Tuesday, June 5, 2012

Beacon with Google AJAXSLT

I was experimenting with XSL transformations on client side so that we could make Beacon embed-able. 
It will be useful if we can pull this off because
  • it will help us concentrate on the Editor itself and leave the authentication and workflow to the services that would like to use Beacon for editing. 
  • This will also help in easier adoption by various organizations as they won't have to set up a different service and look for ways to integrate their workflow with Beacon's. 

jquery.xslt.js (code here)is a JQuery wrapper around the Google AJAXSLT project which fits in in my opinion because Beacon uses JQuery extensively and Google AJAXSLT gives us a browser independent solution that is essential for Beacon to be adopted. 

I looked at the demo code and set up a similar environment in localhost (basically reusing the template XML and  XSL files the docbook plugin folder of Beacon and CSS and media files from Beacon and placing them under /var/www/html along with the test HTML as below:

<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Strict//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-strict.dtd"> <html>
 
<head>
   
<title>XSLT Test</title>
   
<meta http-equiv="Content-Type" content="text/html;charset=utf-8">
   
<link rel="stylesheet" type="text/css" href="docbook.css"/>
 
</head>
 
<body>
   
<div id="output" style="padding: 2px;"> </div>
   
<script type="text/javascript" src="jquery.min.js"> < /script>
   
<script type="text/javascript" src="jquery.xslt.js"> < /script>
   
<script type="text/javascript">
     (function() {

       // The magic

       $('#output').xslt({xmlUrl:'test.xml', xslUrl:
'docbook2html.xsl'});
     });
   
</script>
 
</body>
</html>


The resulting window looked like the screenshot which was really exciting but I was running into the problem that the text in the paragraphs was missing. I added the XSL template for para as below: 

<xsl:template match="para">
 
    <p title="docbookPara">
 
        <xsl:value-of select="." />
 
    </p>
 
    <xsl:apply-templates />
 
</xsl:template>
 

This fixed most of the problems as shown in screenshots here and here. But now I'm stuck at a rather weird issue. The nested tags don't function properly. Trying to come up with the shortest bit of XSL and XML that show the issue, if my XSL is: 
<?xml version="1.0" encoding="ISO-8859-1"?><xsl:stylesheet>
    < xsl:template match="/">
      <xsl:apply-templates select="para"/>
    </xsl:template>
    <xsl:template match="para">
      <p>
        <xsl:value-of select="."/>
        <xsl:apply-templates />
      </p>
    </xsl:template>
    <xsl:template match="bold">
      <b>
        <xsl:value-of select="."/>
      </b>
    <xsl:apply-templates />
</xsl:stylesheet> 


The XML snippet <para> Hello<bold>World</bold>!</para> results in <p>Hello World!<b>World</b></p> 
I tried to find out on #xml on Freenode, and according to them, < ?xml version="1.0" encoding="ISO-8859-1"?> 
<xsl:stylesheet>
  <xsl:template match="/">
      <xsl:apply-templates select="para"/>
  </xsl:template>
  <xsl:template match="para">
    <p>
      <xsl:apply-templates />
    </p>
  </xsl:template>
  <xsl:template match="bold">
    <b>
       <xsl:apply-templates />
    </b>
  </xsl:template>
</xsl:stylesheet> 
should work but removing the <xsl:value-of select="."/> removes the entire text altogether. Any suggestions on how I can fix this? 

Wednesday, November 2, 2011

Talk at FUDCon.in

After working with JBoss and Drools for more than an year at Red Hat, I am giving my first JBoss talk!

I am giving an introduction to JBoss Drools (An open source Business Rules Integration Platform).

As the expected audience is primarily students, I intend to give them a short background of expert systems, give them an idea about inference engines and their applications in the industry using drools as an example.

As the schedule permits only 40 mins for the talk and 10 mins for Q&A, I'll show them how to make a simple application using drools if time permits.

Being a happy drools user for quite sometime, I really hope I am able to convince people to give it a try :)

Oh and now is the time for the uber-cute FUDCon's Happy Guy!

Tuesday, March 15, 2011

mavu.in

Hi All,

My blog has a new home. I've had the domain for quite sometime and I finally decided I'm going to move and be active. See y'all there!

Monday, August 17, 2009

Demo and packaging done

After some re-assessment of our goals for GSoC, we decided to bring out the standalone version of beacon and continue with the integration later.

I got sponsored for sysadmin-test by Fedora-infra and have set the demo up here. To access it, one needs to create an account on http://publictest3.fedoraproject.org/accounts/ and then use that test account.

Also, I finally managed to get over my packaging blues with some kind help from Rakesh Pandit and have submitted a review request here.

On the whole, this GSoC has been fun (I know the project is not finished yet, but the GSoC phase is coming to a close). This year, it was not as much about GSoC as it was about implementing a product which would be guaranteed in use by the community. I interacted with the community more than last time.

Last but not the least, a heartfelt thanks to a very cool mentor, an amazing upstream (esp. Nandeep), and a very supportive community.

Monday, July 13, 2009

Standalone DocBook editor is ready!

Finally finished the DTD for all the tags :) *phew*

Also added a document tree which helps easy navigation, addition and deletion of elements etc. Please check this feature out here. It is really cool.

Checklist: Deliverable 1 done.

For the other two deliverables,test results and documentation, I have an idea. I will write the documentation in DocBook format using Beacon itself.

Note: Since the demo is still in testing do not trust your documents with it yet. We'll make it stable soon. Better yet, it would be great if community could report bugs/suggestions. :)

Saturday, July 11, 2009

XSLs done!

Finally sent my XSLs to Nandeep for review yesterday. He must have committed it by now.
Also found a DocBook CSS which publican uses and replaced the existing one. Using the same CSS might help in integrating later on.

Will finish the DTD today.

Now off to fill the mid term evaluation form...

Tuesday, July 7, 2009

What's Next

Mid-terms evaluations are knocking on door and its been quite a wild ride coding. Web development is a different ball game altogether and I am beginning to understand why they call JavaScript the world's best worst language :D

After I have spruced the Beacon code and finally finished the DocBook DTD with the initial subset, I am going to start the next big task of getting Beacon integrated with Zikula and Publican for a seamless documentation experience.

It looks like Zikula has some text editors and a plug-in for them. I have to find the best way to get Beacon fit in that work flow. Fortunately Beacon already has a feature that allows one to embed it on any page. So it shouldn't be that difficult.

Am still looking for a best way to get Beacon integrated with Publican. It would be nice if the community can suggest a way to do it.

Thursday, July 2, 2009

DocBook Editor 2

Just completed a tiny DTD for DocBook corresponding to the small set of tags for which I had written the XSLs. I am so glad it works! (http://dev.gentooexperimental.org/~n9986/beacon/editor/ works best with firefox)

Now on it is just a matter of adding more elements to get a richer tag-set for DocBook till the mid-term evals. The basic infra is all set up (will try to keep improving usability though).

Also added a brand new and better CSS.

On the other side I think I should try and get more community involvement into developing the editor as the number of DocBook elements is *huge*.

I'll start by explaining what beacon is and how it works.

Beacon is a What you see is what you mean editor which relies heavily on XSL for XML to HTML conversions and vice-versa.

Most of the work involved in making a plugin is writing the two XSLs. One is required for converting the XML to HTML so that it can be displayed in the browser. And the other for converting the HTML back to XML. All these operations are done on the server side using PHP (or Python).

One may wonder, why rewrite XSLs if they already exists, like for Docbook for example. The reason is that Beacon needs some 'hooks' in the generated HTML so it can pick up the nodes using Javascript. This is required to do multitude of tasks like rendering inline editors, maintaining structural sanity, addition/deletion of new nodes, etc. Since HTML and Docbook XML tags are very different, there needs to be some way to map the XML tags over to HTML. These hooks are done using title attribute on tags. So for example, the tag will be rendered in HTML as:

< p title = "docbookPara"> Somethings are best with a title < /p >

Title tags tend to be most unobtrusive here and works well for DOM manipulation which is made even easier thanks to jQuery.

Once the HTML has been rendered, some Javascript magic makes it editable via inline editors of various types. We are still working on the node adding/deleting feature. Once that is done we will have a more or less a complete editor.

The validation of the generated HTML is kept in check via a Javascript based DTD.

So the user cycle of Beacon is:

new document -> XML -> HTML -> XML

All ajax-y communication is handled by JSON (except for file upload which uses an iframe of course).

Since its a web based it relies completely on Ajax for any post first page load. The UI is pretty much like a desktop Application with Tabs, sidebar etc.

To sum it up, a plugin now requires:
  1. Couple of XSLs
  2. Javascript based DTD
  3. Some trivial PHP code
  4. CSS + some template XMLs

I will write a tutorial on how to make a plugin for docbook next. Will be nice if folks chip in to get in as many tags supported as possible.

DocBook Editor 1

It's been very quiet in here. This month has been a marathon of coding and a lot of other events on the home front.

First off, a few non-GSoC related developments:
  • I finished college. (Phew!)
  • I got placed in Red Hat as an Associate Software Engineer. (gleee!)
  • I won the Sun India Code For Freedom contest for my B. Tech Project. (ZOMG!! yes that was COMPLETELY unexpected!)
  • I went out for a small trip this weekend after almost 2 years.
Now to the DocBook Editor...

I started off by learning XSLT (tough!) and making an initial RPM for Beacon and submitted it for review. I created a Feature page and worked on deciding on an initial subset with the help of the Fedora-docs list.

Here are a few links to complement the above:

https://fedoraproject.org/wiki/DocBook_Editor_Documentation
https://fedoraproject.org/wiki/Features/DocBookEditor
https://fedoraproject.org/wiki/DocBook_Editor

Then, I constructed an initial XSL by referring to the example DocBook file in the Fedora Documentation repository and based on the initial tag set that was selected after a community review.

Initially, I fixed a few bugs in Beacon to learn the source code and worked on making it more convenient to add plug-ins by changing the implementation so that beacon needs only a DTD from the plug-in makers (which is us) rather than having us code the whole Javascript. A successful implementation of this would save a lot of time in feature addition and creation of plugins.

Apart from the DTD, I was learning the tools more from implementation perspective like JavaScript and jQuery library used in Beacon.

The last week has been the most exciting. Working with Nandeep Mali, the lead Beacon dev and my point of contact for the Beacon upstream, we got the DTD feature implemented to some extent and its working like a charm. The WYSIWYM is also working and just few more additions need to be done to make this a solid XML editor.

The DTD is actually a giant Javascript object which contains details about every node and its structure (like child, parent, siblings). One may think that why not use an existing WYSIWYG editor like TinyMCE. This is not very feasible because WYSIWYG editors are notorious for
the output they produce. This is what Beacon is trying to avoid by adding a different type of editor. The user will be places with a fill-in-the-blanks type of editor so the generated XML will retain its sanity and conform to Docbook DTD. Makes it easy for the Docs team as well.

Now for the treat: http://dev.gentooexperimental.org/~n9986/beacon/editor/

The above URL is a demo of what we have been working with (I hope the Gentoo URL is not a problem. Beacon was a Gentoo SoC project and I cannot get PHP with JSON and XSLs on my fedora account).

The work till now mainly consisted of writing the XSLs for DocBook and making changes to the beacon core. One final enhancement in the usability of beacon would be a drag and drop feature that has not been committed so far as it is very slow at the moment. Once that is ready it will be very easy to add nodes at any point and still not mess up the corresponding XML.

I have been writing the DTD for DocBook since yesterday and will be able to commit it by very soon. So we can expect another blog post about it very soon.

Has been a fun journey so far. More to come in the next post.

Saturday, May 2, 2009

GSoCer again!

It feels awesome as I type this :) I am a GSoC student for Fedora this year too! I'll be working on Implementing a DocBook plug-in for Beacon. I am being mentored by Yaakov Nemoy and Karsten Wade is my technical guide! I feel honored to work with such amazing people :)

Hope the documentation team finds the editor useful.

Monday, February 23, 2009

What's been up...

Phew! back to blogger after a long time, I must say after trying so hard to get over my hesitation, I am not quite successful in getting myself into blogging.

What has been up recently is that I finally got started with my Red Hat internship. My target is to introduce java profiling in SystemTap. I have been reading the SystemTap code to find out where exactly the tokens are parsed and where is the code translated into the corresponding C code.

To an extent I have been successful thanks to the comments in the main.cxx stating which phase begins when, but I guess the use of kprobes has somewhat hidden how the token representing the probe point is mapped to the location where it is actually found (or maybe it is not and I could not find it).

Also I have found a few C based java application profilers like JMP (TIJMP for the JVMTI supporting versions of java). I am looking at their code to know how to interface with the JVM using such tools.

Looking forward to be able to accomplish the task.

On the Sarai Fellowship front I am still in the process of recording the speech samples for my speech recognition project for the OLPC. Recording speech samples and typing out the text has turned out to be a lot more time consuming and laborious than I initially anticipated (Might also be because I need to arrange my schedule according to the times when the children are available which is becoming difficult because of my classes). I am also very hopeful of getting an XO soon thanks to Sankarshan Sir, Mr. Amit and Mr. Sayamindu.

I am praying it all comes out fine.

Next, I was invited to the National Institute of Technology, Durgapur to talk to the newerbies about how to contribute to FOSS in their annual FOSS festival Mukti.in.

It was an amazing experience. I talked about the philosophy of FOSS, why should they contribute and whats in it for them. Then I moved on to what it takes to be a contributor and how to contribute. The How part was divided into two categories, for those who want to contribute code, and for those who want to contribute by other means like helping out with the artwork, documentation, translations, publicity etc. I also talked about Fedora and OLPC in slight details as I could supplement the information there by my own experience. The presentation slides can be found here.

It was nice to try and answer questions from the audience when they gave me their skill set and asked me where could they fit in. It was also heartening to see people interested in my GSoC project and asking me about system side programming and SystemTap.

Over all, I was satisfied at the end of the talk and felt it was a job well done though I was tensed about screwing up the schedule of the event by extending the talk by about an hour.

Saturday, October 11, 2008

Boot (limn/chart)?

Had drafted a small write-up on bootlimn sometime ago. About what is bootlimn, how does it work, how is it different from bootchart and how to interpret the output of bootlimn.

Just pasting the write-up here for anyone who is interested..

Bootlimn

===========================================

One line description:
It is an analysis and visualization tool for the linux boot process.

===========================================

Working:
Bootlimn uses systemtap[0], a kernel-probing language to extract the
data in an event-based structure where the systemtap scripts probe for
certain functions to be called or a small time-period to elapse before
triggering a corresponding probe handler.
The probe handler contains details as to how to log the information.

This information is stored in XML format for standardization and for
facilitating its use by interested third parties. This information is
parsed using a SAX interface and is used for rendering an SVG image,
whose format is heavily inspired by the svg output of Bootchart[1].

===========================================

Interpretation of results (most important for any user):

An example file is sent along with this text. The XML folder contains
the output as given by the systemtap scripts.

The final bootlimn output consists of an SVG image and five text files.

a) SVG image ( This image has a template similar to Bootchart):

Header: The time shows how long (in seconds) the boot process took.

CPU stats: The first rectangle shows the CPU stats. The pink plot (in
the background) showing the CPU utilization and the blue plot (in the
foreground) showing the CPU throughput.

Disk stats: The second rectangle shows the CPU stats. The pink plot
(in the background) showing the Disk utilization and the green plot
(in the foreground) showing the disk throughput.

The syscalls that have been probed are listed along with their color
coding.

Process Tree: The process tree differs from the classical process tree
in Bootchart in the following ways:

i) The processes are listed in the chronological order of their initial
call and not necessarily as parent child blocks as in bootchart. As the
boot is sequential, a child is never rendered before a parent but the
child and the parent may be separated by a few intermittent processes.
The parent-child relationship is shown by dotted lines connecting the
parent and the child.

ii) All the actions are rendered. But to bring the image to a practical
size, process blocks instead of processes have been used. The processes
with the same name have been merged into a process block (the individual
rectangles in the tree), and all the calls made to the processes in the
process block are rendered sequentially. Hence, one might observe
sys_exit being called more than once on the same process block but the
block might have a sys_clone called before the first exit. The criteria
for trimming the tree can be easily altered to suit various purposes
but changing the condition in the uniqueprocess generator (in the
SVGRenderer.java).

iii) The color code represents the last sys_call that acted upon the
process and not the state directly. This was done because Bootlimn,
unlike Bootchart, does not poll the /proc directory and probes the system
calls instead.


iv) Flexible level of detail: By default, to render the entire image
with manageable dimensions and still be accurate, the timestamps collected
were in milli-seconds. And while rendering each pixel represents 0.1s.
As, no information is discarded while rendering, changing the level of
detail is very easy.
The level of detail in the case of bootlimn is defined by:
The timestamps unit (systemtap offers options to gather timestamps in ns,
ms etc. i.e. by changing the gettimeofday_ms in the systemtap scripts
in stp folder, we can change the level of detail).
The image size and the scale factor in the renderer.java
The scalability of SVG images can be used to keep the image size manageable.

There is no provision to concentrate on a part of boot process and give a
separate detailed view of that part as of now.

b) The text files.
The systemtap scripts are written so as to gather as much information as
possible.As displaying all the details on a graph is not possible, there are
five complementary files that give all the details collected by the systemtap
scripts.

i)The Ioblock.txt gives all the block IO details:
It prints the text output of the ioblock tapset.
It has:
type - whether it was a request for IO or a signal to end
time - timestamp
devname - block device name
ino - i-node number of the mapped file
error - its value is zero on success
sector - beginning sector for the entire bio
flags -
BIO_UPTODATE 0 ok after I/O completion
BIO_RW_BLOCK 1 RW_AHEAD set, and read/write would block
BIO_EOF 2 out-out-bounds error
BIO_SEG_VALID 3 nr_hw_seg valid
BIO_CLONED 4 doesn't own data
BIO_BOUNCED 5 bio is a bounce bio
BIO_USER_MAPPED 6 contains user pages
BIO_EOPNOTSUPP 7 not supported
rw - binary trace for read/write request
vcnt - bio vector count which represents number of array element (page,
offset, length) which make up this I/O request
idx - offset into the bio vector array
phys_segments - number of segments in this bio after physical address
coalescing is performed.
hw_segments - number of segments after physical and DMA remapping
hardware coalescing is performed
size - total size in bytes
bdev - target block device
bdev_contains - points to the device object which contains the
partition (when bio structure represents a partition)
p_start_sect - points to the start sector of the partition
structure of the device

ii) The Perpro.txt gives the per process CPU usage details
It has:
time - timestamp
pid - process id
execname - name of the process
probefunc - the probing function
utime - the user time of the process
stime - the system time of the process

iii) The Process.txt contains the Process details ( the process tree
is derived out of the same XML as this file.
It has:
time - timestamp
pid - process id
ppid - parent process id
execname - process name
probefunc - probing function
pexecname - parent process name
misc - null as of now. any additional information can be added

iv) The Readwrite.txt contains the details of individual system reads
and writes.
It has:
rcount - read count so far
wcount - write count so far
time - timestamp
pid - process id
execname - process name
pexecname - parent process name
type - read or write
file - the file to which data was written or data was read from

v) The Stats.txt contains the CPU and disk statistics. This file is a
direct mapping to the first two rectangles in the image showing CPU and
disk stats.
It has:
time - timestamp
CPUutil - CPU utilization
CPUtput - CPU throughput
diskutil - Disk utilization
disktput - Disk throughput

==============================================
Source:
The svn version of bootlimn can be checked out from [2] and a tarball can
be found at [3].

==============================================

References

[0] http://sourceware.org/systemtap/
[1] http://www.bootchart.org/
[2] http://code.google.com/p/bootlimn/source/checkout
[3] http://code.google.com/p/google-summer-of-code-2008-fedora/downloads/list

============================================================================

Bootlimn on F9

On a hurrah note, finally able to run bootlimn on F9. And am too glad that it required no modification of bootlimn code. I just had to get my systemtap settings right.

The output of the bootlimn from F8 and F9 can be downloaded from here.

Tuesday, September 16, 2008

Project selection

Wanting to continue with systemtap, we came up with this idea of instrumenting XEN or KVM for project 2. Quoting my mentor " writing useful tapsets that one can use to write meaningful scripts to instrument or gather information from the
running guest."

Mr. Masami Hiramatsu Kindly pointed out the VESPER project which is a framework to gather the state of guest kernel.
It looks as if they would like to support systemtap.

The thread on the systemtap mailing list regarding the same can be found here.

I would be really grateful to have some feedback regarding this.

Monday, September 15, 2008

Adieu GSoC and DgpLUG classes; and Hola Red Hat!

Funny, the juggling ended!

I remember trying my best not to miss the DgpLUG classes while struggling to meet my mid term evaluation targets. While I have been blogging about my GSoC project and its updates; I haven't mentioned DgpLUG here till now. Its the Linux Users Group of Durgapur, who came up with this nice initiative for training a few newbies in open source technologies. Thank you folks! It was awesome ^_^

On the GSoC front, its nice to see Bootlimn have so many downloads. I admit that the progress has slowed down a bit given my classes and the recent hunt for a nice and useful project for the Red Hat internship ( yeah! I was as amazed when I was offered and I'm still trying to not sound all stupefied when I am talking about or mailing with regard to it. ) But Bootlimn is far from dead. I'll start working on it again once I settle down with this new routine.

Please do let me know what working well and whats not with bootlimn.

Wednesday, September 3, 2008

oopsey

I am terribly sorry for the multiple bloopers . The code did not get committed last time.
Please checkout the code now.. Revision 18 is the latest.
Sorry again --with an embarrassed look--

checkout : here

Very high on my TODO... port it to Python.. we desperately need that one.. successfully running java the first time seems a miracle now when I get so many head-banging feedbacks saying java not working.

Another development. We were halfway through while trying to run it on Ubuntu today. Though we had to manually modify grub/menu.lst
and had to struggle with java a bit. (JRE ... *bah*)

see this too : http://sourceware.org/systemtap/wiki/SystemtapOnUbuntu

A request: If anyone has any problems or even if you are able to successfully run it, please let me know.

Monday, August 18, 2008

I'm alive

Yes I am :D ... just in case my inactivity here raised any doubts...and I won't surprised if it did because, despite several reminders from my mentor that a blog update has been pending, I have been putting it off for the time that I have something substantial ( or .. was it my laziness? ). Now that the pencils-down date has arrived, I see no further excuse for postponing it.

Since the last time..
  • I got my disk and CPU info without much use of guru mode code (used the queue_stats tapset.. thought had to create my own copy of it where I could change the default time unit to milliseconds instead of microseconds.
The tapset says..
# qstats.stp: Queue statistics gathering tapset

# -------------------------------------------------------------------------

# The default timing function: microseconds. This function could

# go into a separate file (say, qstats_qs_time.stp), so that a user

# script can override it with another definition.

function qs_time () { return gettimeofday_ms () }

# -------------------------------------------------------------------------

Till that is not done.. I might have to stick to my own copy of queue stats.
  • My renderer module is up. Even though it supports only svg for now, I'll extend it to support other formats very soon.
  • I gather per process CPU statistics which show how much system and user time they take (got this idea from bootprobe).
  • I trace sys_open and gather statistics like which process reads/write to what file etc.
  • I also trace the blockIO (in this case I just provide a way to bring out the blockIO information as gathered by the tapset in XML format).
The idea behind tracing points 3, 4 and 5 is to have as much information as possible at least in text format so that even if it cannot be rendered (will terribly clutter the graph if rendered), we can get as much detail as possible.
As all the above information is timestamped, correlation is very easy.
  • The bootlimn (as it has been named tentatively) installs and uninstalls very cleanly.
  • A jar file is packaged along with the source code. It can be run simply by executing ./bootlimn.sh .
  • A build.xml (to be used with ant) is also available.
  • Failure of a part of bootlimn does not crash the entire application. It still tries to give as much output as possible.For example,if one of the XML files cannot be parsed, the others are not affected and neither is the renderer module unless it is *very* critical for the creation of the graph.Even if the XML generated is screwed, the bootlimn still renders till the first occurence of improper entry.
  • The user can specify where to stop by changing the -c option in stpcaller.

Known bugs (Taken care of)):
  • The XML created, sometimes, has negative timestamps. (see update 3)
  • The IOblock parser gives errors at times.(This again is because of the screwed XML) ( see update 1).

What needs more work:
  • All the unique processes are rendered. The user as of now has no control over the degree of detail.
  • The state transitions can be bettered.
  • The CPU wait stats can also be added ( code already present in the stps, just requires slight modification in the XSD and corresponding changes to parser and renderer.. will do it soon)
  • Other formats of images to be supported.( next task at hand after debugging)
  • Header information needs to be added (this will be done soon too)

And anymore that will be suggested when the code is reviewed ( code can be checked out from here ).


Update 1: A temporary workaround is to define the sector, bdev etc fields ( which get some funny values on very rare occasion) as a String type so that just a single instance of screwed up XML does not hinder the parsing of the entire file. Not the best solution but just a temporary workaround as there is no further calculation based on these fields and the only function that uses them is a 'tostring' which converts them to a string anyway.

Update 2: Egads!!! revision 14 is sort of broken.. I am rectifying it.. please don't checkout the code now.

Update 3: The negative timestamps error has been solved ***phew***. The code can now be checked out. Logging has been changed to disk as opposed to in memory (see comments for further details).

Friday, July 4, 2008

Mid term evals!! eeks!

Quoting from my previous post "Now I just need to extend this structure to include the CPU and Disk info."

Well.. it turned out to be anything but just.

For 1, as I am not to use /proc files, that rules out passing those values to my systemtap script using command line arguments.
So, we decided to look into the kernel code and see how those entries are filled in the first place so that we could tap the information from there itself.

From there, I figured out that the files that particularly interest me are
1. /proc/proc_misc.c
2. /proc/array.c
3. /block/genhd.c

Then I mailed the systemtap mailing list and they pointed me to this. It is a very interesting tapset but the following bits left me a little unsure about using it-

Note that blktrace needs to be running in order for these scripts to
have any effect - __blk_add_trace() and therefore the tapset probe
isn't actually called unless tracing is active. Of course, that means
that you need to enable CONFIG_BLK_DEV_IO_TRACE in the kernel.



Things works fine this way, if you're careful - for some reason, if you
don't define overrides for *all* the callback functions, *none* of them
get called. The same thing is apparently true wrt optimization - if any
one of the callback functions gets optimized out, they all do. So in
your script, you need to define handler functions for every event type
whether you use it or not, and furthermore the bodies of the unused
handlers need to contain code that won't be optimized out


Installing blktrace sounded like adding additional baggage to us. Also, I haven't tried out blktrace so I do not know how to use this and whether using it will solve my purpose. I could of course play around with it for sometime but mid term evaluations start on 7th so I probed sys_read and sys_write to count the total reads and writes, and then used guru mode to find out the number of I/O tasks pending at the moment.

Yeah , I know guru mode needs to be used only when there is no other alternative . I couldn't see any. Any suggestion would be very welcome.

Thankfully, it works and I have all the diskstats that I require.So far so good.

But CPU information is turning out to be trickier. The variables that I require are not accessible.

A stap -p2 -e 'probe kernel.function("do_task_stat") {$foo}' -u ( A trick I learnt from fche )
gives the alternatives as : task buffer whole vsize eip esp wchan priority nice tty_pgrp tty_nr sigign sigcatch state res ppid pgid sid num_threads mm start_time cmin_flt cmaj_flt min_flt maj_flt cutime cstime utime stime cgtime gtime rsslim tcomm flags ns

But the variables utime, stime are defined within the function and a
stap -p2 -e 'probe kernel.function("do_task_stat") {$utime}' -u
gives semantic error: not accessible at this address: identifier '$utime' at < input > :1:40

I guess the only variables accessible are the parameters that are passed to the function and the return value as the probe is places at the start of the function definition. and the variables I need have their values worked out inside the function and they get seq.printed from there itself .. so these values are not even returned.

The kernel.statement construct is not what I can use as absolute addresses may be different on different computers.

As of now .. I am trying a very very ugly solution..

Edit 1: License stuff :D
This program is free software; you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation; version 2 of the License.

This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.


%{
#include< linux/kernel_stat.h >
#include< linux/sched.h >
%}

function get_str_info:long(val:long)
%{
struct task_struct *temp = (struct task_struct*)(long)THIS->val;
long x =(long)((temp->state & (0| 1 | 2 | 4| 8))| temp->exit_state);
THIS->__retvalue = x;
%}

function get_u_info_1:long(val:long)
%{
struct task_struct *temp = (struct task_struct*)(long)THIS->val;
cputime_t ut;
long usr;
struct signal_struct *si = temp->signal;
struct task_struct *t = temp;
do {
ut = cputime_add(ut,t->utime);
t = next_thread(t);
} while (t != temp);
ut = cputime_add(ut, si->utime);
usr = (long) ut;
THIS->__retvalue = usr;
%}


function get_u_info_2:long(val:long)
%{
struct task_struct *temp = (struct task_struct*)(long)THIS->val;
cputime_t ut;
long usr;
ut = temp->utime;
usr = (long)ut;
THIS->__retvalue = usr;
%}

function get_s_info_1:long(val:long)
%{
struct task_struct *temp = (struct task_struct*)(long)THIS->val;
long sys;
struct signal_struct *si = temp->signal;
struct task_struct *t = temp;
do {
st = cputime_add(st,t->stime);
t = next_thread(t);
} while (t != temp);
st = cputime_add(st, si->stime);
sys = (long) st;
THIS->__retvalue = sys;
%}

function get_s_info_2:long(val:long)
%{
struct task_struct *temp = (struct task_struct*)(long)THIS->val;
cputime_t st;
long sys;
st = temp->stime;
sys = (long)st;
THIS->__retvalue = sys;
%}

probe kernel.function("do_task_stat")
{
if ($whole)
{
u = get_u_info_1($task)
s = get_s_info_1($task)
state = get_state_info($task)
}
else
{
u = get_u_info_2($task)
s = get_s_info_2($task)
state= get_state_info($task)
}
}


Of course the above code needs to be debugged for successful compilation : function get_state_info:long works but the rest need some more work ( probably due to the calls to other inline functions: next_thread()and cputime_add() within the function ). The reason I am putting up the unfinished code is because I need to know whether its worth spending time debugging it.Is it the right way to proceed?

What I am basically trying to do is.. take the values of task and whole (they are passed as parameters to the function do_task_stat and hence are available. From there I calculate the variables I needed the same way as is being done inside the original function. *very ugly* but am at a loss of ideas (for now at least).

Anyway, I'll just work on this today and tomorrow. If I am unable to find a better way, I guess I'll move on to the cleaning up tasks so that my Bootchart is perfectly ready with the process and disk info at least.

Then I can come back to this after mid-term evaluations.

Saturday, June 21, 2008

Current status of the project

Here is what I have been up to till now..

I have my Bootchart prototype finally running.

What all do I have till now?

1. An install script.
2. A shell script that runs instead of init, runs my systemtap script and then calls init.
3. My systemtap script, of course that does the probing for me and gathers the process info and outputs it in XML format.
4. A script that calls my parser and saves the text output ( the renderer is to be built after mid-term evaluation)
5. A java parser that parses the XML output using the SAX interfaces.

The parsed text-mode output looks like:

No of Processes '1249'.
Process Details - Time1214080512744, Pid451, PPid:450, Execname:sh,Probefunc:sys_execve,Pexecname:stapio.
Process Details - Time1214080516148, Pid1, PPid:0, Execname:init,Probefunc:sys_execve,Pexecname:swapper.
Process Details - Time1214080516167, Pid454, PPid:453, Execname:init,Probefunc:sys_execve,Pexecname:init.
Process Details - Time1214080516203, Pid455, PPid:454, Execname:rc.sysinit,Probefunc:sys_execve,Pexecname:rc.sysinit.

......
Note: Processes are not unique.

Now I just need to extend this structure to include the CPU and Disk info.

Another thing I'd like to mention, mounting a temporary file system is necessary even if I log in memory and dump the output in a file at the end of probing. A 'probe end {//do the logging in file }' results in a read-only file-system error
even though you not writing to the file till the system is up.

Thursday, June 5, 2008

quick update

okay.. so I coded the XML parser to retrieve the Process information using SAX interface.. (suits my purpose better than DOM)..

I need to modify the renderer to use this now.. and then integrate it to bootchartd.. then 1/3 of my project will be done.. or so I hope.

Writing a mail to my mentor now... will put up details later.