Skip to content

Safe and Efficient Message-Passing Control in Pharo with MethodProxies

We recently had accepted our paper about Method proxies on the COLA (Computer Languages) journal!

For those that want the full read, here is the link.

For those that just wants a glimpse, here is a 5-10' read!

The Problem: instrumenting code is difficult

With Sebas we wanted to instrument several parts of the Pharo runtime, including object allocations, deallocations and so on, with the goal of building tools. More specifically, we were interested on building profilers: tools that observe the execution of the program to give you insights about it's behavior. Why is it slow? Is it doing something strange?

The issue is, such instrumentation does not come without a cost. The easiest to see is performance. If you want to track all object allocations, that will make all allocations slower. Now think about tracking all message sends...

Moreover, in Pharo we are typically used to write the instrumentation in Pharo itself, which can lead to some circularity issues. For example, consider the following code: we want to log all sends to the OrderedCollection>>add: method.

logger := SendLogger logDuring: [
    aCollection := OrderedCollection new.
    aCollection add: 17
].

self assert: logger trappedMethods includes: (OrderedCollection>>#new)

The problem is, that the logger itself may use that method too!

SendLogger >> trapSend: aMessage
    trappedMethods add: aMessage method.

And that may provoke a recursion like the following one.

collection>>add: ---trap---> logger>>trapSend: --> collection>>add: --trap--> logger>>trapSend: ...

To handle this, you need to detect you're in the logger, and avoid the recursive trapping. Of course, this is just more expensive checking at runtime!

And worse, with the current techniques in Pharo, except for reflectivity, such recursion needs to be handled by the user.

The Goal: Safe and Efficient by Default

We wanted then to try to avoid such nasty things by default. It's not the user that should specify how to do it.

So we built a stratified framework. Just a fancy way to say "we designed a framework with separation of concerns in mind". That is, at the framework level, what the user uses but does not touch, we handle the tricky stuff. Then the user just defines their own trapping behavior. And that should work.

Concretely, it's as follows: the framework defines method proxies and an abstract handler class. Users inherit from the handler class to define their own trapping behavior. It's the MpMethodProxy class that manages: trapping performance, circularity and recursions, efficient exception handling...

Side effects

How to use it?

Let's say you want to create a callgraph. We then first define a handler class. The idea is that this handler will be invoked for every interesting method invoked. When the handler is invoked we will store the executed method in a stack. When the method returns/exits we will pop it from the stack. And during this process we will link methods as a call chain.

MpHandler << #CallGraphHandler
	slots: { #stack . #instrumentedMethod }; 
	package: #MethodProxyExamples

CallGraphHandler >> beforeMethod
	| childNode |
	childNode := stack top
		addChildFor: instrumentedMethod.
	stack push: childNode.

CallGraphHandler >> afterMethod
	stack pop

CallGraphHandler >> initialize
	super initialize.
	stack := Stack new.

And then this is used just as follows.

  1. We instantiate the handler.
  2. We create a proxies for each method in our application.
  3. We install the proxies
  4. We run our app
  5. We uninstall the proxies
  6. We enjoy our callgraph
handler := CallGraphHandler new.

proxies := MpMethodProxy 
	onMethods: MyApplication definedMethods
	handler: handler.

proxies do: #install.

"Call the instrumented methods"
MyApplication start.

proxies do: #uninstall.

"Program's analysis"
handler stack top. "call graph root node"

A Note on Thread Safety

Important to notice! There is a single handler object, which is shared by all proxies! So that means that if two threads execute a trapped method, both will go into the handler, and magical and wonderful things may happen.

First things first: method proxies is designed to be thread safe. That means, having two threads calling proxies will not break the proxy infrastructure (and if it happens, it is a bug).

However, the handler we defined above is not thread safe. And that may have different consequences such as having a stack that mixes calls from different threads, giving us a wrong callgraph, or in the worst case, the stack getting corrupted because of a race condition.

The first issue may happen because our model is wrong: there is no one call graph. There is a callgraph per thread. Which means that maybe we want to have one stack per thread.

The second issue calls some synchronization in the handler. That is up to now a responsibility of the developer.

Numbers?

Sebas did lots of different benchmarks and studied this in depth. In a gist, the important conclusions are:

  • Practical overhead. Adding proxies in all your package with a handler that does nothing has an overhead ranging between 1.04x and 2.15x depending on the program
  • The main overhead is on the handler. A naïvely written handler may make performance drop by a lot. We have cases were it makes an app 8x slower, because the handler allocates lots of objects and puts a lot of GC pressure

More on the paper

Of course, if you're interested, there is more in the paper.

  • more bench numbers
  • details about safety and unwind/exception handling
  • implementation details