There is a common misconception that is taught in almost every object oriented programming course and book. This misconception is that the real world is composed of macroscopic objects. The common example is a car. It has properties, like make, model, year, color, and so on. It has actions, called methods in OOP, like accelerate, brake, blink left, and more. A car is an object.
Object modeling has a strict definition for objects. If it has properties and actions, then it is an object. Essentially, properties are nouns and adjectives, and actions are verbs. By this definition, wind is an object. It has properties like temperature, direction, and speed. It has actions, like changing speed or direction and knocking things down. In real life terms though, wind is not an object. Using the object model, wind is actually more like a method of the air object. It is an action that air takes. The object model used in object oriented programming is not a reflection of real life applied to programming. It is an abstraction of human reasoning and abstraction applied to programming.
In real life, there is no solid definition of what is an object. The Merriam-Webster definition requires physical objects to be detectable with the senses, which eliminates anything microscopic, no matter how well it fits the object model definition. It also includes things, like wind, which most people would not normally consider objects. Besides all of that, it is too vague to be truly useful.
For most intents, physical objects are defined by their chemical bonds. The difference between two 1oz chunks of steel and one 2oz chunk is that one is coherently bonded in its entirety, while the other is two pieces that are coherently bonded but not to each other. The two pieces could be changed from two objects to one object by welding them together. Unfortunately, this is not a strict rule. A car is composed of many coherent but separate pieces, but it is still considered a single object.
We might define an object based on functional coherence. A board game, composed of a board, game pieces, and a pair of dice, is a single object, because all of the parts serve a single functional purpose: to play the game. A car, composed of an engine, wheels, a chassis, and so on, serves a single purpose of transportation, so it is a functionally coherent object. Except, what about the radio? That does nothing for transportation. It is still considered part of the object. It is not chemically bonded to the object, nor is it a functionally coherent part of the object, so why is the radio part of the car? Obviously functional coherence is not the answer.
Maybe we could add that something which is attached to the object is part of it. This could be viewed as a sort of non-chemical physical continuity. The radio is firmly attached to the car, probably with screws or some kind of snap in tabs. Does this mean that if my son puts super glue on my chair, and I sit on it, I become part of my chair (or maybe my chair becomes part of me)? Again, a dead end.
Even containment does not work. I can be inside my car without being part of it. Putting a CD in a computer does not make it become part of the computer (though putting a hard drive in a computer does...). Even this depends on the specifics.
The problem is that the idea of an object is just that: an idea. Objects are not real things. We perceive things as objects, because we perceive some kind of relational continuity between their parts. This is highly subjective. One person might consider the air filter or battery in the car to be part of the car object, while another might consider them to be separate objects that happen to be used in and by the car. Further, all objects are composed of other objects, until they are not anymore. Current science seems to hold that down to a quantum level all particles have properties and methods (as interactions). In reality though, as you approach sub-atomic levels, the idea of particles seems to break down entirely, giving way to matter that acts like particles but seems to be composed entirely of waves (or one giant wave) with complex harmonics. We cannot say with any degree of certainty that the object model can even be applied to particles smaller than atoms.
The point is, the object model does not reflect real life at all. It merely reflects how most humans tend to perceive real life. Now, this is still very useful, but it exposes some limitations. Because perception is subjective, two people may not always agree on how a specific idea should be modeled as an object or collection of objects. If a battery is part of the car, maybe the car should have a "charge" property to track the charge of the battery. If it is not, maybe we need a whole new "battery" object, with its own "charge" property, and then the car can have a property that points to that battery. Both models have value in specific cases, so there is no common-sense solution. In addition, not all ideas even make sense as objects. Many object oriented programs include objects that are totally incoherent collections of properties and methods that just needed to go somewhere. Besides being bad programming practice (though sometimes necessary in languages that force everything to be part of an object, like Java), this violates the entire purpose of objects. We find this sort of thing in real life all the time though. Many households have a "junk drawer" for holding miscellaneous objects that have nowhere else to go.
While the car analogy might make a good start for helping people to understand objects, very early on it would be a wise idea to stop comparing them to real life objects and start describing them as organizational units. The object model is not about modeling real life. Objects are called objects because it helps people understand the coherency that objects should have. Beyond that though, they resemble real life very poorly, and if the comparison with real life is not left behind, it is likely to result in poor programming practices and absurd arguments over design where real life is used as justification for choosing the worst solutions. In the object model, objects are containers. They are organizational units. Poor organization is poor organization, no matter how people choose to do things in real life. The car analogy is great, but when students start thinking they need to make a "FuelTank" object (and don't forget "FuelPump" and "FuelFilter"), when merely keeping track of current fuel level would be sufficient, the analogy has gone too far.
Object oriented programming is powerful and useful, but it is also far more prone to error than procedural programming. The errors found in object oriented programming are not the kind that break the program, however. They are the kind that waste resources and make maintaining the program very difficult. One of the most effective solutions to this problem would be to reveal objects for what they are, organizational units, instead of comparing them to things that they only marginally resemble, real life objects.
Showing posts with label OOP. Show all posts
Showing posts with label OOP. Show all posts
Wednesday, March 25, 2015
Thursday, February 5, 2015
Inheritance Models
Inheritance models are a controversial subject in Object Oriented Programming language design. The big controversy is multiple inheritance. Python adds a minor controversy on the value of non-public variables and methods.
Multiple inheritance is controversial because multiple parents with methods of the same name implemented can result in naming collisions. Most OOP languages deal with this using priorities. The order of inheritance determines which method is used. The problem with this occurs when the method needs to do different things depending on context. The programmer can always add flexibility, but this is not always desirable. The argument in favor of multiple inheritance is that this is very rare, and when it does occur, it is probably a sign that the program is poorly designed and should be re-factored before continuing.
Java argues that multiple inheritance is rarely necessary, so it eliminates it with a solution that proves that multiple inheritance is indeed frequently necessary. Java solves the multiple inheritance problem by adding a new kind of class to inherit from, called an interface. Any number of interfaces can be attached to a class. Interfaces have some important restrictions. They cannot contain any implemented methods, and all of their methods and member variables must be public. They are essentially fully abstract classes, where all methods are abstract, with a few added restrictions. In reality, Java does support multiple inheritance, it just limits it. Any class can inherit from any number of other classes, on the one condition that only one of those classes is allowed to have implemented code in it. In theory, this solves the naming collision problem. In reality though, it only solves part of it. Two implemented methods with the same name will never be inherited by the same class in Java, so it will never have to determine priority. One implemented method can collide with any number of abstract methods though, and any number of abstract methods can collide with each other. The problem here occurs when two methods with the same name are expected by the interface to do completely different things. This, perhaps the more common problem, is still unresolved. Java not only fails to eliminate multiple inheritance, it fails to eliminate the biggest problem with it.
The evidence that multiple inheritance is necessary and valuable is seen clearly in the fact that Java classes frequently implement one or more interfaces, in addition to inheriting from a parent class. Java's solution to the collision problem does reveal one thing about multiple inheritance: It is probably better to only inherit from one class with implemented methods, and even when violating this, it is almost always better to choose a design that does not require inheriting in ways that cause naming collisions. Java reveals another thing though: There is no reasonable way to entirely eliminate the possibility of naming collisions caused by inheritance. Naming collisions caused by inheritance could be handled by always requiring some kind of prefix for accessing inherited methods and variables, but this would compromise polymorphism, dramatically reducing the value of OOP.
Java's solution to the problems of multiple inheritance is not necessarily bad. It is misleading to claim that it solves all or even most of the problems, but it does solve one, and it encourages the use of naming conventions that are less likely to result in collisions. Overall, it mitigates the problems more than offering a real solution. The problems with multiple inheritance are inherent. They are not related to how it is implemented. Java's mistake was in trying to solve an unsolvable problem, but there is significant learning value in the attempt, so it was not a waste of time or effort.
The value of non-public variables is a different but related matter. What creates the controversy is that it is an entirely artificial limitation. Private and protected variables exist solely to enforce best practices. There is no platform level support for them, and there never will be, because it is already expensive enough to design processors with one protected mode (and sometimes a third access level exists for hardware interrupts). Adding widespread access control to hardware just to enforce good programming practices would be absurd. As such, variable access control in OOP languages is enforced by the compiler or the interpreter. The reason the artificiality matters is that this kind of access control only restricts accessing variables by name. A private variable is no different from any other variable. The only difference is that it can only be accessed by name when it is in scope, and it is only in scope in the methods of the class it belongs to. In any language that allows direct memory manipulation (including through external libraries), accessing private variables directly is fairly simple, if the programmer has a half decent understanding of the object containing it. In reference to the inheritance discussion above, this means that requiring interface methods and variables to be public is merely a semantic thing, not a significant difference between fully abstract methods and interfaces.
The artificiality of variable access restriction by no means devalues it though. Offering programmers some guidance by making it difficult to violate best practices does have value. Access restriction is not the only way to accomplish this though. C++ and Java use access restriction, but Python does not. Nothing in Python is ever strictly private. If you know its name, you can access it. Python does have some ways of making it hard to violate best practices though. A single leading underscore in a variable or method name in Python indicates that it should not be accessed directly, because it is an implementation detail that is subject to change. Two leading underscores, with one or no trailing underscores in Python invokes name mangling that makes direct access from outside the class difficult. Every method or variable reference inside the class definition with the two leading underscores and less than two following will be mangled, which will allow internal references to work properly, but it will break external references. Even this does not make the variable or method private though. The mangling is predictable, so a determined programmer could still access mangled references from outside the class, by using the mangled name. This requires very deliberate action though, and thus will never occur accidentally. A programmer that is willing to go this far to violate best practices would probably also use direct memory manipulation, so the fact that it is easier this way hardly matters.
To be clear, method and instance variable access restriction should never be regarded as a security measure. In any well designed language, there will always be a way around it. Access restriction should be regarded as a means to encourage best practices and provide a warning for attempts to violate them. It should never be treated as anything more than this.
The connection between these two is the fact that access restriction is not a significant difference between two meta data-types. Despite the difference in access restriction between fully abstract classes and interfaces in Java, they are functionally the same thing. Java differentiates to hide the multiple inheritance, and that is it. In fact, considering the value of access restricted variables and methods in fully abstract classes (hint: there is no value), it makes perfect sense to disallow anything but public variables and methods in interfaces. To do anything else would be to restrict or dictate the implementation of the abstract variables and methods, which would directly violate the purpose of fully abstract classes as well as interfaces. This is not a major difference between the two. It is merely another case of the language enforcing best practices.
In the end, Java does not solve the multiple inheritance problem, and access restricted variables and methods are not strictly necessary. In both cases, this does not mean that time or energy was wasted though. They both have value, but they are related in that they are both misunderstood, often even by professionals.
Multiple inheritance is controversial because multiple parents with methods of the same name implemented can result in naming collisions. Most OOP languages deal with this using priorities. The order of inheritance determines which method is used. The problem with this occurs when the method needs to do different things depending on context. The programmer can always add flexibility, but this is not always desirable. The argument in favor of multiple inheritance is that this is very rare, and when it does occur, it is probably a sign that the program is poorly designed and should be re-factored before continuing.
Java argues that multiple inheritance is rarely necessary, so it eliminates it with a solution that proves that multiple inheritance is indeed frequently necessary. Java solves the multiple inheritance problem by adding a new kind of class to inherit from, called an interface. Any number of interfaces can be attached to a class. Interfaces have some important restrictions. They cannot contain any implemented methods, and all of their methods and member variables must be public. They are essentially fully abstract classes, where all methods are abstract, with a few added restrictions. In reality, Java does support multiple inheritance, it just limits it. Any class can inherit from any number of other classes, on the one condition that only one of those classes is allowed to have implemented code in it. In theory, this solves the naming collision problem. In reality though, it only solves part of it. Two implemented methods with the same name will never be inherited by the same class in Java, so it will never have to determine priority. One implemented method can collide with any number of abstract methods though, and any number of abstract methods can collide with each other. The problem here occurs when two methods with the same name are expected by the interface to do completely different things. This, perhaps the more common problem, is still unresolved. Java not only fails to eliminate multiple inheritance, it fails to eliminate the biggest problem with it.
The evidence that multiple inheritance is necessary and valuable is seen clearly in the fact that Java classes frequently implement one or more interfaces, in addition to inheriting from a parent class. Java's solution to the collision problem does reveal one thing about multiple inheritance: It is probably better to only inherit from one class with implemented methods, and even when violating this, it is almost always better to choose a design that does not require inheriting in ways that cause naming collisions. Java reveals another thing though: There is no reasonable way to entirely eliminate the possibility of naming collisions caused by inheritance. Naming collisions caused by inheritance could be handled by always requiring some kind of prefix for accessing inherited methods and variables, but this would compromise polymorphism, dramatically reducing the value of OOP.
Java's solution to the problems of multiple inheritance is not necessarily bad. It is misleading to claim that it solves all or even most of the problems, but it does solve one, and it encourages the use of naming conventions that are less likely to result in collisions. Overall, it mitigates the problems more than offering a real solution. The problems with multiple inheritance are inherent. They are not related to how it is implemented. Java's mistake was in trying to solve an unsolvable problem, but there is significant learning value in the attempt, so it was not a waste of time or effort.
The value of non-public variables is a different but related matter. What creates the controversy is that it is an entirely artificial limitation. Private and protected variables exist solely to enforce best practices. There is no platform level support for them, and there never will be, because it is already expensive enough to design processors with one protected mode (and sometimes a third access level exists for hardware interrupts). Adding widespread access control to hardware just to enforce good programming practices would be absurd. As such, variable access control in OOP languages is enforced by the compiler or the interpreter. The reason the artificiality matters is that this kind of access control only restricts accessing variables by name. A private variable is no different from any other variable. The only difference is that it can only be accessed by name when it is in scope, and it is only in scope in the methods of the class it belongs to. In any language that allows direct memory manipulation (including through external libraries), accessing private variables directly is fairly simple, if the programmer has a half decent understanding of the object containing it. In reference to the inheritance discussion above, this means that requiring interface methods and variables to be public is merely a semantic thing, not a significant difference between fully abstract methods and interfaces.
The artificiality of variable access restriction by no means devalues it though. Offering programmers some guidance by making it difficult to violate best practices does have value. Access restriction is not the only way to accomplish this though. C++ and Java use access restriction, but Python does not. Nothing in Python is ever strictly private. If you know its name, you can access it. Python does have some ways of making it hard to violate best practices though. A single leading underscore in a variable or method name in Python indicates that it should not be accessed directly, because it is an implementation detail that is subject to change. Two leading underscores, with one or no trailing underscores in Python invokes name mangling that makes direct access from outside the class difficult. Every method or variable reference inside the class definition with the two leading underscores and less than two following will be mangled, which will allow internal references to work properly, but it will break external references. Even this does not make the variable or method private though. The mangling is predictable, so a determined programmer could still access mangled references from outside the class, by using the mangled name. This requires very deliberate action though, and thus will never occur accidentally. A programmer that is willing to go this far to violate best practices would probably also use direct memory manipulation, so the fact that it is easier this way hardly matters.
To be clear, method and instance variable access restriction should never be regarded as a security measure. In any well designed language, there will always be a way around it. Access restriction should be regarded as a means to encourage best practices and provide a warning for attempts to violate them. It should never be treated as anything more than this.
The connection between these two is the fact that access restriction is not a significant difference between two meta data-types. Despite the difference in access restriction between fully abstract classes and interfaces in Java, they are functionally the same thing. Java differentiates to hide the multiple inheritance, and that is it. In fact, considering the value of access restricted variables and methods in fully abstract classes (hint: there is no value), it makes perfect sense to disallow anything but public variables and methods in interfaces. To do anything else would be to restrict or dictate the implementation of the abstract variables and methods, which would directly violate the purpose of fully abstract classes as well as interfaces. This is not a major difference between the two. It is merely another case of the language enforcing best practices.
In the end, Java does not solve the multiple inheritance problem, and access restricted variables and methods are not strictly necessary. In both cases, this does not mean that time or energy was wasted though. They both have value, but they are related in that they are both misunderstood, often even by professionals.
Wednesday, October 22, 2014
C Programming: Encapsulation
Encapsulation is perhaps the most valuable aspect of object oriented programming. Without encapsulation, object orientation would not have been useful enough to become popular, in a large part because most of the other useful aspects of object orientation rely heavily on encapsulation. Encapsulation is the intuitive grouping of related data into coherent blocks. In C, the most common form of encapsulation is grouping related functions into libraries. This only brushes the surface of the possibilities though, because a library is only a single instance of an encapsulated entity, and additional instances cannot be made. In object oriented programming, encapsulation is typically used to group related variables together with functions used to manipulate them. Encapsulation's strength is that it can be used to logically order data and operations in a way that is easy to understand and remember.
Encapsulation is not inherently beneficial in programming. It does not improve program efficiency, and it often harms it. Encapsulation can improve development speed and make program source code easier to understand. This can substantially reduce the time required to create a program, and it can also make maintenance far easier. Object oriented programming languages and styles are primarily popular because encapsulation increases development speed, thus increasing potential profits. The benefits of encapsulation are primarily business benefits, not benefits to the program itself.
Encapsulation comes at a heavy cost. While it may be worth the cost to gain the benefits, it is still important to be aware of the costs. Encapsulation almost always results in both memory and performance costs. In object oriented languages, objects must store function pointers in addition to data. These function pointers take up memory, and large numbers of objects or objects with many functions can take up substantial amounts of memory. Contrasted with purely procedural languages, where each function call is explicitly stored in the code only once, object orientation can be very inefficient in memory usage. To add to this, many object oriented languages store additional information about objects, such as object type, even in the compiled program, which uses even more memory. This, however, is not the worst part.
Encapsulation changes how data is stored in memory. On older computers and embedded systems, this effect may be negligible, but on modern systems with advanced memory caching, it is very important. Processor caches are used to avoid slow memory access by storing frequently used data on very fast memory on the processor. This memory is called the cache. Modern processors use additional techniques to further improve cache performance. One of these is to load memory into the cache in chunks. These chunks will always contain the memory that is being accessed, but they also contain nearby sections of memory that are likely to be used by the program in the near future. Since modern processors are many times faster than memory, caching is important to get the full performance of the processor. However, how data is used in a program dramatically affects how efficiently the processor cache is used. A program that jumps around memory, accessing data from widely separated areas, will force the cache to load new data from memory (discarding the old data) very frequently. When data is not found in the cache and must be loaded from memory, this is called a cache miss. Cache misses take a long time, during which the processor will either be idle or executing a different program, reducing the performance of the program waiting for the data to be loaded. When a program loops through a contiguous array of data, cache misses are very infrequent, and performance is maximized. When a program jumps around, performance is dramatically reduced. This is important because organization of data in object oriented programs is different from how data is typically organized in procedural programs. Each object stores its data in a single location. An array of objects, even in contiguous memory (most OOP languages store objects on the heap, which is not guaranteed to be contiguous), that is looped through to access only one or two member variables will still load the entire objects into the cache. This causes a lot of unused data to cycle through the cache, displacing data which would have been used. The result is a dramatic increase in cache misses, which substantially reduces performance. A similar procedural program might store the "object" data in multiple arrays, only accessing (and thus caching) the data in the arrays that are actually used. This adds another benefit. Processor caches are divided into sections, where each section of the cache holds some section of memory. These sections are rotated through (typically the least recently accessed is overwritten by new data being loaded). Looping through two arrays at once can take advantage of this when one cache section holds part of one array and another holds part of another array. Using encapsulation, an array of objects will likely use only one cache section at a time, constantly missing and loading more data. Now, we have only looked at how encapsulation affects cache usage when an array of objects is contiguous in memory. This is not typical. Most object oriented languages dynamically allocate memory for objects, and dynamically allocated memory is often not contiguous. Looping through non-contiguous memory will cause cache misses at almost every iteration, causing severe performance loss. The costs of encapsulation are sometimes justified, but it is important to be aware of them when making design decisions. When using object oriented programming, it is important to understand that OOP is a human invention designed to making interfacing with computers on a low level easier. Computers do not "think" in objects, so there will always be costs to the translation between human ideas of objects and how computers actually work. Understanding the underlying architecture can help mitigate the costs of encapsulation, but it cannot eliminate them entirely.
Now we can discuss how to use encapsulation in C. C already has primitive support for encapsulation, which we can leverage to implement full encapsulation. Note that this is going to be ugly and unwieldy, and it should typically be avoided wherever possible. As an exercise in understanding the C language, however, this may be very useful.
C has a meta data type called a struct. Structs are used to create composite data types that essentially encapsulate data. Structs can contain only data. They cannot contain functions or code. Following is a simple example of a struct.
This code defines a struct of type "item". An instance of this struct could be created and manipulated with the following.
Now that we have a struct for our inventory items, we might want to reuse this idea in our POS system. We might want a struct for transactions. A transaction might contain a pointer to an array of items (we cannot make the array static, since it may have a different size for each transaction). We could use the previous struct for items, but we would use the count element as the quantity purchased instead of inventory. When finalizing a transaction, we might need to calculate sales tax, so a transaction will need a sales tax as well as a function for calculating tax and setting the variable in the struct. Normally, we would just write a function that takes the struct and modifies it. Maybe we want to be able to use different algorithms for sales tax depending on the customer (business customers might have a lower tax rate, while out of state and government customers might be exempt), but we want to be able to treat all transactions the same. This gets into very basic polymorphism, which we will discuss in depth in a different article. For now, however, we need our transaction to use a different tax algorithm for different customers, and we do not want to have to keep track of tons of information to accomplish this. We want to be able to populate most of a transaction, then send it to a function to finalize it that can use the same procedure on all transactions.
While structs cannot contain functions, they can contain function pointers. So long as the functions pointed to have the same signatures (return values and argument lists), they can easily be called in another function that knows how to use them. We could create several functions for calculating tax, and then we could store a pointer to the appropriate function in the struct. Later, when we finalize the transaction, we can just call the function pointed to by the transaction, and it will calculate tax for us. We do not have to care which algorithm is used during finalization. Here is a very simple example of this.
These two functions return void and take transaction pointers, just like the function pointer in the struct definition. The first calculates a 5% sales tax, while the other sets tax to 0 for tax exempt customers. Now, we want to create a transaction. This will be a normal customer, with normal sales tax, who is making a purchase that totals to $10.00 (since we are using ints to store cents, it will be 1000).
There is our transaction. If we were serving a tax exempt customer, we could set gettax to *exemptTax instead. Now we are ready to calculate tax. This can be done with the following line, regardless of what tax algorithm we are using.
This example is obviously contrived, and in real life we probably would have explicitly called a different function for each tax mode. In fact, this would probably be a better way to do it for this situation, but this sort of encapsulation has its strong points in many other applications.
Again, this is an ugly and unwieldy way of implementing encapsulation. If you really need encapsulation, it would probably be better to use an object oriented language like C++. In the rare situation where that is not an option or where you only need very basic encapsulation and only to a very limited degree, this might be appropriate. If you are considering doing this, first ask yourself why you are using C in the first place. It is very likely that the reason you are using C is because object oriented programming is much more expensive in memory and performance. If this is the case, you should probably find a more efficient way of solving your problem.
Here is the source code for a simple C program that implements and uses the transaction example from above:
Encapsulation is not inherently beneficial in programming. It does not improve program efficiency, and it often harms it. Encapsulation can improve development speed and make program source code easier to understand. This can substantially reduce the time required to create a program, and it can also make maintenance far easier. Object oriented programming languages and styles are primarily popular because encapsulation increases development speed, thus increasing potential profits. The benefits of encapsulation are primarily business benefits, not benefits to the program itself.
Encapsulation comes at a heavy cost. While it may be worth the cost to gain the benefits, it is still important to be aware of the costs. Encapsulation almost always results in both memory and performance costs. In object oriented languages, objects must store function pointers in addition to data. These function pointers take up memory, and large numbers of objects or objects with many functions can take up substantial amounts of memory. Contrasted with purely procedural languages, where each function call is explicitly stored in the code only once, object orientation can be very inefficient in memory usage. To add to this, many object oriented languages store additional information about objects, such as object type, even in the compiled program, which uses even more memory. This, however, is not the worst part.
Encapsulation changes how data is stored in memory. On older computers and embedded systems, this effect may be negligible, but on modern systems with advanced memory caching, it is very important. Processor caches are used to avoid slow memory access by storing frequently used data on very fast memory on the processor. This memory is called the cache. Modern processors use additional techniques to further improve cache performance. One of these is to load memory into the cache in chunks. These chunks will always contain the memory that is being accessed, but they also contain nearby sections of memory that are likely to be used by the program in the near future. Since modern processors are many times faster than memory, caching is important to get the full performance of the processor. However, how data is used in a program dramatically affects how efficiently the processor cache is used. A program that jumps around memory, accessing data from widely separated areas, will force the cache to load new data from memory (discarding the old data) very frequently. When data is not found in the cache and must be loaded from memory, this is called a cache miss. Cache misses take a long time, during which the processor will either be idle or executing a different program, reducing the performance of the program waiting for the data to be loaded. When a program loops through a contiguous array of data, cache misses are very infrequent, and performance is maximized. When a program jumps around, performance is dramatically reduced. This is important because organization of data in object oriented programs is different from how data is typically organized in procedural programs. Each object stores its data in a single location. An array of objects, even in contiguous memory (most OOP languages store objects on the heap, which is not guaranteed to be contiguous), that is looped through to access only one or two member variables will still load the entire objects into the cache. This causes a lot of unused data to cycle through the cache, displacing data which would have been used. The result is a dramatic increase in cache misses, which substantially reduces performance. A similar procedural program might store the "object" data in multiple arrays, only accessing (and thus caching) the data in the arrays that are actually used. This adds another benefit. Processor caches are divided into sections, where each section of the cache holds some section of memory. These sections are rotated through (typically the least recently accessed is overwritten by new data being loaded). Looping through two arrays at once can take advantage of this when one cache section holds part of one array and another holds part of another array. Using encapsulation, an array of objects will likely use only one cache section at a time, constantly missing and loading more data. Now, we have only looked at how encapsulation affects cache usage when an array of objects is contiguous in memory. This is not typical. Most object oriented languages dynamically allocate memory for objects, and dynamically allocated memory is often not contiguous. Looping through non-contiguous memory will cause cache misses at almost every iteration, causing severe performance loss. The costs of encapsulation are sometimes justified, but it is important to be aware of them when making design decisions. When using object oriented programming, it is important to understand that OOP is a human invention designed to making interfacing with computers on a low level easier. Computers do not "think" in objects, so there will always be costs to the translation between human ideas of objects and how computers actually work. Understanding the underlying architecture can help mitigate the costs of encapsulation, but it cannot eliminate them entirely.
Now we can discuss how to use encapsulation in C. C already has primitive support for encapsulation, which we can leverage to implement full encapsulation. Note that this is going to be ugly and unwieldy, and it should typically be avoided wherever possible. As an exercise in understanding the C language, however, this may be very useful.
C has a meta data type called a struct. Structs are used to create composite data types that essentially encapsulate data. Structs can contain only data. They cannot contain functions or code. Following is a simple example of a struct.
struct item {
int id;
int price;
int count;
char* name;
};
This code defines a struct of type "item". An instance of this struct could be created and manipulated with the following.
struct item can;The name element would be assigned a pointer to a CString. This struct might be a data type for an inventory system, where id is the inventory id number, price is the price in cents (since it is an int), count is the number of items in stock, and name is a pointer to a character string containing the name of the item. This struct could be passed to and returned from functions like any other variable. We could make an array of this new type to hold all of the different items the store carries. Unlike most object oriented languages, a statically created array of these structs would be contiguous in memory (though, the character arrays pointed to might not be contiguous, and would certainly not be contiguous within the array). This helps keep our data coherent and understandable by human standards. Without the struct, we might instead create a price array, a count array, and a name array, and the ids could be the array indices for each item (we might even avoid storing ids this way, reducing memory costs). For the struct model, if we only looped through the array when we needed to access every element in the struct, we would not get any performance benefits from separate arrays, but since this is unlikely, we will be paying for coherency in performance (in an inventory system, where performance is not that important, this is probably justified).
can.id = 1;
can.price = 120;
can.count = 12;
Now that we have a struct for our inventory items, we might want to reuse this idea in our POS system. We might want a struct for transactions. A transaction might contain a pointer to an array of items (we cannot make the array static, since it may have a different size for each transaction). We could use the previous struct for items, but we would use the count element as the quantity purchased instead of inventory. When finalizing a transaction, we might need to calculate sales tax, so a transaction will need a sales tax as well as a function for calculating tax and setting the variable in the struct. Normally, we would just write a function that takes the struct and modifies it. Maybe we want to be able to use different algorithms for sales tax depending on the customer (business customers might have a lower tax rate, while out of state and government customers might be exempt), but we want to be able to treat all transactions the same. This gets into very basic polymorphism, which we will discuss in depth in a different article. For now, however, we need our transaction to use a different tax algorithm for different customers, and we do not want to have to keep track of tons of information to accomplish this. We want to be able to populate most of a transaction, then send it to a function to finalize it that can use the same procedure on all transactions.
While structs cannot contain functions, they can contain function pointers. So long as the functions pointed to have the same signatures (return values and argument lists), they can easily be called in another function that knows how to use them. We could create several functions for calculating tax, and then we could store a pointer to the appropriate function in the struct. Later, when we finalize the transaction, we can just call the function pointed to by the transaction, and it will calculate tax for us. We do not have to care which algorithm is used during finalization. Here is a very simple example of this.
struct transaction {There is the struct. The pretax element will contain the sum of the prices of items being purchased (in real life, we would have an array of items, and the finalizer would calculate the pretax total from those). The tax element begins empty, and it will be populated by the function pointed to by the gettax element. The gettax element can point to any function that returns void and takes a single argument that is a pointer to a transaction (a pointer because we need to change the original). Now we need some tax functions.
int pretax;
int tax;
void (*gettax)(struct transaction*);
};
void normalTax(struct transaction *t) {
(*t).tax = (*t).pretax * 0.05;
}
void exemptTax(struct transaction *t) {
(*t).tax = 0;
}
These two functions return void and take transaction pointers, just like the function pointer in the struct definition. The first calculates a 5% sales tax, while the other sets tax to 0 for tax exempt customers. Now, we want to create a transaction. This will be a normal customer, with normal sales tax, who is making a purchase that totals to $10.00 (since we are using ints to store cents, it will be 1000).
struct transaction t;
t.pretax = 1000; // $10.00
t.tax = 0; // Initialize to 0
t.gettax = *normalTax; // Normal customer
There is our transaction. If we were serving a tax exempt customer, we could set gettax to *exemptTax instead. Now we are ready to calculate tax. This can be done with the following line, regardless of what tax algorithm we are using.
(*t.gettax)(&t);That will call whichever tax algorithm we selected earlier, passing the transaction in by pointer, so we can set the tax element appropriately. After running the above, we will find that t.tax is equal to 50, which means $0.50. If we had used the tax exempt algorithm, the tax would have been 0.
This example is obviously contrived, and in real life we probably would have explicitly called a different function for each tax mode. In fact, this would probably be a better way to do it for this situation, but this sort of encapsulation has its strong points in many other applications.
Again, this is an ugly and unwieldy way of implementing encapsulation. If you really need encapsulation, it would probably be better to use an object oriented language like C++. In the rare situation where that is not an option or where you only need very basic encapsulation and only to a very limited degree, this might be appropriate. If you are considering doing this, first ask yourself why you are using C in the first place. It is very likely that the reason you are using C is because object oriented programming is much more expensive in memory and performance. If this is the case, you should probably find a more efficient way of solving your problem.
Here is the source code for a simple C program that implements and uses the transaction example from above:
#include <stdio.h>
struct transaction {
int pretax;
int tax;
void (*gettax)(struct transaction*);
};
void normalTax(struct transaction *t);
void exemptTax(struct transaction *t);
void main() {
struct transaction t;
t.pretax = 1000; // $10.00
t.tax = 0;
t.gettax = *normalTax; // *exemptTax for no tax
(*t.gettax)(&t);
printf("Price: %i\n", t.pretax);
printf("Tax: %i\n", t.tax);
}
# For normal 5% sales tax
void normalTax(struct transaction *t) {
(*t).tax = (*t).pretax * 0.05;
}
# For tax exempt customers
void exemptTax(struct transaction *t) {
(*t).tax = 0;
}
Friday, September 26, 2014
C Programming: Singleton Design Pattern
The Singleton design pattern is a common pattern used in object oriented programming. To use the pattern, any constructors of the singleton object must be private. The user must not be able to create new instances of the object explicitly. In this design pattern, the class must only be able to have a single instance. Often this instance is created the first time it is requested, but it may be created at startup time, depending on the programming language. Subsequent requests will be provided with the already existing instance. This design pattern is primarily useful in languages that require object orientation, as a place to collect related global variables and functions, where only one instance of the collection should ever exist. It is less often used in languages that allow object orientation but do not enforce it. There are some cases, however, where it is useful regardless of the language.
One place where the Singleton design pattern is useful regardless of language is the case where a single instance of a global variable is necessary, but it is also necessary to limit how the user may interact with that variable. In 3D graphics, the main camera is one of these global variables. The camera can be stored as a pair of vectors, one representing "up" and the other representing the direction the camera is facing. The third vector, the facing of one of the sides, can easily be calculated from the other two. It is essential, however, that the "up" and "facing" vectors always be perpendicular to each other. If they ever become parallel, the third vector cannot be calculated, and the camera math starts to get zeros and infinities where they do not belong. This makes it impossible for the computer to render graphics that make sense. The Singleton pattern can be used to solve this problem. A single instance of a camera class can be made where the main camera is a private variable of the class. The setter for the camera can ensure that changes to the camera never allow invalid states. Further, methods can be added to the Singleton that allow the user to apply specific transformations to the camera, which removes the burden (and risk) of users trying to do the math for the transforms themselves.
In most cases, the Singleton design pattern is used to hold global things where the language does not provide a better option. In some cases though, this design pattern can be useful in its own right. A problem occurs when the benefits of this design pattern are necessary in a language that does not support object orientation. For example, the C programming language has no object orientation support, but embedded systems often have limited support for languages other than C (or assembly). This may not be true of all non-object oriented languages, but the Singleton design pattern is actually possible in C.
This C Programming series is going to discuss how to use object oriented principles in the C language. In most cases, it is probably a bad idea to use these principles if any other option is available, but in cases like embedded systems, where an object oriented language is not available, it may be necessary, or at least substantially more efficient, to use these principles. The remainder of this article will discuss using the Singleton pattern in C and demonstrate how it can be done.
In C, encapsulation and hiding sensitive data is generally considered impossible. Very basic encapsulation can be accomplished with structs, but the language does not have any built in mechanics for preventing a user from changing any variable that is in scope. This means that protecting a global variable in a getter/setting fashion is impossible. This leads to several difficulties. The first is that it is impossible to enforce data validation. A well designed library might offer setters and getters, but a user of the library might choose to go around them, accessing the variable directly. This puts the burden of correctness on the user, which has proven problematic enough to justify the wide adoption of private and protected variables in object oriented languages.
There is a simple way of making private global variables in external libraries. This is probably nothing new, and has likely been used in many C libraries that use internal state machines. It is not, however, often taught in computer science classes. In C, libraries are contained in separate files from the main program. Each library has at least one source code file as well as a header file. The header file exposes interfaces contained in the library to the program that is using the library. Global variables are exposed with an "extern" statement. If they are not exported, the main program does not even know they exist and thus cannot access them. This does not mean that they do not exist though. The library where the variables are declared can still access them. If this library has exposed functions that can change the hidden variables, then the main program can still access them indirectly. This technique can be used for functions as well. Following is some example code for a C library using what amounts to the Singleton design pattern.
private.c:
private.h
The source code is pretty straight forward. It has a single global variable, a setter, and a getter. For some reason, it is necessary to restrict what the variable is allowed to be, so the setter handles that by ignoring invalid input. The header file is pretty straight forward as well. It exposes the two functions but not the variable. To expose the variable, "extern int private_variable;" could be added to the header file. Notice also that the comment in the header file does not name the variable. If this was distributed as a header and a precompiled object file, the user would not be able to figure out the name of the variable without searching though the object file for intelligible text and then guessing. If the header reveals the name of the variable though, an injudicious user might add an "extern" statement to the header to gain access. Of course, any user that goes to this effort deserves whatever problems it causes, but there is no reason to make it easy. Here is a driver program to test the library with.
main.c
This is not all. Using this same technique, it is possible to put private functions in the library (perhaps for implementation hiding, or maybe just to keep the namespace uncluttered). Any function in the library can be called by other functions in the library, but they can only be called externally if the function prototype is included in the header file. This makes it easy to use the object oriented ideas behind private variables and functions in C. The library represents the object in this case, and the header file determines what is exposed and what is hidden.
This is not one of the object oriented principles that should be avoided if possible. This method of encapsulation and data protection is very straight forward. It is not prone to abuse or errors (and in fact, it is actually designed to reduce the potential for errors). Most of the rest of this series will discuss less stable and manageable techniques that should be used only when absolutely necessary.
One place where the Singleton design pattern is useful regardless of language is the case where a single instance of a global variable is necessary, but it is also necessary to limit how the user may interact with that variable. In 3D graphics, the main camera is one of these global variables. The camera can be stored as a pair of vectors, one representing "up" and the other representing the direction the camera is facing. The third vector, the facing of one of the sides, can easily be calculated from the other two. It is essential, however, that the "up" and "facing" vectors always be perpendicular to each other. If they ever become parallel, the third vector cannot be calculated, and the camera math starts to get zeros and infinities where they do not belong. This makes it impossible for the computer to render graphics that make sense. The Singleton pattern can be used to solve this problem. A single instance of a camera class can be made where the main camera is a private variable of the class. The setter for the camera can ensure that changes to the camera never allow invalid states. Further, methods can be added to the Singleton that allow the user to apply specific transformations to the camera, which removes the burden (and risk) of users trying to do the math for the transforms themselves.
In most cases, the Singleton design pattern is used to hold global things where the language does not provide a better option. In some cases though, this design pattern can be useful in its own right. A problem occurs when the benefits of this design pattern are necessary in a language that does not support object orientation. For example, the C programming language has no object orientation support, but embedded systems often have limited support for languages other than C (or assembly). This may not be true of all non-object oriented languages, but the Singleton design pattern is actually possible in C.
This C Programming series is going to discuss how to use object oriented principles in the C language. In most cases, it is probably a bad idea to use these principles if any other option is available, but in cases like embedded systems, where an object oriented language is not available, it may be necessary, or at least substantially more efficient, to use these principles. The remainder of this article will discuss using the Singleton pattern in C and demonstrate how it can be done.
In C, encapsulation and hiding sensitive data is generally considered impossible. Very basic encapsulation can be accomplished with structs, but the language does not have any built in mechanics for preventing a user from changing any variable that is in scope. This means that protecting a global variable in a getter/setting fashion is impossible. This leads to several difficulties. The first is that it is impossible to enforce data validation. A well designed library might offer setters and getters, but a user of the library might choose to go around them, accessing the variable directly. This puts the burden of correctness on the user, which has proven problematic enough to justify the wide adoption of private and protected variables in object oriented languages.
There is a simple way of making private global variables in external libraries. This is probably nothing new, and has likely been used in many C libraries that use internal state machines. It is not, however, often taught in computer science classes. In C, libraries are contained in separate files from the main program. Each library has at least one source code file as well as a header file. The header file exposes interfaces contained in the library to the program that is using the library. Global variables are exposed with an "extern" statement. If they are not exported, the main program does not even know they exist and thus cannot access them. This does not mean that they do not exist though. The library where the variables are declared can still access them. If this library has exposed functions that can change the hidden variables, then the main program can still access them indirectly. This technique can be used for functions as well. Following is some example code for a C library using what amounts to the Singleton design pattern.
private.c:
// This variable is subject to strict
// requirements.
int private_variable = 5;
// private_variable must be between 5 and 10
// inclusive. Invalid input will be ignored.
void set_private(int input) {
if (input < 5 || input > 10)
return;
else
private_variable = input;
}
// We don't want to expose the variable or its
// memory address, so we use a getter to return
// by value.
int get_private() {
return private_variable;
}
private.h
// The hidden variable must be between 5 and 10
// inclusive. Invalid input will be ignored.
void set_private(int input);
int get_private();
The source code is pretty straight forward. It has a single global variable, a setter, and a getter. For some reason, it is necessary to restrict what the variable is allowed to be, so the setter handles that by ignoring invalid input. The header file is pretty straight forward as well. It exposes the two functions but not the variable. To expose the variable, "extern int private_variable;" could be added to the header file. Notice also that the comment in the header file does not name the variable. If this was distributed as a header and a precompiled object file, the user would not be able to figure out the name of the variable without searching though the object file for intelligible text and then guessing. If the header reveals the name of the variable though, an injudicious user might add an "extern" statement to the header to gain access. Of course, any user that goes to this effort deserves whatever problems it causes, but there is no reason to make it easy. Here is a driver program to test the library with.
main.c
#include <stdio.h>Try adding some code to access private_variable directly. It will not compile. The main program does not even know that variable exists! It can still change and read the variable indirectly through the setter and getter functions though.
#include "private.h"
void main() {
printf("Private = %i\n", get_private());
printf("Setting Private to 10\n");
set_private(10);
printf("Private = %i\n", get_private());
printf("Setting Private to 30\n");
set_private(30);
printf("Private = %i\n", get_private());
printf("Setting Private to 0\n");
set_private(0);
printf("Private = %i\n", get_private());
printf("Setting Private to 7\n");
set_private(7);
printf("Private = %i\n", get_private());
}
This is not all. Using this same technique, it is possible to put private functions in the library (perhaps for implementation hiding, or maybe just to keep the namespace uncluttered). Any function in the library can be called by other functions in the library, but they can only be called externally if the function prototype is included in the header file. This makes it easy to use the object oriented ideas behind private variables and functions in C. The library represents the object in this case, and the header file determines what is exposed and what is hidden.
This is not one of the object oriented principles that should be avoided if possible. This method of encapsulation and data protection is very straight forward. It is not prone to abuse or errors (and in fact, it is actually designed to reduce the potential for errors). Most of the rest of this series will discuss less stable and manageable techniques that should be used only when absolutely necessary.
Thursday, September 18, 2014
Object Oriented Programming
In my studies, work, and research, I have discovered some important things about Object Oriented Programming, Objects, and how each should be used. I want to examine some misconceptions and less well known facts about objects in programming.
Objects are an abstract data type. At the deepest level, an object is a highly flexible template for creating custom data types. This makes objects a data type of data types or a meta data type. Objects are far more than this though. Objects are an amalgam of useful ideas commonly used in programming and programming languages.
Simply put, objects are containers. Objects can contain data and functions. This last part seems pretty novel. An object is a data type that can contain functions. Further, when a contained function is called, it automatically knows which instance of the object it belongs to. These ideas seem very novel. It turns out that they are not.
Objects have some dirty secrets. Objects are hiding places for global variables. In some cases, like the Singleton design pattern, this is easy to see. In other cases it is not. Objects are also containers that often hide the passing of large argument sets to functions. When used properly, this does not usually cause problems, but it can easily hide massive coupling issues. Objects can easily hide poor programming practices, and some common uses for objects would be considered poor programming if they were done without objects.
It turns out that in many cases objects are unnecessary. Because objects have higher overhead than more primitive data types that can be used for the same things, it is important to know where objects will be beneficial and where they may be detrimental. In some cases, it is a matter of trade off between development time and performance, and a judgment call must be made. In many cases, however, objects are unnecessarily used in places where performance is harmed, but no benefits to development time are gained.
Now I want to look at some examples of gaining some of the benefits of objects without actually using objects. This can result in improved performance without sacrificing anything for it.
A few months ago, I was writing a C program where I needed to keep track of a camera in 3D space. It was important that the user be able to create new camera instances, but it was also important that a main camera exist. The main camera would be used for all graphics calculations, and the user could load different camera instances into the main camera. This design had several benefits. One was that the user did not have to pass a camera to the graphics functions every time they were called. Since the graphics functions would typically be called many times per video frame, the overhead of argument passing was an important bottle neck. The other benefit was that the camera required some internal consistency to work properly. The camera consisted of two vectors that were absolutely required to be perpendicular to each other. Allowing the user to directly modify these vectors would make the graphics functions prone to user error, and it would further put a burden on the programmer to ensure that any direct modifications would maintain the vectors properly. In an OOP paradigm, this is an easy problem. The main camera could be made a private member variable in a singleton object, and all access would be controlled with getters and setters. In C, however, objects are not supported. Instead I had to use a novel approach that turned out to be at least as easy as an object but with lower memory and argument passing overhead. The camera handling functions were already contained in a separate file from the main program (to facilitate reuse). So, I put a (global) struct instance for the camera in the .c file for the camera library, but I did not export it in the header file. All of the graphics functions that required a camera were contained in the .c file, so they had direct access to the main camera struct. The main program did not have access to it though. I added some getters and setters to the camera library to allow restricted access to the main camera. The result of this was that I used the Singleton design pattern and I even encapsulated the main camera data, all in a programming language that does not have any support for objects. It also improved program efficiency in several areas.
This experience lead me to another conclusion: Objects are syntactic sugar. In my instance, with the Singleton, I literally avoided passing arguments to frequently used functions. When using large numbers of the same object type, this is impossible. If there had been some benefit to having and using multiple cameras, and if it was normal to use a different camera for each call, passing cameras as arguments would have been more efficient, and in most uses of objects, this is the case. Objects contain references to functions associated with that object type. One benefit of using objects is that the compiler handles the problem of which object instance belongs to which function call. In C, I would have to pass the appropriate data each time I called a function, even if I contained the function references in a struct with the data. The benefit of objects does not, however, improve efficiency. Instead it hides the passing of the argument. This is called "syntactic sugar," because it makes the syntax shorter and faster to type, presumably without making it harder to understand. Syntactic sugar is typically good when it is done well, and in most OOP languages, it is done fairly well. It is important to understand that syntactic sugar does not actually affect program performance though.
Now, as I mentioned before, objects are a meta data type. They are a data type for defining new data types. In this way, they can be very useful. When used properly, they can make programs much easier to develop, read, and understand. Objects in programming are very valuable. High value, however, does not mean that it is appropriate to use objects exclusively. Imagine using structs exclusively in C. We could definitely "encapsulate" all of our functions and variables into structs. We could even define a secondary main function, contain it in a struct, and then call it from main as soon as the program starts (and, in fact, in Java and highly object oriented uses of C++, this pattern is highly recommended). The result would be extremely difficult to read and understand, it would waste substantial amounts of memory, and it would destroy the ability of the compiler to optimize memory usage for cache efficiency (this last one is a problem of all object oriented languages). There are some places where using objects just does not make sense. For some program types, encapsulating most things might work well. For others, it is a waste of time. Selective use of objects can optimize design time where necessary while allowing for optimized performance where it is important. It also turns out that for many tasks, the time spent writing the paper work for the object takes far more time than writing the executable code. In short, objects should not be used where they do not make logical sense. There is no benefit derived from using objects where they are unnecessary and make no sense. Using objects exclusively is like using structs, trees, or any other data structure exclusively. It might be a fun exercise for a challenge, but there is almost no practical application where it is appropriate.
Nearly all of the elements of object oriented programming can be used separately in most modern programming languages. Most newer languages already contain large amounts of syntactic sugar, and many older ones do as well (for instance, the ++ operator in C and C++ is syntactic sugar for simple incrementation). Encapsulation can be attained by grouping things by file (this may not literally make a variable private, but private variables themselves are artificially limited variables that the compiled program knows nothing about). Arrays, structs, tuples, dictionaries, and other data structures can be used to group data, and most languages have some tuple-like mechanic for grouping heterogeneous data. In cases where part of the object paradigm makes sense, but not all of it, it is often possible to get the benefits of the parts you need without actually using objects. This may not always make sense to do, but when it does, it is often better than paying full price for objects when you only need part of them.
Objects and Object Oriented Programming can be very useful in designing and writing applications, but they should never be treated as a complete programming style. Like any other data structure, objects have their place and in their place provide very valuable benefits. Again, like other data structures, overuse of objects results in programs that are inefficient and that do not make logical sense. Objects in programming were designed to model how we imagine real world objects to be. Computers do not think in objects though, and some parts of programming will never fit an object model. Trying to force those things into an object model will ultimately come with extra costs in time, money, and performance. OOP can be very valuable when used properly, but it can cost when it is overused or misused.
Objects are an abstract data type. At the deepest level, an object is a highly flexible template for creating custom data types. This makes objects a data type of data types or a meta data type. Objects are far more than this though. Objects are an amalgam of useful ideas commonly used in programming and programming languages.
Simply put, objects are containers. Objects can contain data and functions. This last part seems pretty novel. An object is a data type that can contain functions. Further, when a contained function is called, it automatically knows which instance of the object it belongs to. These ideas seem very novel. It turns out that they are not.
Objects have some dirty secrets. Objects are hiding places for global variables. In some cases, like the Singleton design pattern, this is easy to see. In other cases it is not. Objects are also containers that often hide the passing of large argument sets to functions. When used properly, this does not usually cause problems, but it can easily hide massive coupling issues. Objects can easily hide poor programming practices, and some common uses for objects would be considered poor programming if they were done without objects.
It turns out that in many cases objects are unnecessary. Because objects have higher overhead than more primitive data types that can be used for the same things, it is important to know where objects will be beneficial and where they may be detrimental. In some cases, it is a matter of trade off between development time and performance, and a judgment call must be made. In many cases, however, objects are unnecessarily used in places where performance is harmed, but no benefits to development time are gained.
Now I want to look at some examples of gaining some of the benefits of objects without actually using objects. This can result in improved performance without sacrificing anything for it.
A few months ago, I was writing a C program where I needed to keep track of a camera in 3D space. It was important that the user be able to create new camera instances, but it was also important that a main camera exist. The main camera would be used for all graphics calculations, and the user could load different camera instances into the main camera. This design had several benefits. One was that the user did not have to pass a camera to the graphics functions every time they were called. Since the graphics functions would typically be called many times per video frame, the overhead of argument passing was an important bottle neck. The other benefit was that the camera required some internal consistency to work properly. The camera consisted of two vectors that were absolutely required to be perpendicular to each other. Allowing the user to directly modify these vectors would make the graphics functions prone to user error, and it would further put a burden on the programmer to ensure that any direct modifications would maintain the vectors properly. In an OOP paradigm, this is an easy problem. The main camera could be made a private member variable in a singleton object, and all access would be controlled with getters and setters. In C, however, objects are not supported. Instead I had to use a novel approach that turned out to be at least as easy as an object but with lower memory and argument passing overhead. The camera handling functions were already contained in a separate file from the main program (to facilitate reuse). So, I put a (global) struct instance for the camera in the .c file for the camera library, but I did not export it in the header file. All of the graphics functions that required a camera were contained in the .c file, so they had direct access to the main camera struct. The main program did not have access to it though. I added some getters and setters to the camera library to allow restricted access to the main camera. The result of this was that I used the Singleton design pattern and I even encapsulated the main camera data, all in a programming language that does not have any support for objects. It also improved program efficiency in several areas.
This experience lead me to another conclusion: Objects are syntactic sugar. In my instance, with the Singleton, I literally avoided passing arguments to frequently used functions. When using large numbers of the same object type, this is impossible. If there had been some benefit to having and using multiple cameras, and if it was normal to use a different camera for each call, passing cameras as arguments would have been more efficient, and in most uses of objects, this is the case. Objects contain references to functions associated with that object type. One benefit of using objects is that the compiler handles the problem of which object instance belongs to which function call. In C, I would have to pass the appropriate data each time I called a function, even if I contained the function references in a struct with the data. The benefit of objects does not, however, improve efficiency. Instead it hides the passing of the argument. This is called "syntactic sugar," because it makes the syntax shorter and faster to type, presumably without making it harder to understand. Syntactic sugar is typically good when it is done well, and in most OOP languages, it is done fairly well. It is important to understand that syntactic sugar does not actually affect program performance though.
Now, as I mentioned before, objects are a meta data type. They are a data type for defining new data types. In this way, they can be very useful. When used properly, they can make programs much easier to develop, read, and understand. Objects in programming are very valuable. High value, however, does not mean that it is appropriate to use objects exclusively. Imagine using structs exclusively in C. We could definitely "encapsulate" all of our functions and variables into structs. We could even define a secondary main function, contain it in a struct, and then call it from main as soon as the program starts (and, in fact, in Java and highly object oriented uses of C++, this pattern is highly recommended). The result would be extremely difficult to read and understand, it would waste substantial amounts of memory, and it would destroy the ability of the compiler to optimize memory usage for cache efficiency (this last one is a problem of all object oriented languages). There are some places where using objects just does not make sense. For some program types, encapsulating most things might work well. For others, it is a waste of time. Selective use of objects can optimize design time where necessary while allowing for optimized performance where it is important. It also turns out that for many tasks, the time spent writing the paper work for the object takes far more time than writing the executable code. In short, objects should not be used where they do not make logical sense. There is no benefit derived from using objects where they are unnecessary and make no sense. Using objects exclusively is like using structs, trees, or any other data structure exclusively. It might be a fun exercise for a challenge, but there is almost no practical application where it is appropriate.
Nearly all of the elements of object oriented programming can be used separately in most modern programming languages. Most newer languages already contain large amounts of syntactic sugar, and many older ones do as well (for instance, the ++ operator in C and C++ is syntactic sugar for simple incrementation). Encapsulation can be attained by grouping things by file (this may not literally make a variable private, but private variables themselves are artificially limited variables that the compiled program knows nothing about). Arrays, structs, tuples, dictionaries, and other data structures can be used to group data, and most languages have some tuple-like mechanic for grouping heterogeneous data. In cases where part of the object paradigm makes sense, but not all of it, it is often possible to get the benefits of the parts you need without actually using objects. This may not always make sense to do, but when it does, it is often better than paying full price for objects when you only need part of them.
Objects and Object Oriented Programming can be very useful in designing and writing applications, but they should never be treated as a complete programming style. Like any other data structure, objects have their place and in their place provide very valuable benefits. Again, like other data structures, overuse of objects results in programs that are inefficient and that do not make logical sense. Objects in programming were designed to model how we imagine real world objects to be. Computers do not think in objects though, and some parts of programming will never fit an object model. Trying to force those things into an object model will ultimately come with extra costs in time, money, and performance. OOP can be very valuable when used properly, but it can cost when it is overused or misused.
Subscribe to:
Posts (Atom)