Wednesday, November 5, 2008

Middleboxes No Longer Considered Harmful

This paper discusses how middle-boxes such as NAT and firewalls violate the two principles of the Internet architecture:
  1. Unique identifier for each Internet entity.
  2. Packets processed by their respective owners.
The authors say that even though the middle-boxes break the rules but they are required for important reasons. Few reasons being security and performance improvement through caching as well as load balancing. The paper then proposes an architecture to include the functionality of middle-boxes without breaking the principles.

The architecture called Delegation Oriented Architecture proposes the following things:
  1. Globally unique ID in flat namespace which is carried by the packets
  2. Sender and Receiver can define the intermediaries that should process the packet.
Without going into the details of the architecture and how it works, I want to list some of the things that concerned me:
  1. After adding the intermediary information in the packet we are still defying the end-to-end principle. What happens if the intermediary crashes?
  2. The unique identifiers are said to be 160 bit long. The packet is supposed to have 2 160-bit identifiers. Isn't this an overhead for small packets?
  3. The idea seems to be interesting but I am concerned about performance. Even though the architecture provides flexibility by allowing the intermediaries to be anywhere and not in the path to the destination. The packet now has to be first traversed to the intermediary and then to the destination. Also, it is required to lookup of the path to the intermediary.
  4. Another question that get raised with such systems is scalability. With so many machines in the Internet, if every machine sends the message to an intermediary and the DHT being used for EID resolving, an important question arises that we are relying on the performance of the DHT for lookup and information retrieval.

Monday, November 3, 2008

DNS Performance and the Effectiveness of Caching

This paper provides analysis of:
  1. DNS performance (latency and failure)
  2. Effect of varying TTL and degree of caching affect
The analysis was performed on three traces that included DNS packets and TCP SYN, FIN and RST packet information. The authors found that over third of all DNS lookups were not answered successfully. 23% of the client lookups in the MIT trace failed to arrive at an answer while 13% lookups gave errors in the answer.

The authors show that name popularity distribution is zipf-like. Effectiveness of caching is determined by finding how useful is it to share DNS caches and the impact of choice of TTL value. Intuitively, I was under the assumption that reducing the TTL will severely affect the performance of DNS but the authors show that reducing TTLs of address(A) records to few hundred seconds has little effect on the hit rate. Also, sharing a forwarding DNS cache does not improve performance.

I found the paper very interesting to read. The analysis of performance of DNS is not dependent on aggressive caching is opposite to the general notion of DNS performance being tightly tied to caching. This observation is directly related to the relationship between TTL values and name popularity. Popular names will be cached effectively even with short TTL while unpopular names may not gain even with long TTL values.

It will be nice to discuss how would the results vary if analyzed with current time traces.

Development of the Domain Name System

This paper discusses the motivation behind the design and need of DNS. Prior to DNS each machine downloaded a file called HOSTS.TXT from the central server on to their system. This file contained the mapping between host names and addresses. This scheme had the inherent problem of distribution and update. With the rise of the number of machines connected to the Internet came the need of building a distributed system for the functionality of HOSTS.TXT.

DNS was designed for this need. It is a hierarchical naming system that had the following design goals:
  1. provide all information of HOSTS.TXT
  2. Allow distributed implementation of the database
  3. Have no size limit issues
  4. Be inter-operable
  5. Provide tolerable performance
DNS contains two main components: name servers and resolvers. Name servers store information and answer queries from the information they possess. Resolvers are the interface to client programs and contain algorithms for querying name servers. DNS name space is organized in the structure of a variable-depth tree. Each node in the tree has a label associated with it and domain name of a node is the concatenation of all the labels on the path from the node to the root of the tree.

One of the main things that makes DNS fast is caching the results of the queries. This overcomes the need to fire a query every time a name lookup has to be performed. Although this optimization comes at the cost of a security concern. People have exploited this feature to perform DNS cache poisoning attacks.

I really enjoyed reading the paper because it gave a good explanation of the ideas behind the design of a system that plays an important role in today's Internet. I would recommend keeping this paper in the syllabus.

Wednesday, October 29, 2008

Lookup Data in P2P Systems

This paper gives a brief summary of lookup operation in P2P networks. It gives an overview of why P2P networks are attractive, discusses the lookup problem, which is basically "Given data X stored in a set of dynamic nodes, how to find it". I really enjoyed reading the last section. The paper discusses talks about CAN, Chord, Kademlia, Pastry, Tapestry and Viceroy.

All these DHTs use different overlay networks (underlying topological structure). Different overlays have different complexity for lookup. Differences between DHTs include distance functions that are combination of being symmetric and unidirectional, cost of node join/failure operations when they happen frequently, network aware routing, malicious nodes etc.

I would recommend keeping this paper.

Chord: A Scalable Peer-to-peer Lookup Service for Internet Applications

This paper talks about the Chord DHT and how it is useful. Instead of giving the complete summary of how it works I want to list the pros and cons of Chord.

The authors in Chord use ring geometry and choice of the DHT geometry greatly affects performance. "The Impact of DHT Routing Geometry on Resilience and Proximity" paper shows that amongst different DHT geometries ring geometry provides greatest flexibility. Another thing I like about Chord is its simplicity.

One place where Chord may not work well is in cases where there is lot of churning. This observation has been made by the Bamboo DHT designers. Also, Chord periodically performs stabilization. This requires that the periodic epoch needs to be tuned as per the need of the network where the nodes reside. Another thing that Chord does not consider is the network topology. DHTs like Tapestry and Pastry take into consideration the network locality and try to minimize the distance traveled by messages.

Monday, October 27, 2008

Active Network Vision and Reality: Lessons from a Capsule-Based System

Active networks allow individual user, or groups of users, to inject customized programs into the nodes of the network. This aids in building a range of applications that can leverage computation within the network. The authors in this paper use ANTS active network toolkit in designing and implementing active networks and share their experience.

The authors express their findings in comparison to the general vision of active networks in three areas:
  1. Capsules: They can be a competitive forwarding mechanism whenever software-based routers are viable. Capsule code can be carried by reference and local on demand. The cost of capsule processing is not high when forwarding is also performed in software.
  2. Accessibility: Each user can handle their packets within the network. Also, code from untrusted users should not do any harm to users of other service.
  3. Applications: Capsules aid in experimenting with new services and deploying them.
The concept of active networks is new to me. The paper was good to read because it talked about the issues active networks have and how the authors experience these issues and resolve them. The only thing that concerns me is the issue that a malicious member of a group can harm other users of the group. The authors talk about this problem and also say that it is not specific to active networks.

Resilient Overlay Networks

This paper proposes an architecture that enables a group of nodes to provide communication in case of failures in the underlying Internet paths. Group of nodes in RON have an application-level protocol for communicating with the other nodes in RON. The protocol ensures that high-speed recovery of path failures, low-latency, and improved loss rate and throughput exist among the nodes participating in the RON.

RON was designed to improve the underlying Internet routing protocols (BGP). BGP can take a long time to converge after link failure. RON aids in recovering from outages and performance issues in seconds. Other goals of RON are tighter integration of routing and path selection with the application and expressive policy routing.

I liked that the authors are proposing an architecture to solve the inherent problem of BGP convergence. Although there are issues that still remain undressed. These include machines behind NAT and violation of BGP policies. The authors discuss these problems as well. The evaluation results are very good but with such techniques one is never sure how they will work in real large scale deployment. Also, it would have been nice see how RON would work in case of congestion, that is how tolerant is RON itself? One more concern I have is related to RONs pushing too aggressively to find alternate routes. This can cause a lot of traffic and be very expensive.

I would recommend keeping this paper because it can stir up a very good discussion in the class regarding the design decisions and their pros and cons.