Monday, November 30, 2009

Secutiry Implications of Cloud Computing

 

 

SECURITY IMPLICATIONS OF CLOUD COMPUTING

Narendran Calluru Rajasekar

November 30th, 2009

 

Supervised by

Dr Chris Imafidon

(formerly Queen Mary University of London)

 

UEL Logo

MSC Internet Systems Engineering

University of East London,

Docklands.

 

Acknowledgements

Oxford Univ Logo

Anne-Marie Imafidon

University of Oxford.

 

Table of Contents

 

1 Abstract

2 Cloud Computing

2.1 Definition

2.2 Understanding Cloud Computing

3 Security Implications

3.1 Security Components

3.1.1 Encryption

3.1.2 Intrusion Detection/Prevention Systems

3.1.3 Antivirus

3.1. 4Firewall

3.2 Security Threat

3.3 Authentication and Access

3.4 Data Security

3.5 Tempting Target for Cybercrime

3.6 Benefit to Risk Ratio

3.7 Legal Issues

4 Conclusion

5 Appendix - Glossary

6 References

 

 

1 Abstract

This paper is focussed on the security implications of cloud computing. Before analysing the security implications, the definition of cloud computing and brief discussion to under cloud computing is presented. The actual analysis of this paper focuses on the basic security components of cloud computing and security threats involved in various aspects of cloud computing.

There are many leading providers in the market like Google, Amazon, Microsoft, HP and IBM. They provide different services and call it with their own names. What exactly is cloud computing? It isn't new; it already exists in different forms such as Virtualisation, Software as a Service, Utility Computing etc. The major aspect in cloud computing is that it is exploited based on the pay per usage charging model.

Though cloud provides many benefits one of which is moving the capital expense to operational expense, the security and legal issues are very high and an organisation can decide to adopt cloud only on based on benefits to risk ratio.

The security implications in cloud computing is discussed in detail in this paper.

 

Keywords: Cloud Computing, Cloud Security, Cloud Legal Issues, Security Implications

 

^table of contents^

 

2 Cloud Computing

2.1 Definition

Cloud computing is an evolving technology and has no concrete definition for it yet. The cloud service providers provide different services based on different capabilities such as SaaS (Software as a Service), PaaS (Platform as a Service), IaaS (Infrastructure as a Service). After analysing definitions from 20 different authors, Vaquero, L., L. Rodero-Merino, et al. (2008) proposed the following definition for cloud computing.

"Clouds are a large pool of easily usable and accessible virtualized resources (such as hardware, development platforms and/or services). These resources can be dynamically reconfigured to adjust to a variable load (scale), allowing also for an optimum resource utilization. This pool of resources is typically exploited by a pay-per-use model in which guarantees are offered by the Infrastructure Provider by means of customized SLA"

- Vaquero, L., L. Rodero-Merino, et al. (2008)

2.2Understanding Cloud Computing

According to Dikaiakos, M., D. Katsaros, et al. (2009), vision of 21st century is accessing Internet services from light weight portable devices, instead of accessing it from a traditional Desktop PC. Cloud computing is a technology which will facilitate companies or organisation to host their services without worrying about IT infrastructure and other supporting services.

The cloud concept draws on the existing technologies which aren't new such as Centralised Computing, Distributed Computing, Utility Computing, SaaS. It is new in the way it integrates all the above and shifts them from a processing unit to a network (Weiss, A., 2007).

The cloud computing facilitates a starting company by moving Capital Expense to Operational Expense (Computing, D. and M. Creeger, 2009). Amazon (EC2, S3), Microsoft Azure, IBM Blue Cloud, HP Cloud Assure are some of the cloud computing services available in the market Kaufman, L. M. (2009).

Organisations can decide upon their operating model either by running their own private cloud or buy it from 3rd party service providers based on their requirements (Grossman, R., 2009). The private cloud is similar to public cloud but it has its own security and compliance needs hosted for and by their own (Rash, W., 2009).

Cloud computing provides extensive computing power for web services but is not mature enough to perform HPC (High Performance Computing). Napper, J. and P. Bientinesi (2009) experimentally shown that the execution speed per dollar spent decreased exponentially with increase computing cores and hence the cost of solving linear systems increased exponentially. Which means cloud computing is in its evolving stage.

 

^table of contents^

 

3 Security Implications

Sloan, K. (2009) has explored and demystified the technologies involved in cloud computing in which he discusses about the challenges posed in security of cloud computing. According to him, security components could be added to the security layer and be delivered as Security as a Service. Figure 1 shows the security architecture of cloud computing.

Cloud Computing Security Architecture

Figure 1: Cloud Computing Security Architecture (Source: Sloan, K., 2009)

To ensure CIA (Confidentiality, Integrity and Availability) of the information, the service provider should offer tested encryption schema, stringent access controls and scheduled data backups (Kaufman, L. M., 2009).

There are many clouds available in the market and the enterprises will start using different clouds for different operations. Eventually there will be a situation where the cloud integration services would be required which again would require a different approach of security implications (Kim, W., 2009). Also there is no single regulatory organisation which regulates the standards for cloud security. Organisation needs to check where the assurance comes from? (Everett, C., 2009).

Although, the basic security components have been identified, the security requirement varies with respect to the domain and business needs. Cloud Security Alliance (2009, April) has identified 15 different domains in cloud computing as shown in the figure 2.

Different domains in cloud computing security

Figure 2: Different domains in cloud computing security Source: Cloud Security Alliance (2009, April)

Since the area is vast and there is no standards clearly defined, cloud security clearly lags and business needs to understand the dangers and weigh them against the benefits (Greene, T., 2009).

For instance, the database service provided by Amazon S3 doesn't support flexible authorisation and granular security (Brantner, M., D. Florescu, et al., 2008).

3.1 Security Components

3.1.1 Encryption

To guarantee the privacy of information hosted on servers in cloud, the information could be encrypted which can only be decrypted at the client level with a key. Again this is only reliable if the data can be quickly decrypted at the client level as it might need high processing power. The multi-core processors which are evolving will make this possible and provide greater integration of information (Hewitt, C., 2008).

A researcher at IBM has cracked a problem with "homomorphic encryption" which is believed boost cloud computing by enabling service providers to analyse the data without actually compromising them (Saran, C., 2009).

In reality, even the leading service providers don't deliver high level of security. For instance, Google services can be used using both http and https. Though by default the service runs using https which is SSL encrypted, it sometime drops back to http which in unencrypted. This will allow attackers to monitor the network traffic and capture the credentials of a specific user (news article: Computer Fraud & Security, 2009). Also when uploading email attachments, Google doesn't use https by default, although the settings could be change to use https always (Herrick, D., 2009).

3.1.2 Intrusion Detection/Prevention Systems

Providing security for cloud computing requires more than authentication using passwords and confidentiality in data transmission. Vieira, K., A. Schulter, et al. (2009) have proposed a solution for intrusion detection in cloud computing. The solution consists of two kinds of analysis behavioural analysis and knowledge analysis. In behavioural analysis, the data mining techniques were used to recognize expected behaviour or a sever deviation of behaviour and in knowledge analysis security policy violations and attack patterns were analysed to detect or prevent intrusion.

3.1.3 Antivirus

Antivirus scanning can be done on the cloud to reduce the risk of malicious activities. It is an expensive operation and doing it once ahead of time for benefit of many could be advantageous, and with the power of cloud more anti-virus engines can be employed to make more efficient. The challenge here is bridging the gap between the threat release and the virus signature release (Walsh, P. J., 2009). Although antivirus scanning is an expensive operation, it should be repeated with the release of new virus signatures.

3.1.4 Firewall

Firewalls could be implemented as a virtual machine image running in its own processing compartment or at the hardware level at each gateway in "out of band" firewall management channels (Sloan, K., 2009).

3.2 Security Threat

The communication between cloud services and consumers can be secured using SSL. Since the technology is too familiar, users usually ignore the warning which can be exploited by attackers. Google has demonstrated such type of exploitation in cloud based services. On the other hand, a flaw in indexing system design of Zoho has resulted in security vulnerability where one user can read others documents. Also there are other XSS and CSRF attacks which were successful on cloud which makes it vulnerable to attacks (Mansfield-Devine, S., 2008).

In SaaS model, the developer should always assume that intruders have full access to the client as anyone including intruders can buy the software. Though they are not supplied with source code, they still have access to binaries using which they can exploit the vulnerabilities. Hence there should always be a verification mechanism to verify client requests before execution (Viega, J., 2009).

3.3 Authentication and Access

There are different authentication mechanisms for different services. The most commonly used mechanisms are Open Id, Open Auth, and User Request Token. The Open Id and Open Auth mechanism is usually used in mobile devises where the authentication information cannot be stored, or have it firewalled as done in regular PC. Yahoo and Google use User Request Token mechanism for authentication where as Amazon AWS uses a custom mechanism which mirrors the Open Id and Open Auth mechanisms and in addition to it, the calling program signs the outbound header elements using HMAC-SHA1 algorithm (Christensen, J., 2009).

2FA (Two Factor Authentication) is one other authentication mechanism which requires two identities or proof which user knows (PIN or Password) / has (Hardware Token, Mobile Phone, Smartcard). Though this mechanism is more secure than the other type of authentication, handling tokens or smartcards could be a burden to users. In this scenario, mobile phones or smart phones can act as a proof if software which generates tokens similar to hardware tokens is installed on it (Abraham, D., 2009).

3.4 Data Security

The organisations using cloud computing should maintain their own data backups even if the providers backs up data for the organisation. This will help continuous access to their data even at the extreme situations such as data providers going bankruptcy or disaster at data center etc (Viega, J., 2009).

Mowbray, M. and S. Pearson (2009) has proposed a client based privacy manager to eliminate the fear of data leakage and loss of privacy in cloud computing. In the paper, they have presented a scenario of salesforce.com which can undergo a security threat; theft of sales data and various ways that an intruder can gain knowledge based on the un-encrypted data. The threats include the collection of personal information and getting inappropriate access to the information. Based on this scenario a set of requirements was derived which include the minimization of personal and sensitive data used in cloud and maximising security protection of data. Finally the overall architecture for client-based privacy data manager has been depicted.

On the other hand, Wang, C., Q. Wang, et al. (2009) says that the model in which public verifiability is enforced can be used where the third party auditor audits the data without intervening with user's time to ensure the data security.

3.5 Tempting Target for Cybercrime

Internet is always a ground of attack for malicious activities. The cloud computing offers a tempting target for cybercrime for various reasons. To maintain data integrity, many providers require 100% of customer's data to be placed in cloud which means that if compromised 100% of data is available to attackers. Leading providers such as Google and Amazon have existing infrastructure to deflect cyber attacks, but this might not be the case with all providers. The cloud architecture is such that it has interlinks with multiple entities and compromise with any one of the weakest links would compromise all the linked entities (Kaufman, L. M., 2009).

The cloud community watching services analyses the cloud activities constantly to detect and prevent newly injected viruses and malicious activities. Active participation of many organisations in this community will help them to curb the malicious activities more effectively (Hawthorn, N., 2009).

3.6 Benefit to Risk Ratio

Viega, J. (2009) presents a scenario of software industry where developers would not have much control over IT Infrastructure. In this scenario, IaaS would be beneficial where the communication between the cloud and local machine is encrypted so that man in middle cannot intercept the traffic. This would be a huge cost saving for the company.

As discussed in section 3.2, in SaaS model the attackers have very less information i.e., the binaries of the software which is quite justifiable to have modest application security program. The cost-effective reality for many organisations is to hire someone to do cheap security testing and skip the cost of training developers on security best practices and review their work (Viega, J., 2009).

3.7 Legal Issues

IT industry's recent focus is on cloud computing due to the 'credit crunch' and a global recession. The key legal issues in cloud with respect to sourcing arrangements are DPA (Data Protection Act 1998), duties of confidentiality and database right. For instance, in the method of storing large volume of data in cloud, the servers could spread across the world. It is debatable whether the informed consent can actually be given in this vague situation. Similarly there are intricacies over confidentiality and database rights as well (Joint, A., E. Baker, et al., 2009).

It is perfectly possible to use cloud-computing in UK in a legal compliant and low risk manner. This would require alteration in operating model which could erode the benefits of cloud computing if not considered in early stages and if contractual or operational management is not properly adopted, there could be significant increase of operational risk (Joint, A., E. Baker, et al., 2009).

A news article published by Computer Fraud & Security (August, 2009) indicates that the data might be subject to search and seizure by government agencies if not specific contracts are made between the service providers. When Google was asked how this situation would be handled, they said that their customers would be notified about any legal order it receives. Hence it is up to the customers to get specific agreements from the service providers.

 

^table of contents^

 

4 Conclusion

The definition of cloud computing is emerging as the various organisations that are developing cloud services are evolving. It is evident that the cloud computing by itself is in evolving stage and hence the security implications in it aren't complete. Even the leading cloud computing providers such as Amazon, Google etc are facing many security issues and are yet to stabilise. Achieving complete solution for legal issues is still a question. With this level of issues in cloud computing, decision to adopt cloud computing in an organisation could be made only based on the benefits to risk ratio.

There is a general assumption at the basic level of all security mechanisms that brute force attack would take considerable time to break it. Considering the power of cloud computing with distributed technology can bring to the computing power, breaking the keys used currently is not far from now! This is a flaw in the low level assumption which could collapse entire security of cloud.

 

^table of contents^

 

5 Appendix - Glossary

Term Definition
2FA Two Factor Authentication
CIA Confidentiality, Integrity and Availability
CRSF Cross Site Request Forgery
DPA Data Protection Act 1998
HMAC Hash based Message Authentication Code
HPC High Performance Computing
HTTP Hyper Text Transfer Protocol
HTTPS Secure Hyper Text Transfer Protocol
IaaS Infrastructure as a Service
PaaS Platform as a Service
SaaS Software as a Service
SHA1 Secure Hash Algorithm
SSL Secure Socket Layer
XSS Cross Site Scripting

 

^table of contents^

 

6 References

[1] Kaufman, L. M. (2009)."Data Security in the World of Cloud Computing." IEEE Security andPrivacy 7(4): 61-64.

[2] Kim, W. (2009). "Cloud Computing: Today and Tomorrow."Journal of object technology 8(1): 65-72.

[3] Grossman, R. (2009). "The Case for Cloud Computing." ITPROFESSIONAL 11(2): 23-27.

[4] Rash, W. (2009). Is cloud computing secure? Prove it. tech in-depth,eWeek. 2009: 8-10.

[5] Computing, D. and M. Creeger (2009). "Cloud Computing: AnOverview." Distributed Computing 7(5).

[6] Weiss, A. (2007). "Computing in the clouds." COMPUTING 16.

[7] Saran, C. (2009). Cryptography breakthrough could secure cloudservices. Computer Weekly. 2009: 20.

[8] Hawthorn, N. (2009). "Finding security in the cloud."Computer Fraud & Security 2009(10): 19-20.

[9] Everett, C. (2009). "Cloud computing - A question oftrust." Computer Fraud & Security 2009(6): 5-7.

[10] (2009). "Data in the cloud might be seized by governmentagencies without you knowing." Computer Fraud & Security 2009(8): 1.

[11] (2009). "Industry to Google: encrypt your cloud." ComputerFraud & Security 2009(6): 3-20.

[12] Hewitt, C. (2008). "ORGsfor scalable, robust, privacy-friendly client cloud computing." IEEEInternet Computing 12(5): 96-99.

[13] Viega, J. (2009). "Cloud Computing and the Common Man."Computer 42(8): 106-108.

[14] Vaquero, L., L. Rodero-Merino, et al. (2008). "A break in theclouds: towards a cloud definition." ACM SIGCOMM Computer CommunicationReview 39(1): 50-55.

[15] Wang, C., Q. Wang, et al. (2009). Ensuring data storage security incloud computing.

[16] Vieira, K., A. Schulter, et al. (2009). "Intrusion DetectionTechniques in Grid and Cloud Computing Environment."

[17] Napper, J. and P. Bientinesi (2009). Can cloud computing reach thetop500?, ACM New York, NY, USA.

[18] Mowbray, M. and S. Pearson (2009). A client-based privacy managerfor cloud computing, ACM.

[19] Herrick, D. (2009). Google this!: using Google apps forcollaboration and productivity, ACM.

[20] de Assunao, M., A. di Costanzo, et al. (2009). Evaluating thecost-benefit of using cloud computing to extend the capacity of clusters, ACMNew York, NY, USA.

[21] Cloud_Security_Alliance (2009, April). "Security Guidance forCritical Areas of Focus in Cloud Computing." Retrieved Nov 25, 2009, from http://www.cloudsecurityalliance.org/guidance/csaguide.pdf

[22] Christensen, J. (2009). Using RESTful web-services and cloud computingto create next generation mobile applications, ACM.

[23] Dikaiakos, M., D. Katsaros, etal. (2009). "Cloud Computing: Distributed Internet Computing for IT andScientific Research." IEEE Internet Computing 13(5): 10-13.

[24] Brantner, M., D. Florescu, et al. (2008). Building a database on S3,ACM.

[25] Greene, T. (2009). "Cloudsecurity fears cast shadow at RSA." Network World 26(16).

[26] Joint, A., E. Baker, et al.(2009). "Hey, you, get off of that cloud?" Computer Law and SecurityReview: The International Journal of Technology and Practice 25(3): 270-274.

[27] Walsh, P. J. (2009). "Thebrightening future of cloud security." Network Security 2009(10): 7-10.

[28] Sloan, K. (2009)."Security in a virtualised world." Network Security 2009(8): 15-18.

[29] Mansfield-Devine, S. (2008)."Danger in the clouds." Network Security 2008(12): 9-11.

[30] Abraham, D. (2009). "Why2FA in the cloud?" Network Security 2009(9): 4-5.

 

^table of contents^

 

Thursday, November 19, 2009

Web Services - Current Trends and Future Opportunities

 

 

WEB SERVICES - CURRENT TRENDS AND FUTURE OPPORTUNITIES

Narendran Calluru Rajasekar

November 19th, 2009

 

Supervised by

Elias Pimenidis

Mike Griffith

 

UEL Logo

MSC Internet Systems Engineering

University of East London,

Docklands.

 

Table of Contents

 

1 Abstract

2 Introduction

3 Web Services

3.1 SOA with Web Services

3.2 Description

3.3 Discovery

3.4 Requirements

3.5 QoS Attributes (Evaluation Criteria)

3.6 Monitoring

3.7 Web Services and Semantic Web

3.8 Web Services in Organisations

3.8.1 Enterprise Application Integration

3.8.2 B2B

4 Evaluating Web Services - Case Study

4.1 Website Architecture

4.2 Experimental Environment

4.3 Quality of Service

4.4 Key Features

4.5 Drawbacks

4.6 Proposals

4.7 Tools used for analysis

5 Conclusion

6 Appendix - Acronyms

7 References

 

^table of contents^

 

1 Abstract

This paper focuses on the current trends and future opportunities in each aspect of web services. Initially the requirements of web services are discussed and the QoS attributes which are the evaluation criteria for web services are defined. This paper also discusses on various standards and technologies involved in web service orchestration, monitoring and the role of web services in enterprise integration and communication. Later in this paper, a case study is presented on a website. The website uses web services provided by a web service provider which is evaluated based on the QoS attributes defined earlier in this paper.

Keywords: Web Services, Semantic Web, QoS Attributes, Amazon Web Service, Zoomii.

 

 

^table of contents^

 

2 Introduction

Today, the cross enterprise business interaction is vital and needs to happen quickly without human intervention. The traditional work flow models are tightly coupled and requires dedicated network between companies for the interaction to happen. This involves high cost for the setup and reduces scalability and reusability (Papazoglou, M., 2003).

Web services are loosely coupled software components which enable interoperability between heterogeneous distributed components. They are available on Internet and are platform independent, thus allowing interaction between different applications (Cerami, E. and S. Laurent, 2002). Hence web services can be used for cross enterprise business interaction with the help of ubiquitous Internet.

 

^table of contents^

 

3 Web Services

Cerami, E. and S. Laurent (2002) defines web service as

"any service that is available over the Internet, uses a standardized XML messaging system, and is not tied to any one operating system or programming language"

- Cerami, E. and S. Laurent (2002)

The web services are self-describing and discoverable. A series of standards such as WSDL, UDDI, SOAP, etc are used to support the activities of web services such as Description, Discovery and Invocation (Lau, R., 2007).

3.1 SOA with Web Services

Service Oriented Architecture (SOA) is a way of defining the communication model with which two different services can talk to each other regardless of the operating system, programming language and IT Infrastructure (Newcomer, E. and G. Lomow, 2005). Web service is a way of implementation of SOA where the components are loosely coupled and interoperable.

The components are loosely coupled and are accessible as individual components rather than as an application. Hence, they are more prone to security attacks (Jensen, M., N. Gruschka, et al., 2007). Also the non functional requirements such as security and integration are difficult since the ways of accessing the components are of wide variety. Yamany, H., M. Capretz, et al., (2009) has proposed metadata consisting of security paradigms and a Web Service is constructed for the consumer to access it.

 

 

3.2 Description

Web services are hosted in a server along with the description. WSDL (Web Service Description Language) is used to describe a web service. WSDL is an XML Grammar based standard accepted by W3C which provides communication level description of protocols and messages. At present keyword based approach is used widely to match description which can be inaccurate, Hu, J., P. Zou, et al. (2005) have proposed QoS based Web Service Description Language and a matching mechanism which is composed of QoS attributes and behaviour restraints and proved it to be efficient based on the experimental results.

Ankolekar, A., M. Burstein, et al. (2002) has provided web service descriptions at application level called DAML-S complementing to WSDL which makes the tasks such as invocation, interoperation, composition, verification, execution and monitoring much more easier with the semantic web.

3.3 Discovery

Web services are being developed everywhere and is available in the Internet, but finding a web service is what matters. UDDI Business Registry (UBR) is a collection of UDDI nodes or servers which consists of web service specification through which a web service can be discovered.

There is no single repository where all the web services are registered and hence finding a specific web service is difficult. Al-Masri, E. and Q. Mahmoud (2007) addressed the issue of searching web services by using an enhanced discovery model using Web Service Crawler Engine and have demonstrated with experimental results of achieving efficient search capabilities. Although there are some limitations such as access restriction to secure content, Al-Masri, E. and Q. Mahmoud (2008) says that "a crawler and a centralized repository for Web services is inevitable".

Discovering a web service is not the end of it! Discovering a web service which is of high quality is also necessary. Current techniques don't allow users to query based on the quality of a web service. For example, it is not possible for a user to select a web service whose response time is between some limits. This can be achieved by Web Service Broker which continuously collects the web services just like a search engine but in addition to it, evaluates it against various quality metrics and stores this information along with the web services (Al-Masri, E. and Q. Mahmoud, 2008). The additional information which the author calls as QWS dataset can be used for the quality driven web service discovery.

Figure 1 shows the classification of QoS parameter by Al-Masri, E. and Q. Mahmoud (2009) based on the quantitative measurements, client's perception and service provider's perception which can be used for discovery purpose.

 

Basic qualities of web service parameters

Figure 1: Basic qualities of web service parameters (Source: Al-Masri, E. and Q. Mahmoud, 2009)

Kritikos, K. and D. Plexousakis (2009) Highlights an approach of selecting QoS based web services where mixed integer programming can be used to select appropriate web services which uses semantic based algorithm to match the description which will eventually increase the accuracy of the match.

3.4 Requirements

The requirement for web services changes according to various factors such as business needs, environments, reusability etc. Web services should be able to adapt or withstand these changes especially in high assurance systems. Hence web services should contain properties which will allow them to be reconfigured either statically or dynamically. Yen, I., H. Ma, et al. (2008) have proposed service transformation framework which allows the web services to be reconfigurable based on the rules defined on properties.

Service Level Agreements (SLA) which is also known as contracts is important for web service which is defined during requirements of a web service. Usually hard contracts like QoS value are some value within limits or predefined values. Soft contracts are something which considers time and the path of the flow as a factor and the thresh hold changes accordingly. The orchestration of a web service is very important to provide high quality service (Rosario, S., A. Benveniste, et al., 2008).

The composition web services i.e., the web services orchestrated by the composition of multiple web services are significant as this is the new way of developing solutions for business (Oh, S., D. Lee, et al., 2008). Thus the complex nature of business requirements sometimes requires dynamic selection of web services based on the environment and current inputs for the web service orchestration (Hwang, S., E. Lim, et al., 2008). The limitations are high in the failure prone environment, moving towards semantic web will help to overcome the limitations.

The dependability of web service can be improved by the mediator framework (Chen, Y. and A. Romanovsky, 2008). According to this framework, the web service mediators accept invocations from client and collect the information on resilience behaviour on web services, which can be used to improve dependability.

General techniques assume the environment and parameter to be static, but in reality this is not true and hence web service should adapt to the change automatically at runtime. This can be achieved by efficiently querying the changing parameters using the value of changed information during runtime (Harney, J. and P. Doshi, 2008). In order to achieve this, requirements should be defined considering all the above factors.

3.5 QoS Attributes (Evaluation Criteria)

QoS attributes of web services may vary for various domains and environment (mobile devices, streaming media, rich content, etc). For instance, the mobile devices requires specific factors for high QoS as they may encounter network problems like, low network bandwidth, loss of connection etc. which may require caching of results to reduce delay (Artail and Saab, 2009).

According to Buccafurri, F., P. De Meo, et al. (2009) the various QoS attributes involved for web services involving real time service provisioning such as streaming media are Response Time, Price, Availability, Reputation, Data quality timeliness, Data quality accuracy and Data quality completeness.

Kritikos, K. and D. Plexousakis (2007) broadly classify QoS parameter into two categories i.e., Domain Dependent QoS attributes and Domain Independent QoS attributes. These QoS attributes are used as evaluation criteria for web services based on their QoS Metrics. Figure 2 shows intuitive representation of this classification.

 

Classification of QoS Attributes

Figure 2: Classification of QoS Attributes (Kritikos, K. and D. Plexousakis, 2007)

3.6 Monitoring

Traditional software monitoring is done parallel to its execution; similarly the web service should be monitored in parallel to check if it is working according to the specification or requirement. Wang, Q., J. Shao, et al. (nd) suggests an approach where constraint specification are predefined and the web service is monitored against the constraints and any anomalies are notified. In this approach a probe is installed at the gateway of web service where the entire client traffic passes through. The probe feeds the central analyser where the QoS attributes are verified against the predefined constraint specification.

The QoS attributes which are to be monitored can be configured statically, but the dynamic nature of Internet sometimes demands it to be configured dynamically and this can be achieved by formalising the specification language. Gan, Y., M. Chechik, et al., (2009) identified that the subset of UML 2.0 can be used as a specification language to formalise and achieve the liveliness of the properties.

3.7 Web Services and Semantic Web

The web service discovery on a keyword based web is inefficient and their matching based on the keyword may not be accurate as the context may differ (Ma, J., J. Cao, et al., 2007). The semantic web services which utilise the power of Ontologies for matching and interchanging information will provide efficient and accurate matching based on the context.

Pathak, J., N. Koul, et al. (2005) has described a framework to discover web services which rely on user supplied Ontology specific mappings to match web services in specific domain to make it meaningful and accurate.

Cuevas-VicenttÌn, V., G. Vargas-Solar, et al. (2008) has presented a design and implementation of web service orchestration engine which provides scalable and robust platform for data management and semantic content across various domains. This will allow web services to automatically adapt to the requirements by discovering required service automatically.

3.8 Web Services in Organisations

The approach of the web services deals with an application integration concept. Web Services Technology is used in Organisations in two broad categories: EAI (Enterprise Application Integration) and B2B (Business-to-Business). ESB (Enterprise Service Bus) Infrastructure enables this integration which uses XML based web services to orchestrate the behaviour of services in distributed process.

3.8.1 Enterprise Application Integration

Enterprise systems integration permits applications to be connected in inter and intra organisational settings. EAI applications define unique data formats and communication protocols. These systems are complicated to change. The approach of the web services suggests a set of technologies which wraps the existing legacy systems as Web Services and integrates with the other systems within the organisation. Moreover, this integration approach permits the reusability of existing applications as well permitting new applications and data to be incorporated. For several organisations, the first and foremost implementations using Web services technology would be internal application integration, since that is the main difficulty for them to deal with IT (Ooi & Su, 2006).

3.8.2 B2B

B2B computing is integrating of business systems of two or more companies to support cross-enterprise business. In a matter of 5 years, the challenges of constructing Business to Business applications collective with the massive market potential initiated innovation that moves the industry from simple business-to-consumer (B2C) applications to SOAP-enabled Web services. Any of the open Internet protocols such as HTTP, SMTP, or FTP, or proprietary networks such as EDI is used by the B2B applications. B2C applications handle data directly over the HTTP protocol. (Cixing & Yunlong, 2009)

 

^table of contents^

 

4 Evaluating Web Services - Case Study

In this case study http://zoomii.com an e-commerce website is evaluated which includes analysis of web services used in the website. This website was developed using Amazon Web SerivesTM and has been analysed based on the paradigms; Quality of Service (attributes listed in Section 3.5), Key Features and Drawbacks.

Zoomii (there after refers to http://zoomii.com) is a book seller website which lists top 25000 books available in amazon.com products list. Its cool interface (similar to Google Maps) allows users to search through the online book store as if they do it in a physical book shelves. It uses Amazon Web Services to get the best out of it. Users can pay for the books purchased online through Amazon FPS web service.

4.1 Website Architecture

Zoomii uses the following Amazon web services:

* Amazon Product Advertising API: formerly known as Amazon E-Commerce Service (ECS) is used to access the books data available in www.amazon.com store using web services (Zoomii Inc, 2009) and (Amazon.com, 2009). Though both REST and SOAP requests can be used to access the web services, Amazon.com (2009) recommends using REST as it is more intuitive and also SOAP requests would need toolkits which is not provided for all platforms and programming languages.

* Amazon Elastic Compute Cloud (Amazon EC2): is used to host the web server and computing infrastructure on which the Zoomii runs.

* Amazon Simple Storage Service (Amazon S3): is used for storing data which is specific to Zoomii such as profile information, images etc.

* Amazon Flexible Payments Service (Amazon FPS): is used for transactions through debit and credit cards.

4.2 Experimental Environment

The experiments were performed on a Laptop with the following specification.

Operating System: Windows XP
Processor: Intel Core 2 Duo - 2.4 GHz
RAM: 2 GB
HDD: 250 GB
Network: 20 Mbps broadband connection

The usability was evaluated in different browsers viz., Fire Fox, Internet Explorer, and Google Chrome. Additionally the usability was also evaluated in the touch interface of iPhone with WiFi connection (Same network used in Laptop).

4.3 Quality of Service

The following attributes (listed in Section 3.5) which has been selected for evaluation is as defined by Kritikos, K. and D. Plexousakis (2007).

4.3.1 Domain Dependent Attributes

Domain dependent attributes are related to the domain to which the web services belong to. Zoomii belongs to book seller domain and hence the attributes which are related to this domain are evaluated in this section.

4.3.1.1 Performance

This quality attribute determines how well the service performs. This is generally determined based on the response time. Due to the limitations and access restrictions, the response time for individual web services used in Zoomii cannot be measured. Hence the overall response time for the page load of each event was captured and analysed. The experimental results showed an average response time of 1.65 seconds which was acceptable for the intuitive look (book shelves).

 

Response time of Zoomii

Figure 2: Response time of Zoomii

4.3.1.2 Dependability

This quality attribute determines if the service can be justifiably trusted. This in turn is sufficed when the attributes Viz., Availability, Reliability and Scalability are sufficed.

Scalability: (Zoomii, 2009) has mentioned that the website can only retrieve top 25000 books from Amazon store. Also there are some limitations to the level of categorisation of books. Hence the scalability of this application is not appreciable. Author is working to improve the scalability. Figure 3 shows the screen shot of a page from Zoomii website showing the limitation.

Availability: The service is available 24×365 except during maintenance activities which could be minimal when planned. Amazon EC2 can be configured such that the web server can scale up or down depending upon the traffic and hence the site never goes down.

Reliability: It can be defined as the ability to perform under stated conditions. The web site retrieved results for all the trials which prove the service to be reliable. Also the server is highly scalable and the services are available 24×7, hence it is reliable.

 

Screen showing the about page of Zoomii

Figure 3: Screen showing the about page of Zoomii

4.3.1.3 Transaction Support Related QoS

Transaction support determines the integrity of the data. Zoomii has features such as Wish List and Cart, where books can be added or deleted and maintained for different sessions. Each session has own instances of wish list and cart which don't interfere with each other. Thus the data integrity is maintained.

https://fps.sandbox.amazonaws.com?
Action=GetTransactionStatus
&AWSAccessKeyId=AKIAIIFXJCFIHITREP4Q
&Signature=2l60qD6%2BDIfVEN7ZiHM0AcUKACZt0GYKFtIryqkCb6g%3D
&SignatureMethod=HmacSHA256
&SignatureVersion=2
&Timestamp=2009-10-06T09%3A12%3A06.921Z
&TransactionId=14GKE3B85HCMF1BTSH5C4PD2IHZL95RJ2LM
&Version=2008-09-17"

The above REST request shows the sample request to get the transaction status. From this it is evident that the transactions is completely taken care by Amazon.

4.3.1.4 Security

Security can be evaluated based on the attributes such as authentication, authorisation, integrity, data encryption etc. Zoomii has a login mechanism (user id and password) using which authentication and authorisation occurs. Users can browse the website without logging in, but can check out the books only after logging in.

The code snippet presented below is a REST request for Item Search operation of the Amazon Product Advertising API. It uses AWSAccessKeyId parameter to authenticate Zoomii to use the web services provided by Amazon. This key pair is provided by Amazon when Zoomii first registered with them.

"http://ecs.amazonaws.com/onca/xml?
Service=AWSECommerceService
&AWSAccessKeyId=[AWS Access Key ID]
&Operation=ItemSearch&
SearchIndex=Books&
Author=Steve%20Davenport&
Version=2006-09-13"

Zoomii maintains user information and it stores password in clear text which is displayed in page source. Hence it is vulnerable to attacks; anyone monitoring network traffic can easily capture the confidential and private data.

 

Page Source Showing Clear Text Password

Figure 3: Page Source Showing Clear Text Password

The payments are taken care by Amazon FPS web service which is a secure service. The communication is done through secure protocol in which the data is encrypted and hence this functionality of the website is secure.

4.3.2 Domain Independent Attributes

The attributes that are related to the technical aspects of the web services are evaluated in this section.

4.3.2.1 Usability

Zoomii's user interface which is similar to Google Maps is easy to use. Instead of traditions tree structure for categorisation, it has used intuitive way of zooming in and out of categories. All these features are good only with a scroll mouse; otherwise it is little bit difficult to navigate between categories.

Website is obsolete when used in a touch interface (Apple iPhone), i.e., the UI is completely not usable as the zoom in and zoom out feature conflicts with touch interface.

4.3.2.2 Data Encryption

Zoomii doesn't use any data encryption. All the payments are carried out through Amazon FPS which uses 128 bit encryption verified by VeriSign Class 3 Secure Server.

 

Security information of www.amazon.com

Figure 4: Security information of www.amazon.com

4.3.2.3 Reputation

The reputation of web services at the Amazon Product Advertising API level cannot be evaluated, but the overall reputation of Amazon Web Services is good as they are the leading web service providers in the market (DatacenterDynamics, 2009).

4.3.2.4 Price

Amazon web services charges only if the service is used. For instance they charge £ 0.08 for one hour usage of Amazon EC2 which is low-priced when compare to the cost of setting up such infrastructure.

4.4 Key Features

* Web Service provides access to several million items with good response time.

* Scaling up and down of servers happens automatically with Amazon EC2.

* Encrypted key value pair is used for authentication with Amazon.

* The user interface gives an intuitive look and is easy to use.

* Complete technical documentation and user support is available.

* E-Commerce design used here is object based which gives real world experience.

4.5 Drawbacks

* Web services doesn't contain enough QoS attributes to make if adapt to changes in environment.

* No synchronisation between Zoomii Cart and Amazon Cart.

* Preview facility is not available in Zoomii but available in Amazon book store.

* Drilling down into sub-categories is limited.

* User information is less secure.

* Only top 25000 ranking books rated in Amazon is available.

4.6 Proposals

* Making both Amazon and Zoomii fully semantic will help overcome the integration problems such as cart synchronisation.

* This could be extended to academic library catalogues.

* Security can be improved by including SSL Certificate.

* Personal book shelf features, where registered users can store and organise their books.

* Mouse hover display of useful information on books in layout mode.

* Drag and drop support for moving books.

 

 

4.7 Tools used for analysis

* Google Chrome, Fire Fox and Internet Explorer - The response time of the website was analysed in various browsers.

* Network Monitor 3.3 - Was used to analyse the network traffic and payloads of various operations in the website.

 

 

^table of contents^

 

5 Conclusion

Computer to computer communication which can happen within companies or across companies, within domains or across domains are not standardised (Davies, N., D. Fensel, et al., 2004). The web services; an implementation of SOA under ESB infrastructure have facilitated a major breakthrough in terms of integration and communication between enterprises. The overall research across the world is oriented towards automating and improving the quality of web services. Semantic web services orchestrated with clearly defined QoS attributes can dynamically adapt and provide high quality of service in a fully semantic web. The case study on Zoomii proved that there are many drawbacks due to the limitations in integrating two websites. This can be facilitated by Semantic Web Services and fully Semantic Web. Hence the vision of web is moving towards fully semantic web which is called Web 3.0. Thus the study on Web Services; its current trends and future opportunities and evaluation of a web services has been performed.

 

 

^table of contents^

 

6 Appendix - Acronyms

Term Definition
B2B Business to Business
B2C Business to Consumer
EAI Enterprise Application Integration
EDI Electronic Data Interchange
ESB Enterprise Service Bus
FTP File Transfer Protocol
HTTP Hyper Text Transfer Protocol
QoS Quality of Service
QWS Quality of Web Service
SLA Service Level Agreement
SMTP Simple Mail Transfer Protocol
SOA Service Oriented Architecture
SOAP Simple Object Access Protocol
W3C World Wide Web Consortium
WS Web Service
WSDi Web Service Discovery
WSDL Web Services Description Language
UBR UDDI Business Registry
XML Extensible Mark-up Language

 

^table of contents^

 

7 References

[1] Acero, A., N. Bernstein, et al. (nd). Live search for mobile: Web services by voice on the cellphone.

[2] Al-Masri, E. and Q. Mahmoud (2007). Crawling multiple UDDI business registries, ACM.

[3] Al-Masri, E. and Q. Mahmoud (2008). "Discovering Web Services in Search Engines." IEEE Internet Computing 12(3): 74-77.

[4] Al-Masri, E. and Q. Mahmoud (2008). "Toward quality-driven web service discovery." IT PROFESSIONAL: 24-28.

[5] Al-Masri, E. and Q. Mahmoud (2009). "Web Service Discovery and Client Goals." Computer 42(1): 104-107.

[6] Amazon.com (2009). "Amazon Elastic Compute Cloud (Amazon EC2)." Retrieved Nov 13, 2009, from http://aws.amazon.com/ec2/.

[7] Amazon.com (2009). "Amazon Flexible Payments Service (Amazon FPS)." Retrieved Nov 13, 2009, from http://aws.amazon.com/fps/.

[8] Amazon.com (2009). "Amazon Simple Storage Service (Amazon S3)." Retrieved Nov 13, 2009, from http://aws.amazon.com/s3/.

[9] Amazon.com (2009). "Case Study - Zoomii." Retrieved Nov 13, 2009, from http://aws.amazon.com/solutions/case-studies/zoomii/.

[10] Amazon.com (2009). "Product Advertising API Developer Guide (API Version 2009-10-01)." Retrieved Nov 13, 2009, from http://docs.amazonwebservices.com/AWSECommerceService/2009-10-01/DG/.

[11] Amazon.com (2009). "Product Advertising API." Retrieved Nov 13, 2009, from https://affiliate-program.amazon.co.uk/gp/advertising/api/detail/main.html?ie=UTF8&pf_rd_t=501&pf_rd_m=A3P5ROKL5A1OLE&pf_rd_p=&pf_rd_s=assoc-right-1&pf_rd_r=&pf_rd_i=assoc_join_menu

[12] Ankolekar, A., M. Burstein, et al. (2002). "DAML-S: Web service description for the semantic web."

[13] Artail, H. and S. Saab (2009). "A Distributed System for Consuming Web Services and Caching Their Responses in MANETs." IEEE TRANSACTIONS ON SERVICES COMPUTING 2(1): 17-33.

[14] Buccafurri, F., P. De Meo, et al. (2009). "A framework for using Web services to enhance QoS for content delivery." IEEE MultiMedia 16(1): 26-35.

[15] Cerami, E. and S. Laurent (2002). Web services essentials, O'Reilly & Associates, Inc. Sebastopol, CA, USA.

[16] Chen, Y. and A. Romanovsky (2008). "Improving the dependability of web services integration." IT PROFESSIONAL: 29-35.

[17] Cuevas-VicenttÌn, V., G. Vargas-Solar, et al. (2008). Web Services Orchestration in the WebContent Semantic Web Framework, IEEE Computer Society.

[18] DatacenterDynamics (2009). "Amazon cloud to extend over Asia-Pacific." Retrieved 2009, from http://www.datacenterdynamics.com/ME2/dirmod.asp?sid=&nm=&type=news&mod=News&mid=9A02E3B96F2A415ABC72CB5F516B4C10&tier=3& nid=28D7AB327DD349FCB3BF1B0D2894AEF9

[19] Davies, N., D. Fensel, et al. (2004). "The future of web services." BT Technology Journal 22(1): 118-130.

[20] Gan, Y., M. Chechik, et al. (2009). "Runtime Monitoring of Web Service Conversations." IEEE TRANSACTIONS ON SERVICES COMPUTING 2(3): 223-244.

[21] Harney, J. and P. Doshi (2008). "Selective Querying for Adapting Web Service Compositions Using the Value of Changed Information." IEEE TRANSACTIONS ON SERVICES COMPUTING 1(3): 169-185.

[22] Hu, J., P. Zou, et al. (2005). "Research on web service description language QWSDL and service matching model." Jisuanji Xuebao(Chin. J. Comput.) 28(4): 504-513.

[23] Hwang, S., E. Lim, et al. (2008). "Dynamic Web Service Selection for Reliable Web Service Composition." IEEE TRANSACTIONS ON SERVICES COMPUTING 1(2): 104-116.

[24] Jensen, M., N. Gruschka, et al. (2007). SOA and web services: New technologies, new standards-new attacks.

[25] Kritikos, K. and D. Plexousakis (2007). Requirements for qos-based web service description and discovery.

[26] Kritikos, K. and D. Plexousakis (2009). "Mixed-Integer Programming for QoS-Based Web Service Matchmaking." IEEE TRANSACTIONS ON SERVICES COMPUTING 2(2): 122-139.

[27] Lau, R. (2007). "Towards a web services and intelligent agents-based negotiation system for B2B eCommerce." Electronic Commerce Research and Applications 6(3): 260-273.

[28] Liu, A., Q. Li, et al. (nd). "FACTS: A Framework for Fault Tolerant Composition of Transactional Web Services."

[29] Ma, J., J. Cao, et al. (2007). A probabilistic semantic approach for discovering web services, ACM.

[30] Newcomer, E. and G. Lomow (2005). Understanding SOA with Web Services, Addison-Wesley, Upper Saddle River, NJ.

[31] Oh, S., D. Lee, et al. (2008). "Effective Web Service Composition in Diverse and Large-Scale Service Networks." IEEE TRANSACTIONS ON SERVICES COMPUTING 1(1): 15-32.

[32] Papazoglou, M. (2003). "Web services and business transactions." World Wide Web 6(1): 49-91.

[33] Pathak, J., N. Koul, et al. (2005). "Discovering Web Services over the Semantic Web." Iowa State University, Dept. of Computer Science Technical Report, ISU-CS-TR: 05-20.

[34] Rosario, S., A. Benveniste, et al. (2008). "Probabilistic QoS and Soft Contracts for Transaction-Based Web Services Orchestrations." IEEE TRANSACTIONS ON SERVICES COMPUTING 1(4): 187-200.

[35] Sedlar, U., L. Zebec, et al. (2008). "Bringing click-to-dial functionality to IPTV users [web services in telecommunications, part II]." IEEE Communications Magazine 46(3): 118-125.

[36] Suo, Y., N. Miyata, et al. (nd). "Open Smart Classroom: Extensible and Scalable Learning System in Smart Space using Web Service Technology." IEEE Transactions on Knowledge and Data Engineering.

[37] Wang, Q., J. Shao, et al. (nd). "An Online Monitoring Approach for Web Service Requirements."

[38] Yamany, H., M. Capretz, et al. (2009). Quality of Security Service for Web Services within SOA, IEEE Computer Society.

[39] Yen, I., H. Ma, et al. (2008). "QoS-Reconfigurable Web Services and Compositions for High-Assurance Systems." Computer 41(8): 48-55.

[40] Zoomii Inc (2009). "About." Retrieved Nov 13, 2009, from http://zoomii.com/files/books/about.html

Friday, May 8, 2009

Data Mining - Classification Algorithm - Evaluation

 

 

DATA MINING - CLASSIFICATION ALGORITHM - EVALUATION

Narendran Calluru Rajasekar

May 8th, 2009

 

Supervised by

Dr. Sin Wee Lee

 

UEL Logo

MSC Internet Systems Engineering

University of East London,

 

Table of Contents

 

1 Abstract

2 Implementation of Classification Algorithm (Part 1)

2.1 Dataset - Network Intrusion Detection

2.1.1 Introduction of the Dataset

2.1.2 Dataset Cleaning

2.1.3 Dataset Transformation

2.1.4 Dataset Reduction

2.1.5 Input Encoding / Input Representation

2.1.6 Input format

2.2 Implementation of Algorithms

2.2.1 Reasons Why the Particular Tool is Chosen

2.2.2 Implementation Procedure Used in Weka

2.2.3 Implementation Procedure Used in SQL Analysis Services

2.2.4 Detailed analysis of the results with the appropriate screen dumps

2.3 Challenges of Implemented Algorithm

2.3.1 Conclusions and comments

2.3.2 Problems encountered and the solutions

3 Research (Part 2)

3.1 Introduction

3.2 Research contents

3.2.1 Decision Trees

3.2.2 Neural Networks

3.2.3 Comparison of Decision Tree (J48) and Neural Networks (Multilayer Perceptron)

3.2.4 Experimental Results of J48 and Multilayer Perceptron

3.3 Conclusions and comments

4 Appendix - Acronyms

5 Bibliography

 

^table of contents^

 

1 Abstract

This paper consists of two parts; the first part consists of implementation of Decision Tree, Naïve Bayes and Neural Networks in Weka and SQL Analysis Services. The attributes such as accuracy and performance are compared among the data before preprocessing and after preprocessing by implementing the above mentioned algorithms. Also screen dumps of tree visualization and implantation steps, tables and charts on the results are provided.

The second part consists of research paper which compare, contrast and evaluates J48 implementation of Decision Trees and Multilayer Perceptron implementation of Neural Networks. It also includes the research topics floating around Decision Trees and Neural Networks.

The above two sections are followed by Appendix (Acronyms) and Bibliography.

Keywords: Data Mining, Decision Trees, Neural Networks, Comparitive Study

 

 

^table of contents^

 

2 Implementation of Classification Algorithm (Part 1)

In this section we will discuss implementation of Decision Tree, Neural Network and Naïve Bayes algorithms to the selected dataset and discuss the results of the implemented algorithms.

2.1 Dataset - Network Intrusion Detection

2.1.1 Introduction of the Dataset

The dataset Network Intrusion Detection was chosen from The UCI KDD Archive [1]. The dataset is the collection of network related information that was captured over a period of time. The dataset downloaded from The UCI KDD Archive consists of 41 attributes, 300,000 instances and 38 different classes which are represented in "csv" format.

2.1.2 Dataset Cleaning

Since the dataset downloaded is corrected already, it doesn't contain any missing value or inconsistent values. Hence data cleaning is not required for the selected dataset.

2.1.3 Dataset Transformation

The "csv" format dataset is imported in to SQL table using the Data Transformation Service (DTS) so that it can be used with SQL Analysis Services and it is also converted to ARFF file format to use it with Weka using the import tool available in Weka software.

2.1.4 Dataset Reduction

To meet the objective of coursework, the training dataset was simplified to 5 attributes, 943 instances and 5 different classes. For simplicity and to understand the domain easily the 5 most common attacks (classes) which involves different combinations of Protocol Type and Service were chosen (Table 2.1) from the dataset. These classes are chosen such that it which may help to understand the implementation of algorithm in real time.

Class Description
back Denial of service attack which targets apache web server by sending requests consists of many back slashes in the url.
guess_passwd Attack by guessing user password.
Ipsweep Attack in which valid IP addresses in a network is found out which can be used for other attacks.
multihop Attack in which access to a machine is gained and is used for attacking other machines.
smurf Denial of service attack in which the ping request is sent to multiple machines with a spoofed host IP address.Hence the target machine would be flooded with ping replies without the ability to distinguish between real traffic and the spoofed ones.

Table 2.1 - Description of classes in the dataset.

Using AttributeSelection filter, 8 attributes were selected out of 41 attributes. Using the domain knowledge another 3 attributes were eliminated which is not related to the selected classes. Finally the 5 attributes (Table 2.2) were selected for the case study.

The 300,000 instances was reduced to 943 by randomly selecting 20,000 instances using SQL Analysis Services and then by filtering down to 5 classes.

Attributes Description Type
Protocol Type Protocol used (tcp, udp, icmp) discrete
Service network service on the destination(http, telnet) discrete
Src Bytes Data bytes transferred from source to destination continuous
Dst Bytes Data bytes transferred from destination to source continuous
Count number of connections to the same host as the current connection in the past two seconds continuous

Table 2.2 - Description of attributes in the dataset.

2.1.5 Input Encoding / Input Representation

The selected attributes consists of both discrete and continuous attribute types. The attributes Protocol Type and Service are of type discrete and the attribute Src Bytes, Dst Bytes, Count are of continuous type. Since the algorithm to be implemented, process the dataset in number of iterations, the continuous type attribute will increase the load on the algorithm and thereby decreasing the performance. Hence they are converted to discrete values by applying Supervised Discretized method available in Weka. This method also helps to remove the outliers to an extent.

2.1.6 Input format

* ARFF - Attribute-Relation File Format is used as input for Weka.

Sample dataset used in this case study

Figure 2.1 - Sample dataset used in this case study.

* SQL table is used for SQL Analysis Server

2.2 Implementation of Algorithms

2.2.1 Reasons Why the Particular Tool is Chosen

The tools chosen for implementation of algorithms were Weka and SQL Analysis services. The objective of selecting these tools is to understand the basic concepts and also application of these algorithms in real time.

Weka is helpful in learning the basic concepts of data mining where we can apply different options and analyze the output that is being produced.

SQL Analysis Services is a professional tool which is used in the market currently. The Mining Accuracy Chart which includes the Lift Chart and Classification Matrix in SQL Analysis services helps to choose the suitable algorithm for the Dataset. This tool will help us understanding the implementation of data mining techniques in real world applications.

2.2.2 Implementation Procedure Used in Weka

.

The transformed ARFF file is fed in to Weka and the classification algorithms are implemented as defined in the following steps:

* In the Preprocess tab, Discretize filter is applied to discretize the attributes Src Bytes, Dst Bytes and Count in the supplied dataset (Figure 2.2).

Choosing Filter

Figure 2.2 - Choosing Filter

* In the Classify tab, choose the classification algorithm to be implemented and start the analysis to get results (Figure 2.3).

Choosing Filter

Figure 2.3 - Choosing Algorithm

* The algorithms implemented in this case study are J48 (Decision Tree), Naïve Bayes and Multilayer Perceptron (Neural Network).

* The parameters of each algorithm are changed and analyzed for improvement in accuracy and performance.

2.2.3 Implementation Procedure Used in SQL Analysis Services

The implementation procedure in SQL Analysis Services is as follows:

* A Data Source View (Figure 2.4) is created using the SQL table in which the data source is loaded.

Data Source View and Mining Structure

Figure 2.4 - Data Source View and Mining Structure

* Using the Data Source View, a Mining Structure (Figure 2.4) is created in with the Mining Models (Microsoft Decision Trees, Microsoft Neural Network, and Microsoft Naïve Bayes).

Mining Models

Figure 2.5 - Mining Models

* The content type is manually selected or is left with the tool for suggestion.

* The Mining Structure and all its related Models are processed.

* Mining Model Viewer (Figure 2.6) allows us to visualize the models.

Mining Model Viewer

Figure 2.6 - Mining Model Viewer

* The Lift Chart provides the Predict Probability for each model and Classification Matrix provides the predicted values against actual value matrix (Figure 2.7).

Classification Matrix

Figure 2.7 - Classification Matrix

2.2.4 Detailed analysis of the results with the appropriate screen dumps

In this section, the results of implemented algorithms i.e., J48, Naïve Bayes and Multilayer Perceptron in Weka and Microsoft Decision Trees, Microsoft Neural Networks and Microsoft Naïve Bayes in Microsoft SQL Analysis Services are evaluated.

2.2.4.1 Evaluation of J48 Decision Tree Results Generated in Weka

The J48 (Binary Tree) algorithm is implemented for both raw data and discretized data and also by changing the Binary Split Option. The following results are achieved.

Data J48 (Binary Split = false)(%) J48 (Binary Split = true)(%)
Raw Data 98.8335 99.1516
Discretized Data 99.2577 99.3637

Table 2.3 - Accuracy Rate of J48 in Weka

Accuracy Rate of J48 in Weka

Chart 2.1 - Accuracy Rate of J48 in Weka

The above chart shows that the accuracy rate of the predicted results is always high for discretized data when compared to raw data (un-preprocessed). This is because the Supervised Discretization used for discretization is based on the clustering of data with the knowledge of class. Also the time taken to build model for raw data is high when compared to the time taken to build model for discretized data. When Binary Split option is set to true, the accuracy is further increased. The visualization of all the binary trees is shown below(Figure 2.8, 2.9, 2.10, 2.11).

J48 Decision Tree with Raw Data and Binary Split = False

Figure 2.8 - J48 Decision Tree with Raw Data and Binary Split = False

J48 Decision Tree with Raw Data and Binary Split = True

Figure 2.9 - J48 Decision Tree with Raw Data and Binary Split = True

J48 Decision Tree with Discretized Data and Binary Split = False

Figure 2.10 - J48 Decision Tree with Discretized Data and Binary Split = False

J48 Decision Tree with Discretized Data and Binary Split = True

Figure 2.11 - J48 Decision Tree with Discretized Data and Binary Split = True

2.2.4.2 Evaluation of Naïve Bayes Results Generated in Weka

The Naïve Bayes algorithm is implemented for both raw data and discretized data and also by changing the Supervised Discretization Option. The following results are achieved.

Data Naïve Bayes (Supervised Discretization = false)(%) Naïve Bayes (Supervised Discretization = true)(%)
Raw Data 99.1516 99.4698
Discretized Data 99.5758 99.5758

Table2.4 - Accuracy Rate of Naïve Bayes in Weka

Accuracy Rate of Naïve Bayes in Weka

Chart 2.2 - Accuracy Rate of Naïve Bayes in Weka

The above results show that the Naïve Bayes accuracy rate has increase when the attributes are discretized, this is for the same reason that the discretization used in Supervised Discretization which clusters data with the knowledge of class to which it belongs to. It also shows that when the Supervised Discretization option is selected for the raw data, the accuracy rate increases drastically but is not up to the level of the preprocessed data. It is also evident that the accuracy rate increases even when Supervised Discretization is applied for the preprocessed data which is already discretized.

2.2.4.3 Evaluation of Multilayer Perceptron (Neural Networks) Generated in SQL Analysis Services

The Multilayer Perceptron algorithm is implemented for both raw data and discretized data and the following results are achieved. It is evident that the accuracy rate has increased when the algorithm is implemented on preprocessed data.

Data Multilayer Perceptron
Raw Data 98.7275
Discretized Data 99.4698

Table 2.5 - Accuracy Rate of Multilayer Perceptron in Weka

2.2.4.4 Evaluation of Results Generated in SQL Analysis Services

The algorithms Microsoft Decision Trees, Microsoft Neural Networks and Microsoft Naïve Bayes are implemented SQL Analysis Services and the following results are achieved. It is evident from the results that the accuracy rate for the Microsoft Neural Networks is high than the Microsoft Decision Trees and Microsoft Naïve Bayes for the analyzed dataset. The preprocessing of data has helped to achieve better accuracy rate for the Microsoft Decision Trees but the accuracy rate for Microsoft Neural Networks has slightly fallen down. Microsoft Naïve Bayes accuracy rate is better than the Microsoft Decision Trees but lower than the Microsoft Neural Networks. The visualization of Microsoft Decision Trees is presented below (Figure 2.12, 2.13).

Data Microsoft Decision Trees(%) Microsoft Neural Networks(%) Microsoft Naïve Bayes(%)
Raw Data 98.95 99.71 99.47
Discretized Data 99.23 99.67

Table2.6 - Accuracy Rate of various algorithms implemented in SQL Analysis Services

Accuracy Rate of various algorithms implemented in SQL Analysis Services

Chart 2.3 - Accuracy Rate of various algorithms implemented in SQL Analysis Services

Microsoft Decision Tree with Raw Data

Figure 2.12 - Microsoft Decision Tree with Raw Data

Microsoft Decision Tree with Discretized Data

Figure 2.13 - Microsoft Decision Tree with Discretized Data

2.3 Challenges of Implemented Algorithm

2.3.1 Conclusions and comments

The table below shows the overall accuracy rate for all the implemented algorithms for the Network Intrusion Detection dataset.

Data Decision Tree(%) Neural Networks(%) Naïve Bayes(%) J48 (Binary Split = ture)(%) Naïve Bayes (Supervised Discretization = true)(%)
Weka - Raw Data 98.8335 98.7275 99.1516 0.9674 0.9948
Weka - Discretized Data 99.2577 99.4698 99.5758 0.9826 0.9956
SQL Analysis Services - Raw 98.95 99.71 99.47
SQL Analysis Services - Discretized 99.23 99.67

Table 2.7 - Percentage of Correctly Classified Instances

Overall accuracy rates of implemented algorithms

Chart 2.4 - Overall accuracy rates of implemented algorithms

From the above chart it is evident that Microsoft Neural Networks has shown the highest accuracy rate for the Network Intrusion Detection dataset when it is applied for the raw data set, whereas all the other algorithms has shown better accuracy rate after preprocessing (Discretization) of data. In general the above result shows drastic accuracy rate changes among algorithms when raw data is used, where as the accuracy rate is consistent for the implemented algorithms after preprocessing of data. Hence the better accuracy rate depends on the choice of algorithm along with the parameters such as selection of attributes along with domain knowledge, preprocessing of data etc.

2.3.2 Problems encountered and the solutions

* The volume of dataset that was downloaded from Internet was huge, and was not able to open in regular document editors or MS Excel. Used Data Transformation Service (DTS) in SQL Server to import the data into the SQL Server database. Later reduced the data and exported to flat file to use it in Weka.

* Microsoft Naïve Bayes algorithm didn't accept the continuous data type. Converting continuous data type to discrete helped to overcome the problem.

* Weka sometimes crashed when the memory exceeded 256MB limit (default). Increased the memory limit by changing the value in batch file which is used to start Weka.

 

^table of contents^

 

3 Research (Part 2)

Compare, contrast and evaluate Decision Tree based algorithms for data mining with the approach taken by Neural Networks.

3.1 Introduction

In real world data is being collected everywhere, and everyone is eager in extracting knowledge out of it and use it to improve their performance. Everyone has different motive for example, business people would need to improve their business and make profit out of it. Physicians may aim at discovering knowledge out of medical records to prevent disease. Network administrators would need to keep their network secure hence they would aim at detecting anomalies to prevent intrusion. It is impractical to manually analyze the data and extract knowledge as the volume of data is very high. Hence, we aim at finding patterns to discover knowledge from the raw data which is called data mining.

There are different definitions floating around for data mining and there are various techniques to find the patterns from the raw data. Two of such techniques to extract pattern from the raw data are Decision Trees and Neural Networks. In this paper, we will compare, contrast and evaluate J48 implementation of C4.5 Decision Tree algorithm and Multilayer Perceptron which is Neural Network based algorithm.

(Simona Despa, 2003) defines "Data Mining is the process of data exploration to extract consistent patterns and relations among variables that can be used to make valid predictions".

3.2 Research contents

3.2.1 Decision Trees

The decision tree consists of three elements, root node, interior node and a leaf node. Top most element is the root node. Leaf node is the terminal element of the structure and the nodes in between is called the interior node. The decision tree is constructed based on "Divide and Conquer" [3]. That is the tree is formed by framing rules which will branch out from the nodes and sub-nodes until the decision is made. There are different methods of forming the decision rules for Decision Trees i.e., the nodes are selected from the top level based on quality attributes such as Information Gain, Gain Ratio, Gini Index etc. The C 4.5 Decision tree uses Gain Ratio to construct the tree, the element with highest gain ratio is taken as the root node and the dataset is split based on the root element values. Again the information gain is calculated for all the sub-nodes individually and the process is repeated until the prediction is completed.

Example of a Decision Tree

Figure 3.1- Example of a Decision Tree

3.2.2 Neural Networks

Multilayer Perceptron (MLP) is a neural network based algorithm which has input layer, hidden layer and output layer [3]. The neurons present in the input layer define the input values. The probabilities of inputs are assigned weights before given to the neurons in the hidden layer. Greater the weight, the neuron favors the result. Finally the neurons in the output layer represent the outcome. Neural networks can be applied to any situation virtually provided that relationship exists between the input and the output. The neurons contain activation function through which the signal is passed to predict the output. There are two types of trainings used in neural networks, supervised and unsupervised training. The supervised training algorithm generally used for MLP is back-propagation algorithm. Back propagation algorithm is where the predicted value is compared with the actual value and if the mean squared error is more, the process is repeated again until the mean squared error is minimized [4].

Example of a Neural Network

Figure 3.2 - Example of a Neural Network

3.2.3 Comparison of Decision Tree (J48) and Neural Networks (Multilayer Perceptron)

The elements of the Decision Tree are called nodes and the elements of Neural Networks are called neurons. The Decision Tree is constructed using "Divide and Conquer" based on the gain ratio i.e., the attribute with the highest information gain in considered as root node and the process is repeated, whereas Neural Networks is constructed using Back Propagation which is done recursively until mean squared error is minimal i.e., the mean squared error is calculated for the predicted value and the actual value and is repeated until it is minimal [3]. In Decision Trees prediction is based on the rules framed by splitting of nodes into branches based on information gain [8], whereas in Neural Networks the prediction is based on the activation function built as a result of back-propagation or learning process [8]. However, both of them are based on the Supervised Learning.

Decision Trees can be used for both numerical and nominal inputs and is easy to interpret where as Neural Networks can be used only for numerical data usually normalized within the range 0 to 1 and is difficult to interpret [7] [8].

Multilayer Perceptron might over fit the training unless the measures like cross validation are used wisely and similarly the drawbacks using J48 is that the splits may be unbalanced i.e., one partition much larger than the other [9]. Also the time taken to construct Mulitlayer Perceptron is very high when compared to the time taken to construct J48 Decision Tree. The data required to train Neural Network is high compared to all other algorithms.

3.2.4 Experimental Results of J48 and Multilayer Perceptron

The experimental results performed on Network Intrusion Detection presented in Section 2 shows that In Weka the accuracy of Multilayer Perceptron is less than J48 algorithm when experimented on un-preprocessed data where as Multilayer Perceptron had shown high accuracy than J48 when run on preprocessed (discretized) data. In SQL Analysis Service Neural Network has performed far high than any other algorithm (Chart 2.8).

Also, for the data which consist of 973 instances, 5 classes and 5 attributes, the training time consumed for building Multilayer Perceptron was 0.07 seconds which is very high when compared to J48 which was almost 0 seconds.

The paper "Neural Networks Vs Decision trees on Intrusion Detection" submitted by Yacine Bouzida for KDD Cup '99 concludes that neural network performs better for generalization but didn't perform well on new type of attacks (unknown classes), whereas Decision trees has performed well on both generalization and new type of attacks.

Though Neural Networks produces high degree of accuracy, the interpretation is very difficult. Hence in the IEEE Journal [10], author proposes a new form of algorithm called artificial neural-network decision tree algorithm (ANN-DT). This algorithm is based on the Neural Network to predict the values and also generates rules similar to Decision Tree which is easy to interpret. Similarly researches are being carried out in many places to extract rules from Neural Networks and they are called with different names such as artificial neural-network decision tree algorithm (ANN-DT) [10], Neural Network Extraction Rules (NNER), Tree Based Neural Networks (TBNN) [11], Rule Generation from Neural Networks [12], Fuzzy Decision Trees [13] etc.

3.3 Conclusions and comments

Though Multilayer Perceptron consumes relatively more time for learning than J48, the outcome (accuracy) is appreciable and the interpretability of Decision tree is intresting. In general, both Decision Trees and Neural Networks has advantages and drawbacks, hence the current research is towards constructing a Hybrid algorithm which encompasses advantages of both the algorithms i.e., accuracy and performance of Neural Networks and the interpretability of Decision Trees.

 

^table of contents^

 

4Appendix - Acronyms

Term Definition
CSV Comma Separated Values
DTS Data Transformation Services
ARFF Attribute-Relation File Format
MLP Multilayer perceptron
ANN-DT Artificial Neural-Network Decision Tree Algorithm
NNER Neural Networks Extraction Rules Algorithm
TBNN Tree based Neural Networks

 

^table of contents^

 

5 Bibliography

[1] Archive, U. K. (1999). "KDD Cup 1999 Data." Retrieved April 20, 2009, from http://kdd.ics.uci.edu/databases/kddcup99/task.html

[2] Corporation, T. C. (1999). Introduction to Data Mining and Knowledge Discovery. Potomac, MD (U.S.A), Two Crows Corporation.

[3] Cscu.cornell.edu, 2003 [Online] Simona Despa, 4 March 2003 Retrieved from http://www.cscu.cornell.edu/news/statnews/stnews55.pdf [Accessed on May 5, 2009]

[4] Fu, L. (1994). "Rule generation from neural networks." IEEE Transactions on Systems, Man and Cybernetics 24(8): 1114-1124.

[5] Halgamuge, S. K. and L. Wang (2005). Classification and clustering for knowledge discovery. Berlin ; New York, Springer.

[6] Han, J. and M. Kamber (2006). Data mining : concepts and techniques. Amsterdam ; Boston San Francisco, CA, Elsevier ; Morgan Kaufmann.

[7] Irena IVANOVA, M. (1995). Initialization of Neural Networks by Means of Decision Tree. Knowledge Based Systems. Sofia, Bulgaria.

[8] Janikow, C. Z. (1998). "Fuzzy decision trees: issues and methods." IEEE Transactions on Systems, Man, and Cybernetics, Part B: Cybernetics, 28(1): 1-14.

[9] José M. Jerez-Aragonésa, J. A. G.-R., Gonzalo Ramos-Jiméneza, José Muñoz-Péreza, Emilio Alba-Conejob (2003). "A combined neural network and decision trees model for prognosis of breast cancer relapse." Artificial Intelligence in Medicine 27(1): 45-63.

[10] Larose, D. T. (2005). Discovering knowledge in data : an introduction to data mining. Hoboken, N.J., Wiley-Interscience.

[11] Schmitz, G. P. J. A., C. Gouws, F.S. (1999). "ANN-DT: an algorithm for extraction of decision trees fromartificial neural networks." IEEE Transactions 10(6): 1392-1401.

[12] Stephen P. Curram, J. M. (1994). "Neural Networks, Decision Tree Induction and Discriminant Analysis: An Empirical Comparison." Operational Research Society 45(4): 440 - 450.

[13] Witten, I. H. and E. Frank (2005). Data mining : practical machine learning tools and techniques. Amsterdam ; Boston, MA, Morgan Kaufman.

[14] Yacine Bouzida (n.d.), "Neural networks vs. decision trees for intrusion detection" Source: Mitsubishi Electric ITE-TCL