Skip to main content

By clicking Submit, you agree to the developerWorks terms of use.

The first time you sign into developerWorks, a profile is created for you. Select information in your profile (name, country/region, and company) is displayed to the public and will accompany any content you post. You may update your IBM account at any time.

All information submitted is secure.

  • Close [x]

The first time you sign in to developerWorks, a profile is created for you, so you need to choose a display name. Your display name accompanies the content you post on developerworks.

Please choose a display name between 3-31 characters. Your display name must be unique in the developerWorks community and should not be your email address for privacy reasons.

By clicking Submit, you agree to the developerWorks terms of use.

All information submitted is secure.

  • Close [x]

developerWorks Community:

  • Close [x]

Background image for big data site banner

Tab navigation

01 October 2013

Top feature

System log analysis using InfoSphere BigInsights and IBM Accelerator for Machine Data Analytics

When isolating performance issues, system logs contain important clues that point to causes. But the task of manually analyzing hundreds of log files is overwhelming without the right tools to help. Learn about IBM solutions for system log analysis.


Featured download

InfoSphere Streams Quick Start Edition

Version 3.1 is available now. Find out what's new.

InfoSphere Streams Quick Start Edition puts real time analytic processing at your fingertips. Now you can analyze massive data volumes quickly (even in real time) and turn data into insight you can use to make better decisions. InfoSphere Streams can quickly ingest, analyze, and correlate information as it arrives from thousands of real-time sources. Try out the newest stream computing software — free to download, quick to start.

Download InfoSphere Streams Quick Start Edition

Highlights

Show descriptions | Hide descriptions

  • Use InfoSphere Streams and BigInsights for real-time Hadoop analytics at scale

    InfoSphere BigInsights and InfoSphere Streams pair up with Apache Hadoop to tame big data analytics in real time. Learn how to build, configure, and deploy real-time analytics to target new revenue streams.

  • SQL to Hadoop and back again, Part 1

    How do you integrate your existing SQL-based data stores with Hadoop to take advantage of different technologies when you need them? Find out in this series of articles that takes a look at a range of methods for integration between Hadoop and traditional SQL databases.

  • Working with Big SQL extended and complex data types

    Sometimes the standard set of SQL data types isn't quite enough. Big SQL -- a new SQL interface introduced in InfoSphere BigInsights -- offers a rich set of extended data types, along with complex data types that make it easier to represent and process semi-structured data. Learn how in this article, complete with code listings and sample queries.

  • Big data architecture and patterns, Part 1

    If you've got big data issues (and who doesn't?), choosing the right solution can feel like a big problem. A good first step is to classify the data according to source, format, and prescribed outcome of the analysis and processing. In Part 1 of this series, take the first step to clarify the factors at work in your big data problem.

  • Get to know the R-project Toolkit in InfoSphere Streams

    InfoSphere Streams addresses a crucial emerging need for platforms and architectures that can process vast amounts of generated streaming data in real time. Learn about the InfoSphere Streams R-project Toolkit, which integrates with the powerful R suite of software facilities and packages.

  • Getting started with real-time stream computing

    InfoSphere Streams is a powerful "platform for real-time analytics on big data" and an important part of IBM's big data architecture. Get up to speed quickly with the tools provided by InfoSphere Streams and take advantage of all the information made available for its use.

  • Do I need to learn R?

    R has proven itself a useful tool within the growing field of big data and has been integrated into several commercial packages, such as IBM SPSS and InfoSphere, as well as Mathematica. This article offers a statistician's perspective on the value of R.

  • ZooKeeper fundamentals, deployment, and applications

    Explore the fundamentals of ZooKeeper, then learn how to set up and deploy a ZooKeeper cluster in a simulated miniature distributed environment. The author also examines the use of ZooKeeper in popular projects.

  • Managing your InfoSphere Streams cluster with IBM Platform Computing

    Managing your big data infrastructure doesn’t have to be challenging. With the appropriate management strategy and tools, multiple large environments can be set up and managed efficiently and effectively. Discover how to use IBM Platform Computing to set up and manage IBM InfoSphere Streams environments that will analyze big data in real time.

  • Deploying and managing scalable web services with Flume

    Explore Flume -- a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log data, using a fault-tolerant architecture. Learn how to deploy and use Flume with a Hadoop cluster and a simple distributed web service.


Featured video

Building confidence in big data through context

Organizations -- and the people in them -- must have confidence in the big data or they simply will not use it to its fullest potential. In this discussion, Claudia Imhoff, president of Intelligent Solutions and founder of Boulder BI Brain Trust, talks with David Corrigan, director of IBM InfoSphere product marketing, about the importance of data governance, integration, security, privacy, and working with Hadoop. (8:23)     |    Watch the video

IBM big data platform capabilities

Key capabilities, at a glance

  • Hadoop-based analytics: Store any data type in the low-cost, scalable Hadoop engine to reduce the cost of processing and analyzing massive volumes of data.

  • Stream computing: Continuously analyze massive volumes of streaming data with sub-millisecond response times to take action in real time.

  • Text analytics: Analyze textual content to uncover hidden meaning and insight in unstructured information.

  • Data warehousing: Store and analyze large volumes of structured information with workload-optimized systems designed for deep and operational analytics.

Supporting capabilities, at a glance

  • Accelerators: Deploy pre-packaged analytical and industry-specific software modules to extract value from big data.

  • Application development: Develop text analytics applications with toolkits and tools, including an extensive library of extractors you can customize and extend.

  • Systems management: Monitor and manage your big data system for secure and optimized performance.

Don't see your topic?

Contribute your perspective

Our award-winning inventory of how-to content is shaped by technical experts like you. We work with new authors every day to help publish their expertise and share it with the worldwide technical community. Send us your content proposal.