As things move forward with the launches of IBM's POWER7 systems this year, there's a flurry of activities on many fronts. In particular, we've been pretty busy with the 2.6.32 kernel in many places. Red Hat has a new RHEL Version 6 in beta, the BlueBioU project of course is using 2.6.32, and Novell has announced the latest service pack for SLES 11 which is based on 2.6.32.
Red Hat's RHEL 6 is already in beta (from last month):
* http://press.redhat.com/2010/04/21/red-hat-enterprise-linux-6-beta-available-today-for-public-download/
Novell just announced their latest service pack for SLES 11 which among other things upgrades the kernel to 2.6.32:
* http://www.novell.com/promo/suse/sle11sp1.html
A key proof point at BlueBioU continues to be worked on collaboratively across many teams. Being based on the 2.6.32 enables the easy availability of a number of latest Linux technologies.
* http://bluebiou.rice.edu/
For some example performance FAQs emerging with the work on the 2.6.32 kernel base, check out:
* http://www.ibm.com/developerworks/wikis/display/LinuxP/Performance+FAQs
To check on current questions being posed see the Linux on Power architecture forum at:
* http://www.ibm.com/developerworks/forums/forum.jspa?forumID=375
Some recent questions address the holes in CPU numbering on POWER7 when SMT=2 or SMT=1 is used, how to control the DSCR settings with the ppc64_cpu command, various tools questions, how to dynamically control the SMT settings on a system, page sizes on Linux, etc. Some of the questions are driving functional updates to the commands or approaches in development, so asking a leading question there is always a good thing.
Various thoughts on the process of improving performance on a Linux system - in a mode of discovering just how much there is to learn. Customers use their systems uniquely - some care passionately about performance, some just want and expect the best "out-of-the-box" experience with no tweaking. I have observed that people in search of performance answers generally want the simple answer, but the practiced answer to any real performance question is: "Well, it depends..." - Bill Buros
Friday, May 21, 2010
Wednesday, April 7, 2010
open-source - patents - and what about performance?
Interesting debates arise around us every day. The latest "patent pledge" excitement on the web is interesting to watch. I tried to dig through and read the details of the various lists of patents, but the eyes glaze and I have to wonder who's pulling who's chain.
With some mild interest, I saw the Linux Foundation blog post this evening re-iterating the pledge from 2005 - quoting a statement from Dan Frye.
Now, I probably don't really count as an unbiased observer, working directly in Dan Frye's organization, but I will observe that in our day-to-day interactions with Linux and customers, our focus is promoting open-source solutions every day. Some of us even take some quiet personal delight in catching up and passing classic IBM proprietary solutions, but our real focus is getting customers up, running, and happy.
Admit'ably, my world is focused primarily on helping customers tune and improve the mostly open-source based deployments of fairly complex applications on Linux on POWER systems. In day to day work, I've been most impressed with the varied partners (both open-source and proprietary) that we implicitly and explicitly work with. The dedication to making things "just work" and "then work nicely" strikes me as the path that customers expect us to embrace.
The process of open-source and improving performance is usually a challenging process. We've got lists of cool performance things I'd love to see implemented. The gate to getting these pieces implemented is not the patents, it's the process of getting consensus and convincing the "community" to adopt something that'll work smoothly across the platforms. Once you have that, we've enabled our customers to have the pieces they need to implement, tune, understand, and optimize their applications and software stacks. And all in all, it's reassuring to see the calm reassurance of IBM's commitments and continued "work nice" community approach being reinforced. It's why many of us greatly prefer working in the Linux space.
Now, back to ganglia and CPU utilization. Something's not quite right there. More on that next week.
With some mild interest, I saw the Linux Foundation blog post this evening re-iterating the pledge from 2005 - quoting a statement from Dan Frye.
Now, I probably don't really count as an unbiased observer, working directly in Dan Frye's organization, but I will observe that in our day-to-day interactions with Linux and customers, our focus is promoting open-source solutions every day. Some of us even take some quiet personal delight in catching up and passing classic IBM proprietary solutions, but our real focus is getting customers up, running, and happy.
Admit'ably, my world is focused primarily on helping customers tune and improve the mostly open-source based deployments of fairly complex applications on Linux on POWER systems. In day to day work, I've been most impressed with the varied partners (both open-source and proprietary) that we implicitly and explicitly work with. The dedication to making things "just work" and "then work nicely" strikes me as the path that customers expect us to embrace.
The process of open-source and improving performance is usually a challenging process. We've got lists of cool performance things I'd love to see implemented. The gate to getting these pieces implemented is not the patents, it's the process of getting consensus and convincing the "community" to adopt something that'll work smoothly across the platforms. Once you have that, we've enabled our customers to have the pieces they need to implement, tune, understand, and optimize their applications and software stacks. And all in all, it's reassuring to see the calm reassurance of IBM's commitments and continued "work nice" community approach being reinforced. It's why many of us greatly prefer working in the Linux space.
Now, back to ganglia and CPU utilization. Something's not quite right there. More on that next week.
Wednesday, March 24, 2010
Busy month for Linux - NCSA picks Linux for Blue Waters
Along the lines of interesting - but not particularly surprising - trends in the industry, we notice that the NCSA has recently endorsed the continuing evolution of HPC workloads towards a Linux base.
Check out the NCSA article titled: "Linux selected as operating system for Blue Waters".
This is particularly encouraging for not only the operating system and in our case the POWER7 platform, but also for the various software stack products that are being developed, improved, enhanced, and deployed in high-demand HPC environments across many industries.
Massive HPC clusters on the scale of projects like Blue Waters are exciting just looking at all of the technologies being worked on. There's a nice Blue Waters project newsletter available which shows the breadth of activities happening around this project.
One of our real-life challenges is to improve the whole food-chain of software products which enable easier deployment of HPC applications on the Linux base and the POWER7 platform. Collaborative projects like the Blue BioU project can provide an expanding community of Linux HPC users access to open-source based POWER7 clusters.
Check out the NCSA article titled: "Linux selected as operating system for Blue Waters".
This is particularly encouraging for not only the operating system and in our case the POWER7 platform, but also for the various software stack products that are being developed, improved, enhanced, and deployed in high-demand HPC environments across many industries.
Massive HPC clusters on the scale of projects like Blue Waters are exciting just looking at all of the technologies being worked on. There's a nice Blue Waters project newsletter available which shows the breadth of activities happening around this project.
One of our real-life challenges is to improve the whole food-chain of software products which enable easier deployment of HPC applications on the Linux base and the POWER7 platform. Collaborative projects like the Blue BioU project can provide an expanding community of Linux HPC users access to open-source based POWER7 clusters.
Monday, March 8, 2010
IBM donates Linux-based POWER7 super-computer
Rice University has deployed a sweet POWER7 based cluster based on Linux. The brand new IBM 750 systems in the cluster are based on Linux, Maui, Torque, openmpi, Infiniband, 10Gb Ethernet, the Advance Toolchain, IBM compilers and IBM's ESSL math libraries.
Check out http://bluebiou.rice.edu/.
Over the coming weeks, we'll describe numerous collaborative efforts underway to develop, deploy, leverage, and execute workloads on the cluster.
One of the exciting aspects of the project is the team at Rice University is well versed in managing and deploying open-source-based production-level HPC clusters used by hundreds of students and researchers.
The initial goal is to integrate the cluster into an existing Maui/Torque infrastructure deployed and in use at Rice University. Extending the infrastructure to allow researchers control of POWER7 SMT hardware threads on each node, the number of 16MB huge pages if desired, energy optimization techniques, and POWER tuning techniques.
Check out http://bluebiou.rice.edu/.
Over the coming weeks, we'll describe numerous collaborative efforts underway to develop, deploy, leverage, and execute workloads on the cluster.
One of the exciting aspects of the project is the team at Rice University is well versed in managing and deploying open-source-based production-level HPC clusters used by hundreds of students and researchers.
The initial goal is to integrate the cluster into an existing Maui/Torque infrastructure deployed and in use at Rice University. Extending the infrastructure to allow researchers control of POWER7 SMT hardware threads on each node, the number of 16MB huge pages if desired, energy optimization techniques, and POWER tuning techniques.
Thursday, February 18, 2010
And here comes POWER7 !
While it's been some time since I last posted much, time flies when you're working on a new generation of hardware and systems.
The latest POWER7 based systems were just recently announced. Naturally, Linux is supported. Linux already exploits the POWER7 technologies and more is coming. For example, SLES 11 has the POWER7 enabling available and was used for numerous standard benchmarks used when we launch systems and operating system updates. For some quick performance data, check out this link.
In the days and weeks coming, several of us will be collaborating together to post insights, hints, and tips on using Linux to exploit the capabilities of IBM's latest POWER-based systems. The system capabilities of what's being delivered and what's coming down the pipeline is pretty impressive.
For more links, see Linux Performance.
For a taste of what's coming, see IBM's Statement of Direction on IBM Power Systems high-end servers.
The latest POWER7 based systems were just recently announced. Naturally, Linux is supported. Linux already exploits the POWER7 technologies and more is coming. For example, SLES 11 has the POWER7 enabling available and was used for numerous standard benchmarks used when we launch systems and operating system updates. For some quick performance data, check out this link.
In the days and weeks coming, several of us will be collaborating together to post insights, hints, and tips on using Linux to exploit the capabilities of IBM's latest POWER-based systems. The system capabilities of what's being delivered and what's coming down the pipeline is pretty impressive.
For more links, see Linux Performance.
For a taste of what's coming, see IBM's Statement of Direction on IBM Power Systems high-end servers.
Wednesday, October 22, 2008
Hey! Who's stealing my CPU cycles?!
I hear this every now and then on the Power systems from customers, programmers, and even peers. In the more recent distro versions, there's a new "st" column in the CPU metrics which tracks the usage of "stolen" CPU cycles, from the perspective of the CPU being measured. This "steal" column has been around for a while, but the most recent service packs of RHEL 5.2 and SLES 10 sp2 have the latest fixes which display the intended values - so the values are getting noticed more.
I believe this "cpu cycle stealing" all came into being when things like Xen were being developed and the programmers wanted a way to account for the CPU cycles which were allocated to another partition. I suspect the programmers were looking at it from the perspective of "my partition", where something devious and nefarious was daring to steal my CPU cycles. Thus the term "stolen CPU cycles". Just guessing though.
This "steal" term is a tad unfortunate. It's been suggested that a more gentle term of "sharing" would be preferred for customers. But digging around the source code I found the term "steal" is fairly pervasive. And what's in the code, tends to end up in the man pages. Ah well.
With Power hardware, there's a mode where the two hardware threads are juggled by the Linux scheduler. This is implemented via cpu pairs (for example, cpu0 and cpu1) which represent the schedule'able individual hardware threads running on the single processor core. This is the SMT mode (simultaneous multi-threaded) on Power.
From a performance perspective, this has tremendous advantages because the processor core can flip between the hardware threads as soon as one thread hits a short-wait for things like memory accesses. Essentially the processor core can fetch the instructions and memory accesses simultaneously for the two hardware threads which improves the efficiency of the core.
In days of old, each CPU's metrics were generally based on the premise that a CPU could get to 100% user busy. Now, the new steal column can account for the processor cycles being shared by the two SMT sibling threads, not to mention additional CPU cycles being shared with other partitions. It's still possible for an individual CPU to go to 100% user busy, while the SMT sibling thread is idle.
For example, in the vmstat output below, the rightmost CPU column is the steal column. On an idle system, this value isn't very meaningful.
In the next example, pushing do-nothing work on every CPU... (in this case a four-core system, SMT was on, so 8 CPUs were available...), we'll see the vmstat "st" column quickly get to the point where the CPU cycles on average are 50% user and 50% steal.
I just wish we could distinguish the SMT sharing of CPU cycles, and the CPU cycles being shared with other partitions.
For more details on the process of sharing the CPU cycles, especially when the CPU cycles are being shared between partitions, check out this page where we dive into more (but not yet all) of the gory details...
I believe this "cpu cycle stealing" all came into being when things like Xen were being developed and the programmers wanted a way to account for the CPU cycles which were allocated to another partition. I suspect the programmers were looking at it from the perspective of "my partition", where something devious and nefarious was daring to steal my CPU cycles. Thus the term "stolen CPU cycles". Just guessing though.
This "steal" term is a tad unfortunate. It's been suggested that a more gentle term of "sharing" would be preferred for customers. But digging around the source code I found the term "steal" is fairly pervasive. And what's in the code, tends to end up in the man pages. Ah well.
With Power hardware, there's a mode where the two hardware threads are juggled by the Linux scheduler. This is implemented via cpu pairs (for example, cpu0 and cpu1) which represent the schedule'able individual hardware threads running on the single processor core. This is the SMT mode (simultaneous multi-threaded) on Power.
- The term "hardware thread" is with respect to the processor core. Each processor core can have two active hardware threads. Software threads and software processes are scheduled on the processor cores by the operating system via the schedule'able CPUs which correspond to the two hardware threads.
From a performance perspective, this has tremendous advantages because the processor core can flip between the hardware threads as soon as one thread hits a short-wait for things like memory accesses. Essentially the processor core can fetch the instructions and memory accesses simultaneously for the two hardware threads which improves the efficiency of the core.
In days of old, each CPU's metrics were generally based on the premise that a CPU could get to 100% user busy. Now, the new steal column can account for the processor cycles being shared by the two SMT sibling threads, not to mention additional CPU cycles being shared with other partitions. It's still possible for an individual CPU to go to 100% user busy, while the SMT sibling thread is idle.
For example, in the vmstat output below, the rightmost CPU column is the steal column. On an idle system, this value isn't very meaningful.
# vmstat 1
procs ---- -------memory------- ---swap-- ---io--- --system-- -----cpu------
r b swpd free buff cache si so bi bo in cs us sy id wa st
0 0 0 14578432 408768 943616 0 0 0 0 2 5 0 0 100 0 0
0 0 0 14578368 408768 943616 0 0 0 0 25 44 0 0 100 0 0
0 0 0 14578432 408768 943616 0 0 0 32 12 44 0 0 100 0 0
0 0 0 14578432 408768 943616 0 0 0 0 21 45 0 0 100 0 0
In the next example, pushing do-nothing work on every CPU... (in this case a four-core system, SMT was on, so 8 CPUs were available...), we'll see the vmstat "st" column quickly get to the point where the CPU cycles on average are 50% user and 50% steal.
- Try using "top", then press the "1" key to see what's happening on a per-CPU basis easier..
while : ; do : ; done &For customers and technical people who were used to seeing their CPUs up to 100% user busy, this can be... disconcerting... but it's now perfectly normal.. even expected..
while : ; do : ; done &
while : ; do : ; done &
while : ; do : ; done &
while : ; do : ; done &
while : ; do : ; done &
while : ; do : ; done &
while : ; do : ; done &
# vmstat 1
procs ---- -------memory------- ---swap-- ---io--- --system-- -----cpu------
r b swpd free buff cache si so bi bo in cs us sy id wa st
8 0 0 14574400 408704 943488 0 0 0 0 26 42 50 0 0 0 50
8 0 0 14574400 408704 943488 0 0 0 0 11 34 50 0 0 0 50
8 0 0 14574400 408704 943488 0 0 0 0 26 42 50 0 0 0 50
8 0 0 14574656 408704 943488 0 0 0 0 10 34 50 0 0 0 50
I just wish we could distinguish the SMT sharing of CPU cycles, and the CPU cycles being shared with other partitions.
For more details on the process of sharing the CPU cycles, especially when the CPU cycles are being shared between partitions, check out this page where we dive into more (but not yet all) of the gory details...
Wednesday, October 15, 2008
Linux on Power.. links and portals..
Several of us were playing around recently counting up how many "portals" we could find for information in and around Linux for Power systems. On the performance side, we were specifically interested in seeing whether there was information "out there" that we could leverage that we weren't really aware of. We actually found a lot of performance information which I'll try to highlight more in the weeks coming.
For the portals, we did find an amazing assortment of web pages available. We hit classic marketing portals, hardware and performance information, generic Linux information, technical portals, a couple of old and outdated portals, various download sites for added-value items, lots of IBM forums on developerWorks, IBM Redbooks (always good information), and pointers to wiki pages spanning a number of subjects.
The marketing and customer teams generally point to the IBM Linux page as the primary entry point (portal). The five web tabs out there (Overview, Getting Started, Solutions, About Linux, and Resources) can get the reader to all sorts of official information.
For our list of web sites and pages, rather than just file the list in another email bucket never to be seen again, we created an index page called Quick Links to keep track of what we wanted to hunt down and get updated and more current. We naturally didn't want to call it another portal. 'Course, now we're hunting down subject-matter experts (aka volunteers) to help update the various wiki pages, especially under the developerWorks Linux for Power architecture wiki. We're particularly interested in providing more of the practical details, one example being the HPC Central - Red Hat page where a series of technical wiki pages are available.
Another interesting observation is seeing IBM's classic reliance on the developerWorks forums which we listed on our Quick Links index page. The Linux community is far more used to mailing lists for interactions, questions, and development issues. Forums are fine for questions and answers, but in our mind many of the forums are rarely used, even if the technology or product covered by each forum is helpful and useful. I would expect that we'll start seeding the forums with answers to questions we get from customers, developers, and peers. Help nudge things along. Which will give us more places to link to and get the practical questions answered.
[edit'ed 10/30/2008 - we made the Quick Links page the LinuxP home page]
For the portals, we did find an amazing assortment of web pages available. We hit classic marketing portals, hardware and performance information, generic Linux information, technical portals, a couple of old and outdated portals, various download sites for added-value items, lots of IBM forums on developerWorks, IBM Redbooks (always good information), and pointers to wiki pages spanning a number of subjects.
The marketing and customer teams generally point to the IBM Linux page as the primary entry point (portal). The five web tabs out there (Overview, Getting Started, Solutions, About Linux, and Resources) can get the reader to all sorts of official information.
For our list of web sites and pages, rather than just file the list in another email bucket never to be seen again, we created an index page called Quick Links to keep track of what we wanted to hunt down and get updated and more current. We naturally didn't want to call it another portal. 'Course, now we're hunting down subject-matter experts (aka volunteers) to help update the various wiki pages, especially under the developerWorks Linux for Power architecture wiki. We're particularly interested in providing more of the practical details, one example being the HPC Central - Red Hat page where a series of technical wiki pages are available.
Another interesting observation is seeing IBM's classic reliance on the developerWorks forums which we listed on our Quick Links index page. The Linux community is far more used to mailing lists for interactions, questions, and development issues. Forums are fine for questions and answers, but in our mind many of the forums are rarely used, even if the technology or product covered by each forum is helpful and useful. I would expect that we'll start seeding the forums with answers to questions we get from customers, developers, and peers. Help nudge things along. Which will give us more places to link to and get the practical questions answered.
[edit'ed 10/30/2008 - we made the Quick Links page the LinuxP home page]
Subscribe to:
Posts (Atom)
Blogs I follow
Bill Buros
Bill leads an IBM Linux performance team in Austin Tx (the only place really to live in Texas). The team is focused on IBM's Power offerings (old, new, and future) working with IBM's Linux Technology Center (the LTC). While the focus is primarily on Power systems, the team also analyzes and improves overall Linux performance for IBM's xSeries products (both Intel and AMD) , driving performance improvements which are both common for Linux and occasionally unique to the hardware offerings.
Performance analysis techniques, tools, and approaches are nicely common across Linux. Having worked for years in performance, there are still daily reminders of how much there is to learn in this space, so in many ways this blog is simply another vehicle in the continuing journey to becoming a more experienced "performance professional". One of several journeys in life.
Performance analysis techniques, tools, and approaches are nicely common across Linux. Having worked for years in performance, there are still daily reminders of how much there is to learn in this space, so in many ways this blog is simply another vehicle in the continuing journey to becoming a more experienced "performance professional". One of several journeys in life.
The Usual Notice
The postings on this site are my own and don't necessarily represent IBM's positions, strategies, or opinions, try as I might to influence them.