An Engineer Looks At Reliability

 

Introduction

As an engineer, I have always paid attention to reliability in everything that I design, use and recommend. Since my engineering experience is largely from the automotive world, most of what I will be stating here will have a flavour from that background. I will mention other areas as well as I move along.

Reliability is one of those words we use constantly without necessarily thinking about what it really means.

We say that a car is reliable. A washing machine is reliable. An aircraft is reliable. A company is reliable. Even a person can be described as reliable.

But what does reliability actually mean?

From an engineer's perspective, reliability is much more than simply asking whether something works.

It is about whether something will continue to perform its intended function, under specified conditions, for the required period of time.

That distinction is important.

A product that works perfectly for one day is not necessarily reliable. A product that fails every year but is quickly repaired may be convenient to own, but I would hesitate to call it highly reliable.

A truly reliable system is one that does what it is supposed to do, when it is supposed to do it, repeatedly and predictably.



Reliability begins with design

One of the biggest misconceptions about reliability is that it is primarily a manufacturing or maintenance issue.

It isn't.

Reliability begins on the drawing board.

An engineer has to ask:

  • What can go wrong?
  • How often might it go wrong?
  • Under what conditions will it fail?
  • What happens if it fails?
  • Can the failure be prevented?
  • Can the failure be detected before it becomes serious?
  • Can the component be repaired or replaced?
  • What happens when several failures occur together?

These questions should be asked before the product reaches the customer.

This is why good engineering involves concepts such as Failure Mode and Effects Analysis (FMEA), design reviews, testing, validation, redundancy, tolerances, material selection and environmental testing.

The objective isn't to make failure impossible.

That would often be unrealistic and prohibitively expensive.

The objective is to make failure sufficiently unlikely, adequately predictable and largely manageable.




The weakest component can determine the reliability of the whole system

Consider a modern automobile.

It might contain an excellent engine, transmission, suspension system and braking system.

But suppose a small electronic module repeatedly fails.

The vehicle may then become unreliable despite the fact that 99% of its components are functioning perfectly.

This illustrates an important engineering principle:

System reliability depends on the reliability of its components and, critically, on how those components interact.

As systems become more complicated, there are simply more opportunities for something to go wrong.

This doesn't mean that sophisticated technology is inherently unreliable.

It means that complexity has to be justified by corresponding improvements in design, validation, manufacturing and quality control.





More features don't necessarily mean better engineering

Modern products are often marketed by counting features.

More cameras.

More screens.

More sensors.

More software.

More motors.

More connectivity.

More automated functions.

From a consumer's perspective, this can be attractive.

From an engineer's perspective, however, every additional component potentially introduces another failure mode.

Imagine a simple mechanical switch.

It has a relatively straightforward function.

Now replace it with:

switch → electronic module → communication network → software → control module → actuator.

You may have gained functionality, convenience and programmability.

But you have also created additional interfaces and dependencies.

The engineering question therefore shouldn't be:

"Can we add this feature?"

It should be:

"What additional value does this feature provide, and is that value worth the additional complexity and potential failure modes?"

That is a very different way of thinking.






Reliability is not the same as durability

These two terms are often confused and typically used interchangeably. But, they are not the same!

Durability generally concerns how well something withstands use, wear, loads and environmental conditions over time.

Reliability concerns the probability that it will perform its required function without failure for a specified period and under specified conditions.

They overlap, but they aren't identical.

A component could be extremely durable but still be unreliable if it occasionally suffers an unpredictable electronic failure.

Conversely, a component could have a relatively short service life but be highly reliable during that service life.

For example, an aircraft component may be deliberately replaced after a certain number of operating hours—not because it has necessarily failed, but because engineers have determined that replacing it at that interval provides an acceptable level of reliability.

That is an important lesson:

Preventing failure is often better engineering than waiting for failure.



Maintenance can improve reliability — but it cannot compensate for bad design

Maintenance is essential. The best equipment ins the world can become useless without maintainenance.

Oil changes, inspections, lubrication, replacement of wear components and software updates can all contribute to reliable operation.

But there is a dangerous temptation to use maintenance as a substitute for good design.

Suppose a component repeatedly fails because it is exposed to excessive heat.

We could tell the owner:

"Replace it every two years."

Or we could ask:

"Why is the component getting so hot in the first place?"

The second question is the engineering question.

This is where root-cause analysis becomes important.

Replacing a failed component may restore the system.

Finding and eliminating the reason it failed may prevent the failure from occurring again.





Fixing the symptom isn't the same as fixing the problem

Imagine a machine repeatedly shuts down because a sensor reports excessive temperature.

An inexperienced approach might be:

Replace the sensor.

If the new sensor fails in exactly the same circumstances, we have learned something.

Perhaps the sensor wasn't the problem.

Maybe:

  • the machine is genuinely overheating;
  • airflow is inadequate;
  • the sensor is incorrectly positioned;
  • vibration is damaging the wiring;
  • electrical noise is affecting the signal;
  • or the control software is interpreting the signal incorrectly.

The failed sensor was merely the symptom.

Good engineering tries to discover the root cause.

This is one reason engineers frequently ask:

"Why?"

And then ask it again.

And again.

Until the apparent problem leads to an underlying physical, electrical, software, manufacturing or human cause.








Reliability is also about simplicity

One of the most powerful lessons I have learned from engineering is that simplicity can be an enormous advantage.

A simple system isn't necessarily primitive.

It can be highly sophisticated in its design philosophy.

If two systems perform the same function, and one requires twice as many components, twice as many interfaces and considerably more complicated software, I would naturally ask:

What am I gaining for that additional complexity?

If the answer is substantial additional capability, the complexity may be worthwhile.

If the answer is merely a feature that looks impressive in a brochure, I would be less enthusiastic.

This is particularly relevant to modern consumer products.

Sometimes the best engineering solution isn't the one with the most technology.

It is the one with the least technology necessary to accomplish the objective reliably.




Reliability must be designed into the entire system

A common mistake is to think about reliability component by component.

But components don't operate in isolation.

A reliable system requires:

Reliable design + reliable components + reliable interfaces + reliable manufacturing + reliable software + appropriate maintenance + appropriate human interaction.

Consider something as simple as a connector.

The electrical component may be perfectly reliable.

But if the connector is poorly sealed against water, the system can still fail.

Similarly, an excellent mechanical component can become unreliable if it is incorrectly assembled.

This is why reliability engineering has to consider the whole system, not merely individual parts.







Reliability has a cost

There is another uncomfortable truth:

You can almost always spend more money trying to increase reliability.

A thicker material may last longer.

A higher-grade bearing may last longer.

A more sophisticated sealing system may reduce contamination.

A redundant system may continue operating after a failure.

More testing may uncover more problems before production.

But all of these things cost money.

Therefore engineering is ultimately about trade-offs.

The objective isn't necessarily maximum reliability at any cost.

It is:

The appropriate level of reliability for the intended application, at an acceptable cost.

The required reliability of a household toaster is obviously different from that of an aircraft engine.

The consequences of failure determine how much reliability we should demand.











The consequences of failure matter

This is perhaps the most important point.

Not all failures are equal.

If my toaster stops working, I may be annoyed.

If my washing machine stops halfway through a cycle, I have an inconvenience.

If my car loses an important safety function, the consequences could be considerably more serious.

If an aircraft component fails, the consequences could potentially be catastrophic.

Therefore reliability engineering is closely connected to risk management.

Engineers don't merely ask:

"How likely is this component to fail?"

They also ask:

"What happens if it does?"

A relatively unlikely failure can demand enormous attention if the consequences are severe.





The Lesson From The Automotive/Car Industry

One of the engineering philosophies I particularly admire is the emphasis placed by automotive industries, especially the Japanese, on identifying problems rather than simply avoiding or hiding them.

The underlying philosophy is powerful:

A problem is an opportunity to improve the system.

If a production line repeatedly produces a defect, the objective shouldn't simply be to increase inspection and remove defective parts.

The better question is:

Why is the process producing the defect?

If the root cause can be eliminated, the inspection requirement may eventually become less important.

This is the essence of continuous improvement.

Reliability isn't achieved once.

It is continuously engineered.








Reliability and trust

There is also a human dimension to reliability.

We trust things that behave predictably.

We trust a car that starts every morning.

We trust a bridge to support us.

We trust an aircraft to take off , fly, land and arrive safely.

We trust a financial institution to protect our money.

We trust a person who consistently does what they say they will do.

In that sense, reliability creates trust.

And trust is ultimately built through repeated evidence.

One successful event proves very little.

Thousands of successful events create confidence.

That is why reliability is fundamentally about consistency over time.

 




Conclusion - The engineer's definition of a good product

If I were asked to define a truly well-engineered product, I wouldn't necessarily choose the product with the most features.

I would choose the product that:

  • performs its intended function consistently;
  • is appropriately designed for its environment;
  • has sufficient safety margins;
  • contains no unnecessary complexity;
  • can be manufactured consistently;
  • can be maintained economically;
  • fails predictably when failure eventually occurs;
  • can be repaired without unnecessary difficulty;
  • and provides the required performance throughout its intended service life.

In other words:

Good engineering isn't about preventing every possible failure. It's about understanding failure, controlling risk and designing a system that continues to perform reliably throughout its useful life.

And perhaps that is the most important lesson of reliability engineering.






Further Reading

1. Design for Reliability: Developing Assets That Meet The Needs of Owners - Daniel T. Daley features

Engineering Maintainability : How to Design for Reliability and Engineering 


Post a Comment

0 Comments