Bioreactor Operation Control Using Reinforcement Learning Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for manufacturing systems, such as bioreactors, fail to consistently achieve suitable operation content for efficient cell culture, leading to suboptimal yield and production outcomes.

Innovation Solution

A system comprising a bioreactor and an apparatus that utilizes reinforcement learning to determine optimal operation content by acquiring state parameters, calculating reward values, and adjusting control models to maximize production efficiency, even in varying external environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional manufacturing methods are used for cell culture, then the manufacturing process is simple and easy to operate, but the yield and production efficiency are suboptimal

Engineering Contradiction:
Improveyield and production efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements feedback mechanisms where the learning unit continuously receives operation content and state parameter data from the bioreactor system, processes this information through reinforcement learning, and generates optimized operation content that is fed back to the control unit. This closed-loop feedback system enables the system to learn from past operations and continuously improve yield and production efficiency without requiring complex manual intervention

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system employs self-service through automated reinforcement learning where the learning unit independently analyzes state parameters, determines optimal operation content, and adjusts control parameters without external intervention. The bioreactor system serves itself by automatically optimizing its own operation based on accumulated learning data, improving productivity while maintaining operational simplicity

Inventive Principle:
Principle #25Self-service

2Productivity

If trial and error methods are used to determine operation content, then the system can find suitable parameters, but the process is time-consuming and inefficient

Engineering Contradiction:
Improvedetermination efficiencyVSAvoidtime for parameter optimization
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-storing multiple types of state parameter data and operation content in the learning unit's storage system. When determining optimal operation content, the system retrieves and analyzes this pre-collected data through reinforcement learning algorithms, enabling rapid determination without time-consuming trial and error. The system prepares learning data in advance, allowing quick optimization when needed

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces manual trial-and-error mechanical adjustment with automated reinforcement learning algorithms. The learning unit uses computational methods to analyze state parameters and determine optimal operation content, substituting the time-consuming mechanical process of manual parameter adjustment with efficient computational optimization that significantly reduces determination time

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If fixed operation parameters are used, then the control process is simple, but the system cannot adapt to changing external environments

Engineering Contradiction:
Improveenvironmental adaptationVSAvoidcontrol system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamics by enabling the operation content to dynamically adapt to changing external environments. The learning unit continuously monitors state parameters and uses reinforcement learning to adjust operation content in real-time based on current conditions. This dynamic adjustment mechanism allows the system to respond flexibly to environmental changes while maintaining controlled complexity through automated learning algorithms

Inventive Principle:
Principle #15Dynamics

4Productivity

If manual optimization of operation content is performed, then the system requires less computational resources, but the manufacturing efficiency and yield are reduced

Engineering Contradiction:
Improvemanufacturing efficiencyVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by stationary object

Solution Approach 1:

The patent applies partial action by implementing reinforcement learning selectively for critical optimization tasks rather than continuously optimizing all parameters. The learning unit focuses computational resources on determining optimal operation content for key state parameters that most significantly impact yield and production efficiency. This selective approach achieves high manufacturing efficiency while avoiding excessive computational resource consumption on less critical parameters

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3816735B1Apparatus, method, and program
Publication Date: 2023.04.05 YOKOGAWA ELECTRIC CORP
  • EP3816735B1 patent drawingFigure 1
  • EP3816735B1 patent drawingFigure 2
  • EP3816735B1 patent drawingFigure 3

AI summary

An apparatus is provided, which includes a setting unit for setting an operation content for a manufacturing system configured to manufacture an object to be manufactured, a first acquisition unit for acquiring a posterior state parameter set indicating a state of at least one of the manufacturing system or the object to be manufactured after the operation content is set, and a learning processing unit for executing, by using learning data including the operation content and the posterior state parameter set, a learning process of a control model of the manufacturing system configured to output the operation content that increases a reward value determined by a preset reward function in response to input of a state parameter set indicating a state of at least one of the manufacturing system or the object to be manufactured.