High-strength steel rolling box optimization method based on reinforcement learning cellular algorithm

Through the reinforcement learning cellular algorithm, the plate thickness of the high-strength steel rolling box is optimized, and the limitations of material selection is solved, efficient and flexible box design is achieved, and performance and material utilization is improved.

CN120337744APending Publication Date: 2025-07-18GUANGXI JINJUSHI NEW ENERGY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510406133.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing high-strength steel rolling box has limitations in material selection. The roller rolling material thickness of each side beam and horizontal beam is single, making it difficult to find the optimal thickness for different structures, resulting in redundancy in performance and waste of materials.

Method used

The reliability model and optimization model are constructed based on reinforcement learning cellular algorithms, and the design parameters are transferred and optimized to optimize the plate thickness distribution of each functional beam by learning performance under different working conditions.

Benefits of technology

It realizes rapid and precise optimization of steel rolling boxes in different structures, reduces material waste, improves optimization efficiency and performance maximization, and reduces production costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337744A_ABST
    Figure CN120337744A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning cellular algorithm-based high-strength steel rolling box body optimization method, which comprises the following steps of: obtaining structure parameters and material parameters of a high-strength steel rolling box body, and constructing a high-strength steel rolling box body reliability model and a high-strength steel rolling box body optimization model; a reinforcement learning cellular model is constructed, and the high-strength steel rolling box reliability model learns the to-be-optimized high-strength steel rolling box according to the reinforcement learning cellular model; the high-strength steel rolling box optimization model carries out migration optimization on a to-be-optimized high-strength steel rolling box; and outputting the plate thickness distribution of each functional beam of the optimized high-strength steel rolling box body. The method has the migration optimization capacity, optimization can be conducted on another model through an existing learning model, and the most appropriate plate thickness of each beam can be rapidly solved and summarized for steel rolling box bodies of different structures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric vehicle production, and particularly relates to an optimization method for a high-strength steel roll-formed box body based on a reinforcement learning cellular algorithm. Background Art

[0002] As a power battery box body, the roll-formed high-strength steel box body uses high-strength steel and is processed into parts with complex shapes through a roll-forming process to achieve structural lightweighting, with a weight reduction of 15%-20%. The roll-formed high-strength steel box body has a lower material cost. In addition, the roll-forming process has high production efficiency and is suitable for large-scale production, further reducing the unit cost. Compared with traditional aluminum box bodies, the roll-formed high-strength steel box body is not only more economical in terms of material cost, but also shows higher efficiency in large-scale production, can effectively reduce the unit production cost, and improve market competitiveness.

[0003] At the same time, with the continuous progress of new energy vehicle technology, the concept of CTB (Cell to Body) structure battery pack-body integration design has gradually emerged. This design concept directly integrates the battery cells into the bottom of the vehicle body, sharing the battery pack and the vehicle body structural parts, thereby achieving the lightweight goal of the battery box body. The roll-formed high-strength steel box body, with its high strength and lightweight characteristics, can well meet this design requirement and provide strong support for the lightweight development of new energy vehicles.

[0004] However, there are certain limitations in the material selection of high-strength steel box bodies currently on the market. Specifically, the thickness of the roll-formed materials for each side beam and transverse and longitudinal beams is relatively single, and it is difficult to find the optimal steel plate thickness for steel roll-formed box bodies with different structures. This single material thickness selection often leads to performance redundancy in the box body design, unable to maximize performance, and thus causing unnecessary material waste, increasing production costs and resource consumption. Summary of the Invention

[0005] The present invention provides an optimization method for a high-strength steel roll-formed box body based on a reinforcement learning cellular algorithm, which has the ability of migration optimization and can optimize on another model through an existing learning model. For steel roll-formed box bodies with different structures, it can quickly solve and summarize the most suitable plate thickness for each beam. The specific technical solutions are as follows:

[0006] An optimization method for a high-strength steel roll-formed box body based on a reinforcement learning cellular algorithm, comprising the following steps:

[0007] Obtain the structural parameters and material parameters of the high-strength steel roll-formed box body, and construct a reliability model and an optimization model for the high-strength steel roll-formed box body;

[0008] Build a reinforcement learning cellular model. The reliability model of the high-strength steel roll-formed box body learns about the high-strength steel roll-formed box body to be optimized according to the reinforcement learning cellular model;

[0009] The optimization model of the high-strength steel roll-formed box body performs migration optimization on the high-strength steel roll-formed box body to be optimized;

[0010] Output the plate thickness distribution of each functional beam of the optimized high-strength steel roll-formed box body.

[0011] Preferably, the reliability model of the high-strength steel roll-formed box body constructs a side column extrusion working condition model and a simulated collision working condition model.

[0012] Preferably, the reliability model of the high-strength steel roll-formed box body learns about the high-strength steel roll-formed box body to be optimized according to the reinforcement learning cellular model, including the following steps:

[0013] Set the number of learning times, optimization goal, and initialize the value function;

[0014] Initialize the design variables and obtain the cellular state;

[0015] The cell selects a cell action according to the decision-making strategy to obtain the next state of the cell;

[0016] Calculate the cell value function and record the cell state, cell action, and next state;

[0017] Use the next state of the cell as the input for iteration until the learning result converges and reaches the set number of learning times.

[0018] Preferably, setting the number of learning times, optimization goal, and initializing the value function includes the following steps:

[0019] Take each unit in the structural grid as an agent cell, take the unit density and field variables as the state functions describing the agent, and define the neighborhood range through the von Neumann type neighborhood;

[0020] Describe the state of the agent cell through variable values, strain energy, and neighborhood state;

[0021] Construct the value function of the agent cell;

[0022] Set the number of learning times and optimization goal of the reliability model of the high-strength steel roll-formed box body.

[0023] Preferably, the optimization model of the high-strength steel roll-formed box body performs migration optimization on the high-strength steel roll-formed box body to be optimized, including the following steps:

[0024] Set the optimization parameters and read the value function;

[0025] Initialize the design variables and obtain the cellular state;

[0026] The cells select cell actions according to the decision-making strategy to obtain the next state of the cells;

[0027] Take the next state of the cells as the input for iteration until the learning result converges.

[0028] Preferably, the setting of the optimization parameters and the reading of the value function include the following steps:

[0029] Map the density parameter to the elastic modulus by the variable density method, and use the elastic modulus as the optimization parameter;

[0030] Maximize the stiffness and minimize the volume of each rolling beam of the rolling box as the optimization goal;

[0031] Read the value function of the reinforcement learning cell model.

[0032] Preferably, the value function of the agent cell is:

[0033] R(S,a) = αR1 + βR2 + γR .3

[0034] Among them, R(S,a) represents the reward obtained by the cell when selecting action a in state s; α, β, and γ are the weight coefficients of the three benefit sub-functions respectively; R1 represents the benefit sub-item of the design variable; R2 represents the benefit sub-item of the strain energy; R3 represents the benefit sub-item of the unit neighborhood state.

[0035] Preferably, the agent cells use the same value function for actions.

[0036] An optimization system for a high-strength steel rolling box based on a reinforcement learning cell algorithm, including a first calculation unit for obtaining the structural parameters and material parameters of the high-strength steel rolling box, and constructing a reliability model and an optimization model of the high-strength steel rolling box;

[0037] A second calculation unit for constructing a reinforcement learning cell model, and the reliability model of the high-strength steel rolling box learns the high-strength steel rolling box to be optimized according to the reinforcement learning cell model;

[0038] A third calculation unit for performing migration optimization on the high-strength steel rolling box to be optimized by the high-strength steel rolling box optimization model;

[0039] A fourth calculation unit for outputting the plate thickness distribution of each functional beam of the optimized high-strength steel rolling box.

[0040] An optimization device for a high-strength steel roll-formed box body based on a reinforcement learning cellular algorithm, the device comprising a processor and a memory; the memory is used to store program code and transmit the program code to the processor;

[0041] The processor is configured to execute the steps of the above-mentioned high-strength steel roll-formed box body optimization method based on the reinforcement learning cellular algorithm according to the instructions in the program code.

[0042] A computer-readable storage medium for storing program code for executing the steps of the above-mentioned high-strength steel roll-formed box body optimization method based on the reinforcement learning cellular algorithm.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] The present invention combines topology optimization with reinforcement learning theory and proposes a reinforcement learning cellular optimization algorithm for optimizing each side beam and transverse and longitudinal beams of a high-strength steel roll-formed box body. It includes a learning mode and an optimization mode. Through the transfer optimization method of learning in one mode and optimizing in another model, compared with the traditional SIMP method and BESO method, it saves the model reconstruction time and greatly improves the solution efficiency. Using this method, for high-strength steel roll-formed box bodies with different structures, the optimal plate thickness can be quickly solved on the basis of the original learning mode, breaking the limitation of the single roll-formed steel plate thickness of the existing high-strength steel roll-formed box body, realizing the maximization of performance, and reducing material waste. Specifically, through the reinforcement learning cellular model, the reliability model can learn the performance of the high-strength steel roll-formed box body under different working conditions, thereby providing data support for the optimization model. The optimization model then uses the learned knowledge to adjust the design parameters through the transfer optimization method to achieve rapid and precise optimization of the high-strength steel roll-formed box body. This method not only improves the optimization efficiency but also can find the optimal plate thickness distribution for box bodies with different structures, avoiding the time waste caused by model reconstruction in traditional methods, and providing a more efficient and flexible solution for the design of high-strength steel roll-formed box bodies. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally denoted by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0046] Figure 1 It is a flowchart of the method of the present invention.

[0047] Figure 2 It is a flowchart of the learning method of the present invention.

[0048] Figure 3 This is the flow chart of the optimization method of the present invention. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0051] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0052] It should be further understood that the term " / and / " used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0053] Embodiment 1

[0054] An optimization method for a high-strength steel roll-formed box based on a reinforcement learning cellular algorithm as shown in the figure includes the following steps:

[0055] Obtain the structural parameters and material parameters of the high-strength steel roll-formed box, and construct a reliability model and an optimization model for the high-strength steel roll-formed box;

[0056] Construct a reinforcement learning cellular model, and the reliability model of the high-strength steel roll-formed box learns about the high-strength steel roll-formed box to be optimized according to the reinforcement learning cellular model;

[0057] The optimization model of the high-strength steel roll-formed box performs migration optimization on the high-strength steel roll-formed box to be optimized;

[0058] Output the plate thickness distribution of each functional beam of the optimized high-strength steel roll-formed box.

[0059] This optimization method first obtains the structural parameters and material parameters of the high-strength steel roll-formed box body, constructs a reliability model and an optimization model. Then, a reinforcement learning cell model is constructed, and the reliability model learns about the high-strength steel roll-formed box body based on this model, learning its performance and characteristics under different working conditions to understand the reliability and performance of the structure. Next, the optimization model uses the learned knowledge to perform migration optimization on the high-strength steel roll-formed box body, adjusting the design parameters according to the learned characteristics, such as the plate thickness distribution of the functional beams, to achieve the optimization goal. Finally, the plate thickness distribution of each functional beam after optimization is output, realizing the optimized design of the box body structure, improving its performance and reliability, and at the same time possibly increasing the material utilization rate and reducing costs.

[0060] When performing the topology optimization of the structure, first, the original finite element model is initialized. This process involves converting each element in the model into an independent cell, and each cell has unique attributes, materials, and identification numbers. Subsequently, this transformed model is submitted to the solver LS-DYNA for computational analysis, with the focus on extracting key response data such as strain energy density. To further optimize the structure, based on the preset neighborhood definition and cell evolution criterion, the relative density of each cell is dynamically adjusted. This process is achieved through multiple iterations until the entire system reaches a stable state, that is, an effective structure topology optimization cycle is completed. By initializing the finite element model, decomposing it into independently operable cells, using the solver to analyze and extract important information, and then adjusting the cell attributes according to specific rules, and iterating repeatedly until the optimization goal is achieved, the precise topology optimization of the structure is realized.

[0061] Example 2

[0062] The difference between this example and Example 1 is that the reliability model of the high-strength steel roll-formed box body constructs a side column extrusion working condition model and a simulated collision working condition model.

[0063] These two working condition models respectively simulate the side column extrusion and collision situations that the box body may encounter during actual use. By adding these two models to the reliability model, the reliability and performance of the high-strength steel roll-formed box body under different working conditions can be more comprehensively evaluated, so that these actual working conditions can be fully considered during the optimization process, making the optimized box body design more in line with the actual use requirements and improving its reliability and safety in actual applications.

[0064] For the side column extrusion condition model: The purpose of the extrusion analysis of the power battery box is to evaluate the ability of the power battery Pack to withstand extrusion during the vehicle collision, ensuring that no accidents such as fire or explosion occur to the battery pack under the vehicle collision condition. It is mainly carried out in two directions: X (the vehicle driving direction) and Y (the horizontal direction perpendicular to the driving direction). First, it is necessary to select a suitable extrusion plate according to the structural characteristics of the object under study. There are mainly two specifications used, both of which are semi-cylinders with a radius of 75 mm, and their length (L) is greater than the size of the battery pack to be extruded (such as Figure 1 as shown), and at the same time, there is a given standard for the speed requirement of the extrusion experiment, defined as not greater than 2 mm / s. Furthermore, use the selected semi-cylinders to extrude the battery pack system in the X and Y directions respectively, and then solve the deformation amount of the power battery box through finite element software. When the extrusion deformation amount reaches 30% of the overall size in the extrusion direction or the extrusion force reaches 100 kN, stop the extrusion, and the corresponding displacement-time curves in the X and Y directions. Finally, by observing the extrusion deformation amount of the battery pack box structure and the damage degree of the internal modules and other phenomena, to judge whether the structure meets the extrusion requirements.

[0065] For the simulated collision condition model: The simulated collision analysis of the power battery box mainly tests the ability of the battery pack to withstand impact loads during the vehicle collision, ensuring the mechanical structure safety of the battery pack under the vehicle collision condition, not squeezing the battery cells, and no accidents such as fire or explosion occur. Usually, for the extreme collision condition suffered by the battery pack system, its excitation is generally reflected by the magnitude of the acceleration pulse. The test needs to install the battery pack on the vehicle frame, and then combine the vehicle body weight to apply acceleration pulses to the battery pack system in two directions: X (the vehicle driving direction) and Y (the horizontal direction perpendicular to the driving direction) at the same time, and use finite element software to solve the position and magnitude of the maximum plastic strain of the power battery box to judge whether the structure meets the collision requirements.

[0066] Example 3

[0067] The difference between this example and Example 2 is that the reliability model of the high-strength steel roll-formed box learns the high-strength steel roll-formed box to be optimized according to the reinforcement learning cell model, including the following steps:

[0068] Set the number of learning times, optimization objectives, and initialize the value function;

[0069] Initialize the design variables and obtain the cell state;

[0070] The cell selects the cell action according to the decision-making strategy to obtain the next state of the cell;

[0071] Calculate the cell value function, and record the cell state, cell action, and next state;

[0072] Iterate with the next state of the cell as the input until the learning result converges and the set number of learning times is reached.

[0073] First, set the number of learning times, optimization objective, and initialize the value function to clarify the scope and objective of learning, as well as the initial value function for evaluation. Then, initialize the design variables and obtain the initial state of the cells. These cells represent the individual units in the structure, and their states reflect the current situation of the units. Next, the cells select actions according to the decision-making strategy, and this action determines the next state of the cells, simulating the changes of the structure under different conditions. Subsequently, calculate the value function of the cells to evaluate the quality of the current state and action, and record the relevant information. Finally, iterate with the next state of the cells as the input, continuously repeat the above process until the learning result converges and the set number of learning times is reached. Through this process, the model can learn the performance and reliability characteristics of the high-strength steel roll-formed box under different working conditions and conditions, providing a more accurate basis for subsequent optimization, thereby achieving more precise structural optimization.

[0074] Embodiment 4

[0075] The difference between this embodiment and Embodiment 3 is that setting the number of learning times, optimization objective, and initializing the value function includes the following steps:

[0076] Take each unit in the structural grid as an agent cell, take the unit density and field variables as state functions describing the agent, and define the neighborhood range through the von Neumann-type neighborhood.

[0077] Describe the state of the agent cell through variable values, strain energy, and neighborhood state.

[0078] Construct the value function of the agent cell.

[0079] Set the number of learning times and optimization objective of the reliability model of the high-strength steel roll-formed box.

[0080] The two most important objects in the reinforcement learning model are the agent and the environment. To solve the topology optimization problem using the reinforcement learning method, the first step is to define the agent in the problem. By breaking down the structure into smaller parts, a single unit in the structural grid is considered as an agent. Since the shapes of individual units in the grid are basically the same, and during the solution process, only the changes in its own density value x and field variables such as displacement stress occur for a single unit, the unit density and field variables can be used as the state functions to describe the agent, while the change in density is regarded as the action of the agent. The cellular automaton method provides a way to establish a connection between the individual and the whole. By combining the concept of cells in cellular automata with that in finite elements, each structural grid unit exists as an agent. By integrating the cell sensing behavior and the agent learning behavior, they are called agent cells. Each agent cell only senses the state of its own neighborhood and makes decisions independently. In this way, the overall action space and state space can be reduced to the action space and state space of a single cell, greatly reducing the number of state and action variables described and making the problem solvable.

[0081] Regarding the selection of agents in cellular automata: For each cell in the system, regardless of its specific location, the update rule they follow is unified. The core of this rule is to collect and analyze the state information of other cells in the neighborhood of each cell, so that the cell can adjust its own state based on this information. The concept of the von Neumann neighborhood is used to collect and analyze the neighborhood state. By quantifying the central distance between the cells in the neighborhood and the central cell, the cells whose centers are within the neighborhood radius of a certain cell can be regarded as the neighborhood of that cell. The corresponding neighborhood radius of the von Neumann type is r = a, and the diameter of an equal-area circle or an equal-volume sphere is used to equivalent the side length of the cell, and its calculation formula is as follows:

[0082]

[0083] After introducing the neighborhood radius, the concept of the cell neighborhood can be extended to irregular grids.

[0084] When the strain energy density of the unit rises to the peak value allowed by the material properties, the density of the unit will tend to be stable and will no longer change with the optimization process. At this time, the relative density of the unit is locked at 1, indicating that it has been transformed into a fully filled, void-free solid material state. Therefore, the optimization objective function is expressed as:

[0085]

[0086] In the formula: They are respectively the average strain energy density and the strain energy density target value of unit i, and N is the total number of units; It is the minimum value of the design variable, taking 0.001 to avoid a singular matrix.

[0087] Description of agent state: The state of the agent cell needs to record the feedback from the environment to the agent and the perception of the neighborhood state. Use n variables to describe the state, and record the state of the cell in an n - dimensional array:

[0088] s i =[α1 α2 α3 ... α n (4)

[0089] where Si represents the state of cell i, and α i represents a variable.

[0090] To improve the learning efficiency, only three variables are used to describe the cell state, namely the design variable value x, the strain energy U, and the neighborhood state s. Since the strain energy comprehensively considers the strain and stress states of the cell, and the energy is a scalar, it is easy to compare, analyze, and calculate. The calculation method of the neighborhood state is as follows:

[0091]

[0092] In the formula, s e is the neighborhood state of cell e: n is the number of cells in the neighborhood; η i is the weight factor of cell i in the neighborhood; U i is the strain energy of cell i in the neighborhood; w is a function of the cell center - to - center distance r ie This function can adjust its weight factor according to the distance of each cell in the neighborhood from the central cell.

[0093] Example 5

[0094] The difference between this example and Example 4 is that the optimization model of the high - strength steel roll - pressed box body performs migration optimization on the high - strength steel roll - pressed box body to be optimized, including the following steps:

[0095] Set the optimization parameters and read the value function;

[0096] Initialize the design variables and obtain the cell state;

[0097] The cell selects the cell action according to the decision strategy to obtain the next state of the cell;

[0098] Use the next state of the cell as the input for iteration until the learning result converges.

[0099] First, set the optimization parameters and read the value function, which clarifies the parameters to be adjusted during the optimization process and the basis for evaluation. Then, initialize the design variables and obtain the initial state of the unit cells to prepare for the optimization process. Next, the unit cells select actions according to the decision-making strategy to obtain the next state of the unit cells, simulating the changes in the structure during the optimization process. Finally, use the next state of the unit cells as the input for iteration and continuously optimize until the learning result converges. Through this process, the optimization model can gradually adjust the design variables based on the previously learned knowledge and the evaluation of the value function, optimize the structure of the high-strength steel roll-formed box body, and make its performance and reliability reach the expected optimization goal.

[0100] Example 6

[0101] The difference between this example and Example 5 is that the setting of the optimization parameters and the reading of the value function include the following steps:

[0102] Map the density parameter to the elastic modulus by the variable density method, and use the elastic modulus as the optimization parameter;

[0103] Maximize the stiffness and minimize the volume of each roll-formed beam of the roll-formed box body as the optimization goal;

[0104] Read the value function of the reinforcement learning unit cell model.

[0105] In the variable density method, an artificial assumption is made about a density-variable material unit that does not exist in actual engineering. According to the discretized topology optimization modeling idea, the density of this material unit is set as a continuous variable between [0, 1], and the same density is assigned to each discretized unit, and this is used as the optimization variable. The cellular automaton describes the state by assigning a density parameter x i to each unit, and assigns a density parameter to each unit to characterize its state. Through the penalty function formula in the variable density method, this abstract density parameter is mapped to specific physical properties and thus can be reflected in numerical simulations. Given that the elastic modulus has a significant impact on the stiffness of the unit, it is used as the target for mapping the density parameter, so as to effectively simulate the "void" or "solid" state of the unit at the mechanical level, that is, whether the unit exists in the structure. Map x i to the elastic modulus of the unit. The calculation formula for the unit elastic modulus is as follows:

[0106] E i (x i ) = x i p E0(7)

[0107] Considering the use of the elastoplastic material constitutive model in the structural design of each functional beam of the roll-formed battery box body, on this basis, the initial yield stress σ y0 and the strain hardening modulus Eh0 Thus, the yield stress and strain hardening modulus of element i can be obtained respectively:

[0108] σ yi (x i ) = x i p σ y0 (8)

[0109] E hi (x i ) = x i p E h0 (9)

[0110] Therefore, the optimization of the structural topology shape is transformed into the optimization of parameter x i , and parameter x i is the optimization variable.

[0111] When setting the optimization goal, the maximization of the stiffness of each rolling beam of the rolling box while minimizing the volume is taken as the optimization goal. The strain energy density, as a local rigidity index, can achieve the optimization criterion of uniform distribution of strain energy, and the stiffness can be quantified by the strain energy. Therefore, the maximization of stiffness is equivalent to the minimization of strain energy. The calculation formula of the structural strain energy is:

[0112] C(X) = U T KU(10)

[0113] The calculation formula of the structural volume is:

[0114]

[0115] Where: C(X) is the structural strain energy function; X = {x i} is the element density vector; K is the total stiffness matrix of the structure; U is the structural nodal displacement vector.

[0116] Regarding the constraint conditions, two parameters are selected to evaluate whether the structure meets the performance requirements: the morphological stability and the maximum stress level of the structure. Selecting the structural configuration state and the maximum stress of the structure to evaluate whether the element density xi meets the value limit and mechanical property requirements. The morphological stability is evaluated by monitoring the changes of the structure during the optimization process, and observing whether the structure evolves from the original rigid body state to the mechanism state. Once the structure transforms into a mechanism, it means that it has lost its original load-bearing capacity and is therefore regarded as a failure state. This transformation can be intuitively reflected by monitoring the maximum displacement of the structure.

[0117] The maximum stress of the structure must be controlled within the stress range allowed by the material to ensure the safety and reliability of the structure. Considering the stresses that the structure may bear in different directions, evaluate the effects of these stresses to ensure that none of them exceed the limit of the material. The maximum mechanical abuse that the battery box can withstand should meet the requirement of not causing internal short-circuit failure.

[0118] Example 7

[0119] The difference between this example and Example 6 is that the value function of the intelligent cell is as follows:

[0120] R(S,a) = αR1 + βR2 + γR .3

[0121] Wherein, R(S,a) represents the reward obtained by the cell when selecting action a in state s; α, β, and γ are the weight coefficients of the three sub-reward functions respectively; R1 represents the sub-reward item of the design variable; R2 represents the sub-reward item of the strain energy; R3 represents the sub-reward item of the neighborhood state of the cell.

[0122] The intelligent cell uses the same value function for actions.

[0123] Because the objectives and constraints of topology optimization need to be reflected in the reward function of the intelligent agent to have an impact on the decision-making of the intelligent agent. In general topology optimization problems, it is desired to maximize the structural stiffness while minimizing the volume as much as possible, that is, to reduce the amount of material used as much as possible.

[0124] To meet the above requirements, several sub-reward functions are constructed:

[0125]

[0126] R1 represents the sub-reward item of the design variable; R2 represents the sub-reward item of the strain energy; R3 represents the sub-reward item of the neighborhood state of the cell. x′ e ,U′ e ,s′ e , respectively represent the design variable value, strain energy, and neighborhood state of the next state after executing the action. Since the goal is to minimize the structural strain energy as much as possible, when constructing the reward function, the above-mentioned rewards are processed negatively to ensure the consistency of the goal and the correctness of the direction, and the sub-reward functions are unified to obtain the reward R(S,a) of the cell after performing action a in state S;

[0127] R(S,a) = αR1 + βR2 + γR .3 (13)

[0128] In the formula, R(S,a) represents the reward obtained by the cell when selecting action a in state s; α, β, and γ are the weight coefficients of the three sub-reward functions respectively.

[0129] Regarding the action mode of the agent: During the optimization process of the reinforcement learning cellular method, all units of the structured grid are agent cells, and all cells share the same action-value function. However, in each iteration step, all cells independently select actions and obtain rewards. Since each cell makes independent decisions and takes independent actions, a large amount of feedback data can be obtained by the value function in one learning process. This method solves the problem of too many design variables and can accelerate the learning speed.

[0130] It can be found from the definition of the agent's state that only the design variable x and strain U of the agent are stored in the state, and the information related to structural loads and constraints is not saved. Therefore, the action-value function of the agent is independent of the structure form. An agent that has completed learning under one structure is very likely to be used to optimize another different structure. This method is also called transfer optimization, which makes this optimization method highly versatile.

[0131] Embodiment 8

[0132] A high-strength steel roll-formed box optimization system based on the reinforcement learning cellular algorithm, including a first calculation unit for obtaining the structural parameters and material parameters of the high-strength steel roll-formed box, and constructing a reliability model and an optimization model of the high-strength steel roll-formed box;

[0133] A second calculation unit for constructing a reinforcement learning cellular model, and the reliability model of the high-strength steel roll-formed box learns about the high-strength steel roll-formed box to be optimized according to the reinforcement learning cellular model;

[0134] A third calculation unit for performing transfer optimization on the high-strength steel roll-formed box to be optimized by the optimization model of the high-strength steel roll-formed box;

[0135] A fourth calculation unit for outputting the plate thickness distribution of each functional beam of the optimized high-strength steel roll-formed box.

[0136] Embodiment 9

[0137] A high-strength steel roll-formed box optimization device based on the reinforcement learning cellular algorithm, the device includes a processor and a memory; the memory is used to store program codes and transmit the program codes to the processor;

[0138] The processor is used to execute the steps of the high-strength steel roll-formed box optimization method based on the reinforcement learning cellular algorithm described in Embodiments 1-7 according to the instructions in the program codes.

[0139] Embodiment 10

[0140] A computer-readable storage medium for storing program code for performing the steps of the high-strength steel roll-formed box body optimization method based on the reinforcement learning cellular algorithm described in Embodiments 1-7.

[0141] In summary, the present invention combines topology optimization with reinforcement learning theory, and proposes a reinforcement learning cellular optimization algorithm for optimizing the side beams and transverse and longitudinal beams of a high-strength steel roll-formed box body. It includes a learning mode and an optimization mode. Through the transfer optimization method of learning in one mode and optimizing in another mode, compared with the traditional SIMP method and BESO method, it saves the model reconstruction time and greatly improves the solution efficiency. By using this method, for high-strength steel roll-formed box bodies with different structures, the optimal sheet thickness can be quickly solved on the basis of the original learning mode, breaking the limitation of the single roll-formed steel sheet thickness of the existing high-strength steel roll-formed box body, maximizing the performance, and reducing material waste. Specifically, through the reinforcement learning cellular model, the reliability model can learn the performance of the high-strength steel roll-formed box body under different working conditions, thereby providing data support for the optimization model. The optimization model then uses the learned knowledge to adjust the design parameters through the transfer optimization method to achieve rapid and accurate optimization of the high-strength steel roll-formed box body. This method not only improves the optimization efficiency, but also can find the optimal sheet thickness distribution for box bodies with different structures, avoiding the time waste caused by model reconstruction in the traditional method, and providing a more efficient and flexible solution for the design of high-strength steel roll-formed box bodies.

[0142] Those of ordinary skill in the art can realize that the units of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components of each example have been generally described according to their functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0143] In the embodiments provided by the present invention, it should be understood that the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units can be combined into one unit, one unit can be split into multiple units, or some features can be ignored, etc.

[0144] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0145] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0146] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.

Claims

1. An optimization method for a high-strength steel roll-formed box body based on a reinforcement learning cellular algorithm, characterized in that, It includes the following steps: Obtain the structural parameters and material parameters of the high-strength steel roll-formed box body, and construct a reliability model and an optimization model for the high-strength steel roll-formed box body; Construct a reinforcement learning cellular model, and the reliability model of the high-strength steel roll-formed box body learns the high-strength steel roll-formed box body to be optimized according to the reinforcement learning cellular model; The optimization model of the high-strength steel roll-formed box body performs migration optimization on the high-strength steel roll-formed box body to be optimized; Output the plate thickness distribution of each functional beam of the optimized high-strength steel roll-formed box body.

2. According to the method for optimizing a high-strength steel roll-formed box body based on a reinforcement learning cellular algorithm described in claim 1, the reliability model of the high-strength steel roll-formed box body constructs a side column extrusion working condition model and a simulated collision working condition model.

3. A high-strength steel roll-formed box body optimization method based on a reinforcement learning cellular algorithm according to claim 1, characterized in that The reliability model of the high-strength steel roll-formed box body learns the high-strength steel roll-formed box body to be optimized according to the reinforcement learning cellular model, including the following steps: Set the number of learning times, optimization objectives, and initialize the value function; Initialize the design variables and obtain the cellular state; The cell selects a cell action according to the decision-making strategy to obtain the next state of the cell; Calculate the cell value function, and record the cell state, cell action, and the next state; Use the next state of the cell as the input for iteration until the learning result converges and reaches the set number of learning times.

4. The optimized method for a high-strength steel roll-formed box body based on a reinforcement learning cellular algorithm according to claim 3, characterized in that The setting of the number of learning times, optimization objectives, and initialization of the value function includes the following steps: Regard each unit in the structural grid as an agent cell, regard the unit density and field variables as state functions describing the agent, and define the neighborhood range through the von Neumann type neighborhood; Describe the state of the agent cell through variable values, strain energy, and neighborhood state; Construct the value function of the agent cell; Set the number of learning times and optimization objectives of the reliability model of the high-strength steel roll-formed box body.

5. A high-strength steel roll-formed box body optimization method based on a reinforcement learning cellular algorithm according to claim 1, characterized in that, The optimization model of the high-strength steel roll-formed box body performs migration optimization on the high-strength steel roll-formed box body to be optimized, including the following steps: Set the optimization parameters and read the value function; Initialize the design variables and obtain the cellular state; The cell selects a cell action according to the decision-making strategy to obtain the next state of the cell; Use the next state of the cell as the input for iteration until the learning result converges.

6. A high-strength steel roll-formed box body optimization method based on a reinforcement learning cellular algorithm according to claim 5, characterized in that The setting of the optimization parameters and reading of the value function includes the following steps: Map the density parameter to the elastic modulus through the variable density method, and use the elastic modulus as the optimization parameter; Maximize the stiffness and minimize the volume of each roll-formed beam of the roll-formed box body as the optimization objective; Read the value function of the reinforcement learning cellular model.

7. A high-strength steel roll-formed box body optimization method based on a reinforcement learning cellular algorithm according to claim 1, characterized in that The value function of the agent cell is: R(S,a) = αR1 + βR2 + γR .3 Among them, R(S,a) represents the reward obtained by the cell when selecting action a in state s; α, β, γ are the weight coefficients of the three benefit sub-functions respectively; R1 represents the benefit sub-function of the design variable; R2 represents the benefit sub-function of the strain energy; R3 represents the benefit sub-function of the unit neighborhood state.

8. An optimization system for a high-strength steel roll-formed box body based on a reinforcement learning cellular algorithm, including a first calculation unit for obtaining the structural parameters and material parameters of the high-strength steel roll-formed box body, and constructing a reliability model and an optimization model for the high-strength steel roll-formed box body; A second computing unit for constructing a reinforcement learning cellular model, and a high-strength steel roll-formed box reliability model learns from the high-strength steel roll-formed box to be optimized according to the reinforcement learning cellular model; A third computing unit for performing migration optimization on the high-strength steel roll-formed box to be optimized using a high-strength steel roll-formed box optimization model; A fourth computing unit for outputting the plate thickness distribution of each functional beam of the optimized high-strength steel roll-formed box.

9. A high-strength steel roll-formed box optimization device based on a reinforcement learning cellular algorithm, the device comprising a processor and a memory; the memory is used for storing program codes and transmitting the program codes to the processor; The processor is used for executing the steps of the high-strength steel roll-formed box optimization method based on the reinforcement learning cellular algorithm according to the instructions in the program codes as claimed in claims 1-7.

10. A computer-readable storage medium for storing program codes for executing the steps of the high-strength steel roll-formed box optimization method based on the reinforcement learning cellular algorithm according to claims 1-7.