Power battery disassembly path decision-making method based on mixed attention and reinforcement learning
By adopting a hybrid attention and reinforcement learning method in the decision-making of power battery disassembly paths, the decoder structure and multimodal model are optimized, and the problems of inefficient power battery disassembly and environmental impact in the existing technology are solved, achieving a more efficient and resource-saving disassembly process.
Patent Information
- Application Number
- CN202510637230.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing power battery dismantling path decision-making methods are inefficient and susceptible to the environment. Traditional reinforcement learning methods face problems such as slow convergence speed, dimensional disasters and empirical sample quality dependence when dealing with complex tasks.
Using a power battery disassembly path decision-making method based on hybrid attention and reinforcement learning, a multi-modal model of multi-decoding layers is constructed, combining one-dimensional convolution, residual connection and multi-layer normalization strategies, the decoder structure is optimized to improve timing data processing capabilities and training stability.
It significantly improves disassembly efficiency, shortens task execution time, reduces resource consumption, enhances the system's flexibility and real-time adaptability, and solves the problems of inefficiency and environmental impact.
Smart Images

Figure CN120163300A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field, and in particular to a decision-making method for the disassembly path of power batteries based on hybrid attention and reinforcement learning. Background Art
[0002] With the rapid development of battery recycling technology, the disassembly and treatment of retired power batteries have become an important part of the environmental protection and resource recycling fields. However, traditional power battery disassembly methods usually rely on manual operation and single equipment scheduling, facing problems such as low efficiency, high cost, and complex operation. Therefore, how to optimize the battery disassembly process through an intelligent decision-making system, improve disassembly efficiency, reduce costs, and reduce environmental impacts has become the focus of current research.
[0003] Most existing disassembly methods adopt scheduling methods based on heuristic rules or task ranking methods based on optimization algorithms (such as genetic algorithms and particle swarm optimization). These traditional methods usually only find an approximate optimal solution under given environments and constraints. However, due to the strong dependence between tasks and the intertwining of various factors involved in the disassembly process (such as resource consumption, execution time, task priority, etc.), the performance of traditional methods in complex environments is not satisfactory.
[0004] Traditional reinforcement learning methods learn optimal policies by interacting with the environment and have been widely applied to task scheduling and resource allocation. However, traditional reinforcement learning methods still face several significant problems when dealing with complex tasks.
[0005] Firstly, the slow convergence speed is a major problem of traditional reinforcement learning. Since traditional algorithms require a large amount of interaction data and training time, especially when the task space is large, the convergence speed significantly slows down, resulting in low training efficiency. Secondly, the high-dimensionality problem is also a challenge faced by traditional RL methods. When facing high-dimensional data and complex environments, traditional RL methods are prone to the curse of dimensionality, leading to increased training difficulty. Finally, traditional RL methods usually rely on experience, and their performance highly depends on the experience samples obtained during the training process. If the quality of experience sampling is poor or the training environment is unstable, it may lead to unsatisfactory training effects and affect the final decision-making quality.
[0006] In view of the deficiencies of traditional methods, the present invention proposes an intelligent decision-making system for power battery disassembly based on hybrid attention and reinforcement learning. Summary of the Invention
[0007] The present invention provides a decision-making method for the disassembly path of power batteries based on hybrid attention and reinforcement learning to solve the problems of low efficiency and susceptibility to environmental impacts of existing battery disassembly path decision-making methods.
[0008] To achieve the above object, the present invention is implemented through the following technical solutions: The present invention provides a decision-making method for the disassembly path of power batteries based on hybrid attention and reinforcement learning, including the following steps: Step 1: Construct a decoder with multiple normalizations and residual connections based on hybrid attention and one-dimensional convolution. According to the decoder, construct a multi-modal model with multiple decoding layers, and obtain historical disassembly data of previous battery disassembly. Combine reinforcement learning to train the multi-modal model to obtain a battery disassembly model; Step 2: Obtain the disassembly data of the battery to be disassembled, input it into the battery disassembly model, obtain the disassembly action and score, and execute the current disassembly action according to the comparison result of the score and the predetermined threshold; Step 3: Input the disassembly action and disassembly data in Step 2 into the battery disassembly model together, obtain the next disassembly action, and execute the next disassembly action according to the comparison result of the score and the predetermined threshold until the battery to be disassembled is completely disassembled.
[0009] Among them, the disassembly data includes the disassembly process of the battery pack, disassembly time, and disassembly cost. After the first disassembly action is executed, the obtained disassembly action is also used as an input and input into the battery disassembly model together with the disassembly data.
[0010] Further, the decoder constructed based on hybrid attention and one-dimensional convolution with multiple normalizations and residual connections includes: constructing a first residual connection between the original input and the first residual connection after normalization and hybrid attention processing according to hybrid attention, and then constructing a second residual connection between the one-dimensional convolution output and the second residual connection after one-dimensional convolution and normalization processing based on one-dimensional convolution. Construct a decoder based on the first residual connection and the second residual connection.
[0011] Further, the decoder includes a first residual connection unit and a second residual connection unit, and the first residual connection unit is connected to the second residual connection unit; The first residual connection unit includes an input layer, a first normalization layer, a hybrid attention block, and a first residual accumulation layer, and is connected in sequence according to the order of the input layer, the first normalization layer, the hybrid attention block, and the first residual accumulation layer. The input layer is residually connected to the first residual accumulation layer; The second residual connection unit includes a one-dimensional convolution layer, a second normalization layer, a linear layer, a second residual accumulation layer, and a first output layer, and is connected in sequence according to the order of the one-dimensional convolution layer, the second normalization layer, the linear layer, the second residual accumulation layer, and the first output layer. The one-dimensional convolution layer is residually connected to the second residual accumulation layer, and the first residual accumulation layer is connected to the one-dimensional convolution layer.
[0012] By optimizing the decoder structure and combining one-dimensional convolutional branches, residual connections, and multi-layer normalization strategies, the ability to process time-series data and training stability are further improved. The convolutional branch enhances the ability to extract local features, the residual connection ensures the effective transmission of information, and the normalization strategy accelerates the network convergence process. In the prior art, most decoders use traditional neural network structures and lack a dedicated design optimized for time-series data. However, the present invention effectively improves the ability to handle complex disassembly tasks and the stability of network training through a multi-level optimized structure.
[0013] Furthermore, the hybrid attention block obtains the query matrix, key matrix, and value matrix corresponding to the input features through three independent linear transformations, calculates the local attention weights for the query matrix and key matrix through the attention mechanism of the one-dimensional convolutional block, performs weighted summation on the value matrix according to the local attention weights to obtain the first sequence, calculates the second sequence of the query matrix, key matrix, and value matrix according to the multi-head attention, aligns the first sequence and the second sequence through linear interpolation, and then obtains the output features through the normalization and fusion layer.
[0014] Among them, the input feature is the input feature after normalization processing in the first residual connection unit, and the feature output by the hybrid attention block is used for the residual connection of the subsequent first residual connection unit.
[0015] Through the above operations, the adopted hybrid attention mechanism combines convolutional attention and multi-head self-attention, and can capture both global dependencies and local information simultaneously. Convolutional attention enhances the ability to capture short-term temporal dependencies through the local receptive field mechanism, while multi-head self-attention helps the model capture long-term dependencies. This combination method significantly improves the feature extraction ability in time-series data, especially in handling disassembly tasks with time-dependence and variability, and has good performance. The self-attention mechanism in the prior art usually only focuses on global dependency information and lacks effective capture of local information.
[0016] Furthermore, the one-dimensional convolutional block converts the dimensions of the query matrix and key matrix through two independent one-dimensional convolutional layers, then calculates the preliminary attention weights of the query matrix and key matrix through matrix multiplication, and finally adjusts the weight probability distribution through normalization to obtain the final attention weights.
[0017] Furthermore, the one-dimensional convolutional block includes a first input branch, a second input branch, a multiplication layer, and a third normalization layer; Both the first input branch and the second input branch are connected to the multiplication layer, and the multiplication layer is connected to the third normalization layer; Both the first input branch and the second input branch are one-dimensional convolutional layers with a convolutional kernel set to 3 and a stride set to 2.
[0018] Further, the multimodal model constructs an incentive function based on execution time, execution cost, and delay penalty; The incentive function is represented by the following formula: ; where represents the execution cost; represents the execution time; represents the delay penalty; , and are all weight coefficients; The delay penalty is represented by the following formula: ; where represents the actual completion time of battery disassembly; represents the specified completion time of battery disassembly.
[0019] Further, the training of the multimodal model by combining reinforcement learning includes: model initialization, action selection exploration strategy, experience replay storage, and model optimization; The model initialization is to set the network structure and initial parameters; The action selection exploration strategy is to balance exploration and action selection by combining the greedy strategy; The experience replay storage is to put the current state, selected action, reward, next state, and task completion flag into the experience replay pool after interacting with the environment; The model optimization is to calculate the reward based on the task execution cost, execution time, and delay penalty and update.
[0020] Reinforcement learning is used to optimize the execution order and resource allocation of disassembly tasks, and an incentive function is combined to balance the time and cost of task execution. The incentive function not only considers the execution time and cost of tasks, but also comprehensively considers multiple factors such as waiting time, task completion reward, and delay penalty. The design of this incentive function can dynamically adjust decisions during task scheduling to ensure that each task is completed at the lowest cost and shortest time, while avoiding task delays and resource waste. In the prior art, reinforcement learning is often used to optimize a single objective (such as time or cost), and the innovation of the present invention lies in balancing multiple objectives through the design of a comprehensive incentive function, improving the ability of multi-objective optimization.
[0021] Furthermore, the battery disassembly model maps the discrete sequence of disassembly data into a continuous vector through an embedding layer, converts the continuous vector into a dimensional vector of a unified decoder dimension through a linear mapping, performs feature fusion on the dimensional vector through a fusion layer to obtain a fusion feature, then extracts the fusion feature through a multi-decoder layer to obtain an output feature, finally removes the sequence dimension through dimensional compression, and performs a linear mapping through a fully connected layer to obtain disassembly actions and scores.
[0022] Embed and fuse different types of data such as the battery pack disassembly process, disassembly time, and disassembly cost. Existing technologies usually rely on single-modal data for decision-making, while the present invention can more comprehensively understand the global information of the task and more accurately capture local features by making full use of the complementarity between different modal data. This data fusion method has not been widely applied in task scheduling by existing technologies and is a key innovation point of the present invention.
[0023] Beneficial effects: The power battery disassembly path decision-making method based on hybrid attention and reinforcement learning provided by the present invention introduces an autoregressive mechanism, uses the disassembly action output by the current task as the input of the next task, and gradually optimizes each step in the disassembly process. This innovative method significantly improves the stability, accuracy, and adaptability of decision-making when dealing with long-term dependencies in task scheduling, and solves the problems of low efficiency and susceptibility to environmental influence of existing battery disassembly path decision-making methods.
[0024] By optimizing the decoder structure and combining a one-dimensional convolution branch, residual connection, and multi-layer normalization strategy, the time-series data processing ability and training stability are further improved, and the problem of easily encountering the curse of dimensionality when facing high-dimensional data is solved.
[0025] Through the combination of autoregressive strategy, reinforcement learning, and hybrid attention mechanism, real-time decision adjustment can be performed in a complex and dynamically changing task scheduling environment. Compared with the static decision-making system in the prior art, the present invention can dynamically adjust task scheduling and resource allocation according to environmental changes and task requirements, reduces the dependence on empirical samples, and greatly enhances the flexibility and real-time adaptability of the system. Brief description of the drawings
[0026] Figure 1 It is a schematic diagram of the overall network structure of the battery disassembly model according to an embodiment of the present invention; Figure 2 It is a schematic diagram of the network structure of the decoder according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the network structure of the hybrid attention block according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the network structure of the one-dimensional convolution block according to an embodiment of the present invention. Detailed implementation mode
[0027] The technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.
[0028] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "a" or "one" do not indicate a quantity limitation, but indicate that there is at least one. The terms "connected" or "linked" and the like are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right" and the like are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship also changes accordingly.
[0029] Please refer to Figure 1 , the embodiment of the present application provides a power battery disassembly path decision-making method based on hybrid attention and reinforcement learning, including the following steps: Step 1: Construct a decoder with multiple normalizations and residual connections based on hybrid attention and one-dimensional convolution. According to the decoder, construct a multi-modal model with multiple decoding layers, and obtain historical disassembly data of previous battery disassembly. Combine reinforcement learning to train the multi-modal model to obtain a battery disassembly model; Among them, constructing a decoder with multiple normalizations and residual connections based on hybrid attention and one-dimensional convolution includes: constructing a first residual connection between the original input and the first residual connection after normalization and hybrid attention processing according to hybrid attention, and then constructing a second residual connection between the one-dimensional convolution output and the one-dimensional convolution and normalization processing based on one-dimensional convolution. Construct a decoder based on the first residual connection and the second residual connection; Specifically, please refer to Figure 2 , the decoder includes a first residual connection unit and a second residual connection unit, and the first residual connection unit is connected to the second residual connection unit; The first residual connection unit includes an input layer, a first normalization layer, a hybrid attention block, and a first residual accumulation layer, and is connected in sequence according to the input layer, the first normalization layer, the hybrid attention block, and the first residual accumulation layer. The input layer is residually connected to the first residual accumulation layer; The second residual connection unit includes a one-dimensional convolutional layer, a second normalization layer, a linear layer, a second residual accumulation layer, and a first output layer, and they are connected in sequence according to the order of the one-dimensional convolutional layer, the second normalization layer, the linear layer, the second residual accumulation layer, and the first output layer. The one-dimensional convolutional layer is in residual connection with the second residual accumulation layer, and the first residual accumulation layer is connected to the one-dimensional convolutional layer.
[0030] Here, the first normalization layer uses the Norm normalization function, and the second normalization layer uses layer normalization.
[0031] In the decoder, by fusing the Hybrid Attention mechanism with one-dimensional convolution (the size of the convolution kernel is 1*3), it aims to improve the feature extraction ability of time series data. Compared with the traditional Transformer decoder layer, the present invention adds a one-dimensional convolution branch in the structure and introduces multiple normalizations and residual connections in key calculation steps, thereby effectively improving the stability and generalization performance of the model in complex tasks.
[0032] When the decoder receives the input data, it first normalizes the input data to adjust it to the same size, then obtains the attention-weighted feature fusion representation through hybrid attention calculation, improves the overall learning effect through residual links, extracts local time series features through one-dimensional convolution, adjusts the size again through normalization and linear transformation, and finally enhances information transmission and optimizes gradients through residual links.
[0033] Please refer to Figure 3 , when the hybrid attention block receives the input data, it first obtains the query matrix, key matrix, and value matrix corresponding to the input features through three independent linear transformations, calculates the local attention weights for the query matrix and key matrix through the attention mechanism of the one-dimensional convolution block, performs weighted summation on the value matrix according to the local attention weights to obtain the first sequence, calculates the second sequence of the query matrix, key matrix, and value matrix according to multi-head attention, aligns the first sequence and the second sequence through linear interpolation, and then obtains the output features through normalization and the fusion layer.
[0034] Among them, please refer to Figure 4 , the one-dimensional convolution block transforms the dimensions of the query matrix and key matrix through two independent one-dimensional convolutional layers, then calculates the preliminary attention weights of the query matrix and key matrix through matrix multiplication, and finally adjusts the weight probability distribution through normalization to obtain the final attention weights. The one-dimensional convolution block includes a first input branch, a second input branch, a multiplication layer, and a third normalization layer; Both the first input branch and the second input branch are connected to the multiplication layer, and the multiplication layer is connected to the third normalization layer; Both the first input branch and the second input branch are one-dimensional convolutional layers with a convolution kernel set to 3 and a stride set to 2.
[0035] In this embodiment, when converting the dimension of the feature, from is converted to , that is, where and represent the converted matrix, and the third normalization layer adopts Softmax normalization processing.
[0036] The multi-modal model constructs an incentive function based on execution time, execution cost, and delay penalty; The incentive function is represented by the following formula: ; where represents the execution cost, which is the resources consumed during the execution of the task. The goal is to minimize this value to reduce the resource consumption of disassembly; represents the execution time, which refers to the time required for the task to start and complete. The goal is to shorten the disassembly time as much as possible to improve the execution efficiency of the task; represents the delay penalty. If the task fails to be completed on time, a penalty will be imposed according to the overtime part. The delay penalty increases quadratically with the increase of the task delay time to ensure that the task is completed on time and avoid additional costs caused by overtime; , and are all weight coefficients, is the weight of the disassembly cost. The lower the disassembly cost, the better. Therefore, this coefficient should be set to a relatively large positive value to ensure effective control of the task cost, is the weight of the disassembly time. The shorter the time, the better. Shortening the disassembly time helps to improve the disassembly efficiency. Therefore, this coefficient should be set to a relatively large positive value, is the weight of the delay penalty to ensure that the task can be completed on time. If the task is overtime, the delay penalty will increase quadratically, thus avoiding additional costs caused by task delays; In other embodiments, the relative importance of time, cost, and delay penalty can be flexibly balanced to optimize the scheduling effect.
[0037] The delay penalty is represented by the following formula: ; where represents the actual completion time of battery disassembly; represents the specified completion time of battery disassembly.
[0038] Training the multi-modal model in combination with reinforcement learning includes: model initialization, action selection exploration strategy, experience replay storage, and model optimization; The model is initialized by setting the network structure and initial parameters. Specifically, the vocabulary size is set to vocab_size = 128, the embedding dimension is embedding_dim = 256, the decoder dimension is decoder_dim = 512, and the number of decoder layers is num_decoder_layers = 6. The learning rate of reinforcement learning is lr = 0.001, the discount factor gamma = 0.95, the batch size batch_size = 32, and the experience replay pool size buffer_size = 10,000. The reasonable setting of these parameters ensures the stability and efficiency of model training.
[0039] The action selection exploration strategy is to balance exploration and action selection by combining the greedy strategy. Specifically, the -greedy strategy is adopted. In the initial stage, is set to 0.1, that is, actions are randomly selected with a probability of 10% (exploration), and the remaining 90% are selected based on the predictions of the current model (exploitation). As the training progresses, the value gradually decays, enhancing the exploitation ability of the model and reducing unnecessary random selections; After each interaction with the environment, the current state, the selected action, the obtained reward, the next state, and the task completion flag are stored in the experience replay pool. The introduction of the experience replay pool helps to break the correlation of training samples, ensuring data diversity and stability during the training process. During each training, batch_size = 32 samples are randomly sampled from the experience pool for update, strengthening the generalization ability of the model and accelerating the convergence process; The optimized model calculates and updates the reward based on the task execution cost, execution time, and delay penalty. The specific update operation is to update the Q value according to the reward value calculated by the incentive function. The Q value is calculated by the following formula: ; Among them, the target Q value ( ) is calculated by combining the current reward ( ) with the future Q value ( ). The future Q value is adjusted by the discount factor ( ), and the term ensures that the future Q value is only considered when the task is not completed. The model continuously adjusts the strategy by minimizing the difference between the target Q value and the current Q value, optimizing the task decomposition decision; In this embodiment, the training process is set to 500 times. After each training session ends, the total reward is obtained, and the model is evaluated based on this total reward. Specifically, after each training iteration ends, the model calculates the total reward by accumulating the reward values in each step. The calculation formula for the total reward is: ; where represents the reward for the step, and is the total number of steps in the current training round. After the training process ends, the performance of the model is evaluated based on the total reward value of each round. By continuously optimizing the reward signal, the model can gradually adjust its strategy, thereby improving its performance in task execution. During the evaluation process, some performance metrics are usually adopted, such as average reward, variance of the reward, etc., to comprehensively understand the learning effect and robustness of the model.
[0040] Step 2: Obtain the disassembly data of the battery to be disassembled, input it into the battery disassembly model, obtain the disassembly action and the score, and execute the current disassembly action according to the comparison result between the score and the predetermined threshold; The battery disassembly model maps the discrete sequence of disassembly data into a continuous vector through the embedding layer, converts the continuous vector into a dimensional vector of a unified decoder dimension through a linear mapping, performs feature fusion on the dimensional vector through the fusion layer to obtain the fusion feature, then extracts the output feature from the fusion feature through the multi-decoding layer, finally removes the sequence dimension through dimensional compression, and obtains the disassembly action and the score through a linear mapping of the fully connected layer.
[0041] By performing embedding, mapping, and fusion on the multi-modal input data of the disassembly data, and combining with the layer-by-layer refinement of the decoder layer, an efficient battery disassembly model for the disassembly sequence decision of retired power batteries is constructed. Using multiple information sources to achieve accurate prediction of the disassembly operation and its score, and at the same time ensuring the stability and generalization performance of model training through techniques such as normalization and residual connection.
[0042] Step 3: Input the disassembly action and the disassembly data in Step 2 into the battery disassembly model together, obtain the next disassembly action, and execute the next disassembly action according to the comparison result between the score and the predetermined threshold until the battery to be disassembled is completely disassembled.
[0043] Finally, combine the genetic algorithm, particle swarm optimization, traditional reinforcement learning with the method and battery disassembly model proposed in the present invention for the planning of the battery disassembly path, and mainly compare from the disassembly time and disassembly cost. Please refer to Table 1; Table 1: Comparison results of the battery disassembly model with the genetic algorithm, particle swarm optimization, and traditional reinforcement learning.
[0044]
[0045] The battery disassembly model proposed by the present invention significantly improves the disassembly efficiency through reinforcement learning and a hybrid attention mechanism. Compared with traditional methods, namely genetic algorithms, particle swarm optimization, and traditional reinforcement learning, the execution time is reduced by approximately 30%. Especially when dealing with complex task scheduling, the model can dynamically adjust the strategy, reducing unnecessary waiting time and idle time, thereby significantly shortening the execution time of tasks.
[0046] By optimizing task scheduling and resource allocation, approximately 20% of resource consumption is saved during the disassembly process. Traditional methods often result in higher costs due to the lack of fine control over resource consumption. The battery disassembly model can dynamically adjust the execution order and resource allocation of each task through the time and cost objectives in the incentive function, thereby effectively reducing the resource consumption of disassembly tasks.
[0047] Through comparison with traditional methods, the power battery disassembly path decision-making method based on hybrid attention and reinforcement learning proposed by the present invention and the corresponding battery disassembly model demonstrate their obvious advantages in optimizing disassembly time and disassembly cost. Especially in terms of optimizing disassembly efficiency and resource consumption, this system is significantly superior to traditional methods.
[0048] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A power battery disassembly path decision method based on hybrid attention and reinforcement learning, characterized in that: The steps include: Step 1: Based on hybrid attention and one-dimensional convolution, a decoder with multiple normalization and residual connection is constructed. A multimodal model with multiple decoding layers is constructed based on the decoder. The historical disassembly data of previous battery disassembly is obtained and combined with reinforcement learning to train the multimodal model to obtain a battery disassembly model. Step 2: Obtain disassembly data of the battery to be disassembled, input the battery disassembly model, obtain the disassembly action and score, and execute the current disassembly action according to the comparison result between the score and the predetermined threshold; Step 3: Input the disassembly action and disassembly data of step 2 into the battery disassembly model, obtain the next disassembly action, and execute the next disassembly action according to the comparison result between the score and the predetermined threshold until the disassembly of the battery to be disassembled is completed.
2. The power battery disassembly path decision method based on hybrid attention and reinforcement learning according to claim 1 is characterized in that: The decoder based on hybrid attention and one-dimensional convolution to construct multiple normalization and residual connections includes: constructing the original input and the first residual connection after normalization and hybrid attention processing according to the hybrid attention, and then constructing the one-dimensional convolution output and the second residual connection after one-dimensional convolution and normalization processing based on the one-dimensional convolution, and constructing a decoder based on the first residual connection and the second residual connection.
3. The power battery disassembly path decision method based on hybrid attention and reinforcement learning according to claim 2 is characterized in that: The decoder comprises a first residual connection unit and a second residual connection unit, wherein the first residual connection unit is connected to the second residual connection unit; The first residual connection unit includes an input layer, a first normalized layer, a hybrid attention block, and a first residual accumulation layer, and is connected in sequence in the order of the input layer, the first normalized layer, the hybrid attention block, and the first residual accumulation layer, and the input layer is residually connected to the first residual accumulation layer; The second residual connection unit includes a one-dimensional convolutional layer, a second normalized layer, a linear layer, a second residual accumulation layer and a first output layer, and is connected in sequence in the order of the one-dimensional convolutional layer, the second normalized layer, the linear layer, the second residual accumulation layer and the first output layer. The one-dimensional convolutional layer is residually connected to the second residual accumulation layer, and the first residual accumulation layer is connected to the one-dimensional convolutional layer.
4. The power battery disassembly path decision method based on hybrid attention and reinforcement learning according to claim 3 is characterized in that: The hybrid attention block obtains the query matrix, key matrix and value matrix corresponding to the input features through three independent linear transformations, calculates the local attention weights for the query matrix and the key matrix through the attention mechanism of the one-dimensional convolution block, obtains the first sequence by weighted summation of the value matrix according to the local attention weights, calculates the second sequence of the query matrix, key matrix and value matrix according to multi-head attention, aligns the first sequence with the second sequence in combination with linear interpolation, and obtains the output features through normalization and fusion layers.
5. The power battery disassembly path decision method based on hybrid attention and reinforcement learning according to claim 4 is characterized in that: The one-dimensional convolution block converts the dimensions of the query matrix and the key matrix through two independent one-dimensional convolution layers, and then calculates the preliminary attention weights of the query matrix and the key matrix by combining matrix multiplication, and finally obtains the final attention weight by normalizing and adjusting the weight probability distribution.
6. The power battery disassembly path decision method based on hybrid attention and reinforcement learning according to claim 5 is characterized in that: The one-dimensional convolution block includes a first input branch, a second input branch, a multiplication layer and a third normalization layer; The first input branch and the second input branch are both connected to the multiplication layer, and the multiplication layer is connected to a third normalization layer; The first input branch and the second input branch are both one-dimensional convolution layers with a convolution kernel set to 3 and a step size set to 2.
7. The power battery disassembly path decision method based on hybrid attention and reinforcement learning according to any one of claims 1 to 6, characterized in that: The multimodal model constructs an incentive function based on execution time, execution cost, and delay penalty; The activation function is expressed by the following formula: ; in, represents the execution cost; Indicates execution time; Indicates delayed punishment; , as well as All are weight coefficients; The delay penalty is expressed by the following formula: ; in, Indicates the actual completion time of battery disassembly; Indicates the specified completion time for battery disassembly.
8. The power battery disassembly path decision method based on hybrid attention and reinforcement learning according to claim 7 is characterized in that: The training of the multimodal model in combination with reinforcement learning includes: model initialization, action selection exploration strategy, experience replay storage and model optimization; The model initialization includes setting the network structure and setting the initial parameters; The action selection exploration strategy is to balance exploration and action selection in combination with a greedy strategy; The experience replay storage is to put the current state, selected action, reward, next state and task completion flag into the experience replay pool after interacting with the environment; The optimization model calculates rewards based on task execution cost, execution time and delay penalty, and updates them.
9. The power battery disassembly path decision method based on hybrid attention and reinforcement learning according to claim 7, characterized in that: The battery disassembly model maps the discrete sequence of disassembly data into a continuous vector through an embedding layer, converts the continuous vector into a dimensional vector of a unified decoder dimension through linear mapping, fuses the dimensional vector through a fusion layer to obtain fusion features, extracts the fusion features through multiple decoding layers to obtain output features, finally removes the sequence dimension through dimensional compression, and obtains the disassembly action and score through linear mapping through a fully connected layer.
Citation Information
Patent Citations
Optimization method of decommissioned power battery disassembling process
CN117728064A
Man-machine collaborative disassembly retired power battery task sequence optimization method based on reinforcement learning
CN118627826A
Disassembly sequence planning method and device for retired power battery considering cascade failure
CN118761764A
Detired battery dynamic disassembly path decision-making method based on deep search
CN119849331A
Retired battery disassembly scheduling method based on reinforcement learning
CN119849887A
Cited By
Automobile retired power battery disassembly sequence optimization method and related device
CN120724867A
A method and related device for optimizing disassembly sequence of retired power battery of an automobile
CN120724867B
Waste electric energy meter disassembling method and system
CN122133676A