Shared physical memory allocation method, electronic device, and storage medium
Patent Information
- Application Number
- CN202611171753.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-03
- Publication Date
- 2026-09-29
AI Technical Summary
不同处理组件之间通过总线或通信接口传输处理数据,增加了数据传输开销,进而增加了安全控制时延
Smart Images

Figure CN122838129A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to methods for allocating shared physical memory, electronic devices, and storage media. Background Technology
[0002] Safety control systems for electronic devices (such as robots and robotic arms) typically involve multiple processing stages (e.g., motion model inference, safety distance field processing, model predictive control solution, and safety residual gradient backpropagation). In related technologies, these processing stages are usually distributed across multiple processing components, with each component capable of executing one or more of them. The transmission of processed data between different processing components via bus or communication interfaces increases data transmission overhead, thereby increasing safety control latency. Summary of the Invention
[0003] This application provides a shared physical memory allocation method, electronic device, and storage medium, designed to reduce data transfer overhead and security control latency.
[0004] A first aspect provides a shared physical memory allocation method applied to an electronic device, the electronic device including multiple processing components, the multiple processing components being used to execute multiple processing stages, each of the multiple processing components being used to execute at least one of the multiple processing stages; the method includes: A shared physical memory pool is allocated in the electronic device for the common access of the multiple processing components; Based on the access attribute information corresponding to the multiple processing stages, the shared physical memory pool is divided into multiple logical partitions, and each logical partition is used to store the associated data of at least one of the multiple processing stages. A partition description information is established for each logical partition. The partition description information is used by the multiple processing components to locate and access the data stored in the corresponding logical partition.
[0005] This technical solution employs a three-layer design of "physical memory pool sharing + logical partitioning + partition description information location," transforming cross-component data transmission, which originally relied on bus / communication interfaces, into direct access based on shared memory without altering the independent execution capabilities of each processing component. This significantly reduces data transmission overhead and security control latency. Simultaneously, partitioned management ensures data organization efficiency and scalability in multi-stage, multi-component concurrent access scenarios, making it particularly suitable for electronic devices with high real-time security control requirements, such as robots and robotic arms.
[0006] In conjunction with the first aspect, in one possible implementation, the method further includes: A unified data description format is pre-defined. The data description format is used to establish data description information corresponding to data stored in any of the logical partitions. The data description information is used by the multiple processing components to locate and read the corresponding data.
[0007] In conjunction with the first aspect, in one possible implementation, the multiple processing stages include an action model inference stage, a model prediction and control solution stage, a safety distance field processing stage, and a safety residual gradient backpropagation stage; the method further includes: By executing the processing component of the action model inference stage, action model inference is performed to obtain action embeddings, and the action embeddings are written into the first logical partition among the plurality of logical partitions. A first data description information for the action embedding is established based on the data description format; By executing the processing component of the model predictive control solution stage, the action embedding is read based on the first data description information, the optimal motion trajectory is obtained by performing model predictive control solution based on the action embedding, and the optimal motion trajectory is written into the second logical partition among the multiple logical partitions. A second data description information for the optimal motion trajectory is established based on the data description format; By executing the processing component of the safe distance field processing stage, the optimal motion trajectory is read based on the second data description information of the optimal motion trajectory, the safe residual is determined based on the optimal motion trajectory, and the safe residual is written into the third logical partition among the plurality of logical partitions. A third data description information for the safety residual is established based on the data description format; By executing the processing component of the backpropagation phase of the security residual gradient, the security residual is read based on the third data description information of the security residual, the security residual gradient embedded by the security residual for the action is determined, and the security residual gradient is written into the fourth logical partition among the plurality of logical partitions. The fourth data description information of the safety residual gradient is established based on the data description format.
[0008] In conjunction with the first aspect, in one possible implementation, the execution of action model reasoning to obtain the action embedding includes: reasoning about the input data through the action model to obtain the action embedding; The step of obtaining the optimal motion trajectory based on the action embedding execution model predictive control includes: decoding the action embedding into model predictive control parameters; constructing a constrained optimization problem based on the model predictive control parameters; iteratively updating the optimization variables of the optimization problem and the dual variables corresponding to the constraints in the optimization problem until a preset convergence condition is met to obtain the optimal solution of the optimization variables; and determining the optimal motion trajectory based on the optimal solution of the optimization variables. The step of determining the safety residual based on the optimal motion trajectory includes: calculating the safety distance corresponding to each of multiple trajectory points in the optimal motion trajectory; determining the minimum safety distance from the multiple safety distances; and determining the safety residual based on the comparison result between the minimum safety distance and a safety distance threshold. Determining the safety residual gradient of the safety residual with respect to the action embedding includes: taking the derivative of the safety residual with respect to the optimal motion trajectory as the derivative variable to obtain a first gradient; taking the derivative of the optimal motion trajectory with respect to the model prediction control parameters as the derivative variable to obtain a second gradient; taking the derivative of the model prediction control parameters with respect to the action embedding as the derivative variable to obtain a third gradient; and determining the safety residual gradient based on the first gradient, the second gradient, and the third gradient.
[0009] In conjunction with the first aspect, in one possible implementation, the method further includes: Obtain the first state of the electronic device in the current processing cycle and the second state of the electronic device in the previous processing cycle; The state change amount of the electronic device is determined based on the first state and the second state; When the state change satisfies the preset conditions, the initial value of the optimization variable for the current processing cycle is determined based on the optimal motion trajectory of the previous processing cycle, and the dual variable when the previous processing cycle satisfies the preset convergence condition is used as the initial value of the dual variable for the current processing cycle.
[0010] In conjunction with the first aspect, in one possible implementation, the method further includes: The security residual gradient is read based on the fourth data description information, the action embedding correction amount is determined based on the security residual gradient, and the action embedding correction amount is written into the fourth logical partition; A fifth data description information for the action embedding correction amount is established based on the data description format; In the next processing cycle, by executing the processing component of the action model inference stage, the action embedding correction amount is read based on the fifth data description information, and the action embedding is corrected based on the action embedding correction amount.
[0011] In conjunction with the first aspect, in one possible implementation, the method further includes: For each processing stage, a corresponding numerical precision is configured based on the numerical stability requirements and computational efficiency requirements of that processing stage.
[0012] In conjunction with the first aspect, in one possible implementation, the method further includes: Independent computation execution queues are established for the action model inference stage, the model prediction and control solution stage, the safety distance field processing stage, and the safety residual gradient backpropagation stage, respectively. If the safety residual does not meet the preset safety event triggering conditions, the computation execution queues corresponding to the action model inference stage, the model prediction and control solution stage, and the safety distance field processing stage are scheduled in parallel. When the safety residual meets the preset safety event triggering condition, a safety event flag is generated, and the computation execution queue corresponding to the backpropagation stage of the safety residual gradient is scheduled based on the safety event flag.
[0013] In a second aspect, a shared physical memory allocation device is provided, applied to an electronic device, the electronic device including a plurality of processing components, the plurality of processing components being configured to execute a plurality of processing stages, each of the plurality of processing components being configured to execute at least one of the plurality of processing stages; the device includes: An allocation module is used to allocate a shared physical memory pool in the electronic device for common access by the multiple processing components; The partitioning module is used to divide the shared physical memory pool into multiple logical partitions based on the access attribute information corresponding to the multiple processing stages. Each logical partition is used to store the associated data of at least one of the multiple processing stages. A module is established to create partition description information for each logical partition. The partition description information is used by the multiple processing components to locate and access the data stored in the corresponding logical partition.
[0014] Thirdly, an electronic device is provided, including a memory and a processor, the memory being connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, wherein when the processor executes the one or more computer programs, the electronic device enables the shared physical memory allocation method of the first aspect described above.
[0015] Fourthly, a computer-readable storage medium is provided, which stores a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the shared physical memory allocation method of the first aspect.
[0016] Fifthly, a computer program product is provided that, when run on an electronic device, causes any of the methods provided in the first aspect to be executed. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating a shared physical memory allocation method provided in an embodiment of this application; Figure 2 This is a schematic diagram of a shared physical memory allocation device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0020] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.
[0021] The technical solution of this application is described in detail below, and the technical solution of this application can be applied to electronic devices.
[0022] In this context, electronic devices refer to devices capable of performing functions such as perception, decision-making, and control, and having the ability to interact with the external environment. Their safety control systems typically involve stages such as motion model inference, safety distance field processing, model predictive control solution, and safety residual gradient backpropagation. For example, electronic devices may include robots, robotic arms, etc., but this application does not limit them to these categories.
[0023] The electronic device includes multiple processing components for performing multiple processing stages, each processing component for performing at least one of the multiple processing stages, and the processing components are devices applied to the electronic device for performing multiple processing stages in a safety control system of the electronic device.
[0024] A processing component is a functional unit in an electronic device that has data processing capabilities and is used to execute at least one of the multiple processing stages. It can be one or more hardware processing units such as a central processing unit (CPU), a graphics processing unit (GPU), or a neural network processing unit (NPU).
[0025] The multiple processing stages and the multiple processing components can have a flexible correspondence, with each processing component executing at least one of the multiple processing stages. For example, the action model inference stage can be executed by an NPU or GPU; the safe distance field processing stage can be executed by an NPU, GPU, etc.; the model prediction control solution stage can be executed by a CPU, GPU, etc.; and the safe residual gradient backpropagation stage can be executed by a GPU, NPU, etc. The same processing stage can be executed by multiple different types of processing components, and the same processing component can execute multiple processing stages; this application does not impose any limitations on this.
[0026] See Figure 1 , Figure 1 This is a flowchart illustrating a shared physical memory allocation method provided in an embodiment of this application, as shown below. Figure 1 As shown, this method is applied to an electronic device and includes the following steps: S101, a shared physical memory pool is allocated in the electronic device for the common access of the plurality of processing components.
[0027] In this process, a shared physical memory pool is allocated on the single-chip unified memory accelerator of the electronic device for the multiple processing components to access together.
[0028] S102, based on the access attribute information corresponding to the multiple processing stages, the shared physical memory pool is divided into multiple logical partitions, and each logical partition is used to store the associated data of at least one of the multiple processing stages.
[0029] Each processing stage corresponds to an access attribute, which includes a data access mode and a data lifecycle. The data access mode characterizes the read / write method used by the processing stage to access its associated data, such as read-only, write-only, or read / write access. The data lifecycle characterizes the duration of the associated data in the processing stage from its creation to its expiration or when it can be overwritten or released.
[0030] In one embodiment, dividing the shared physical memory pool into multiple logical partitions based on the access attribute information corresponding to the multiple processing stages includes: mapping processing stages with similar access attribute information to the same logical partition, and mapping processing stages with different access attribute information to different logical partitions, thereby dividing the shared physical memory pool into the multiple logical partitions. For example, the weight data and key-value cache data corresponding to the action model inference stage, since both are generated and used by the action model inference stage, have similar data access patterns and data lifecycles, and therefore can be mapped to the same logical partition. The network parameters and encoding table data corresponding to the safe distance field processing stage, since their data access patterns and data lifecycles differ from those of the action model inference stage, are mapped to different logical partitions. The associated data corresponding to the model predictive control solution stage, since its data access patterns and data lifecycles differ from those of the aforementioned processing stages, are mapped to different logical partitions. The associated data corresponding to the safe residual gradient backpropagation stage, since its data access patterns and data lifecycles differ from those of the aforementioned processing stages, are mapped to different logical partitions.
[0031] In one embodiment, the plurality of logical partitions include: The first logical partition is used to store the model parameters and key-value cache data of the action model; The second logical partition is used to store the model parameters and encoding table data of the safe distance field model; The third logical partition is used to store the working data for model predictive control solutions; The fourth logical partition is used to store safe residual gradients and context data.
[0032] S103, establish partition description information for each logical partition, the partition description information being used by the plurality of processing components to locate and access the data stored in the corresponding logical partition.
[0033] Each logical partition establishes a partition description, which includes the physical starting address, data shape, step size, and data type.
[0034] The physical starting address refers to the starting position of the byte offset of the logical partition in the shared physical memory pool.
[0035] Data shape refers to the dimensional information of the data stored in a logical partition.
[0036] Step size refers to the number of bytes that a logical partition stores along each dimension to locate any element in that data.
[0037] Data type refers to the numerical type of the data stored in a logical partition, which is used to determine the parsing format used when accessing the data.
[0038] In this embodiment, a three-layer design of "physical memory pool sharing + logical partitioning + partition description information positioning" is adopted. Without changing the independent execution capability of each processing component, the cross-component data transmission that originally relied on bus / communication interface is transformed into direct access based on shared memory. This significantly reduces data transmission overhead and security control latency. At the same time, partitioned management ensures data organization efficiency and scalability in multi-stage and multi-component concurrent access scenarios. It is especially suitable for electronic equipment scenarios such as robots and robotic arms with high requirements for real-time security control.
[0039] In one implementation, the method further includes: A unified data description format is pre-defined. The data description format is used to establish data description information corresponding to data stored in any of the logical partitions. The data description information is used by the multiple processing components to locate and read the corresponding data.
[0040] The data description format is a pre-defined, fixed structure format used to describe the attribute information of the data stored in the logical partition. The data description format includes the starting address, data shape, step array, data type, and memory layout flags.
[0041] The starting address represents the physical memory offset of the data within the shared physical memory pool.
[0042] Data shape represents the dimensional array of the data.
[0043] The step size array represents the number of bytes spanned in each dimension of the data.
[0044] Data types represent the numerical type of data, including FP16 or FP32, etc.
[0045] Memory layout flags indicate how data is arranged in memory, including row-major or column-major order.
[0046] Data description information is a description record established for data stored in any logical partition according to the above data description format, including the specific values of the starting address, data shape, step array, data type, and memory layout flags corresponding to the data.
[0047] In this embodiment of the application, by pre-setting a unified data description format and establishing data description information corresponding to data stored in any logical partition based on the data description format, multiple processing components based on different inference frameworks can access the same segment of data in the shared physical memory pool from a unified perspective, thus avoiding data parsing obstacles caused by differences in the underlying implementation frameworks of the processing components.
[0048] In addition, when any of the multiple processing components needs to access data, it only needs to locate and read the corresponding data through the data description information, without having to copy the data, thus reducing the additional overhead caused by data transmission and copying.
[0049] In one implementation, the plurality of processing stages includes an action model inference stage; the method further includes: By executing the processing component of the action model inference stage, action model inference is performed to obtain action embeddings, and the action embeddings are written into the first logical partition among the plurality of logical partitions. The first data description information for the action embedding is established based on the data description format.
[0050] In one embodiment, the action model reasoning to obtain the action embedding includes: reasoning about the input data through the action model to obtain the action embedding.
[0051] The input data includes current visual scene data and language command data.
[0052] The action model can be a Vision-Language-Action Model (VLA), which is used to perform forward reasoning on the input data to achieve a joint understanding of the current visual scene data and language instruction data, and generate the corresponding action embedding.
[0053] Action embedding is a high-dimensional vector generated by the output layer after the action model completes forward inference. Action embedding encodes the expected action semantics obtained by jointly understanding the current visual scene data and language instruction data.
[0054] It should be noted that the step "establishing the first data description information of the action embedding based on the data description format" can be executed by the processing component that performs the action model inference stage, or by other processing components of the electronic device, and is not limited here.
[0055] In this embodiment, the action model inference process is executed by the processing component in the action model inference stage to obtain the action embedding, which is then written to the first logical partition. This eliminates the need to transmit the action embedding to other processing components via a bus or communication interface, reducing data transmission overhead and security control latency. A first data description information corresponding to the action embedding is established based on the data description format, enabling other processing components to locate and read the action embedding through this information without needing to copy it, thus reducing the additional overhead associated with data copying.
[0056] In one implementation, the plurality of processing stages further includes a model predictive control solution stage; the method further includes: By executing the processing component of the model predictive control solution stage, the action embedding is read based on the first data description information, the optimal motion trajectory is obtained by performing model predictive control solution based on the action embedding, and the optimal motion trajectory is written into the second logical partition among the multiple logical partitions. The second data description information of the optimal motion trajectory is established based on the data description format.
[0057] The optimal trajectory is the optimal solution obtained under constraints during the model predictive control solution stage.
[0058] In one embodiment, obtaining the optimal motion trajectory based on the predictive control solution of the action embedding execution model includes: The action is embedded and decoded into model predictive control parameters; a constrained optimization problem is constructed based on the model predictive control parameters; the optimization variables and the dual variables corresponding to the constraints in the optimization problem are iteratively updated until a preset convergence condition is met to obtain the optimal solution of the optimization variables; the optimal motion trajectory is determined based on the optimal solution of the optimization variables.
[0059] One specific implementation method for decoding the action embedding into model predictive control parameters is as follows: the decoder of the electronic device translates the action embedding to obtain the model predictive control parameters. The decoder can be a lightweight decoder, such as a small multilayer perceptron. The translation processing performed by the decoder on the action embedding refers to the process by which the decoder converts the expected action semantics encoded by the action embedding into the structured parameters required for solving the model predictive control problem.
[0060] The model predictive control parameters include at least one of the following: state weighting matrix, control input weighting matrix, terminal cost matrix, prediction time domain length, and target state reference value.
[0061] The state weighting matrix is used to characterize the degree of penalty imposed on state variables when they deviate from the target state during the model predictive control solution process, indicating which states are penalized more severely for deviating from the target.
[0062] The control input weighting matrix is used to characterize the degree of penalty imposed on the control input during the model predictive control solution process, indicating which joints consume more energy and are penalized.
[0063] The terminal cost matrix is used to characterize the accuracy requirements of the predicted end state in the time domain during the model predictive control solution process, i.e., the accuracy requirements of the trajectory end.
[0064] The prediction time domain length is used to characterize the range of prediction time steps considered when constructing the optimal control problem in the finite time domain during the model predictive control solution process.
[0065] The target state reference value is used to characterize the target state value that is expected to be achieved during the model predictive control solution process, and serves as a reference benchmark for measuring the degree of deviation of the state variables at each prediction time step.
[0066] The specific implementation of constructing a constrained optimization problem based on the model predicting control parameters includes: reading the current state of the electronic device and concatenating the state variables and control inputs of each prediction time step within the prediction time domain into an optimization vector; expanding the state deviation term and control input term in the optimization vector into quadratic penalty terms based on the state weighting matrix and the control input weighting matrix; constructing the terminal cost term at the end of the prediction time domain based on the terminal cost matrix and the target state reference value to obtain the objective function of the optimization problem; and expanding the kinematic constraints into linear equality constraints and inequality constraints as constraints of the optimization problem, thereby constructing a constrained optimization problem in the form of a standard quadratic programming problem.
[0067] The specific implementation method for iteratively updating the optimization variables and the dual variables corresponding to the constraints in the optimization problem until a preset convergence condition is met to obtain the optimal solution of the optimization variables includes: iteratively solving the optimization problem using the interior point method or the activity set method; calculating the current search direction and step size in each iteration, and updating the optimization variables and the dual variables based on the search direction and the step size; checking the constraint satisfaction degree to determine whether the updated optimization variables and the dual variables meet the preset convergence condition; if not, continuing to the next iteration until the preset convergence condition is met to obtain the optimal solution of the optimization variables.
[0068] The optimization variable in the optimization problem refers to the optimization vector formed by concatenating the state variables and control inputs at each prediction time step within the prediction time domain length.
[0069] In optimization problems, constraints refer to linear equality constraints and inequality constraints obtained by expanding kinematic constraints. These constraints are used to limit the range of values or relationships that the optimization variables need to satisfy during the solution process.
[0070] The dual variable corresponding to the constraint condition refers to the variable associated with the constraint condition during the iterative solution of the constrained optimization problem using the interior point method or the activity set method. The dual variable and the optimization variable are updated together based on the search direction and step size in each iteration to cooperate in checking the constraint satisfaction.
[0071] The preset convergence condition refers to the pre-set judgment condition used to determine whether the iterative solution process can be stopped during the iterative solution of the constrained optimization problem. The preset convergence condition is determined by checking the constraint satisfaction of the updated optimization variable and the dual variable.
[0072] The optimal solution for the optimization variable refers to the value of the optimization variable obtained by iteratively updating the optimization variable and the dual variable until the preset convergence condition is met, and is used to determine the optimal motion trajectory.
[0073] The specific implementation of determining the optimal motion trajectory based on the optimal solution of the optimization variables includes: decomposing the optimal solution of the optimization variables to obtain the joint angle sequence and torque sequence corresponding to each prediction time step; and determining the joint angle sequence and torque sequence as the optimal motion trajectory in tensor form.
[0074] It should be noted that the step "establishing the second data description information of the optimal motion trajectory based on the data description format" can be executed by the processing component that performs the model predictive control solution stage, or by other processing components of the electronic device, and is not limited here.
[0075] In this embodiment, by executing the processing component of the model prediction control solution stage, the action embedding is read based on the first data description information, eliminating the need to copy the action embedding and reducing the additional overhead caused by data copying.
[0076] In addition, the optimal motion trajectory is obtained by predictive control based on the motion embedding execution model and written into the second logical partition. This eliminates the need to transmit the optimal motion trajectory to other processing components via a bus or communication interface, thus reducing data transmission overhead and security control latency.
[0077] In one implementation, the plurality of processing stages further includes a safe distance field processing stage; the method further includes: By executing the processing component of the safe distance field processing stage, the optimal motion trajectory is read based on the second data description information of the optimal motion trajectory, the safe residual is determined based on the optimal motion trajectory, and the safe residual is written into the third logical partition among the plurality of logical partitions. A third data description information for the safety residual is established based on the data description format.
[0078] In one embodiment, determining the safety residual based on the optimal motion trajectory includes: calculating the safety distance corresponding to each of the multiple trajectory points in the optimal motion trajectory; determining the minimum safety distance from the multiple safety distances; and determining the safety residual based on the comparison result between the minimum safety distance and a safety distance threshold.
[0079] Determining the minimum safe distance from a plurality of safe distances means determining the safe distance with the smallest value from a plurality of safe distances, wherein the minimum safe distance corresponds to the distance value of the most dangerous point in the optimal motion trajectory.
[0080] The step of determining the safety residual based on the comparison result between the minimum safety distance and the safety distance threshold includes: if the minimum safety distance is greater than the safety distance threshold, determining the safety residual to be zero; if the minimum safety distance is less than or equal to the safety distance threshold, determining the difference between the minimum safety distance and the safety distance threshold as the safety residual.
[0081] The calculation of the safe distance corresponding to each trajectory point in the optimal motion trajectory includes: Multiple sampling trajectory points are selected from the optimal motion trajectory; Calculate the safety distance and distance gradient corresponding to each of the sampling trajectory points; Based on the safety distance and distance gradient corresponding to adjacent sampling trajectory points, the safety distance and distance gradient corresponding to the trajectory points between the adjacent sampling trajectory points are determined by interpolation.
[0082] Selecting multiple sampling trajectory points from the optimal motion trajectory means selecting some or all of the trajectory points from the multiple trajectory points in the optimal motion trajectory as the sampling trajectory points.
[0083] The calculation of the safety distance and distance gradient corresponding to each of the sampling trajectory points includes: after converting the sampling trajectory points from joint configuration to end effector spatial coordinates, inputting them into a safety distance field network to obtain the safety distance and distance gradient corresponding to the sampling trajectory point, wherein the safety distance is the signed distance to the surface of the nearest obstacle, with positive values indicating safety and negative values indicating penetration, and the distance gradient is a vector pointing to the fastest direction away from the obstacle.
[0084] Specifically, determining the safety distance and distance gradient between adjacent sampled trajectory points by interpolation based on the safety distance and distance gradient between adjacent sampled trajectory points means that for trajectory points in the optimal motion trajectory that are located between two adjacent sampled trajectory points but were not selected as sampled trajectory points, the safety distance and distance gradient corresponding to the trajectory point are determined by interpolation based on the safety distance and distance gradient corresponding to each of the two adjacent sampled trajectory points.
[0085] It should be noted that the step "establishing the third data description information of the safety residual based on the data description format" can be performed by the processing component that performs the safety distance field processing stage, or by other processing components of the electronic device, and is not limited here.
[0086] In this embodiment, by executing the processing component of the safe distance field processing stage, the optimal motion trajectory is read based on the second data description information of the optimal motion trajectory, eliminating the need to copy the optimal motion trajectory and reducing the additional overhead caused by data copying. The safety residual is determined based on the optimal motion trajectory and written to the third logical partition, eliminating the need to transmit the safety residual to other processing components via a bus or communication interface, thus reducing data transmission overhead and safety control latency.
[0087] In one implementation, the plurality of processing stages further includes a safety residual gradient backpropagation stage; the method further includes: By executing the processing component of the backpropagation phase of the security residual gradient, the security residual is read based on the third data description information of the security residual, the security residual gradient embedded by the security residual for the action is determined, and the security residual gradient is written into the fourth logical partition among the plurality of logical partitions. The fourth data description information of the safety residual gradient is established based on the data description format.
[0088] In one embodiment, determining the safety residual gradient of the safety residual with respect to the action embedding includes: taking the derivative of the safety residual with respect to the optimal motion trajectory as the derivative variable to obtain a first gradient; taking the derivative of the optimal motion trajectory with respect to the model prediction control parameters as the derivative variable to obtain a second gradient; taking the derivative of the model prediction control parameters with respect to the action embedding as the derivative variable to obtain a third gradient; and determining the safety residual gradient based on the first gradient, the second gradient, and the third gradient.
[0089] The first gradient is used to characterize the sensitivity of the safety residual to the optimal motion trajectory.
[0090] The second gradient is used to characterize the sensitivity of the optimal motion trajectory to the model's predictive control parameters.
[0091] The third gradient is used to characterize the sensitivity of the model's predictive control parameters to the action embedding.
[0092] Determining the safety residual gradient based on the first gradient, the second gradient, and the third gradient includes: sequentially multiplying the first gradient, the second gradient, and the third gradient to obtain the safety residual gradient. The safety residual gradient is a vector used to characterize the sensitivity of the safety residual to each dimension of the action embedding.
[0093] It should be noted that the step "establishing the fourth data description information of the security residual gradient based on the data description format" can be executed by the processing component that performs the backpropagation stage of the security residual gradient, or by other processing components of the electronic device, and is not limited here.
[0094] In this embodiment, by executing the processing component of the backpropagation phase of the security residual gradient, the security residual is read based on the third data description information of the security residual, eliminating the need to copy the security residual and reducing the additional overhead of data copying. The security residual gradient embedded in the action is determined, and the security residual gradient is written to the fourth logical partition in the shared physical memory pool. This eliminates the need to transmit the security residual gradient to other processing components via a bus or communication interface, reducing data transmission overhead and security control latency.
[0095] In one implementation, the method further includes: Obtain the first state of the electronic device in the current processing cycle and the second state of the electronic device in the previous processing cycle; The state change amount of the electronic device is determined based on the first state and the second state; When the state change satisfies the preset conditions, the initial value of the optimization variable for the current processing cycle is determined based on the optimal motion trajectory of the previous processing cycle, and the dual variable when the previous processing cycle satisfies the preset convergence condition is used as the initial value of the dual variable for the current processing cycle.
[0096] The current processing cycle refers to a processing flow that the electronic device security control system is currently executing, which is the next processing flow that follows the previous processing cycle.
[0097] The first state refers to the joint angle, angular velocity, and end effector position of the electronic device in the current processing cycle.
[0098] The second state refers to the joint angle, angular velocity, and end effector position of the electronic device in the previous processing cycle.
[0099] The step of determining the state change amount of the electronic device based on the first state and the second state includes: calculating a first change amount between the joint angles corresponding to the first state and the second state, calculating a second change amount between the angular velocities corresponding to the first state and the second state, and calculating a third change amount between the end effector positions corresponding to the first state and the second state. The first change, the second change, and the third change are defined as the state change.
[0100] The state change satisfies the preset conditions, which means that the first change is lower than the preset angle threshold, the second change is lower than the preset angular velocity threshold, and the third change is lower than the preset displacement threshold.
[0101] Determining the initial values of the optimization variables for the current processing cycle based on the optimal motion trajectory of the previous processing cycle includes: using the optimal motion trajectory obtained in the previous processing cycle as the initial values of the optimization variables in the constrained optimization problem constructed in the current processing cycle.
[0102] The dual variable that met the preset convergence condition in the previous processing cycle refers to the value of the dual variable that met the preset convergence condition during the iterative update of the dual variable in the previous processing cycle, and this value is used as the initial value of the dual variable in the current processing cycle.
[0103] If the state change does not meet the preset conditions, the preset optimization variable is determined as the initial value of the optimization variable in the current processing cycle, and the preset dual variable is determined as the initial value of the dual variable in the current processing cycle.
[0104] Based on the initial values of the optimization variables and the dual variables in the current processing cycle, the first state of the electronic device in the current processing cycle is read, the constrained optimization problem of the current processing cycle is constructed, the optimization variables and the dual variables of the optimization problem are iteratively updated, the search direction and step size are calculated in each iteration, and the constraint satisfaction is checked to see if the preset convergence condition is met. If not, the next iteration continues. If it is met, the optimal motion trajectory of the current processing cycle is output.
[0105] When the state change amount meets the preset condition, the key-value pairs corresponding to the visual instruction class Token in the action model are cached, and the key-value pairs are recalculated only for the position of the modified action Token.
[0106] When the state change satisfies the preset condition, sparse sampling is performed on the safe distance field query, and interpolation is used between adjacent sampling points to estimate the safe distance and the distance gradient.
[0107] In this embodiment, the first state of the electronic device in the current processing cycle and the second state of the electronic device in the previous processing cycle are obtained, and the state change of the electronic device is determined based on the first state and the second state. When the state change satisfies a preset condition, the initial value of the optimization variable in the current processing cycle is determined based on the optimal motion trajectory of the previous processing cycle, and the initial value of the dual variable in the current processing cycle when the preset convergence condition is met is used as the initial value of the dual variable. This allows the iterative solution of the current processing cycle to start based on initial values closer to the optimal solution, reducing the number of iterations required to reach the preset convergence condition, shortening the processing time of the model predictive control solution stage, and reducing the overall latency of the safety control system. Simultaneously, the solution results of the previous processing cycle are reused only when the state change satisfies the preset condition, avoiding iterative solution errors caused by sudden changes in the state of the electronic device and ensuring the reliability of the solution results.
[0108] In one implementation, the method further includes: The security residual gradient is read based on the fourth data description information, the action embedding correction amount is determined based on the security residual gradient, and the action embedding correction amount is written into the fourth logical partition; A fifth data description information for the action embedding correction amount is established based on the data description format; In the next processing cycle, by executing the processing component of the action model inference stage, the action embedding correction amount is read based on the fifth data description information, and the action embedding is corrected based on the action embedding correction amount.
[0109] The step of determining the action embedding correction amount based on the safety residual gradient includes: scaling the safety residual gradient along a direction opposite to the safety residual gradient by a preset step size to obtain the action embedding correction amount. The direction of the action embedding correction amount is opposite to the direction of the safety residual gradient, and it is used to correct the action embedding to reduce the safety residual corresponding to the corrected action embedding.
[0110] The next processing cycle refers to the next processing cycle that follows the current processing flow, which consists of the action model inference stage, the model prediction and control solution stage, the safety distance field processing stage, and the safety residual gradient backpropagation stage.
[0111] Correcting the action embedding based on the action embedding correction amount means that the processing component executing the action model inference stage adds the action embedding correction amount to the action embedding, so that the optimal motion trajectory generated based on the corrected action embedding shifts in a safer direction.
[0112] In the embodiments of this application, in the next processing cycle, the motion embedding is corrected based on the motion embedding correction amount, so that the optimal motion trajectory generated based on the corrected motion embedding shifts in a safer direction, forming a safety feedback closed loop across processing cycles, thereby improving the safety of motion control of electronic devices.
[0113] In one implementation, the method further includes: For each processing stage, a corresponding numerical precision is configured based on the numerical stability requirements and computational efficiency requirements of that processing stage.
[0114] Among them, the numerical stability requirement is used to characterize the accuracy requirements that need to be met during the numerical calculation process in the processing stage to avoid abnormal calculation results due to insufficient numerical accuracy.
[0115] Computational efficiency requirements are used to characterize the efficiency requirements that need to be considered to improve computational throughput during the numerical computation process in the processing stage.
[0116] In one embodiment, configuring the corresponding numerical precision for the processing stage based on the numerical stability requirements and computational efficiency requirements of the processing stage includes: If the data stability requirement of the processing stage is low and the computational efficiency requirement of the processing stage is high, then the corresponding numerical precision for the processing stage is configured to be low (e.g., FP16). If the data stability requirement for the processing stage is high, then the corresponding numerical precision for the processing stage is configured to be high numerical precision (such as FP32).
[0117] In one embodiment, the safe distance field processing stage includes a safe distance field forward calculation process and a safe distance field online update process; The action model inference stage uses a first numerical precision, while the online update process of the safe distance field and the backpropagation stage of the safe residual gradient use a second numerical precision, which is higher than the first numerical precision.
[0118] In one embodiment, for two processing stages with a data transmission relationship, if the numerical precision of the data output by the previous processing stage is different from the numerical precision used in the subsequent processing stage, the data output by the previous processing stage is converted to the numerical precision used in the subsequent processing stage.
[0119] In this context, two processing stages with a data transfer relationship refer to two processing stages that are sequentially connected, where the associated data output by the preceding processing stage is read by the following processing stage and used to execute the corresponding computational task. For example, the action model inference stage and the model predictive control solution stage, the model predictive control solution stage and the safe distance field processing stage, and the safe distance field processing stage and the safe residual gradient backpropagation stage are all examples of two processing stages with a data transfer relationship.
[0120] The step of converting the data output from the previous processing stage into the numerical precision used in the next processing stage includes: at the interface between the processing component corresponding to the previous processing stage and the processing component corresponding to the next processing stage, identifying, through the data type marked in the data description information, that the numerical precision of the data output from the previous processing stage is different from the numerical precision used in the next processing stage; and performing precision conversion on the data output from the previous processing stage through a hardware format conversion unit to obtain data that conforms to the numerical precision used in the next processing stage. The conversion process requires no additional software overhead.
[0121] In this embodiment of the application, for each processing stage, based on the numerical stability requirements and computational efficiency requirements of the processing stage, a corresponding numerical precision is configured for the processing stage. This allows processing stages with lower numerical stability requirements and higher computational efficiency requirements to improve computational throughput with lower numerical precision, and allows processing stages with higher numerical stability requirements to ensure the numerical stability of the computation process with higher numerical precision. This improves the overall computational efficiency of the system while ensuring the correctness of the computation.
[0122] In one implementation, the method further includes: Independent computation execution queues are established for the action model inference stage, the model prediction and control solution stage, the safety distance field processing stage, and the safety residual gradient backpropagation stage, respectively. If the safety residual does not meet the preset safety event triggering conditions, the computation execution queues corresponding to the action model inference stage, the model prediction and control solution stage, and the safety distance field processing stage are scheduled in parallel. When the safety residual meets the preset safety event triggering condition, a safety event flag is generated, and the computation execution queue corresponding to the backpropagation stage of the safety residual gradient is scheduled based on the safety event flag.
[0123] The computation execution queue is an independent task scheduling unit created on the unified memory accelerator of the electronic device for each processing stage. The kernel functions corresponding to different computation execution queues run in parallel on different subsets of hardware computing units without blocking each other. The hardware scheduler allocates computing resources to each computation execution queue according to priority and time constraints.
[0124] The safety residual does not meet the preset safety event triggering conditions, which means that the safety distance of each predicted point in the optimal motion trajectory is not lower than the safety distance threshold, and the safety residual does not exceed the preset threshold.
[0125] The safety residual satisfies the preset safety event triggering condition, which means that the safety distance of any predicted point in the optimal motion trajectory is lower than the safety distance threshold, and the safety residual exceeds the preset threshold.
[0126] Generating a security event flag means writing the event flag into the event register to trigger a security interrupt when the security residual meets the preset security event triggering conditions.
[0127] The step of scheduling the computation execution queue corresponding to the security residual gradient backpropagation stage based on the security event flag includes: responding to the security event flag written in the event register, additionally starting the computation execution queue corresponding to the security residual gradient backpropagation stage, increasing the scheduling priority of the computation execution queue, and correspondingly reducing the resources of other computation execution queues.
[0128] It should be noted that if the safety residual does not meet the preset safety event triggering conditions, the redundant computation of the computation execution queue corresponding to the safety residual gradient backpropagation stage is closed, and only the computation execution queues corresponding to the action model inference stage, the model prediction control solution stage, and the safety distance field processing stage are scheduled in parallel.
[0129] After the security incident is resolved, the computation execution queue corresponding to the security residual gradient backpropagation stage is automatically closed, and the state is restored to where only the computation execution queues corresponding to the action model inference stage, the model prediction and control solution stage, and the security distance field processing stage are scheduled in parallel.
[0130] In this embodiment, independent computation execution queues are established for the action model inference stage, the model prediction and control solution stage, the safety distance field processing stage, and the safety residual gradient backpropagation stage, respectively. This allows the kernel functions corresponding to each processing stage to run in parallel on different subsets of hardware computing units without blocking each other, thus improving overall processing efficiency. When the safety residual does not meet the preset safety event triggering condition, the computation execution queues corresponding to the action model inference stage, the model prediction and control solution stage, and the safety distance field processing stage are scheduled in parallel, preventing the computation execution queue corresponding to the safety residual gradient backpropagation stage from occupying computing resources for extended periods, thus saving hardware resource overhead. When the safety residual meets the preset safety event triggering condition, a safety event flag is generated. Based on the safety event flag, the computation execution queue corresponding to the safety residual gradient backpropagation stage is scheduled, and its scheduling priority is increased. This ensures that the safety correction process can be completed with the shortest possible latency, improving the response speed and reliability of the electronic device safety control system in response to sudden safety events.
[0131] The method of this application has been described above; the apparatus of this application will be described below.
[0132] See Figure 2 , Figure 2 This is a schematic diagram of a shared physical memory allocation device provided in an embodiment of this application. The shared physical memory allocation device is applied to an electronic device, which includes multiple processing components. The multiple processing components are used to execute multiple processing stages, and each processing component is used to execute at least one of the multiple processing stages, such as... Figure 2 As shown, the shared physical memory allocation device 20 includes: Allocation module 201 is used to allocate a shared physical memory pool in the electronic device for the common access of the multiple processing components; The partitioning module 202 is used to divide the shared physical memory pool into multiple logical partitions based on the access attribute information corresponding to the multiple processing stages. Each logical partition is used to store the associated data of at least one of the multiple processing stages. The module 203 is used to establish partition description information for each logical partition. The partition description information is used by the multiple processing components to locate and access the data stored in the corresponding logical partition.
[0133] It should be noted that the shared physical memory allocation device 20 described above can execute the shared physical memory allocation method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the embodiments can be found in the shared physical memory allocation method provided in the embodiments of this application.
[0134] See Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 30 includes a processor 301 and a memory 302. The memory 302 is connected to the processor 301, for example, via a bus.
[0135] Processor 301 is configured to support the electronic device 30 in performing the corresponding functions in the methods described in the above method embodiments. Processor 301 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0136] Memory 302 is used to store program code, etc. Memory 302 may include volatile memory (VM), such as random access memory (RAM); memory 302 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory 302 may also include combinations of the above types of memory.
[0137] The memory 302 is used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the shared physical memory allocation method in the embodiments of this application. The processor executes various functional applications and data processing of the shared physical memory allocation method by running the non-volatile software programs, instructions, and modules stored in the memory, thereby realizing the functions of the shared physical memory allocation method provided in the above method embodiments.
[0138] Memory 302 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function. The data storage area may store data created based on the use of the shared physical memory allocation device. In some embodiments, the memory may include memory remotely located relative to the processor, which can be connected to the shared physical memory allocation device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0139] The one or more modules are stored in the memory. When executed by the one or more processors, they perform the shared physical memory allocation method in any of the above method embodiments. For example, they perform the method steps described in the above method embodiments to realize the functions of the modules described in the above device embodiments.
[0140] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the method described in the foregoing embodiments.
[0141] This application also provides a computer program product that, when run on an electronic device, causes any of the above methods to be executed.
[0142] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0143] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A method for allocating shared physical memory, characterized in that, Applied to an electronic device, the electronic device including a plurality of processing components, the plurality of processing components being configured to perform a plurality of processing stages, each of the plurality of processing components being configured to perform at least one of the plurality of processing stages; the method comprising: A shared physical memory pool is allocated in the electronic device for the common access of the multiple processing components; Based on the access attribute information corresponding to the multiple processing stages, the shared physical memory pool is divided into multiple logical partitions, and each logical partition is used to store the associated data of at least one of the multiple processing stages. A partition description information is established for each logical partition. The partition description information is used by the multiple processing components to locate and access the data stored in the corresponding logical partition.
2. The method according to claim 1, characterized in that, The method further includes: A unified data description format is pre-defined. The data description format is used to establish data description information corresponding to data stored in any of the logical partitions. The data description information is used by the multiple processing components to locate and read the corresponding data.
3. The method according to claim 2, characterized in that, The multiple processing stages include an action model inference stage, a model prediction and control solution stage, a safe distance field processing stage, and a safe residual gradient backpropagation stage; the method further includes: By executing the processing component of the action model inference stage, action model inference is performed to obtain action embeddings, and the action embeddings are written into the first logical partition among the plurality of logical partitions. A first data description information for the action embedding is established based on the data description format; By executing the processing component of the model predictive control solution stage, the action embedding is read based on the first data description information, the optimal motion trajectory is obtained by performing model predictive control solution based on the action embedding, and the optimal motion trajectory is written into the second logical partition among the plurality of logical partitions. A second data description information for the optimal motion trajectory is established based on the data description format; By executing the processing component of the safe distance field processing stage, the optimal motion trajectory is read based on the second data description information of the optimal motion trajectory, the safe residual is determined based on the optimal motion trajectory, and the safe residual is written into the third logical partition among the plurality of logical partitions. A third data description information for the safety residual is established based on the data description format; By executing the processing component of the backpropagation phase of the security residual gradient, the security residual is read based on the third data description information of the security residual, the security residual gradient embedded by the security residual for the action is determined, and the security residual gradient is written into the fourth logical partition among the plurality of logical partitions. The fourth data description information of the safety residual gradient is established based on the data description format.
4. The method according to claim 3, characterized in that, The action model reasoning to obtain the action embedding includes: reasoning about the input data through the action model to obtain the action embedding; The step of obtaining the optimal motion trajectory based on the action embedding execution model predictive control includes: decoding the action embedding into model predictive control parameters; constructing a constrained optimization problem based on the model predictive control parameters; iteratively updating the optimization variables of the optimization problem and the dual variables corresponding to the constraints in the optimization problem until a preset convergence condition is met to obtain the optimal solution of the optimization variables; and determining the optimal motion trajectory based on the optimal solution of the optimization variables. The step of determining the safety residual based on the optimal motion trajectory includes: calculating the safety distance corresponding to each of multiple trajectory points in the optimal motion trajectory; determining the minimum safety distance from the multiple safety distances; and determining the safety residual based on the comparison result between the minimum safety distance and a safety distance threshold. Determining the safety residual gradient of the safety residual with respect to the action embedding includes: taking the derivative of the safety residual with respect to the optimal motion trajectory as the derivative variable to obtain a first gradient; taking the derivative of the optimal motion trajectory with respect to the model prediction control parameters as the derivative variable to obtain a second gradient; taking the derivative of the model prediction control parameters with respect to the action embedding as the derivative variable to obtain a third gradient; and determining the safety residual gradient based on the first gradient, the second gradient, and the third gradient.
5. The method according to claim 4, characterized in that, The method further includes: Obtain the first state of the electronic device in the current processing cycle and the second state of the electronic device in the previous processing cycle; The state change amount of the electronic device is determined based on the first state and the second state; When the state change satisfies the preset conditions, the initial value of the optimization variable for the current processing cycle is determined based on the optimal motion trajectory of the previous processing cycle, and the dual variable when the previous processing cycle satisfies the preset convergence condition is used as the initial value of the dual variable for the current processing cycle.
6. The method according to claim 3, characterized in that, The method further includes: The security residual gradient is read based on the fourth data description information, the action embedding correction amount is determined based on the security residual gradient, and the action embedding correction amount is written into the fourth logical partition; A fifth data description information for the action embedding correction amount is established based on the data description format; In the next processing cycle, by executing the processing component of the action model inference stage, the action embedding correction amount is read based on the fifth data description information, and the action embedding is corrected based on the action embedding correction amount.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: For each processing stage, a corresponding numerical precision is configured based on the numerical stability requirements and computational efficiency requirements of that processing stage.
8. The method according to claim 3, characterized in that, The method further includes: Independent computation execution queues are established for the action model inference stage, the model prediction and control solution stage, the safety distance field processing stage, and the safety residual gradient backpropagation stage, respectively. If the safety residual does not meet the preset safety event triggering conditions, the computation execution queues corresponding to the action model inference stage, the model prediction and control solution stage, and the safety distance field processing stage are scheduled in parallel. When the safety residual meets the preset safety event triggering condition, a safety event flag is generated, and the computation execution queue corresponding to the backpropagation stage of the safety residual gradient is scheduled based on the safety event flag.
9. An electronic device, characterized in that, The device includes a memory and a processor, the memory being connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, the processor causing the electronic device to perform the method as described in any one of claims 1-8 when executing the one or more computer programs.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-8.