A traveling body control method and device

By generating multiple sets of candidate action sequences and performing weighted fusion, the control error problem of a single control sequence in nonlinear environments for autonomous driving is solved, achieving higher precision and robust autonomous driving control.

CN122111011APending Publication Date: 2026-05-29JINGDONG KUNPENG (JIANGSU) TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINGDONG KUNPENG (JIANGSU) TECH CO LTD
Filing Date
2026-02-26
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing autonomous driving technologies, predictive control methods based on a single control sequence are difficult to accurately reflect the nonlinear characteristics of dynamics and the external environment, leading to the accumulation of control errors and affecting control accuracy and robustness.

Method used

By generating multiple sets of candidate action sequences and performing weighted fusion, the target action sequence is determined. Considering various control methods and their effects on vehicle motion, the path cost is quantified and weighted averaged to generate the final control command.

Benefits of technology

It improves the accuracy and robustness of autonomous driving control, reduces control deviation and error in nonlinear scenarios, and enhances trajectory tracking capability and stability in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122111011A_ABST
    Figure CN122111011A_ABST
Patent Text Reader

Abstract

The application provides a driving main body control method and device, and the method comprises the following steps: determining an initial action sequence for controlling a target main body to converge to a reference track; generating a plurality of groups of candidate action sequences of the target main body according to the initial action sequence, and determining a path cost corresponding to each candidate action sequence; and weighting and fusing the plurality of groups of candidate action sequences based on the path cost to obtain a target action sequence for controlling the movement of the target main body. According to the embodiment, the target action sequence for controlling the movement of the target main body is obtained by weighting and fusing the plurality of groups of candidate action sequences, nonlinear loss in the driving process of the main body is avoided, the precision of the driving main body control is improved, and the robustness of the control system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and more specifically, to a method and apparatus for controlling a driving vehicle. Background Technology

[0002] In the field of autonomous driving and intelligent driving vehicle control, vehicle motion control typically requires generating corresponding control commands based on a reference trajectory to guide the vehicle along a target path. Existing technologies often employ predictive control methods based on a single control sequence to control the vehicle. These methods usually rely on linear assumptions or local approximations to generate control commands. However, in actual driving processes, dynamic characteristics and external environmental factors exhibit significant nonlinear features. A single control sequence cannot accurately reflect the impact of different control actions on the vehicle's motion state, easily leading to the accumulation of control errors under complex conditions, thus affecting control accuracy and robustness. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide at least one driving subject control method, device, electronic device and storage medium, which can obtain a target action sequence for controlling the movement of the target subject by weighted fusion of multiple sets of candidate action sequences, avoid nonlinear losses during the subject's driving process, improve the accuracy of driving subject control and improve the robustness of the control system.

[0004] In a first aspect, embodiments of the present invention provide a driving vehicle control method, comprising: Determine the initial sequence of actions used to control the convergence of the target subject to the reference trajectory; Based on the initial action sequence, generate multiple sets of candidate action sequences for the target subject and determine the path cost corresponding to each candidate action sequence; The target action sequence is obtained by weighted fusion of multiple candidate action sequences based on path cost to control the movement of the target subject.

[0005] Optionally, an initial sequence of actions for controlling the convergence of the target subject to the reference trajectory is determined, including: Based on the state information of the target subject and the reference trajectory, determine the deviation information between the target subject and the reference trajectory; Determine the linear relationship between deviation information and the historical command data of the target entity; Based on the linear relationship, the initial action sequence used to control the target body to converge to the reference trajectory is predicted.

[0006] Optionally, based on the initial action sequence, multiple sets of candidate action sequences for the target subject are generated, including: Based on the initial action sequence and the preset constraint range, perturbations that satisfy the preset probability distribution are applied to the steering control and acceleration control quantities in the initial action sequence to generate multiple sets of perturbation control quantities. Based on the disturbance control quantity and the state information of the target entity, multiple sets of candidate action sequences for the target entity are generated.

[0007] Optionally, the path cost corresponding to each candidate action sequence is determined, including: Each candidate action sequence is input into the motion model of the target subject to obtain the predicted trajectory of the target subject under the control of each candidate action sequence; Based on the deviation between the predicted trajectory and the reference trajectory, as well as the preset cost term, the path cost corresponding to each candidate action sequence is determined.

[0008] Optionally, the motion model is trained based on the following steps: Collect historical command data and historical state data of the target subject, and use the historical command data as input data so that the motion model can output the predicted state data of the target subject under the historical command data; The motion model is obtained by minimizing the error between the predicted state data and the historical state data.

[0009] Optionally, multiple candidate action sequences are weighted and fused based on path cost to obtain a target action sequence for controlling the motion of the target subject, including: Based on the path cost corresponding to each candidate action sequence, the weight value corresponding to each candidate action sequence is determined according to the preset weight allocation rule; among them, the candidate action sequence with the larger path cost corresponds to the smaller weight value. Based on the weight values ​​corresponding to each candidate action sequence, multiple candidate action sequences are weighted and fused to obtain the target action sequence.

[0010] In a second aspect, embodiments of the present invention provide a driving body control device, comprising: The determination module is used to determine the initial sequence of actions for controlling the target body to converge toward the reference trajectory; The generation module is used to generate multiple sets of candidate action sequences for the target subject based on the initial action sequence, and to determine the path cost corresponding to each candidate action sequence. The control module is used to perform weighted fusion of multiple candidate action sequences based on path cost to obtain the target action sequence used to control the movement of the target subject.

[0011] Thirdly, embodiments of the present invention also provide an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the first aspect or any optional implementation of the first aspect are performed.

[0012] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the first aspect or any optional implementation thereof.

[0013] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the method of any of the above embodiments.

[0014] In any of the above aspects or any implementation methods, the determined initial action sequence can be used to provide a basic control direction consistent with the reference trajectory, so that the generated candidate action sequences are concentrated within a reasonable control range. Based on this, multiple sets of candidate action sequences are generated around the initial action sequence, enabling the control system to simultaneously consider multiple possible control methods and their corresponding vehicle motion results, thereby covering the performance of vehicle dynamics nonlinear characteristics under different control conditions. By calculating the path cost for each candidate action sequence, the advantages and disadvantages of different candidate action sequences in terms of trajectory tracking error, driving smoothness, and constraint satisfaction can be quantified, and different weights can be assigned to each candidate action sequence accordingly. Furthermore, based on the path cost, multiple sets of candidate action sequences are weighted and fused, so that the final target action sequence is not derived from a single path or a single control decision, but rather is a weighted average of multiple control results, thereby effectively reducing the control deviation and error amplification problems that a single action sequence may introduce in nonlinear scenarios. This improves the stability and continuity of the control output, and thus achieves more accurate trajectory tracking in complex driving environments. Meanwhile, by weighted fusion of multiple candidate action sequences, control performance can be maintained when facing external interference, thereby improving the robustness of the driving body control.

[0015] The beneficial effects of the aforementioned vehicle control device, electronic equipment, and storage medium are described in the description of the aforementioned vehicle control method, and will not be repeated here. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to the present invention and, together with the specification, serve to explain the technical solutions of the present invention. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 A flowchart of a driving body control method provided by an embodiment of the present invention is shown; Figure 2 A schematic diagram of the disturbance control quantity provided in an embodiment of the present invention is shown; Figure 3 A schematic diagram of a driving body control device provided in an embodiment of the present invention is shown; Figure 4 An exemplary system architecture in which embodiments of the present invention can be applied is shown; Figure 5 A schematic diagram of the structure of a computer system used to implement embodiments of the present invention is shown. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0019] It should be noted that the collection, use, storage, sharing and transfer of user personal information involved in the technical solution of the present invention all comply with the provisions of relevant laws and regulations, and require notification to users and obtaining their consent or authorization. When applicable, user personal information is subjected to de-identification and / or anonymization and / or encryption technical processing.

[0020] The above problems and solutions are the result of the inventor's practice and careful research. The discovery process of the above problems and the solutions proposed for the above problems should be considered as the inventor's contribution to the invention.

[0021] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of this invention. The components of this invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0022] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0023] To facilitate understanding of this embodiment, a detailed description of the driving entity control method disclosed in this invention will be provided first. The driving entity control method provided in this invention generally executes a computer device with a certain computing capability. This computer device may include, for example, a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, this driving entity control method can be implemented by a processor calling computer-readable instructions stored in memory.

[0024] See Figure 1 The diagram shows a flowchart of a driving entity control method provided in an embodiment of the present invention. The method includes steps S101 to S103, wherein: S101: Determine the initial sequence of actions used to control the convergence of the target subject to the reference trajectory.

[0025] In this embodiment of the invention, the target entity can be an execution object with autonomous movement capabilities, such as an autonomous vehicle, mobile robot, unmanned delivery equipment, or other device that requires motion control along a predetermined path. The reference trajectory can be a target driving path pre-generated according to task requirements, environmental constraints, or planning algorithms, used to indicate the position, posture, and driving direction that the target entity expects to achieve during movement. The reference trajectory typically reflects control objectives such as safety, efficiency, or comfort. Since the target entity is affected by factors such as dynamic nonlinearity, external disturbances, and state measurement errors during actual movement, its current motion state often deviates from the reference trajectory. Therefore, it is necessary to adjust the motion state of the target entity through control actions. To this end, an initial action sequence converging towards the reference trajectory is introduced during the control process. This provides a basic control direction consistent with the reference trajectory for subsequent control, allowing the target entity's motion trend to approach the reference trajectory in the initial stage, thereby reducing the deviation range.

[0026] In this embodiment of the invention, determining the initial action sequence for controlling the convergence of the target subject to the reference trajectory includes: determining the deviation information between the target subject and the reference trajectory based on the state information of the target subject and the reference trajectory; determining the linear relationship between the deviation information and the historical command data of the target subject; and predicting the initial action sequence for controlling the convergence of the target subject to the reference trajectory based on the linear relationship.

[0027] In practical implementation, the current state information of the target entity can be obtained first. This state information can include the target entity's current position, velocity, heading angle, heading angular velocity, and other motion-related state parameters. Simultaneously, a reference trajectory corresponding to the target entity's driving task can be obtained. By comparing the target entity's current state with the expected state at the corresponding position in the reference trajectory, the deviation information of the target entity in terms of position, attitude, or motion state relative to the reference trajectory can be calculated. This deviation information can reflect the degree of difference between the target entity's current motion state and its expected motion state. Then, the relationship between the deviation information and control commands can be analyzed by combining the command data recorded by the target entity during its historical operation. Since the influence of control commands on changes in the target entity's motion state near its current state and within a short-term prediction range has the characteristics of local continuity and approximate descriptibility, the mapping relationship between the deviation information and historical command data can be considered as an approximate linear relationship. This approximate linear relationship is not a precise description of the overall motion characteristics of the target entity, but rather a linear approximation of the state deviation changes caused by changes in control commands near the current state and the reference trajectory, used to simplify the control calculation process. An approximate linear relationship can be obtained by linear fitting or local regression analysis of historical command data and corresponding deviation changes, thus obtaining a linear correspondence between deviation information and control commands under the current state conditions. Based on this linear relationship, the impact of different control commands on deviation information can be estimated in the prediction time domain, and a sequence of control commands that can gradually reduce deviation information can be calculated, thereby forming an initial action sequence to guide the target subject to converge to the reference trajectory.

[0028] S102: Based on the initial action sequence, generate multiple sets of candidate action sequences for the target subject and determine the path cost corresponding to each candidate action sequence.

[0029] In this embodiment of the invention, after obtaining the initial action sequence, multiple candidate action sequences can be searched using the initial action sequence as the center. Specifically, generating multiple candidate action sequences for the target subject based on the initial action sequence includes: applying disturbances satisfying a preset probability distribution to the steering control and acceleration control quantities in the initial action sequence based on the initial action sequence and a preset constraint range, thereby generating multiple sets of disturbance control quantities; and generating multiple candidate action sequences for the target subject based on the disturbance control quantities and the state information of the target subject.

[0030] In practical implementation, after obtaining the initial action sequence used to guide the target subject to converge to the reference trajectory, multiple candidate action sequences can be generated around this initial action sequence to fully explore the impact of different control actions on the target subject's motion outcome. The initial action sequence can be used to characterize a relatively reasonable control direction in the current state. However, since the nonlinear relationship between the control command and the subject's state was not considered when obtaining the initial action sequence, this initial action sequence cannot achieve precise control of the target subject. Therefore, multiple candidate action sequences can be obtained based on this initial action sequence by making certain amplitude changes to the control quantity to explore the nonlinear relationship between the control command and the target subject's motion state. For example... Figure 2 As shown, the variation range of the candidate action sequence can be limited by preset constraints. These constraints can be pre-set based on the physical characteristics of the target entity, the capabilities of the actuator, and safety requirements. For example, the maximum turning angle range and rate of change of the steering control quantity, as well as the upper and lower limits of the acceleration control quantity, can be limited to ensure the feasibility and safety of the generated control action sequence during actual execution. Furthermore, to introduce a certain degree of control diversity within the constraints, perturbations satisfying a preset probability distribution can be applied to the steering and acceleration control quantities in the initial action sequence. This preset probability distribution can be used to describe the variation pattern of the control quantities near the initial action sequence. For example, it can make smaller-amplitude perturbations have a higher probability of occurrence, while larger-amplitude perturbations have a relatively lower probability of occurrence, thus ensuring both stability and exploratory nature in the generated candidate action sequence. Specifically, random perturbation values ​​corresponding to the length of the initial action sequence can be generated according to the preset probability distribution. Under the premise of satisfying the constraints, these perturbation values ​​are superimposed on the steering and acceleration control quantities in the initial action sequence to obtain multiple sets of perturbation control quantities. Then, by combining the current state information of the target subject, the disturbance control quantity can be arranged into multiple sets of control command sequences in chronological order, thereby generating multiple sets of candidate action sequences for the target subject.

[0031] In this embodiment of the invention, after obtaining multiple sets of candidate action sequences, the path cost for the target subject to execute each candidate action sequence can be determined. Specifically, determining the path cost corresponding to each candidate action sequence includes: inputting each candidate action sequence into the motion model of the target subject to obtain the predicted trajectory of the target subject under the control of each candidate action sequence; and determining the path cost corresponding to each candidate action sequence based on the deviation between the predicted trajectory and the reference trajectory and a preset cost term.

[0032] In practice, for each candidate action sequence, the sequence is combined with the target's current state information and input into a pre-trained motion model describing the target's motion characteristics. This predicts the target's motion state changes over a future period under the control of the candidate action sequence, thus obtaining the corresponding predicted trajectory. The training method for the motion model is detailed below and will not be elaborated here. The predicted trajectory typically consists of multiple consecutive state points, reflecting the changes in the target's spatial position, posture, and velocity over time. After obtaining the predicted trajectory, it is compared with a reference trajectory. The degree to which the candidate action sequence conforms to the target's motion along the reference trajectory is measured by the deviation between the two in terms of position, posture, or motion state. Based on this, a pre-defined cost term can be introduced to comprehensively evaluate the candidate action sequences. The cost items may include trajectory tracking error cost, used to quantify the deviation between the predicted trajectory and the reference trajectory; control smoothness cost, used to reflect the magnitude of the change in control action over time to avoid overly drastic control commands; control amplitude cost, used to limit the magnitude of steering control and acceleration control quantities to ensure that control commands meet the capabilities and safety requirements of the actuator; and stability or comfort-related cost items, used to constrain changes in the motion state of the target subject. By combining one or more of the above cost items according to preset weights, the corresponding path cost can be calculated for each set of candidate action sequences. The obtained path cost can be used to reflect the quality of each candidate action sequence in the corresponding dimension. It should be noted that the above cost items are only illustrative examples of feasible implementation methods in the embodiments of the present invention and do not constitute an improper limitation of the present invention. In practical applications, cost items and calculation methods for each cost item can be set according to actual needs and circumstances. The embodiments of the present invention do not make specific limitations in this regard, and the functionality shall prevail.

[0033] In this embodiment of the invention, the motion model mentioned above can be trained based on the historical command data and historical state data of the target subject. Specifically, the motion model is trained based on the following steps: collecting the historical command data and historical state data of the target subject, using the historical command data as input data, so that the motion model outputs the predicted state data of the target subject under the historical command data; and using minimizing the error between the predicted state data and the historical state data as the training objective to obtain the motion model.

[0034] In practical implementation, historical command data and corresponding historical state data can be continuously collected during the actual or simulated operation of the target entity. Historical command data can include steering control quantities, acceleration control quantities, or other control commands issued at each control moment to drive the target entity's motion. Historical state data can include the target entity's position, velocity, heading angle, and heading angular velocity after executing control commands. These historical command and state data correspond to each other in time and can be used to reflect the relationship between control commands and the target entity's motion results. During model training, historical command data can be used as input data to the motion model to be trained, allowing the model to output predicted state data of the target entity under the corresponding control commands. The predicted state data can then be compared with historical state data collected at the same time scale. The error between the two can be calculated to measure the degree of fit between the current motion model and the actual motion behavior of the target entity. The error can be calculated based on the difference between state variables, and corresponding weights can be set for different state dimensions to reflect the importance of each state variable in the motion description. Based on this, the parameters of the motion model can be adjusted and updated with the goal of minimizing the overall error between the predicted state data and the historical state data. This allows the model to gradually learn the intrinsic mapping relationship between control commands and changes in the target subject's motion state. By repeatedly training on a large amount of historical data, the resulting motion model can accurately represent the motion patterns of the target subject under different control commands. This enables the model to predict the target subject's trajectory under each candidate action sequence based on the input target subject's current motion state and candidate action sequences.

[0035] S103: Based on path cost, multiple candidate action sequences are weighted and fused to obtain the target action sequence used to control the movement of the target subject.

[0036] In this embodiment of the invention, a target action sequence for controlling the movement of a target subject is obtained by weighted fusion of multiple candidate action sequences based on path cost. This includes: determining the weight value corresponding to each candidate action sequence according to the path cost of each candidate action sequence and a preset weight allocation rule; wherein the candidate action sequence with a larger path cost has a smaller weight value; and weighted fusion of multiple candidate action sequences based on the weight value of each candidate action sequence to obtain the target action sequence.

[0037] In practical implementation, each candidate action sequence can first be assigned a corresponding weight value according to the path cost of each candidate action sequence and a preset weight allocation rule. The weight allocation rule describes the correspondence between path cost and weight value. The smaller the path cost, the better the overall performance of the candidate action sequence in terms of trajectory tracking accuracy, control smoothness, and feasibility, and the larger the corresponding weight value can be set; conversely, the larger the path cost, the smaller the corresponding weight value. In one feasible implementation, the reciprocal of the path cost of each candidate action sequence can be calculated or the weight can be generated based on an exponential decay function. The obtained weights are then normalized so that the sum of all weight values ​​is 1, thereby ensuring that the influence of each candidate action sequence on the final control result is controllable. After determining the weight value of each candidate action sequence, multiple sets of candidate action sequences can be aligned according to the time dimension, and the control quantity can be weighted and calculated based on their respective weight values, thereby obtaining a comprehensive control command at each control moment. Through the weighted fusion method described above, the final target action sequence integrates the advantages of multiple candidate action sequences under different control schemes. Its control result does not depend on a single candidate action sequence, but rather balances various possible control paths. This effectively reduces the negative impact of a single candidate action sequence on the control result in the presence of model errors or environmental disturbances, enabling the target action sequence to maintain good tracking ability of the reference trajectory while possessing higher stability and continuity, thereby improving the robustness and reliability of the target's motion control.

[0038] According to a second aspect of the embodiments of the present invention, such as Figure 3 As shown, a driving body control device 300 is provided, including: The determination module 301 is used to determine the initial action sequence for controlling the target body to converge toward the reference trajectory; The generation module 302 is used to generate multiple sets of candidate action sequences for the target subject based on the initial action sequence, and determine the path cost corresponding to each candidate action sequence. The control module 303 is used to perform weighted fusion of multiple candidate action sequences based on path cost to obtain a target action sequence for controlling the movement of the target subject.

[0039] Optionally, the determining module 301 is specifically used for: Based on the state information of the target subject and the reference trajectory, determine the deviation information between the target subject and the reference trajectory; Determine the linear relationship between deviation information and the historical command data of the target entity; Based on the linear relationship, the initial action sequence used to control the target body to converge to the reference trajectory is predicted.

[0040] Optionally, the generation module 302 is specifically used for: Based on the initial action sequence and the preset constraint range, perturbations that satisfy the preset probability distribution are applied to the steering control and acceleration control quantities in the initial action sequence to generate multiple sets of perturbation control quantities. Based on the disturbance control quantity and the state information of the target entity, multiple sets of candidate action sequences for the target entity are generated.

[0041] Optionally, the generation module 302 is specifically used for: Each candidate action sequence is input into the motion model of the target subject to obtain the predicted trajectory of the target subject under the control of each candidate action sequence; Based on the deviation between the predicted trajectory and the reference trajectory, as well as the preset cost term, the path cost corresponding to each candidate action sequence is determined.

[0042] Optionally, the generation module 302 is also used for: Collect historical command data and historical state data of the target subject, and use the historical command data as input data so that the motion model can output the predicted state data of the target subject under the historical command data; The motion model is obtained by minimizing the error between the predicted state data and the historical state data.

[0043] Optionally, the control module 303 is specifically used for: Based on the path cost corresponding to each candidate action sequence, the weight value corresponding to each candidate action sequence is determined according to the preset weight allocation rule; among them, the candidate action sequence with the larger path cost corresponds to the smaller weight value. Based on the weight values ​​corresponding to each candidate action sequence, multiple candidate action sequences are weighted and fused to obtain the target action sequence.

[0044] According to a third aspect of the present invention, an electronic device for controlling a driving entity is provided, comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the first aspect of the present invention.

[0045] According to a fourth aspect of the present invention, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method provided in the first aspect of the present invention.

[0046] According to a fifth aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method of any of the above embodiments.

[0047] Figure 4An exemplary system architecture 400 is shown that can be applied to the driving vehicle control method or driving vehicle control device implemented in this invention.

[0048] like Figure 4 As shown, system architecture 400 may include terminal devices 401, 402, and 403, a network 404, and a server 405. Network 404 serves as the medium for providing communication links between terminal devices 401, 402, and 403 and server 405. Network 404 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0049] Users can use terminal devices 401, 402, and 403 to interact with server 405 via network 404 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 401, 402, and 403, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0050] Terminal devices 401, 402, and 403 can be various electronic devices with displays that support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0051] Server 405 can be a server that provides various services, such as a backend management server that supports shopping websites browsed by users using terminal devices 401, 402, and 403 (for example only). The backend management server can process the received vehicle control requests and send the processing results (for example only) back to the terminal devices.

[0052] It should be noted that the driving vehicle control method provided in this embodiment of the invention is generally executed by server 405, and correspondingly, the driving vehicle control device is generally located in server 405. The driving vehicle control method provided in this embodiment of the invention can also be executed by terminal devices 401, 402, and 403, and correspondingly, the driving vehicle control device can be located in terminal devices 401, 402, and 403.

[0053] It should be understood that Figure 4 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0054] The following is for reference. Figure 5 It shows a schematic diagram of the structure of a computer system 500 suitable for implementing a terminal device of the present invention. Figure 5 The terminal device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0055] like Figure 5 As shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 502 or programs loaded from storage section 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the system 500. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0056] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 510 as needed so that computer programs read from it can be installed into storage section 508 as needed.

[0057] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, it performs the functions defined above in the system of this invention.

[0058] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0059] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0060] The modules described in the embodiments of the present invention can be implemented in software or hardware. The described modules can also be located in a processor. For example, a processor includes a determining module, a generating module, and a controlling module. The names of these modules do not necessarily limit the module itself. For example, the determining module can also be described as "a module for determining the initial action sequence for controlling the target subject to converge to the reference trajectory".

[0061] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, which, when executed by the device, implement the following method: determining an initial action sequence for controlling the convergence of a target subject to a reference trajectory; generating multiple sets of candidate action sequences for the target subject based on the initial action sequence, and determining the path cost corresponding to each candidate action sequence; and performing weighted fusion of the multiple sets of candidate action sequences based on the path cost to obtain a target action sequence for controlling the movement of the target subject.

[0062] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for controlling a driving vehicle, characterized in that, include: Determine the initial sequence of actions used to control the convergence of the target subject to the reference trajectory; Based on the initial action sequence, multiple candidate action sequences for the target subject are generated, and the path cost corresponding to each candidate action sequence is determined. The multiple candidate action sequences are weighted and fused based on the path cost to obtain the target action sequence used to control the movement of the target subject.

2. The method according to claim 1, characterized in that, Determine the initial sequence of actions used to control the convergence of the target subject to the reference trajectory, including: Based on the state information of the target entity and the reference trajectory, the deviation information between the target entity and the reference trajectory is determined; Determine the linear relationship between the deviation information and the historical command data of the target entity; Based on the linear relationship, an initial action sequence for controlling the target body to converge toward the reference trajectory is predicted.

3. The method according to claim 1, characterized in that, Based on the initial action sequence, multiple candidate action sequences for the target subject are generated, including: Based on the initial action sequence and the preset constraint range, perturbations that satisfy a preset probability distribution are applied to the steering control quantity and acceleration control quantity in the initial action sequence to generate multiple sets of perturbation control quantities. Based on the disturbance control quantity and the state information of the target entity, multiple sets of candidate action sequences for the target entity are generated.

4. The method according to claim 1, characterized in that, Determine the path cost corresponding to each candidate action sequence, including: The candidate action sequences are input into the motion model of the target subject to obtain the predicted trajectory of the target subject under the control of each candidate action sequence. Based on the deviation between the predicted trajectory and the reference trajectory, and the preset cost term, the path cost corresponding to each candidate action sequence is determined.

5. The method according to claim 4, characterized in that, The motion model was trained based on the following steps: Collect historical command data and historical state data of the target subject, and use the historical command data as input data so that the motion model outputs the predicted state data of the target subject under the historical command data; The motion model is obtained by minimizing the error between the predicted state data and the historical state data as the training objective.

6. The method according to claim 1, characterized in that, The multiple candidate action sequences are weighted and fused based on the path cost to obtain a target action sequence for controlling the motion of the target subject, including: Based on the path cost corresponding to each candidate action sequence, the weight value corresponding to each candidate action sequence is determined according to a preset weight allocation rule; wherein, the candidate action sequence with the larger path cost has a smaller weight value. Based on the weight values ​​corresponding to each candidate action sequence, the multiple sets of candidate action sequences are weighted and fused to obtain the target action sequence.

7. A driving body control device, characterized in that, include: The determination module is used to determine the initial sequence of actions for controlling the target body to converge toward the reference trajectory; The generation module is used to generate multiple sets of candidate action sequences for the target subject based on the initial action sequence, and to determine the path cost corresponding to each candidate action sequence. The control module is used to perform weighted fusion of the multiple candidate action sequences based on the path cost to obtain a target action sequence for controlling the movement of the target subject.

8. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.

10. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.