Vehicle control method and device, electronic equipment and storage medium

By generating environmental description information and predicting driving trajectories, the problem of low vehicle control accuracy is solved, and high-precision control is achieved in sudden dynamic scenarios.

CN122009209APending Publication Date: 2026-05-12CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHERY AUTOMOBILE CO LTD
Filing Date
2026-01-20
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing multi-sensor fusion technology has failed to effectively solve the problem of data error propagation in sudden dynamic scenarios, resulting in low vehicle control accuracy.

Method used

By acquiring sensor data from multiple sensor devices, environmental description information is generated. Based on this information and the first sensor data, the driving trajectory is predicted. Image analysis, clustering and fitting techniques are used to generate the vehicle's predicted driving trajectory. The trajectory is then adjusted in conjunction with external input data to control the vehicle's movement.

Benefits of technology

This effectively avoids the propagation of data errors in sudden dynamic scenarios and improves the control accuracy of the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122009209A_ABST
    Figure CN122009209A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle control method and device, electronic equipment and a storage medium, a vehicle is deployed with multiple sensing devices, the method comprises the steps that in the current driving process of the vehicle, sensing data from the multiple sensing devices are obtained, multiple pieces of sensing data are obtained, and the sensing data are used for displaying the environment where the vehicle is located; based on the multiple pieces of sensing data, environment description information of the vehicle is generated, and the environment description information is used for describing the environment through a preset information format; on the basis of the environment description information and first sensing data in the multiple pieces of sensing data, a predicted driving track of the vehicle is generated, the first sensing data is used for displaying the environment through the image, the predicted driving track is the driving track of the vehicle in the future driving process, and the future driving process is the driving process after the current driving process; and in the environment, controlling the vehicle to run according to the predicted running track. The technical problem that the control precision of the vehicle is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicles, and more specifically, to a vehicle control method, apparatus, electronic device, and storage medium. Background Technology

[0002] Currently, existing multi-sensor fusion technologies often treat the sensing module as an independent module, without coupling it with the vehicle's control system. This can lead to data error propagation in sudden dynamic scenarios, making it difficult to provide accurate data for subsequent decisions and resulting in low vehicle control precision.

[0003] There is currently no effective solution to the technical problem of low control precision in the aforementioned vehicles. Summary of the Invention

[0004] This application provides a vehicle control method, apparatus, electronic device, and storage medium to at least solve the technical problem of low vehicle control accuracy.

[0005] According to one aspect of the embodiments of this application, a vehicle control method is provided. The vehicle is equipped with multiple sensing devices. The method includes: acquiring sensing data from the multiple sensing devices during the current driving process of the vehicle to obtain multiple sensing data, wherein the sensing data is used to display the environment in which the vehicle is located; generating environmental description information of the vehicle based on the multiple sensing data, wherein the environmental description information is used to describe the environment through a preset information format; generating a predicted driving trajectory of the vehicle based on the environmental description information and first sensing data from the multiple sensing data, wherein the first sensing data is used to display the environment through an image, the predicted driving trajectory is the driving trajectory of the vehicle in the future driving process, and the future driving process is the driving process after the current driving process; and controlling the vehicle to drive according to the predicted driving trajectory in the environment.

[0006] Furthermore, the multiple sensor data include: second sensor data and third sensor data. The first sensor data is sensor data from a first sensor among multiple sensor devices, the second sensor data is sensor data from a second sensor among multiple sensor devices, and the third sensor data is sensor data from a third sensor among multiple sensor devices. The first sensor device, the second sensor device, and the third sensor device are of different device types. Based on the multiple sensor data, generating environmental description information for the vehicle includes: performing image analysis on the first sensor data to obtain image analysis results, where the image analysis results are used to represent different types of objects in the environment; performing clustering and fitting on the second sensor data to obtain fitting results, where the fitting results are used to represent the size of different types of objects; determining the predicted state information of different types of objects based on the third sensor data, where the predicted state information is used to represent the position and size of different types of objects; and generating environmental description information based on the image analysis results, the fitting results, and the predicted state information.

[0007] Further, image analysis is performed on the first sensing data to obtain image analysis results, including: inputting the first sensing data into an image analyzer; extracting first features of different types of objects from the first sensing data in the image analyzer; performing semantic segmentation on the first sensing data in the image analyzer to obtain semantic segmentation results, wherein the semantic segmentation results are used to represent the semantic labels to which different types of objects belong; performing visual description on the first sensing data in the image analyzer to obtain visual description results, wherein the visual description results are used to describe the state of different types of objects under visual conditions; and determining the first features, semantic segmentation results, and visual description results as the image analysis results.

[0008] Furthermore, the second sensing data is clustered and fitted to obtain fitting results, including: clustering the second sensing data to obtain cluster data of different types of objects, wherein the cluster data is used to represent the cluster category to which the object belongs; and fitting the cluster data to obtain fitting results.

[0009] Furthermore, based on the third sensor data, the predicted state information of different types of objects is determined, including: extracting the second features of different types of objects from the third sensor data, and generating feature maps corresponding to the second features; using the state prediction model to predict the feature maps to obtain the predicted state information, wherein the state prediction model is constructed based on the center point algorithm.

[0010] Furthermore, based on the image analysis results, fitting results, and predicted state information, environmental description information is generated, including: performing time correction on the image analysis results and fitting results according to the timestamp of the third sensor data; transforming the corrected image analysis results, corrected fitting results, and predicted state information to the coordinate system of the vehicle to obtain the transformed image analysis results, transformed fitting results, and transformed predicted state information; and integrating the transformed image analysis results, transformed fitting results, and transformed predicted state information into environmental description information.

[0011] Furthermore, the method also includes: acquiring input data from outside the vehicle, wherein the input data includes: environmental meteorological data and scene data, and vehicle driving data; generating a predicted driving trajectory of the vehicle based on environmental description information and first sensor data from multiple sensor data, including: generating driving trajectory points of the vehicle within a future driving time based on the input data, environmental description information, and first sensor data; fitting the driving trajectory points to obtain a fitted driving trajectory; and generating a predicted driving trajectory based on the fitted driving trajectory.

[0012] Furthermore, based on the fitted driving trajectory, a predicted driving trajectory is generated, including: generating a buffer region along the side of the fitted trajectory; adjusting the fitted driving trajectory in response to different types of objects in the environment being in the buffer region, and determining the adjusted fitted driving trajectory as the predicted driving trajectory; and determining the fitted driving trajectory as the predicted driving trajectory in response to different types of objects in the environment not being in the buffer region.

[0013] Furthermore, in the environment, controlling vehicle driving according to the predicted driving trajectory includes: determining initial predicted driving data corresponding to the predicted driving trajectory, wherein the initial predicted driving data is used to represent the initial position and initial speed of the vehicle in the future driving process; and controlling vehicle driving according to the initial predicted driving data in the environment.

[0014] Furthermore, the initial predicted driving data includes: initial predicted position data and initial predicted speed data. The initial predicted position data represents the initial position of the vehicle during future driving, and the initial predicted speed data represents the initial speed of the vehicle during future driving. In the given environment, controlling the vehicle's driving according to the initial predicted driving data includes: determining target predicted position data corresponding to the predicted driving trajectory, wherein the target predicted position data represents the target position of the vehicle during future driving, and the distance between the target position and the initial position is less than or equal to a preset distance; adjusting the initial predicted speed data to obtain target predicted speed data, wherein the target predicted speed data represents the target speed of the vehicle during future driving, and the speed difference between the target speed and the initial speed is less than or equal to a preset speed difference; and in the given environment, controlling the vehicle's driving according to the target predicted position data and the target predicted speed data.

[0015] According to another aspect of the embodiments of this application, a vehicle control device is also provided. The vehicle is equipped with multiple sensing devices. The device includes: a first acquisition unit, configured to acquire sensing data from the multiple sensing devices during the current driving process of the vehicle, thereby obtaining multiple sensing data, wherein the sensing data is used to display the environment in which the vehicle is located; a first generation unit, configured to generate environmental description information of the vehicle based on the multiple sensing data, wherein the environmental description information is used to describe the environment through a preset information format; a second generation unit, configured to generate a predicted driving trajectory of the vehicle based on the environmental description information and the first sensing data from the multiple sensing data, wherein the first sensing data is used to display the environment through images, the predicted driving trajectory is the driving trajectory of the vehicle in the future driving process, and the future driving process is the driving process after the current driving process; and a control unit, configured to control the vehicle to drive according to the predicted driving trajectory in the environment.

[0016] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.

[0017] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0018] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0019] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.

[0020] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.

[0021] According to another aspect of the embodiments of this application, a vehicle is also provided, which includes the electronic equipment described in this application.

[0022] In this embodiment, when controlling the vehicle, during the vehicle's current driving process, sensor data from multiple sensing devices are acquired to obtain multiple sensor data. Based on the multiple sensor data, environmental description information of the vehicle is generated. Based on the environmental description information and the first sensor data among the multiple sensor data, a predicted driving trajectory of the vehicle is generated. In the environment, the vehicle is controlled to drive according to the predicted driving trajectory. Since this embodiment, when the vehicle is in a driving state, can generate environmental description information describing the vehicle's environment in a preset information format based on the acquired multiple sensor data, and combine the generated environmental description information with the first sensor data among the acquired multiple sensor data, a predicted driving trajectory of the vehicle can be generated. That is, the driving trajectory of the vehicle in the future driving process can be generated. Then, in the environment where the vehicle is located, the vehicle can be controlled to drive according to the generated predicted driving trajectory. This achieves the purpose of avoiding data error propagation in sudden dynamic scenarios, thereby solving the technical problem of low vehicle control accuracy and achieving the technical effect of improving vehicle control accuracy. Attached Figure Description

[0023] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0024] Figure 1(a) is a schematic diagram of an application scenario of a vehicle control method according to an embodiment of this application;

[0025] Figure 1(b) is a flowchart of a vehicle control method according to an embodiment of this application;

[0026] Figure 2 This is a flowchart of a data fusion method for a multimodal sensor according to an embodiment of this application;

[0027] Figure 3This is a flowchart of a trajectory generation method based on an enhanced vision-language-action model according to an embodiment of this application;

[0028] Figure 4 This is a structural block diagram of a vehicle control device according to an embodiment of this application;

[0029] Figure 5 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0030] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0032] According to an embodiment of this application, an embodiment of a vehicle control method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0033] As an optional implementation, the vehicle control method described above can be applied, but is not limited to, the application scenario shown in Figure 1(a). Figure 1(a) is a schematic diagram of an application scenario of a vehicle control method according to an embodiment of this application. As shown in Figure 1(a), in the application scenario, the terminal device 10 can communicate with the server 13 via the network 11, but is not limited to. The server 13 can perform operations on the database, such as writing or reading data. The terminal device 10 can include, but is not limited to, a human-machine interface screen, a processor, and a memory. The human-machine interface screen can be used, but is not limited to, to display a virtual machine on the mobile terminal 10. The vehicle 12 can be used, but is not limited to, to respond to the human-machine interface operation, execute the corresponding operation, or generate the corresponding instruction and send the generated instruction to the server 13.

[0034] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here. The vehicle control method of this application may include: step S102, acquiring sensing data from multiple sensing devices during the current driving process of the vehicle to obtain multiple sensing data; step S104, generating environmental description information of the vehicle based on the multiple sensing data; step S106, generating a predicted driving trajectory of the vehicle based on the environmental description information and the first sensing data among the multiple sensing data; and step S108, controlling the vehicle to drive according to the predicted driving trajectory in the environment.

[0035] It should be noted that all relevant information (including but not limited to environmental description information) and data (including but not limited to sensor data) involved in this application are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0036] Figure 1(b) is a flowchart of a vehicle control method according to an embodiment of the present application. The vehicle is equipped with multiple sensing devices. As shown in Figure 1(b), the method may include the following steps.

[0037] Step S112: During the current driving process of the vehicle, sensor data from multiple sensor devices are acquired to obtain multiple sensor data, which are used to display the environment in which the vehicle is located.

[0038] In the technical solution provided in step S112 of this application, the aforementioned sensing data can be used to display the environment in which the vehicle is located. The multiple sensing data can each display the environment in which the vehicle is located through different data formats. For example, the aforementioned sensing device can also be called a sensor perception device. If the multiple sensor perception devices include: a camera device, millimeter-wave radar, and lidar, then the multiple sensing data can include: image data, millimeter-wave radar point cloud data, and lidar point cloud data.

[0039] In this embodiment, the aforementioned image data can display the vehicle's environment in image form, the aforementioned millimeter-wave radar point cloud data can display the vehicle's environment in millimeter-wave radar point cloud form, and the aforementioned laser point cloud data can display the vehicle's environment in laser radar point cloud form. That is, the aforementioned multiple sensor data can respectively display the vehicle's environment in image form, millimeter-wave radar point cloud form, and laser radar point cloud form.

[0040] In this embodiment, the aforementioned environment may include: natural environment and driving environment. For example, the aforementioned natural environment may be any one of the following: sunny environment, rainy / snowy environment, and windy environment, etc.; the aforementioned driving environment may be any one of the following: congested driving environment, smooth driving environment, and random driving environment, etc.; the random driving environment may be a driving environment that randomly occurs during the current driving process of the vehicle, which is only an example and is not specifically limited.

[0041] In this embodiment, during the vehicle's current driving process, sensor data from multiple sensing devices are acquired, resulting in multiple sensor data sets. Optionally, this embodiment performs vehicle state detection. If the vehicle is detected to be in motion, during the vehicle's current driving process, sensor data collected by each of the multiple sensing devices is acquired, resulting in multiple sensor data sets. For example, during the vehicle's current driving process, image data is acquired from a camera device, millimeter-wave radar point cloud data is acquired from a millimeter-wave radar device, and lidar point cloud data is acquired from a lidar device. The aforementioned camera devices can be panoramic cameras or multi-view cameras; this is merely an example and not a specific limitation.

[0042] Step S114: Based on multiple sensor data, generate environmental description information for the vehicle, wherein the environmental description information is used to describe the environment using a preset information format.

[0043] In the technical solution provided in step S114 of this application, the aforementioned environment description information can be used to describe the environment through a preset information format, and the environment description information can be embodied in the form of an environment description file. For example, the preset information format can be a lightweight data exchange format (JavaScript Object Notation, abbreviated as JSON), and the environment description file can be a structured text scene description file. This is only an example and is not specifically limited.

[0044] In this embodiment, during the vehicle's current driving process, sensor data from multiple sensing devices are acquired. After obtaining the multiple sensor data, environmental description information of the vehicle is generated based on the multiple sensor data. Optionally, this embodiment performs different types of data processing operations on the multiple sensor data based on the acquired data, and then fuses the multiple processed sensor data to obtain environmental description information, thereby achieving the purpose of describing the environment through a preset information format.

[0045] Optionally, based on the acquired multiple sensor data, different types of data processing operations are performed on each of the multiple sensor data, then the processed multiple sensor data are corrected, the corrected multiple sensor data are transformed, and the transformed multiple sensor data are fused to obtain environmental description information. For example, transforming the corrected multiple sensor data to the same coordinate system and merging the multiple sensor data in the same coordinate system can obtain environmental description information.

[0046] Step S116: Based on the environmental description information and the first sensor data among multiple sensor data, a predicted driving trajectory of the vehicle is generated. The first sensor data is used to display the environment through images, and the predicted driving trajectory is the driving trajectory of the vehicle in the future driving process. The future driving process is the driving process after the current driving process.

[0047] In the technical solution provided by step S116 of this application, the first sensing data can be used to display the environment through images. Optionally, the first sensing data can be panoramic image data, and the panoramic image data can be stored in the first sensing device in an image file format (Joint Photographic Experts Group, abbreviated as JPEG). For example, the first sensing device can be a camera device suitable for vehicles.

[0048] In this embodiment, the predicted driving trajectory can be the driving trajectory of the vehicle in the future, and the future driving process is the driving process after the current driving process. Optionally, the predicted driving trajectory can be a continuous trajectory of the vehicle within a future driving time, and the future driving time can be used to reflect the duration of the future driving process. For example, the future driving time can be, but is not limited to, 5 minutes or 10 minutes, etc. The values ​​here are only illustrative and are not specifically limited.

[0049] In this embodiment, after generating environmental description information of the vehicle based on multiple sensor data, a predicted driving trajectory of the vehicle is generated based on the environmental description information and the first sensor data among the multiple sensor data. Optionally, after generating the environmental description information of the vehicle, this embodiment can generate a predicted driving trajectory of the vehicle by combining the generated environmental description information and the first sensor data among the acquired multiple sensor data, thereby achieving the purpose of determining the driving trajectory of the vehicle in the future driving process.

[0050] Optionally, by combining the aforementioned environmental description information and the first sensing data, the vehicle's trajectory points over the future driving time can be generated. Based on the generated trajectory points, a predicted driving trajectory for the vehicle can be generated. For example, by performing trajectory fitting on the generated trajectory points, the predicted driving trajectory for the vehicle can be obtained.

[0051] It should be noted that the method described above for generating a vehicle's predicted driving trajectory by combining environmental description information and first sensor data is merely illustrative and is not intended to impose specific limitations. Any method that can generate a vehicle's predicted driving trajectory based on the aforementioned environmental description information and the aforementioned first sensor data is within the protection scope of this application's embodiments, and will not be described in detail here.

[0052] Step S118: In the environment, control the vehicle to drive according to the predicted driving trajectory.

[0053] In the technical solution provided in step S118 of this application, after generating a predicted driving trajectory for the vehicle based on environmental description information and first sensor data from multiple sensor data, the vehicle is controlled to drive according to the predicted driving trajectory in the environment. Optionally, this embodiment uses the generated predicted driving trajectory to predict the vehicle's driving data, thereby obtaining initial predicted driving data corresponding to the predicted driving trajectory. The initial predicted driving data can be used to represent the initial position and initial speed of the vehicle during future driving. Then, in the aforementioned environment where the vehicle is located, the vehicle is controlled to drive according to the initial predicted driving data, thereby achieving the goal of avoiding data error propagation in sudden dynamic scenarios.

[0054] Optionally, the initial predicted driving data is adjusted to obtain target predicted driving data. This target predicted driving data can represent the vehicle's target position and target speed during future driving. The distance between the target position and the initial position is less than or equal to a preset distance, and the speed difference between the target speed and the initial speed is less than or equal to a preset speed difference. For example, the preset distance can be, but is not limited to, 0.3m, and the preset speed difference can be, but is not limited to, 2km / h. These values ​​are merely illustrative and not specifically limited. Subsequently, in the aforementioned environment, the vehicle is controlled according to the adjusted target predicted driving data, thereby avoiding the propagation of data errors in sudden dynamic scenarios.

[0055] In steps S112 to S118 of this application, when controlling the vehicle, during the vehicle's current driving process, sensor data from multiple sensing devices are acquired to obtain multiple sensor data; based on the multiple sensor data, environmental description information of the vehicle is generated; based on the environmental description information and the first sensor data among the multiple sensor data, a predicted driving trajectory of the vehicle is generated; and in the environment, the vehicle is controlled to drive according to the predicted driving trajectory. Since in this embodiment, when the vehicle is in a driving state, based on the acquired multiple sensor data, environmental description information describing the vehicle's environment in a preset information format can be generated, and combined with the generated environmental description information and the first sensor data among the acquired multiple sensor data, a predicted driving trajectory of the vehicle can be generated, that is, the driving trajectory of the vehicle in the future driving process can be generated. Then, in the environment where the vehicle is located, the vehicle can be controlled to drive according to the generated predicted driving trajectory, thereby achieving the purpose of avoiding data error propagation in sudden dynamic scenarios, thus solving the technical problem of low vehicle control accuracy, and thereby achieving the technical effect of improving vehicle control accuracy.

[0056] The steps for generating vehicle environmental description information based on multiple sensor data in this embodiment will be further described below.

[0057] As an optional embodiment, the multiple sensor data includes: second sensor data and third sensor data, where the first sensor data is sensor data from a first sensor among multiple sensor devices, the second sensor data is sensor data from a second sensor among multiple sensor devices, and the third sensor data is sensor data from a third sensor among multiple sensor devices. The types of the first sensor device, the second sensor device, and the third sensor device are different device types. Step S114, generating environmental description information for the vehicle based on the multiple sensor data, includes: performing image analysis on the first sensor data to obtain image analysis results, wherein the image analysis results are used to represent different types of objects in the environment; performing clustering and fitting on the second sensor data to obtain fitting results, wherein the fitting results are used to represent the size of different types of objects; determining predicted state information for different types of objects based on the third sensor data, wherein the predicted state information is used to represent the position and size of different types of objects; and generating environmental description information based on the image analysis results, the fitting results, and the predicted state information.

[0058] In this embodiment, the aforementioned multiple sensing data may include: first sensing data, second sensing data, and third sensing data. The first sensing data may be panoramic image data or non-panoramic image data, the second sensing data may be Doppler data, and the third sensing data may be three-dimensional point cloud data.

[0059] In this embodiment, the first sensing data is sensing data from a first sensing device among multiple sensing devices, the second sensing data is sensing data from a second sensing device among multiple sensing devices, and the third sensing data is sensing data from a third sensing device among multiple sensing devices. For example, the first sensing data may be image data from a camera device, the second sensing data may be point cloud data from a millimeter-wave radar, and the third sensing data may be point cloud data from a lidar.

[0060] In this embodiment, the aforementioned plurality of sensing devices may include a first sensing device, a second sensing device, and a third sensing device. The first sensing device, the second sensing device, and the third sensing device may be of different types. For example, the first sensing device may be a camera device suitable for vehicles, the second sensing device may be a millimeter-wave radar suitable for vehicles, and the third sensing device may be a lidar suitable for vehicles. This is merely an example and not a specific limitation.

[0061] In this embodiment, the image analysis may include: feature extraction, semantic segmentation, and visual description of the first sensing data. For example, the feature extraction of the first sensing data may be feature extraction of image data, the semantics of the first sensing data may be semantic segmentation of image data, and the visual description of the first sensing data may be a 360° visual description.

[0062] In this embodiment, the image analysis results described above can be used to represent different types of objects in the environment. For example, these different types of objects can be any one or any combination of the following: lane lines, traffic lights, pedestrians, and vehicles.

[0063] In this embodiment, image analysis is performed on the first sensing data to obtain image analysis results. Optionally, after acquiring multiple sensing data, this embodiment determines the first sensing data from the multiple sensing data, then performs feature extraction on the first sensing data to obtain first features of different types of objects, performs semantic segmentation on the first sensing data to obtain semantic segmentation results, and performs visual description on the first sensing data to obtain visual description results. Finally, the first features, the semantic segmentation results, and the visual description results are determined as the image analysis results.

[0064] Optionally, image data is determined from multiple raw datasets. Feature extraction is then performed on this image data to obtain feature vectors for different types of objects. Semantic segmentation is performed on the image data to obtain multiple semantic labels, and visual description is performed on the image data to obtain a natural language description. Finally, the feature vectors, semantic labels, and natural language description are combined to form the image analysis result.

[0065] In this embodiment, the fitting results described above can be used to represent the dimensions of different types of objects, where the dimensions can be described using the length and width of the object. For example, if the different types of objects are different types of obstacles, then the dimensions of the different types of objects can be described using the length and width of the different types of obstacles, respectively.

[0066] In this embodiment, the second sensor data is clustered and fitted to obtain the aforementioned fitting result. Optionally, after acquiring multiple sensor data, this embodiment determines the millimeter-wave radar data from the multiple raw data, and then clusters and fits the millimeter-wave radar data to obtain the length and width of different types of objects, for example, the length and width of different types of obstacles.

[0067] In this embodiment, the predicted state information can be used to represent the position and size of different types of objects. The position can be described using the object's center and heading angle. For example, if the different types of objects are different types of obstacles, the position of each type of object can be described using the center and heading angle of the respective obstacle, and the size of each type of object can be described using the length and width of the respective obstacle.

[0068] In this embodiment, after determining the predicted state information of different types of objects based on third sensor data, environmental description information can be generated based on image analysis results, fitting results, and predicted state information. Optionally, after acquiring multiple sensor data, this embodiment determines the third sensor data from multiple raw data sets. Using the third sensor data, state prediction is performed on different types of objects to obtain predicted state information for different types of objects. Subsequently, the image analysis results obtained from image analysis, the fitting results obtained from fitting, and the determined predicted state information are fused to obtain environmental description information. This achieves the goal of describing the environment using a preset information format, thereby improving the accuracy of environmental description information.

[0069] The following section further explains the steps of performing image analysis on the first sensing data to obtain image analysis results in this embodiment.

[0070] As an optional embodiment, image analysis is performed on the first sensing data to obtain image analysis results, including: inputting the first sensing data into an image analyzer; extracting first features of different types of objects from the first sensing data in the image analyzer; performing semantic segmentation on the first sensing data in the image analyzer to obtain semantic segmentation results, wherein the semantic segmentation results are used to represent the semantic labels to which different types of objects belong; performing visual description on the first sensing data in the image analyzer to obtain visual description results, wherein the visual description results are used to describe the visual states of different types of objects; and determining the first features, semantic segmentation results, and visual description results as the image analysis results.

[0071] In this embodiment, the image analyzer can be an image analyzer based on DeepSeek-Version3 (DeepSeek-V3).

[0072] In this embodiment, the first feature can be represented by a two-dimensional (2D) feature vector.

[0073] In this embodiment, during the vehicle's current driving process, sensor data from multiple sensing devices are acquired. After obtaining multiple sensor data, the first sensor data is input into an image analyzer. Then, in the image analyzer, first features of different types of objects are extracted from the first sensor data. Optionally, in this embodiment, image data is input into the image analyzer, where visual feature extraction is performed on the image data to obtain the first features of different types of objects. For example, two-dimensional feature vectors of different types of objects can be obtained.

[0074] In this embodiment, the semantic segmentation results described above can be used to represent the semantic tags to which different types of objects belong. Optionally, the semantic tags to which different types of objects belong can be multiple semantic tags. For example, multiple semantic tags can include: the semantic tag of "construction guardrail", the semantic tag of "zebra crossing" and the semantic tag of "no left turn sign", etc. This is only an example and is not specifically limited.

[0075] In this embodiment, the first sensing data is semantically segmented in the image analyzer to obtain the semantic segmentation result. Then, the first sensing data is visually described in the image analyzer to obtain the visual description result. Optionally, in this embodiment, the image data is segmented at the pixel level in the DeepSeek-V3-based image analyzer to obtain multiple semantic labels; and in the DeepSeek-V3-based image analyzer, multiple sub-image data from the image data are stitched together to form panoramic image data, and then a 360° visual description is performed on the panoramic image data to obtain a natural language description. The aforementioned multiple sub-image data can be 8-channel sub-image data; this value is only illustrative and not specifically limited.

[0076] In this embodiment, the visual description results described above can be used to describe the state of different types of objects under visual perception. Optionally, the visual description results can be natural language descriptions, and these natural language descriptions are associated with the coordinate system in which the vehicle is located. For example, the coordinate system in which the vehicle is located can also be referred to herein as the vehicle coordinate system or the target coordinate system.

[0077] In this embodiment, after obtaining the semantic segmentation result and the visual description result, the first feature, the semantic segmentation result, and the visual description result are determined as the image analysis result. Optionally, this embodiment uses the extracted two-dimensional feature vector, multi-class semantic labels, and natural language description as the specific processing result of the image data, thereby achieving the goal of describing the environment through a preset information format, and thus realizing the technical effect of improving the accuracy of environmental description information.

[0078] The following section further explains the steps of clustering and fitting the second sensing data to obtain the fitting result in this embodiment.

[0079] As an optional embodiment, clustering and fitting the second sensing data to obtain a fitting result includes: clustering the second sensing data to obtain cluster data of different types of objects, wherein the cluster data is used to represent the cluster category to which the object belongs; and fitting the cluster data to obtain a fitting result.

[0080] In this embodiment, the clustering data described above can be used to represent the cluster category to which an object belongs. For example, the clustering data described above can be cluster data.

[0081] In this embodiment, after clustering the second sensing data to obtain cluster data for different types of objects, the cluster data is fitted to obtain a fitting result. Optionally, this embodiment performs Euclidean clustering on the millimeter-wave radar data according to spatial distance to obtain cluster data. Then, geometric fitting is performed on the cluster data to obtain a fitting result, thereby achieving the purpose of determining the size of different types of objects, and thus realizing the technical effect of improving the fitting accuracy of cluster data.

[0082] For example, with a clustering threshold of 0.5m, millimeter-wave radar data is grouped according to spatial distance, merging discrete detection points of the same obstacle, and outputting the cluster center coordinates and size range of the obstacle. By performing minimum bounding rectangle fitting on the above cluster data, the length and width of the obstacle can be estimated (the error in the length and width of the obstacle is ≤10%). Thus, from the vehicle's environment, different types of objects such as "vehicles," "pedestrians," and "static obstacles" can be distinguished.

[0083] The following description further explains the steps of determining the predicted state information of different types of objects based on third sensor data in this embodiment.

[0084] As an optional implementation method, based on third sensing data, the predicted state information of different types of objects is determined, including: extracting second features of different types of objects from the third sensing data, and generating feature maps corresponding to the second features; using a state prediction model to predict the feature maps to obtain predicted state information, wherein the state prediction model is constructed based on the center point algorithm.

[0085] In this embodiment, the aforementioned third sensing data can be lidar point cloud data, which can also be simply referred to as lidar data.

[0086] In this embodiment, the second feature can be a pillar feature.

[0087] In this embodiment, the feature map can be a feature map displayed in a bird's-eye view (BEV), for example, the feature map has dimensions of 64×512×512.

[0088] In this embodiment, the state prediction model described above can be constructed based on the CenterPoint algorithm.

[0089] In this embodiment, during the vehicle's current driving process, sensor data from multiple sensing devices are acquired. After obtaining the multiple sensor data, second features of different types of objects are extracted from the third sensor data, and feature maps corresponding to the second features are generated. Then, a state prediction model is used to predict the generated feature maps to obtain predicted state information. Optionally, after acquiring the original data, this embodiment can determine the LiDAR point cloud data from the original data, extract columnar features from the LiDAR point cloud data, obtain columnar features of different types of objects, and generate feature maps corresponding to the columnar features. Then, the generated feature maps are input into the state prediction model. In the state prediction model, state prediction is performed on different types of objects to obtain predicted state information for different types of objects. This achieves the goal of determining the position and size of different types of objects, thereby improving the technical effect of improving the accuracy of predicted state information.

[0090] Optionally, columnar feature extraction can be performed on the LiDAR point cloud data to obtain columnar features for different types of objects. For example, by dividing the LiDAR point cloud data into columns and inputting the divided LiDAR point cloud data into a 3D convolutional network for columnar feature extraction, columnar features for different types of objects can be obtained.

[0091] The steps for generating environmental description information based on image analysis results, fitting results, and predicted state information in this embodiment will be further explained below.

[0092] As an optional embodiment, environmental description information is generated based on image analysis results, fitting results, and predicted state information, including: performing time correction on the image analysis results and fitting results according to the timestamp of the third sensor data; converting the corrected image analysis results, corrected fitting results, and predicted state information to the coordinate system of the vehicle to obtain the converted image analysis results, converted fitting results, and converted predicted state information; and integrating the converted image analysis results, converted fitting results, and converted predicted state information into environmental description information.

[0093] In this embodiment, the timestamp of the third sensing data can be the timestamp of the lidar point cloud data.

[0094] In this embodiment, the transformed image analysis results, the transformed fitting results, and the transformed predicted state information are all located in the vehicle's coordinate system. The vehicle's coordinate system can also be referred to as the vehicle coordinate system or the target coordinate system in this document.

[0095] In this embodiment, after obtaining the image analysis results, fitting results, and predicted state information, time correction is performed on the image analysis results and fitting results according to the timestamp of the third sensor data. Optionally, this embodiment performs time correction on the image analysis results and fitting results using interpolation based on the timestamp of the lidar point cloud data. For example, the data delay between the camera (±1 frame) and the millimeter-wave radar (±10ms) is corrected using interpolation based on the timestamp of the lidar point cloud data, thereby ensuring that the time deviation is ≤20ms.

[0096] In this embodiment, the corrected image analysis results, corrected fitting results, and predicted state information are respectively transformed into the coordinate system of the vehicle to obtain the transformed image analysis results, transformed fitting results, and transformed predicted state information. These transformed image analysis results, transformed fitting results, and transformed predicted state information are then integrated into environmental description information. Optionally, after obtaining the corrected image analysis results, corrected fitting results, and predicted state information, this embodiment transforms them into the coordinate system of the vehicle by calibration parameters to obtain the transformed image analysis results, transformed fitting results, and transformed predicted state information. Then, the transformed image analysis results, transformed fitting results, and transformed predicted state information are integrated to obtain an integrated result. This integrated result is then used to generate an environmental description file, which can be used to store environmental description information. This achieves the goal of describing the environment using a preset information format, thereby improving the accuracy of environmental description.

[0097] The steps for generating a predicted driving trajectory of a vehicle based on environmental description information and the first sensor data from multiple sensor data in this embodiment will be further explained below.

[0098] As an optional embodiment, the method further includes: acquiring input data from outside the vehicle, wherein the input data includes: environmental meteorological data and scene data, and vehicle driving data; step S116, generating a predicted driving trajectory of the vehicle based on environmental description information and first sensor data from multiple sensor data, including: generating driving trajectory points of the vehicle within a future driving time based on the input data, environmental description information and first sensor data; fitting the driving trajectory points to obtain a fitted driving trajectory; and generating a predicted driving trajectory based on the fitted driving trajectory.

[0099] In this embodiment, the input data can be used to represent natural language input into the vehicle. The input data may include: environmental meteorological data and scene data, as well as vehicle driving data, etc. Specifically, the meteorological data can be used to represent the weather conditions, the scene data can be used to represent the driving scenario in that environment, and the driving data can be used to represent driving tasks from the navigation system or user voice commands.

[0100] In this embodiment, input data from outside the vehicle is acquired. Optionally, this embodiment acquires meteorological data from an onboard meteorological sensor, scene data from a road speed limit and vehicle dynamics model, and driving data from a navigation system, and then uses the aforementioned meteorological data, scene data, and driving data as input data.

[0101] In this embodiment, the aforementioned driving trajectory point may include multiple sub-driving trajectory points. For example, the multiple sub-driving trajectory points may be, but are not limited to, 10 sub-driving trajectory points. The number here is only for illustrative purposes and is not specifically limited.

[0102] In this embodiment, after generating environmental description information of the vehicle based on multiple sensor data and acquiring input data from outside the vehicle, the vehicle's driving trajectory points for the future driving time are generated based on the input data, environmental description information, and first sensor data. The generated driving trajectory points are then fitted to obtain a fitted driving trajectory, and a predicted driving trajectory is generated based on the fitted driving trajectory. Optionally, this embodiment inputs the aforementioned input data, environmental description information, and first sensor data to a Visual-Language-Action (VLA) inference unit to generate trajectory points, obtaining the vehicle's driving trajectory points for the future driving time. Then, a fifth-order polynomial is used to curve-fit the generated driving trajectory points to obtain a fitted driving trajectory. The fitted driving trajectory is then adjusted to obtain a predicted driving trajectory, or the fitted driving trajectory is determined as the predicted driving trajectory. This achieves the goal of determining the vehicle's driving trajectory during future driving, thereby improving the technical effect of predicting the driving trajectory.

[0103] The steps for generating a predicted driving trajectory based on the fitted driving trajectory described in this embodiment will be further explained below.

[0104] As an optional embodiment, generating a predicted driving trajectory based on the fitted driving trajectory includes: generating a buffer region along the side of the fitted trajectory; adjusting the fitted driving trajectory in response to different types of objects in the environment being in the buffer region; and determining the adjusted fitted driving trajectory as the predicted driving trajectory; and determining the fitted driving trajectory as the predicted driving trajectory in response to different types of objects in the environment not being in the buffer region.

[0105] In this embodiment, the side of the fitted trajectory may include a first side of the fitted trajectory and a second side of the fitted trajectory, and the direction of the first side is opposite to the direction of the second side.

[0106] In this embodiment, the aforementioned buffer area can be a safety buffer zone with a preset width. For example, the preset width can be, but is not limited to, 0.5m.

[0107] In this embodiment, a buffer region is generated along the side of the fitted trajectory. Optionally, this embodiment generates a safety buffer of a preset width along the first and second sides of the fitted trajectory. For example, a safety buffer of less than 0.45m can be generated along the left and right sides of the fitted trajectory.

[0108] In this embodiment, after generating a buffer region along the side of the fitted trajectory, the fitted driving trajectory is adjusted in response to different types of objects in the environment being located within the buffer region, and the adjusted fitted driving trajectory is determined as the predicted driving trajectory. Optionally, this embodiment identifies whether different types of objects in the environment are located within the buffer region, and an identification result can be obtained. If the identification result indicates that the different types of objects are located within the buffer region, the local path in the fitted driving trajectory is adjusted, and the adjusted fitted driving trajectory is determined as the predicted driving trajectory. This achieves the goal of determining the vehicle's driving trajectory in the future, thereby improving the accuracy of the predicted driving trajectory.

[0109] For example, if the above identification results indicate that the obstacle is in the safety buffer zone, then the local path in the continuous trajectory is adjusted, and the adjusted continuous trajectory is determined as the final trajectory that the vehicle needs to refer to.

[0110] In this embodiment, after generating a buffer region along the side of the fitted trajectory, the fitted driving trajectory is determined as the predicted driving trajectory in response to different types of objects in the environment not being located within the buffer region. Optionally, this embodiment identifies whether different types of objects in the environment are within the buffer region, and an identification result can be obtained. If the identification result indicates that the aforementioned different types of objects are not within the buffer region, then there is no need to adjust the fitted driving trajectory; the fitted driving trajectory is directly determined as the predicted driving trajectory. This achieves the goal of determining the vehicle's driving trajectory in the future, thereby improving the technical effect of predicting the driving trajectory.

[0111] For example, if the above identification results indicate that the obstacle is not in the safety buffer zone, then there is no need to adjust the local path in the continuous trajectory, and the generated continuous trajectory is directly determined as the final trajectory that the vehicle needs to refer to.

[0112] The steps for controlling vehicle movement according to a predicted trajectory in the above-described environment of this embodiment will be further explained below.

[0113] As an optional embodiment, step S118, in the environment, controls the vehicle to drive according to the predicted driving trajectory, including: determining initial predicted driving data corresponding to the predicted driving trajectory, wherein the initial predicted driving data is used to represent the initial position and initial speed of the vehicle in the future driving process; and controlling the vehicle to drive according to the initial predicted driving data in the environment.

[0114] In this embodiment, the aforementioned initial predicted driving data can be used to represent the initial position and initial speed of the vehicle during future driving. The initial position can be described using the vehicle's three-dimensional or two-dimensional coordinates, and the initial speed can be a predicted speed to be adjusted.

[0115] In this embodiment, after generating a predicted driving trajectory for the vehicle based on environmental description information and first sensor data from multiple sensor data sources, initial predicted driving data corresponding to the predicted driving trajectory is determined. Then, in the environment where the vehicle is located, the vehicle is controlled to drive according to the initial predicted driving data. Optionally, this embodiment uses the generated predicted driving trajectory to predict the vehicle's driving data, thereby obtaining initial predicted driving data corresponding to the predicted driving trajectory. Then, in the aforementioned environment where the vehicle is located, the vehicle is controlled to drive according to the initial predicted driving data. This achieves the goal of avoiding data error propagation in sudden dynamic scenarios, thus improving the technical effect of vehicle control accuracy.

[0116] For example, by using the generated final trajectory to predict the vehicle's driving data, the initial position and initial speed corresponding to the final trajectory can be obtained. Then, in the aforementioned environment where the vehicle is located, the vehicle's driving is controlled according to the aforementioned initial position and initial speed.

[0117] The steps for controlling vehicle driving according to initial predicted driving data in the above-described environment of this embodiment will be further explained below.

[0118] As an optional embodiment, the initial predicted driving data includes: initial predicted position data and initial predicted speed data. The initial predicted position data represents the initial position of the vehicle during future driving, and the initial predicted speed data represents the initial speed of the vehicle during future driving. In the given environment, controlling the vehicle's driving according to the initial predicted driving data includes: determining target predicted position data corresponding to the predicted driving trajectory, wherein the target predicted position data represents the target position of the vehicle during future driving, and the distance between the target position and the initial position is less than or equal to a preset distance; adjusting the initial predicted speed data to obtain target predicted speed data, wherein the target predicted speed data represents the target speed of the vehicle during future driving, and the speed difference between the target speed and the initial speed is less than or equal to a preset speed difference; and in the given environment, controlling the vehicle's driving according to the target predicted position data and the target predicted speed data.

[0119] In this embodiment, the initial predicted driving data may include: initial predicted position data and initial predicted speed data. The initial predicted position data can be used to represent the initial position of the vehicle during future driving, and the initial predicted speed data can be used to represent the initial speed of the vehicle during future driving. The initial position can be described using the vehicle's three-dimensional or two-dimensional coordinates, and the initial speed can be a predicted speed to be adjusted.

[0120] In this embodiment, the aforementioned predicted target location data can be used to represent the target location of the vehicle during future travel, and the distance between the target location and the initial location is less than or equal to a preset distance. For example, the preset distance can be, but is not limited to, 0.3m or 0.31m.

[0121] In this embodiment, after generating a predicted driving trajectory for the vehicle based on environmental description information and first sensor data from multiple sensor data sources, the predicted target position data corresponding to the predicted driving trajectory is determined. Optionally, this embodiment utilizes a pure tracking algorithm to calculate the vehicle's steering angle. By using the steering angle to predict the vehicle's position during future driving, the predicted target position data corresponding to the predicted driving trajectory can be obtained, thereby achieving the goal of determining the target position of the vehicle during future driving.

[0122] In this embodiment, the aforementioned target predicted speed data can be used to represent the target speed that the vehicle will possess during future driving, and the speed difference between the target speed and the initial speed is less than or equal to a preset speed difference. For example, the preset speed difference can be, but is not limited to, 2 km / h or 2.5 km / h.

[0123] In this embodiment, after generating the predicted driving trajectory of the vehicle based on environmental description information and first sensor data from multiple sensor data sources, the initial predicted speed data is adjusted to obtain target predicted speed data. Then, in the vehicle's environment, the vehicle is controlled according to the target predicted position data and the target predicted speed data. Optionally, this embodiment uses a proportional-integral-derivative (PID) controller to adjust the initial predicted speed data to obtain the target predicted speed data. Then, in the vehicle's environment, the vehicle is controlled according to the determined target predicted position data and the adjusted target predicted speed data. This achieves the goal of avoiding data error propagation in sudden dynamic scenarios, thereby improving the technical effect of vehicle control accuracy.

[0124] For example, by using a proportional-integral-derivative controller, the initial speed can be adjusted according to a preset speed difference to obtain the target speed. Then, in the environment where the vehicle is located, the vehicle can be controlled to move according to the determined target position and the adjusted target speed.

[0125] In this embodiment, when controlling the vehicle, during the vehicle's current driving process, sensor data from multiple sensing devices are acquired to obtain multiple sensor data. Based on the multiple sensor data, environmental description information of the vehicle is generated. Based on the environmental description information and the first sensor data among the multiple sensor data, a predicted driving trajectory of the vehicle is generated. In the environment, the vehicle is controlled to drive according to the predicted driving trajectory. Since this embodiment, when the vehicle is in a driving state, can generate environmental description information describing the vehicle's environment in a preset information format based on the acquired multiple sensor data, and combine the generated environmental description information with the first sensor data among the acquired multiple sensor data, a predicted driving trajectory of the vehicle can be generated. That is, the driving trajectory of the vehicle in the future driving process can be generated. Then, in the environment where the vehicle is located, the vehicle can be controlled to drive according to the generated predicted driving trajectory. This achieves the purpose of avoiding data error propagation in sudden dynamic scenarios, thereby solving the technical problem of low vehicle control accuracy and achieving the technical effect of improving vehicle control accuracy.

[0126] The technical solutions of the embodiments of this application will be illustrated below with reference to preferred embodiments.

[0127] Currently, existing multi-sensor fusion technologies often treat the sensing module as an independent module, without coupling it with the vehicle's control system. This can lead to data error propagation in sudden dynamic scenarios, making it difficult to provide accurate data for subsequent decisions and resulting in low vehicle control precision.

[0128] However, this application proposes a vehicle control method. When the vehicle is in motion, based on multiple sensor data acquired, environmental description information describing the vehicle's environment in a preset information format can be generated. Combining the generated environmental description information with the first sensor data among the acquired multiple sensor data, a predicted driving trajectory of the vehicle can be generated. That is, the driving trajectory of the vehicle in the future driving process can be generated. Then, in the environment where the vehicle is located, the vehicle can be controlled to drive according to the generated predicted driving trajectory. This achieves the purpose of avoiding the propagation of data errors in sudden dynamic scenarios, thereby solving the technical problem of low vehicle control accuracy and achieving the technical effect of improving vehicle control accuracy.

[0129] In this embodiment, by executing a multimodal sensor data fusion method, sensor data collected by multiple sensors can be specifically processed and fused. For example, Figure 2 This is a flowchart of a data fusion method for a multimodal sensor according to an embodiment of this application, such as... Figure 2 As shown, the data fusion method for this multimodal sensor may include the following steps.

[0130] Step S201: Acquire red, green, and blue (RGB) image data using a camera.

[0131] After acquiring RGB image data through the camera, step S202 is executed to input the RGB image data into the image analyzer.

[0132] In the technical solution provided by step S202 of this application, the image analyzer can be an image analyzer based on DeepSeek-V3.

[0133] After the RGB image data is input into the image analyzer, step S203 is performed to extract visual features, perform semantic segmentation, and perform 360° visual description on the RGB image data.

[0134] In the technical solution provided in step S203 of this application, for visual feature extraction, the visual branch of DeepSeek-V3 is used to extract 2D feature vectors of targets such as lane lines, traffic lights, pedestrians, and vehicles in the image. For example, the dimension of the feature vector is 512. For semantic segmentation, the image is segmented at the pixel level through a network transformation architecture, which can output multiple semantic labels. For example, 32 semantic labels can be output (such as "construction guardrail", "zebra crossing", and "no left turn sign"). For 360° visual description, eight images are stitched together to form a panoramic image. DeepSeek-V3 is used to generate a natural language description and associate it with the target coordinates. For example, the natural language description can be presented as follows: There is a red traffic light 30m ahead, there is a construction area on the right side of the intersection, and the guardrail occupies 1 / 3 of the lane. The target coordinates can be coordinates relative to the vehicle coordinate system.

[0135] Step S204: Collect the first point cloud data using millimeter-wave radar.

[0136] After acquiring the first point cloud data using millimeter-wave radar, step S205 is executed to perform Euclidean clustering on the first point cloud data.

[0137] In the technical solution provided by step S205 of this application, the radar target list is grouped by spatial distance, discrete detection points of the same obstacle are merged, and the cluster center coordinates and size range of the obstacle are output. For example, the spatial distance is the clustering threshold, and the value of the clustering threshold can be, but is not limited to, 0.5m.

[0138] After performing Euclidean clustering on the point cloud data, step S206 is executed to perform geometric target fitting on the cluster data.

[0139] In the technical solution provided in step S206 of this application, the cluster data is fitted with the minimum bounding rectangle to estimate the length and width of obstacles (the error in the length and width of obstacles is ≤10%), and to distinguish types such as "vehicles", "pedestrians", and "static obstacles". For example, based on the following speed thresholds: pedestrians ≤5km / h and vehicles ≥0km / h, types such as "vehicles" and "pedestrians" can be distinguished.

[0140] Step S207: Collect second point cloud data using lidar.

[0141] After acquiring the second point cloud data using LiDAR, step S208 is executed to extract features from the second point cloud data.

[0142] In the technical solution provided in step S208 of this application, the point cloud is divided into "pillars". Through a three-dimensional (3D) convolutional network, the pillar features can be extracted from the second point cloud data to generate a feature map displayed in a bird's-eye view (BEV). For example, the dimension of the feature map is 64×512×512.

[0143] After feature extraction from the second point cloud data, step S209 is executed to enter the target detection stage.

[0144] In the technical solution provided in step S209 of this application, the obstacle center, size and heading angle can be predicted based on the BEV feature map and using the CenterPoint algorithm.

[0145] After entering the target detection stage, step S210 is executed to perform bounding box regression on obstacles in the environment.

[0146] In the technical solution provided by step S210 of this application, a 3D bounding box is output, wherein the 3D bounding box can be used to display the length, width, height and / or heading angle of the obstacle.

[0147] After performing geometric target fitting on the cluster data and bounding box regression on obstacles in the environment, step S211 is performed to perform delay correction on velocity and coordinates.

[0148] In the technical solution provided in step S211 of this application, regarding time synchronization, the data delay between the camera (±1 frame) and the millimeter-wave radar (±10ms) is corrected by interpolation based on the lidar timestamp, thereby ensuring a time deviation of ≤20ms. Regarding spatial calibration, based on the vehicle coordinate system (origin at the rear axle center), the data from each sensor is transformed to a unified coordinate system through calibration parameters (e.g., extrinsic parameter matrix), thereby eliminating installation position deviations.

[0149] After performing delay correction on the velocity and coordinates, step S212 is executed to generate a structured text scene description file.

[0150] In the technical solution provided by step S212 of this application, the processing results are integrated to generate a standardized lightweight data exchange (JavaScript Object Notation, or JSON) format file. The JSON format file may include: vehicle status: speed (km / h), acceleration (m / s²), heading angle (°), and coordinates of the Global Positioning System (GPS); obstacle list: identification (id), type, relative distance (m), relative speed (m / s), 3D bounding box, semantic label, and weighted values ​​of confidence scores for each sensor. For example, the weight of the LiDAR confidence score is 0.6, the weight of the camera confidence score is 0.3, and the weight of the radar confidence score is 0.1; environmental description: weather (e.g., sunny / rainy / foggy), lighting (e.g., day / night / tunnel), road type (e.g., urban road / highway), etc.

[0151] In this embodiment, a trajectory can be generated by executing a trajectory generation method based on an enhanced vision-language-action model. For example, Figure 3 This is a flowchart of a trajectory generation method based on an enhanced vision-language-action model according to an embodiment of this application, such as... Figure 3 As shown, the trajectory generation method based on the enhanced vision-language-action model may include the following steps.

[0152] Step S301: Collect red, green and blue image data through the camera.

[0153] When red, green and blue image data are collected through the camera, step S302 is executed, and a range Doppler image can be collected through millimeter-wave radar.

[0154] When red, green and blue image data are collected through the camera, step S303 is executed to collect three-dimensional point cloud data through the LiDAR.

[0155] After acquiring red-green-blue image data, range Doppler images, and 3D point cloud data, step S304 is executed to input the red-green-blue image data, range Doppler images, and 3D point cloud data into the multimodal sensor feature encoder.

[0156] Step S305: Obtain environmental conditions.

[0157] When environmental conditions are obtained, step S306 is executed to determine the driving task.

[0158] When the environmental conditions are obtained, step S307 is executed to obtain the scene constraints.

[0159] After obtaining environmental conditions, determining the driving task, and obtaining scene constraints, step S308 is executed, inputting the environmental conditions, driving task, and scene constraints into the text segmenter.

[0160] After inputting red-green-blue image data, distance Doppler maps, and 3D point cloud data into the multimodal sensor feature encoder, and inputting environmental conditions, driving tasks, and scene constraints into the text segmenter, step S309 is executed, in which the structured text file, panoramic camera image, and natural language input from the outside are input into the Large Language Model (LLM) based on DeepSeek-V3.

[0161] In the technical solution provided by step S309 of this application, the structured text file can be in JSON format, and the panoramic image from the camera can be in JPEG format.

[0162] After inputting the structured text file, panoramic camera image, and natural language input from the outside into the DeepSeek-V3-based LLM model, steps S310 and S311 are executed to enter the scene planning stage and the action planning stage.

[0163] In the technical solution provided in step S310 of this application, the construction of prompt words can be achieved by automatically generating structured prompts. An example of generating structured prompts is as follows: "Known scene: {structured text}, image: {panoramic image features}, task: {driving task}, constraint: {scene constraint}. Please analyze the risks (e.g., the impact of construction area), output: 1. Driving instructions (speed, steering angle); 2. Reasoning basis (e.g., 'Due to construction occupying the road, you need to slow down and avoid it'); 3. Trajectory point in the next 1 second." For example, the above trajectory point can be represented by x-coordinates and y-coordinates, with a time interval of 0.1 seconds for each trajectory point.

[0164] In this embodiment, a high-risk factor can be identified using an LLM model based on DeepSeek-V3. For example, the high-risk factor can be any one or any combination of the following: construction area guardrails encroaching on the lane, pedestrians crossing the road, and oncoming vehicles crossing the line. Furthermore, the risk level and response priority can be output. For example, the risk level can be represented by levels 1-5, with higher numbers indicating higher risk levels.

[0165] In this embodiment, based on risk analysis, specific instructions are output, as shown in the following example: "Current speed is 20km / h, need to decelerate by 5km / h, and maintain the speed at 15km / h; steering angle: +5° (shift 0.5m to the right); reasoning: a construction guardrail 15m ahead occupies 1 / 3 of the left lane, and there are no obstacles in the right lane, so it is necessary to swerve to the right to maintain a safe distance."

[0166] In this embodiment, the output driving instructions can be in JSON format, and these instructions may include: target speed value, acceleration limit, target steering angle value, and steering rate limit. The output trajectory visualization results may include: the absolute coordinates of 10 trajectory points within the next 1 second, where the absolute coordinates are based on GPS positioning. The reasoning basis can be natural language text, which can be used for log recording and human-computer interaction interpretation.

[0167] After entering the scenario planning stage and the action planning stage, step S312 is executed to post-process the generated continuous trajectory.

[0168] In the technical solution provided by step S312 of this application, the continuous trajectory can be generated by performing fifth-order polynomial curve fitting on the trajectory points output by the Visual-Language-Action (VLA) inference unit. The curvature of the continuous trajectory is continuous (to avoid sharp turns), and the continuous trajectory satisfies the constraint of the vehicle's minimum turning radius. For example, if the minimum turning radius is 5m, then the continuous trajectory is less than or equal to 5m.

[0169] In this embodiment, a safety boundary is superimposed on the continuous trajectory. For example, a 0.5m wide safety buffer zone is generated along both sides of the continuous trajectory. If an obstacle intrudes into the buffer zone, secondary planning is triggered to adjust the local path corresponding to the continuous trajectory.

[0170] After post-processing the generated continuous trajectory, step S313 is executed to generate the predicted trajectory of the vehicle.

[0171] In this embodiment, the Model Predictive Control (MPC) algorithm is used to convert the trajectory tracking error into a control variable. The trajectory tracking error may include position deviation and velocity deviation.

[0172] Optionally, a proportional-integral-derivative (PID) controller can be used to ensure that the deviation between the vehicle's actual speed and the target speed is ≤2 km / h. Based on the tracking algorithm, the steering angle can be calculated, which can be used to ensure that the distance deviation between the vehicle's actual trajectory and the predicted trajectory is ≤0.3m.

[0173] Optionally, in response to driving commands from the vehicle, a continuous trajectory is generated, and then the real-world scenario is reproduced through high-fidelity digital twin simulation to verify the safety and efficiency of the continuous trajectory. Finally, signals that directly control the vehicle's movement are generated; for example, these signals may include at least one of the following: accelerator signal, brake signal, and steering wheel control signal.

[0174] Optionally, the control signal is output via a transmission bus and the vehicle's execution results are fed back to the vehicle's control system in real time for closed-loop correction. The response delay of the aforementioned vehicle chassis actuator is ≤50ms, and the execution results may include actual steering angle and speed, etc.

[0175] In this embodiment, when controlling the vehicle, during the vehicle's current driving process, sensor data from multiple sensing devices are acquired to obtain multiple sensor data. Based on the multiple sensor data, environmental description information of the vehicle is generated. Based on the environmental description information and the first sensor data among the multiple sensor data, a predicted driving trajectory of the vehicle is generated. In the environment, the vehicle is controlled to drive according to the predicted driving trajectory. Since this embodiment can generate environmental description information describing the environment in which the vehicle is located using a preset information format based on the acquired multiple sensor data when the vehicle is in a driving state, and can generate a predicted driving trajectory of the vehicle by combining the generated environmental description information and the first sensor data among the acquired multiple sensor data, that is, can generate the driving trajectory of the vehicle in the future driving process, and then, in the environment in which the vehicle is located, the vehicle can be controlled to drive according to the generated predicted driving trajectory, thereby achieving the purpose of avoiding the propagation of data errors in sudden dynamic scenarios, thus solving the technical problem of low vehicle control accuracy, and thus achieving the technical effect of improving vehicle control accuracy.

[0176] According to another aspect of the embodiments of this application, corresponding to the embodiments of the above-described vehicle control method, the embodiments of this application also provide a vehicle control device, wherein the vehicle is equipped with multiple sensing devices. Figure 4 This is a structural block diagram of a vehicle control device according to an embodiment of this application, such as... Figure 4 As shown, the vehicle control device 400 may include: a first acquisition unit 402, a first generation unit 404, a second generation unit 406, and a control unit 408.

[0177] The first acquisition unit 402 is used to acquire sensing data from multiple sensing devices during the current driving process of the vehicle, and obtain multiple sensing data, wherein the sensing data is used to display the environment in which the vehicle is located.

[0178] The first generation unit 404 is used to generate environmental description information of the vehicle based on multiple sensor data, wherein the environmental description information is used to describe the environment through a preset information format.

[0179] The second generation unit 406 is used to generate a predicted driving trajectory of the vehicle based on environmental description information and first sensor data from multiple sensor data. The first sensor data is used to display the environment through images, and the predicted driving trajectory is the driving trajectory of the vehicle in the future driving process. The future driving process is the driving process after the current driving process.

[0180] Control unit 408 is used to control the vehicle's movement according to a predicted driving trajectory in the given environment.

[0181] Optionally, the multiple sensing data include: second sensing data and third sensing data, where the first sensing data is sensing data from a first sensing device among multiple sensing devices, the second sensing data is sensing data from a second sensing device among multiple sensing devices, and the third sensing data is sensing data from a third sensing device among multiple sensing devices. The types of the first sensing device, the second sensing device, and the third sensing device are different device types. The first generation unit 404 may include: an analysis module for performing image analysis on the first sensing data to obtain image analysis results, wherein the image analysis results are used to represent different types of objects in the environment; a clustering and fitting module for performing clustering and fitting on the second sensing data to obtain fitting results, wherein the fitting results are used to represent the size of different types of objects; a first determination module for determining the predicted state information of different types of objects based on the third sensing data, wherein the predicted state information is used to represent the position and size of different types of objects; and a first generation module for generating environmental description information based on the image analysis results, the fitting results, and the predicted state information.

[0182] Optionally, the analysis module may include: an input submodule for inputting first sensing data into an image analyzer; an extraction submodule for extracting first features of different types of objects from the first sensing data in the image analyzer; a semantic segmentation submodule for performing semantic segmentation on the first sensing data in the image analyzer to obtain a semantic segmentation result, wherein the semantic segmentation result is used to represent the semantic labels to which different types of objects belong; a visual description submodule for performing visual description on the first sensing data in the image analyzer to obtain a visual description result, wherein the visual description result is used to describe the visual states of different types of objects; and a first determination submodule for determining the first features, the semantic segmentation result, and the visual description result as the image analysis result.

[0183] Optionally, the clustering and fitting module may include: a clustering submodule for clustering the second sensing data to obtain cluster data for different types of objects, wherein the cluster data is used to represent the cluster category to which the objects belong; and a fitting submodule for fitting the cluster data to obtain fitting results.

[0184] Optionally, the first determining module may include: a first generating submodule, used to extract second features of different types of objects from the third sensing data, and generate feature maps corresponding to the second features; and a prediction submodule, used to predict the feature maps using a state prediction model to obtain predicted state information, wherein the state prediction model is constructed based on the center point algorithm.

[0185] Optionally, the first generation module may include: a correction submodule, used to perform time correction on the image analysis results and fitting results according to the timestamp of the third sensor data; a transformation submodule, used to transform the corrected image analysis results, the corrected fitting results, and the predicted state information to the coordinate system where the vehicle is located, respectively, to obtain the transformed image analysis results, the transformed fitting results, and the transformed predicted state information; and an integration submodule, used to integrate the transformed image analysis results, the transformed fitting results, and the transformed predicted state information into environmental description information.

[0186] Optionally, the vehicle control device 400 may further include: a second acquisition unit for acquiring input data from outside the vehicle, wherein the input data includes: environmental meteorological data and scene data, and vehicle driving data; the second generation unit 406 may include: a second generation module for generating driving trajectory points of the vehicle within a future driving time based on the input data, environmental description information and first sensor data; a fitting module for fitting the driving trajectory points to obtain a fitted driving trajectory; and a third generation module for generating a predicted driving trajectory based on the fitted driving trajectory.

[0187] Optionally, the third generation module may include: a second generation submodule for generating a buffer region along the side of the fitted trajectory; a second determination submodule for adjusting the fitted driving trajectory in response to different types of objects in the environment being in the buffer region, and determining the adjusted fitted driving trajectory as the predicted driving trajectory; and a third determination submodule for determining the fitted driving trajectory as the predicted driving trajectory in response to different types of objects in the environment not being in the buffer region.

[0188] Optionally, the control unit 408 may include: a second determining module, used to determine initial predicted driving data corresponding to the predicted driving trajectory, wherein the initial predicted driving data is used to represent the initial position and initial speed of the vehicle during future driving; and a control module, used to control the vehicle to drive according to the initial predicted driving data in the environment.

[0189] Optionally, the initial predicted driving data includes: initial predicted position data and initial predicted speed data. The initial predicted position data represents the initial position of the vehicle during future driving, and the initial predicted speed data represents the initial speed of the vehicle during future driving. The control module may include: a fourth determining submodule, used to determine target predicted position data corresponding to the predicted driving trajectory, wherein the target predicted position data represents the target position of the vehicle during future driving, and the distance between the target position and the initial position is less than or equal to a preset distance; an adjusting submodule, used to adjust the initial predicted speed data to obtain target predicted speed data, wherein the target predicted speed data represents the target speed of the vehicle during future driving, and the speed difference between the target speed and the initial speed is less than or equal to a preset speed difference; and a control submodule, used to control the vehicle's driving according to the target predicted position data and the target predicted speed data in the given environment.

[0190] In this embodiment, the vehicle control device includes the following units: a first acquisition unit, used to acquire sensing data from multiple sensing devices during the current driving process of the vehicle, obtaining multiple sensing data, wherein the sensing data is used to display the environment in which the vehicle is located; a first generation unit, used to generate environmental description information of the vehicle based on the multiple sensing data, wherein the environmental description information is used to describe the environment through a preset information format; a second generation unit, used to generate a predicted driving trajectory of the vehicle based on the environmental description information and the first sensing data from the multiple sensing data, wherein the first sensing data is used to display the environment through an image, the predicted driving trajectory is the driving trajectory of the vehicle in the future driving process, and the future driving process is the driving process after the current driving process; and a control unit, used to control the vehicle to drive according to the predicted driving trajectory in the environment, thereby achieving the purpose of avoiding the propagation of data errors in sudden dynamic scenarios, thus solving the technical problem of low vehicle control accuracy, and thereby achieving the technical effect of improving the vehicle control accuracy.

[0191] Embodiments of this application also provide an electronic device, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.

[0192] Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0193] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0194] Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.

[0195] Embodiments of this application also provide a computer program that, when executed by a processor, implements the methods described in the various embodiments of this application.

[0196] Embodiments of this application also provide a vehicle that includes the electronic devices described in this application.

[0197] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0198] According to an embodiment of this application, an electronic device is also provided. Figure 5 This is a schematic diagram of an electronic device according to an embodiment of this application, such as... Figure 5 As shown, the electronic device 500 may include a memory 510 and a processor 520, wherein the memory 510 is used to store an executable program; and the processor 520 is used to run the program stored in the memory 510, and the program executes the method of this application when it runs.

[0199] In this application, "multiple" refers to two or more.

[0200] In this application, unless otherwise expressly defined, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0201] The terms “first,” “second,” “third,” “fourth,” etc., in this application (if present) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0202] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0203] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided. The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the device control method for the vehicle in the embodiment.

[0204] Computer-readable storage media, also known as computer storage media, may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. These propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable storage media can transmit, propagate, or transfer programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0205] The program code contained in a computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, radio frequency, or any suitable combination thereof.

[0206] Optionally, when the above-mentioned computer program is executed by the processor, the program code implements the following steps: during the current driving process of the vehicle, acquiring sensor data from multiple sensor devices to obtain multiple sensor data, wherein the sensor data is used to display the environment in which the vehicle is located; based on the multiple sensor data, generating environmental description information of the vehicle, wherein the environmental description information is used to describe the environment through a preset information format; based on the environmental description information and the first sensor data among the multiple sensor data, generating a predicted driving trajectory of the vehicle, wherein the first sensor data is used to display the environment through an image, the predicted driving trajectory is the driving trajectory of the vehicle in the future driving process, and the future driving process is the driving process after the current driving process; and controlling the vehicle to drive according to the predicted driving trajectory in the environment.

[0207] Optionally, when the above computer program is executed by the processor, the program code implements the following steps: performing image analysis on the first sensing data to obtain image analysis results, wherein the image analysis results are used to represent different types of objects in the environment; performing clustering and fitting on the second sensing data to obtain fitting results, wherein the fitting results are used to represent the size of different types of objects; determining the predicted state information of different types of objects based on the third sensing data, wherein the predicted state information is used to represent the position and size of different types of objects; and generating environmental description information based on the image analysis results, fitting results, and predicted state information.

[0208] Optionally, when the above-mentioned computer program is executed by the processor, the program code implements the following steps: inputting first sensing data into an image analyzer; in the image analyzer, extracting first features of different types of objects from the first sensing data; in the image analyzer, performing semantic segmentation on the first sensing data to obtain a semantic segmentation result, wherein the semantic segmentation result is used to represent the semantic labels to which different types of objects belong; in the image analyzer, performing visual description on the first sensing data to obtain a visual description result, wherein the visual description result is used to describe the state of different types of objects under visual conditions; and determining the first features, semantic segmentation result, and visual description result as the image analysis result.

[0209] Optionally, when the above computer program is executed by the processor, the program code implements the following steps: clustering the second sensing data to obtain cluster data of different types of objects, wherein the cluster data is used to represent the cluster category to which the objects belong; fitting the cluster data to obtain the fitting result.

[0210] Optionally, when the above computer program is executed by the processor, the program code implements the following steps: extracting second features of different types of objects from the third sensing data, and generating feature maps corresponding to the second features; using a state prediction model to predict the feature maps and obtain predicted state information, wherein the state prediction model is constructed based on the center point algorithm.

[0211] Optionally, when the above computer program is executed by the processor, the program code implements the following steps: performing time correction on the image analysis results and fitting results according to the timestamp of the third sensor data; converting the corrected image analysis results, corrected fitting results, and predicted state information to the coordinate system where the vehicle is located, respectively, to obtain the converted image analysis results, converted fitting results, and converted predicted state information; and integrating the converted image analysis results, converted fitting results, and converted predicted state information into environmental description information.

[0212] Optionally, when the above computer program is executed by the processor, the program code implements the following steps: acquiring input data from outside the vehicle, wherein the input data includes: environmental meteorological data and scene data, and vehicle driving data; generating driving trajectory points of the vehicle within the future driving time based on the input data, environmental description information and first sensor data; fitting the driving trajectory points to obtain a fitted driving trajectory; and generating a predicted driving trajectory based on the fitted driving trajectory.

[0213] Optionally, when the above computer program is executed by the processor, the program code implements the following steps: generating a buffer region along the side of the fitted trajectory; adjusting the fitted driving trajectory in response to different types of objects in the environment being in the buffer region, and determining the adjusted fitted driving trajectory as the predicted driving trajectory; and determining the fitted driving trajectory as the predicted driving trajectory in response to different types of objects in the environment not being in the buffer region.

[0214] Optionally, when the above computer program is executed by the processor, the program code implements the following steps: determining initial predicted driving data corresponding to the predicted driving trajectory, wherein the initial predicted driving data is used to represent the initial position and initial speed of the vehicle in the future driving process; and controlling the vehicle to drive according to the initial predicted driving data in the environment.

[0215] Optionally, when the above computer program is executed by the processor, the program code implements the following steps: determining the target predicted position data corresponding to the predicted driving trajectory, wherein the target predicted position data is used to represent the target position of the vehicle during future driving, and the distance between the target position and the initial position is less than or equal to a preset distance; adjusting the initial predicted speed data to obtain target predicted speed data, wherein the target predicted speed data is used to represent the target speed that the vehicle will have during future driving, and the speed difference between the target speed and the initial speed is less than or equal to a preset speed difference; and controlling the vehicle to drive according to the target predicted position data and the target predicted speed data in the environment.

[0216] In the embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0217] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0218] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0219] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0220] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for controlling a vehicle, characterized in that, The vehicle is equipped with multiple sensing devices, and the method includes: During the current driving process of the vehicle, sensor data from multiple sensor devices are acquired to obtain multiple sensor data, wherein the sensor data is used to display the environment in which the vehicle is located. Based on multiple sensor data, environmental description information of the vehicle is generated, wherein the environmental description information is used to describe the environment in a preset information format; Based on the environmental description information and the first sensor data among the multiple sensor data, a predicted driving trajectory of the vehicle is generated, wherein the first sensor data is used to display the environment through images, and the predicted driving trajectory is the driving trajectory of the vehicle in the future driving process, and the future driving process is the driving process after the current driving process. In the aforementioned environment, the vehicle is controlled to drive according to the predicted driving trajectory.

2. The method according to claim 1, characterized in that, The plurality of sensor data includes: second sensor data and third sensor data, wherein the first sensor data is sensor data from a first sensor among the plurality of sensor devices, the second sensor data is sensor data from a second sensor among the plurality of sensor devices, and the third sensor data is sensor data from a third sensor among the plurality of sensor devices, wherein the first sensor device, the second sensor device, and the third sensor device are of different device types, wherein, based on the plurality of sensor data, generating environmental description information of the vehicle includes: Image analysis is performed on the first sensing data to obtain image analysis results, wherein the image analysis results are used to represent different types of objects in the environment; Clustering and fitting are performed on the second sensing data to obtain fitting results, wherein the fitting results are used to represent the size of the different types of objects; Based on the third sensor data, predictive state information of the different types of objects is determined, wherein the predictive state information is used to represent the position and size of the different types of objects; Based on the image analysis results, the fitting results, and the predicted state information, environmental description information is generated.

3. The method according to claim 2, characterized in that, Image analysis is performed on the first sensing data to obtain image analysis results, including: The first sensor data is input into the image analyzer; In the image analyzer, the first features of the different types of objects are extracted from the first sensing data; In the image analyzer, the first sensing data is semantically segmented to obtain a semantic segmentation result, wherein the semantic segmentation result is used to represent the semantic labels to which the different types of objects belong; In the image analyzer, the first sensing data is visually described to obtain a visual description result, wherein the visual description result is used to describe the state of the different types of objects under visual conditions. The first feature, the semantic segmentation result, and the visual description result are determined as the image analysis result.

4. The method according to claim 2, characterized in that, Clustering and fitting are performed on the second sensing data to obtain the fitting results, including: Clustering is performed on the second sensing data to obtain cluster data of the different types of objects, wherein the cluster data is used to represent the cluster category to which the object belongs; The clustering data is fitted to obtain the fitting result.

5. The method according to claim 2, characterized in that, Based on the third sensor data, the predicted state information of the different types of objects is determined, including: From the third sensing data, the second features of the different types of objects are extracted, and a feature map corresponding to the second features is generated; The feature map is predicted using a state prediction model to obtain the predicted state information, wherein the state prediction model is constructed based on the center point algorithm.

6. The method according to claim 2, characterized in that, Based on the image analysis results, the fitting results, and the predicted state information, environmental description information is generated, including: The image analysis results and the fitting results are time-corrected according to the timestamp of the third sensor data; The corrected image analysis results, the corrected fitting results, and the predicted state information are respectively transformed into the coordinate system of the vehicle to obtain the transformed image analysis results, the transformed fitting results, and the transformed predicted state information; The converted image analysis results, the converted fitting results, and the converted predicted state information are integrated into the environmental description information.

7. The method according to claim 1, characterized in that, The method further includes: Acquire input data from outside the vehicle, wherein the input data includes: meteorological data and scene data of the environment, and driving data of the vehicle; Based on the environmental description information and the first sensor data from the plurality of sensor data, a predicted driving trajectory of the vehicle is generated, including: generating driving trajectory points of the vehicle within a future driving time based on the input data, the environmental description information, and the first sensor data; fitting the driving trajectory points to obtain a fitted driving trajectory; and generating the predicted driving trajectory based on the fitted driving trajectory.

8. The method according to claim 7, characterized in that, Based on the fitted driving trajectory, the predicted driving trajectory is generated, including: A buffer region is generated along the side of the fitted trajectory; In response to different types of objects in the environment being in the buffer area, the fitted driving trajectory is adjusted, and the adjusted fitted driving trajectory is determined as the predicted driving trajectory; In response to different types of objects in the environment not being in the buffer area, the fitted driving trajectory is determined as the predicted driving trajectory.

9. The method according to claim 1, characterized in that, In the stated environment, controlling the vehicle's movement according to the predicted driving trajectory includes: Determine initial predicted driving data corresponding to the predicted driving trajectory, wherein the initial predicted driving data is used to represent the initial position and initial speed of the vehicle during the future driving process; In the aforementioned environment, the vehicle is controlled to drive according to the initial predicted driving data.

10. The method according to claim 9, characterized in that, The initial predicted driving data includes: initial predicted position data and initial predicted speed data. The initial predicted position data represents the initial position of the vehicle during the future driving process, and the initial predicted speed data represents the initial speed of the vehicle during the future driving process. In the given environment, controlling the vehicle's driving according to the initial predicted driving data includes: Determine the target predicted position data corresponding to the predicted driving trajectory, wherein the target predicted position data is used to represent the target position of the vehicle during the future driving process, and the distance between the target position and the initial position is less than or equal to a preset distance; The initial predicted speed data is adjusted to obtain target predicted speed data, wherein the target predicted speed data is used to represent the target speed that the vehicle will have during the future driving process, and the speed difference between the target speed and the initial speed is less than or equal to a preset speed difference; In the aforementioned environment, the vehicle is controlled to move according to the predicted target location data and the predicted target speed data.

11. A vehicle control device, characterized in that, The vehicle is equipped with multiple sensing devices, the devices including: The first acquisition unit is used to acquire sensing data from multiple sensing devices during the current driving process of the vehicle, thereby obtaining multiple sensing data, wherein the sensing data is used to display the environment in which the vehicle is located. The first generation unit is configured to generate environmental description information of the vehicle based on multiple sensor data, wherein the environmental description information is used to describe the environment through a preset information format; The second generation unit is used to generate a predicted driving trajectory of the vehicle based on the environmental description information and the first sensor data among the multiple sensor data, wherein the first sensor data is used to display the environment through an image, and the predicted driving trajectory is the driving trajectory of the vehicle in the future driving process, and the future driving process is the driving process after the current driving process. A control unit is configured to control the vehicle to drive according to the predicted driving trajectory in the environment.

12. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the method according to any one of claims 1 to 10.