Vehicle control method, control device and storage medium

By preprocessing sensor data, aligning features and preliminary fusion, optimizing communications, and adjusting data weights, the problem of poor system performance caused by multi-sensor data transmission and fusion is solved, efficient and accurate vehicle control is achieved, and transmission delays and accident risks are reduced.

CN120756518APending Publication Date: 2025-10-10YOUDI ROBOT (WUXI) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510855561.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

The existing technology of multi-sensor data transmission and fusion leads to low system performance and data transmission efficiency, high communication bandwidth requirements, increased system complexity, increased costs and reduced real-time performance, which may cause delays in vehicle control information transmission and accidents.

Method used

The sensor data is preprocessed through the preprocessing module, the feature fusion module performs feature alignment and preliminary fusion of multimodal data, the communication optimization module processes the data according to the communication bandwidth requirements, and the decision control module adjusts the sensor data weight according to the scene mode for fusion to generate environmental perception data to control vehicle driving.

Benefits of technology

It improves system performance and data transmission efficiency, ensures efficient and accurate sensor data fusion, reduces the transmission delay of vehicle control information, and reduces the occurrence of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120756518A_ABST
    Figure CN120756518A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of vehicle control, in particular to a vehicle control method, a control device and a storage medium, and the method comprises the steps: carrying out the preprocessing, feature alignment and preliminary fusion of first data obtained by all sensors, obtaining third data, and processing the third data according to a communication bandwidth demand, according to the invention, the system performance and the data transmission efficiency are improved, and it is ensured that the fourth data of each sensor are efficiently and accurately fused to generate the environmental perception data, so that the vehicle driving is controlled, the transmission delay of the vehicle control information is reduced, and the occurrence of traffic accidents is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of vehicle control, in particular to a vehicle control method, a control device and a storage medium. BACKGROUND

[0002] In the automatic driving technology, multi-sensor data fusion is a key link to realize high-precision environment perception. However, in the traditional multi-sensor data fusion scheme, the communication bandwidth demand is high, which leads to the increase of system complexity, the rise of cost and the decline of real-time performance, etc., which seriously reduces the system performance and data transmission efficiency, causes the delay of vehicle control information transmission, and even may lead to more serious accidents. SUMMARY

[0003] Therefore, an embodiment of the present application aims to provide a vehicle control method, a control device and a storage medium to solve the technical problem of low system performance and data transmission efficiency caused by multi-sensor data transmission and fusion in the prior art.

[0004] To solve the above technical problems, an embodiment of the present application provides the following technical solutions: In a first aspect, an embodiment of the present application provides a vehicle control method, comprising: A preprocessing module performs data preprocessing on first data obtained by each sensor to obtain second data; A feature fusion module performs feature alignment and preliminary fusion of multi-modal data on the second data to obtain third data; A communication optimization module processes the third data according to the communication bandwidth demand to obtain fourth data; A decision control module adjusts the data weight of each sensor according to the scene mode, fuses the fourth data of each sensor to generate environment perception data, and controls the vehicle to travel according to the environment perception data.

[0005] In some embodiments, the first data includes initial point cloud data and initial image data, and the preprocessing module performs data preprocessing on the first data obtained by each sensor to obtain second data, including: The preprocessing module removes noise points in the initial point cloud data to obtain candidate point cloud data; Determine the key feature points in the initial image data, and mark the key feature points as feature data in the initial image data to obtain candidate image data; The second data includes the candidate point cloud data and the candidate image data.

[0006] In some embodiments, the feature fusion module performs multimodal data feature alignment and preliminary fusion on the second data to obtain third data, including: The feature fusion module generates image description text based on the candidate image data; The visual features in the candidate image data are aligned and fused with the text features in the image description text, and the visual features in the candidate image data are aligned and fused with the point cloud features in the candidate point cloud data to obtain standard image data and standard point cloud data after feature alignment and fusion, and the third data includes the standard image data and the standard point cloud data.

[0007] In some embodiments, the communication optimization module processes the third data according to the communication bandwidth requirement to obtain fourth data, including: The communication optimization module compresses the standard image data and the standard point cloud data according to data compression requirements, dynamically adjusts the data resolution and update frequency according to the scene mode, and obtains target image data and target point cloud data, wherein the fourth data includes the target image data and the target point cloud data.

[0008] In some embodiments, the method further comprises: The third data is interpreted in real time through the AI ​​perception algorithm embedded in the sensor hardware to determine the scene mode that is adapted to the current environmental conditions.

[0009] In some embodiments, the sensor includes a laser radar and a visual sensor; The decision control module adjusts the data weight of each sensor according to the scene mode, fuses the fourth data of each sensor to generate environmental perception data, and controls the vehicle driving according to the environmental perception data, including: The decision control module determines a target scene mode according to the target image data and the target point cloud data; Adjusting and determining the data weight of the laser radar and the data weight of the visual sensor according to the target scene mode; fusing the target point cloud data of the laser radar and the target image data of the visual sensor according to the data weight of the laser radar and the data weight of the visual sensor to generate environmental perception data; Adjust the speed and driving direction of the vehicle according to the environmental perception data.

[0010] In some embodiments, determining the target scene mode according to the target image data and the target point cloud data includes: Determining, based on the target image data and the target point cloud data, the current environmental conditions of the vehicle, the environmental conditions including the number of obstacles, traffic flow, and weather conditions; Based on the number of obstacles, the traffic flow and the weather conditions, a target scene mode adapted to the environmental conditions is determined, the target scene mode including a first mode or a second mode, wherein the complexity of the environmental conditions corresponding to the first mode is higher than the complexity of the environmental conditions corresponding to the second mode.

[0011] In some embodiments, adjusting and determining the data weight of the laser radar and the data weight of the visual sensor according to the target scene mode includes: If the target scene mode is the first mode, adjusting the data weight of the laser radar to the first weight, and adjusting the data weight of the visual sensor to the second weight, wherein the first weight is greater than the second weight; If the target scene mode is the second mode, the data weight of the lidar is adjusted to the third weight, and the data weight of the visual sensor is adjusted to the fourth weight, and the third weight is smaller than the fourth weight.

[0012] In a second aspect, an embodiment of the present invention provides a control device, including: A preprocessing module, a feature fusion module, a communication optimization module, and a decision control module, wherein the feature fusion module is communicatively connected to the preprocessing module and the communication optimization module respectively, and the communication optimization module is also communicatively connected to the decision control module; The preprocessing module, the feature fusion module, the communication optimization module and the decision control module are used to execute any one of the vehicle control methods proposed in the first aspect.

[0013] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer program instructions executable by a processor are stored. When the computer program instructions are executed by the processor, the computer executes any one of the vehicle control methods proposed in the first aspect.

[0014] The embodiments of the present invention have the following beneficial effects: Different from the existing technology, the vehicle control method provided by the embodiments of the present invention includes: a preprocessing module performs data preprocessing on the first data acquired by each sensor to obtain second data, a feature fusion module performs multimodal data feature alignment and preliminary fusion on the second data to obtain third data, a communication optimization module processes the third data according to the communication bandwidth requirement to obtain fourth data, a decision control module adjusts the data weight of each sensor according to the scene mode, fuses the fourth data of each sensor to generate environmental perception data, and controls vehicle driving according to the environmental perception data.

[0015] The embodiment of the present invention obtains third data by preprocessing, feature aligning and preliminary fusion of the first data acquired by each sensor, and processes the third data according to the communication bandwidth requirement to obtain fourth data that meets the communication bandwidth requirement, thereby improving system performance and data transmission efficiency, ensuring efficient and accurate fusion of the fourth data of each sensor to generate environmental perception data, thereby controlling vehicle driving, reducing the transmission delay of vehicle control information, and reducing the occurrence of traffic accidents. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the prior art or embodiments. Obviously, the drawings described below only illustrate certain embodiments of the present invention and should not be construed as limiting the scope of protection. Those skilled in the art can, without inventive effort, derive other relevant drawings based on these drawings.

[0017] Figure 1 is a schematic diagram of an application scenario of a vehicle control method provided by some embodiments of the present invention; Figure 2 is a schematic structural diagram of a vehicle in some embodiments of the present invention; Figure 3 is a schematic structural diagram of a control device in a vehicle provided by some embodiments of the present invention; Figure 4 is a schematic structural diagram of a control device in a vehicle provided by other embodiments of the present invention; Figure 5 is a flow chart of a vehicle control method provided by some embodiments of the present invention; Figure 6 yes Figure 5 A schematic diagram of a sub-flowchart of step S400 in the vehicle control method shown in the embodiment. DETAILED DESCRIPTION

[0018] In order to make the purposes and advantages of the embodiments of the present invention easier to understand, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. The detailed description of the embodiments of the present invention in the drawings below does not limit the scope of protection claimed by the present invention, but only represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0019] It should be noted that, if no conflict is constituted, the various technical features involved in the embodiments of the present invention described below can be combined with each other and are all within the scope of protection of the present invention. In addition, although the functional modules are divided in the device or structural diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different module division than in the device or in an order different from that in the flow chart. In addition, the "first", "second", "third" and other similar expressions used herein do not limit the data and execution order, but are only for the purpose of convenience of explanation and to distinguish between the same items or similar items with basically the same functions and effects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features.

[0020] Unless otherwise defined, the technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art within the technical field of the present invention. The terms used in this specification are intended solely to describe specific embodiments and are not intended to limit the present invention. It should be understood that the term "and / or" as used in this specification includes any and all combinations of one or more of the listed items.

[0021] See also Figure 1 and Figure 2 , Figure 1 The following schematically illustrates an application scenario of the vehicle control method provided by some embodiments of the present invention. Figure 2 The structural diagram of a vehicle in some embodiments of the present invention is schematically shown.

[0022] like Figure 1 As shown, the application scenario includes vehicles 11, 12, 13, 14, 15, 16, trees 17, and 18. Vehicles 11, 12, 13, 14, 15, and 16 are traveling on a road in a direction N, and trees 17 and 18 are located on the side of the road.

[0023] Specifically, taking vehicle 12 as an example, please refer to Figure 2The vehicle 12 comprises a body 121, and a first sensor 122, a second sensor 123 and a control device 124 arranged on the body 121. The control device 124 is in communication connection with the first sensor 122 and the second sensor 123 respectively, the first sensor 122 is configured to acquire image data of an environment in front of the vehicle 12, and the second sensor 123 is configured to acquire point cloud data of the environment in front of the vehicle 12.

[0024] It is easy to understand that the first sensor 122 can be any suitable type of component such as a camera, an infrared camera, a fisheye camera, etc., and the second sensor 123 can be any suitable type of component such as a laser radar, a structured light scanner, a ToF camera, a stereo vision camera, etc.

[0025] Referring to Figure 3 The control device 124 comprises a preprocessing module 1241, a feature fusion module 1242, a communication optimization module 1243, and a decision control module 1244. The feature fusion module 1242 is in communication connection with the preprocessing module 1241 and the communication optimization module 1243 respectively, and the communication optimization module 1243 is also in communication connection with the decision control module 1244.

[0026] For example, after the first sensor 122 acquires the image data and the second sensor 123 acquires the point cloud data, and the image data and the point cloud data are transmitted to the control device 124, the preprocessing module 1241 pre-processes the image data (i.e., first data) acquired by the first sensor 122 and the point cloud data (i.e., first data) acquired by the second sensor 123, to obtain pre-processed image data and point cloud data, which are second data.

[0027] Then, the feature fusion module 1242 aligns and fuses the features in the pre-processed image data and the point cloud data according to the characteristics of the image data and the point cloud data, to achieve feature alignment and preliminary fusion of multi-modal data, and obtain image data and point cloud data after feature alignment and preliminary fusion, which are third data.

[0028] Next, the communication optimization module 1243 processes the third data according to the communication bandwidth requirement, to obtain fourth data. For example, when the traffic is heavy and the weather condition is poor, it is necessary to improve the data transmission efficiency to achieve efficient transmission of the image data and the point cloud data required for driving control of the vehicle 12. At this time, the communication bandwidth requirement is a high bandwidth requirement level, and the third data is processed by compression, removal of redundant data, adjustment of data resolution and update frequency, etc. according to the communication bandwidth requirement, to obtain target image data and target point cloud data, which are fourth data.

[0029] Finally, decision control module 1244 adjusts the data weights of each sensor based on the scene mode, fuses the fourth data from each sensor to generate environmental perception data, and controls vehicle travel based on the environmental perception data. For example, if vehicle 12 is currently in a scene mode where the environment undergoes a sudden change, the point cloud data acquired by second sensor 123 is prioritized for controlling vehicle travel. Specifically, the data weight of second sensor 123 is adjusted to be greater than the data weight of first sensor 122. Based on the data weights of first sensor 122 and second sensor 123, the target image data from first sensor 122 and target point cloud data from second sensor 123 are fused to generate environmental perception data, and vehicle travel is controlled based on the environmental perception data.

[0030] The embodiment of the present invention obtains third data by preprocessing, feature aligning and preliminary fusion of the first data obtained by each sensor, and processes the third data according to the communication bandwidth requirement to obtain fourth data that meets the communication bandwidth requirement, thereby improving system performance and data transmission efficiency, ensuring efficient and accurate fusion of the fourth data of each sensor to generate environmental perception data, thereby controlling vehicle driving, reducing the transmission delay of vehicle control information, and reducing the occurrence of traffic accidents.

[0031] It should be understood that Figure 1 and Figure 2 The above examples are merely illustrative of the vehicle driving scenarios and vehicle structures in some embodiments of the present invention and do not limit the structure, type, quantity, or other aspects of the vehicles in other embodiments. For example, in other embodiments, the vehicle may further include a display device, or the vehicle may further travel in other scenarios.

[0032] To facilitate understanding of the vehicle control method provided by the embodiment of the present invention, the control device provided by the embodiment of the present invention is first introduced in detail.

[0033] See also Figure 3 , Figure 3 The structural diagram of the control device provided by some embodiments of the present invention is schematically shown.

[0034] like Figure 3 As shown, the control device 124 includes a preprocessing module 1241, a feature fusion module 1242, a communication optimization module 1243, and a decision control module 1244. The feature fusion module 1242 is respectively communicated with the preprocessing module 1241 and the communication optimization module 1243, and the communication optimization module 1243 is also communicated with the decision control module 1244.

[0035] Specifically, the preprocessing module 1241 is used to preprocess the first data acquired by each sensor to obtain second data. The feature fusion module 1242 is used to perform multimodal feature alignment and preliminary fusion on the second data to obtain third data. The communication optimization module 1243 is used to process the third data according to communication bandwidth requirements to obtain fourth data. The decision control module 1244 is used to adjust the data weight of each sensor based on the scene mode, fuse the fourth data of each sensor to generate environmental perception data, and control vehicle driving based on the environmental perception data.

[0036] In some embodiments, the first data includes initial point cloud data and initial image data, and the preprocessing module 1241 is specifically used to: remove noise points in the initial point cloud data to obtain candidate point cloud data, determine key feature points in the initial image data, mark the key feature points as feature data in the initial image data, and obtain candidate image data. The second data includes candidate point cloud data and candidate image data.

[0037] In some embodiments, the feature fusion module 1242 is specifically used to: generate image description text based on candidate image data, align and fuse the visual features in the candidate image data with the text features in the image description text, and align and fuse the visual features in the candidate image data with the point cloud features in the candidate point cloud data to obtain standard image data and standard point cloud data after feature alignment and fusion, and the third data includes standard image data and standard point cloud data.

[0038] In some embodiments, the communication optimization module 1243 is specifically used to: compress standard image data and standard point cloud data according to data compression requirements, dynamically adjust data resolution and update frequency according to scene mode, and obtain target image data and target point cloud data, wherein the fourth data includes target image data and target point cloud data.

[0039] In some embodiments, see Figure 4 The control device 124 also includes an interpretation and matching module 1245, which is respectively communicated with the communication optimization module 1243 and the decision control module 1244. The interpretation and matching module 1245 is used to interpret the third data in real time through the AI ​​perception algorithm embedded in the sensor hardware to determine a scene mode adapted to the current environmental conditions.

[0040] In some embodiments, the sensors include a laser radar and a vision sensor, and the decision control module 1244 is specifically configured to: determine a target scene mode according to the target image data and the target point cloud data, adjust the data weight of the laser radar and the data weight of the vision sensor according to the target scene mode, fuse the target point cloud data of the laser radar and the target image data of the vision sensor according to the data weight of the laser radar and the data weight of the vision sensor to generate environment perception data, and adjust the speed and the driving direction of the vehicle according to the environment perception data.

[0041] In some embodiments, the decision control module 1244 is further specifically configured to: determine an environmental condition in which the vehicle currently locates based on the target image data and the target point cloud data, the environmental condition including the number of obstacles, the traffic flow and the weather condition, determine a target scene mode adapted to the environmental condition according to the number of obstacles, the traffic flow and the weather condition, and the target scene mode including a first mode or a second mode, wherein the complexity of the environmental condition corresponding to the first mode is higher than the complexity of the environmental condition corresponding to the second mode.

[0042] In some embodiments, the decision control module 1244 is further specifically configured to: if the target scene mode is the first mode, adjust the data weight of the laser radar to a first weight and adjust the data weight of the vision sensor to a second weight, the first weight being greater than the second weight, and if the target scene mode is the second mode, adjust the data weight of the laser radar to a third weight and adjust the data weight of the vision sensor to a fourth weight, the third weight being less than the fourth weight.

[0043] According to the above, it can be understood that the implementation execution subject of any vehicle control method provided by the embodiments of the present application can be any suitable type of control device with certain computing and control capabilities, for example, can be implemented by the above-mentioned control device 124.

[0044] The vehicle control method provided by the embodiments of the present application will be described in detail below in combination with the exemplary application and implementation of the control device provided by the embodiments of the present application.

[0045] Please refer to Figure 5 , Figure 5 The flowchart schematically shows the vehicle control method provided by some embodiments of the present application.

[0046] As Figure 5 shown, the vehicle control method includes but is not limited to the following steps S100-S400: S100: The preprocessing module performs data preprocessing on the first data acquired by each sensor to obtain second data.

[0047] In this embodiment of the present invention, various sensors are installed at any appropriate location on the vehicle, such as an image sensor (e.g., a camera, infrared camera, or fisheye camera) for acquiring image data, and a lidar, structured light scanner, time-of-flight camera, or stereo vision camera for acquiring point cloud data. This embodiment uses a vehicle equipped with a camera and a lidar as an example. While the vehicle is in motion, the camera captures and acquires image data of the environment in front of the vehicle, while the lidar acquires point cloud data of the environment in front of the vehicle.

[0048] Specifically, after each sensor obtains corresponding data, each sensor directly transmits the data obtained to the control device, or the control device directly obtains corresponding data from each sensor, or each sensor transmits the data obtained to a storage medium, and the control device obtains the data of each sensor from the storage medium, so that the control device obtains the image data obtained by camera photography and the point cloud data obtained by laser radar sensing. The image data obtained by camera photography and the point cloud data obtained by laser radar sensing are the first data.

[0049] After obtaining the first data, the preprocessing module of the control device preprocesses the image data and the point cloud data respectively, such as performing operations such as denoising, image enhancement and format unification on the image data, and removing noise points and abnormal points in the point cloud data, to obtain preprocessed image data and point cloud data, wherein the preprocessed image data and point cloud data are the second data.

[0050] In some embodiments, the preprocessing module performs data preprocessing on the first data acquired by each sensor to obtain second data, specifically including but not limited to the following steps S110-S120: S110: The pre-processing module removes noise points from the initial point cloud data to obtain candidate point cloud data.

[0051] In this step, the image data captured by the camera is the initial point cloud data, and the point cloud data captured by the lidar sensing is the initial point cloud data, that is, the first data includes the initial point cloud data and the initial image data.

[0052] Specifically, the preprocessing module identifies and removes noise points from the initial point cloud data to obtain candidate point cloud data. For example, a statistical method can be used to calculate the local density of each point in the initial point cloud data, identify points with a density below a preset density threshold as noise points, and remove the identified noise points to obtain candidate point cloud data.

[0053] S120: Determine key feature points in the initial image data, mark the key feature points as feature data in the initial image data, and obtain candidate image data.

[0054] Exemplarily, the initial image data is processed to identify and determine key feature points in the initial image data, such as edges, corners, and object contours, and the identified key feature points are marked as feature data in the initial image data to obtain candidate image data.

[0055] After preprocessing the initial point cloud data and the initial image data, preprocessed candidate point cloud data and candidate image data are obtained, and the candidate point cloud data and candidate image data are the second data, that is, the second data includes the candidate point cloud data and the candidate image data.

[0056] S200: The feature fusion module performs multimodal data feature alignment and preliminary fusion on the second data to obtain third data.

[0057] Specifically, after the preprocessing module performs data preprocessing on the first data to obtain second data, the second data is transmitted to the feature fusion module. The feature fusion module performs multimodal data feature alignment and preliminary fusion on the second data to obtain third data.

[0058] In some embodiments, the feature fusion module performs feature alignment and preliminary fusion of multimodal data on the second data of each sensor to obtain third data in the following specific steps: interpolating, synchronizing, or resampling the data based on the timestamp of each sensor; if the sensor frequencies are different, uniformly aligning them based on the master clock; projecting or transforming the data of all modalities into a unified coordinate system (e.g., a vehicle coordinate system or an image coordinate system) based on the sensor calibration parameters (external parameters); extracting the spatial semantic features of the image data using CNN or Transformer; extracting the spatial structural features of the point cloud data using a point cloud network (e.g., PointNet, VoxelNet, etc.); performing comparative learning or embedding mapping on the spatial semantic features of the image data and the spatial structural features of the point cloud data using projection, attention mechanism, or multimodal encoder, aligning the spatial semantic features of the image data and the spatial structural features of the point cloud data; and finally, preliminarily fusing the spatial semantic features of the image data and the spatial structural features of the point cloud data through splicing, weighted fusion, or adaptive fusion using a multi-head attention mechanism, thereby obtaining the image data and point cloud data after feature alignment and preliminary fusion. The image data and point cloud data after feature alignment and preliminary fusion are the third data.

[0059] In some embodiments, the feature fusion module performs multimodal data feature alignment and preliminary fusion on the second data to obtain third data, specifically including but not limited to the following steps S210-S220: S210: The feature fusion module generates image description text based on the candidate image data.

[0060] Specifically, an image coding network (e.g., CNN, ViT) is used to encode the candidate image data and extract the image's visual feature vector. Visual features include, but are not limited to, object category information, positional relationships, color, action, and scene. The extracted visual features are converted into a sequence format to adapt it to a text generation model. A language generation model (e.g., LSTM, GRU, Transformer Decoder) is then used to decode the visual features and output a natural language sentence describing the image content (e.g., "There is a white car parked on the zebra crossing ahead"). This generates an image description text, which is used to represent the semantic information in the image. By converting image content into image description text, subsequent multimodal feature alignment and semantic fusion are facilitated.

[0061] S220: Aligning and fusing the visual features in the candidate image data with the text features in the image description text, and aligning and fusing the visual features in the candidate image data with the point cloud features in the candidate point cloud data, to obtain standard image data and standard point cloud data after feature alignment and fusion.

[0062] Specifically, visual features such as object categories, bounding boxes, and semantic segmentation information are extracted from the candidate image data; text features such as semantic embeddings and keywords are extracted from the image description text; spatial structural features (i.e., point cloud features) such as point density, depth, and object contours are extracted from the candidate point cloud data; and based on semantic alignment strategies (e.g., multimodal contrastive learning, attention mechanisms, or cross-modal projections), the visual features are aligned and fused with the text features to generate semantically enhanced image features. Furthermore, the visual features are aligned and fused with the point cloud features, for example, by mapping the visual features to a point cloud coordinate system based on the extrinsic parameters of the camera and lidar. The complementary features are enhanced using a multimodal fusion network (e.g., FusionNet, BEVFusion, etc.), resulting in feature-aligned and fused standard image data (including aligned semantic visual features) and standard point cloud data (including enhanced spatial structural features). The feature-aligned and fused standard image data and standard point cloud data constitute the third data, i.e., the third data includes the standard image data and the standard point cloud data.

[0063] In some embodiments, the alignment and fusion of visual features in the candidate image data with text features in the image description text, and the alignment and fusion of visual features in the candidate image data with point cloud features in the candidate point cloud data, may be performed using any one or more of the following methods: Feature alignment: 1) Transformer-based alignment: The Transformer model integrates a self-attention mechanism and an encoder-decoder structure, automatically learning the attention distribution between data of different modalities, thereby achieving implicit alignment. In the image description text generation task, the Transformer model can dynamically generate weight vectors between images and text, achieving alignment and weighted fusion of cross-modal information; 2) Local alignment based on optimal transfer (OT). Optimal transfer theory provides a basic framework for comparing and aligning probability distributions, finding the optimal way to transform one distribution into another while minimizing the transfer cost. In multimodal data alignment, optimal transfer can be used to learn token-level correspondences between data from different modalities, treating feature sequences as discrete distributions, thereby achieving fine-grained cross-modal alignment and fusion of features. 3) Global alignment based on Maximum Mean Difference (MMD). MMD measures the distribution differences of data from different modalities by comparing their statistical differences in the high-dimensional Reproducing Kernel Hilbert Space (RKHS). A system for minimizing the MMD loss is proposed to ensure the consistency of the global distribution of features from different modalities. 4) Contrastive learning alignment: Alignment is achieved by maximizing the similarity between different modal data. The ALBEF model aligns visual and textual features using a contrastive learning loss function before inputting them into the multimodal Transformer. This significantly improves the multimodal Transformer's ability to learn cross-modal relationships. Feature fusion method: 1) Multi-stream methods: equip each modality with a separate encoder and decoder, and use cross-modal Transformers to facilitate information exchange between different modal data. Models such as LXMERT and ViLBERT use co-attention Transformer layers to model the bidirectional relationship between visual and textual features to achieve the fusion of visual and textual features. 2) The single-stream method combines features from different modal data into a unified sequence, which is then processed through a shared Transformer layer. The VisualBERT model embeds and concatenates visual and textual features into tokens and then feeds them into the Transformer for fusion. 3) Semantic alignment based on graph neural networks (GNNs). Graph neural networks (GNNs) build semantic graphs between images and text, learn semantic relationships between nodes (modal data), and achieve semantic alignment between visual features and text features. 4) Pre-training model combined with contrast learning, pre-training model CLIP is trained on a large number of image-text pairs through contrast learning, so that the model learns the corresponding relationship between image and text at the semantic level, thereby realizing efficient implicit semantic alignment, i.e. semantic alignment of visual features and text features; 5) Combination of local and global alignment, the combination of local and global alignment methods more comprehensively captures the relationship of cross-modal data, and the AlignMamba method realizes comprehensive alignment of features from Token level to distribution level by introducing OT-based local alignment module and MMD-based global alignment loss.

[0064] Of course, those skilled in the art can also align and fuse visual features and text features, and align and fuse visual features and point cloud features according to actual needs by using any other suitable method or way, and the embodiments of the present application do not make any specific limitation thereto.

[0065] S300: The communication optimization module processes the third data according to the communication bandwidth requirement to obtain fourth data.

[0066] Specifically, after the feature fusion module aligns and preliminarily fuses the multi-modal data features of the second data to obtain the third data, the third data is transmitted to the communication optimization module, and the communication optimization module compresses, filters or encodes the third data according to the current communication bandwidth requirement to obtain the fourth data. For example, when the vehicle is in an environment with heavy traffic and poor weather conditions, the data transmission efficiency needs to be improved to realize efficient transmission of image data and point cloud data required for vehicle driving control. At this time, the communication bandwidth requirement is a high bandwidth requirement level, and the communication optimization module processes the third data according to the communication bandwidth requirement, such as compression, removal of redundant data, adjustment of data resolution and update frequency, to obtain processed image data and point cloud data, which is the fourth data.

[0067] For example, in some embodiments, the communication optimization module processes the third data according to the communication bandwidth requirement to obtain fourth data, which specifically includes but is not limited to the following steps S310: S310: The communication optimization module compresses the standard image data and the standard point cloud data according to the data compression requirement, dynamically adjusts the data resolution and update frequency according to the scene mode, and obtains the target image data and the target point cloud data, wherein the fourth data includes the target image data and the target point cloud data.

[0068] The communication optimization module processes the image data and the point cloud data before the multi-modal perception data transmission, especially under the limited communication bandwidth, by compressing and adjusting the image data and the point cloud data to ensure the transmission efficiency and integrity of the key data.

[0069] Specifically, the communication optimization module compresses standard image data using standard compression algorithms such as JPEG or H.264 according to data compression requirements. By adjusting the compression parameters, it ensures that while compressing the data volume, sufficient feature information in the standard image data can be retained.

[0070] In some embodiments, the compression parameters include quality factor, quantization table, bit rate control parameters and GOP structure parameters. Process of adjusting compression parameters: Select a suitable bit rate control mode according to the application scenario and transmission conditions of the image data. For real-time image data transmission, CBR or dynamic bit rate can be used, and for image data storage or on-demand, VBR can be used. Adjust the image frame rate and GOP structure according to the complexity of the image data content and the image quality requirements. For images with more dynamic scenes, appropriately increase the frame rate and shorten the GOP interval. Select a suitable encoding profile and level based on the compatibility of the target device and the compression efficiency requirements. Under the premise of ensuring image quality, the quantization parameter (QP) value can be appropriately increased to reduce the file size.

[0071] Specifically, efficient point cloud data compression algorithms, such as octree-based compression methods, are used to compress standard point cloud data. These methods include voxel downsampling, inter-frame redundancy elimination, and sparse coding. By dividing standard point cloud data into multiple octree nodes and transmitting information from key nodes, the amount of transmitted data is reduced. Data compression prioritizes retaining key feature information, such as object edges, moving targets, and dense areas, to ensure accurate environmental perception.

[0072] In this embodiment, the communication optimization module dynamically adjusts the data resolution and update frequency (such as image size scaling, point cloud data precision downsampling, image data and point cloud data upload update frequency or frame rate) according to the current scene mode of the vehicle (such as urban roads, high-speed driving, congestion, complex interactions, etc.) to adapt to the communication bandwidth limitation. Finally, the image data and point cloud data after compression and adjustment of resolution and update frequency are obtained. The image data after compression and adjustment of resolution and update frequency is the target image data, and the point cloud data after compression and adjustment of resolution and update frequency is the target point cloud data. The target image data and target point cloud data are the fourth data, that is, the fourth data includes the target image data and the target point cloud data.

[0073] In some embodiments, the vehicle control method further includes but is not limited to the following steps S301: S301: Through the AI ​​perception algorithm embedded in the sensor hardware, the third data is interpreted in real time to determine the scene mode that is adapted to the current environmental conditions.

[0074] In this embodiment, an AI perception algorithm is embedded or deployed on the sensor hardware (i.e., camera, lidar, etc.). The AI ​​perception algorithm takes the third data output by the feature fusion module (i.e., standard image data and standard point cloud data after feature alignment and fusion) as input, and uses a deep learning model (such as multimodal Transformer, lightweight CNN, etc.) to interpret and identify environmental features in the third data in real time, and determines the scene mode corresponding to the current driving environment of the vehicle based on the environmental features, that is, determines the scene mode that is adapted to the current environmental conditions of the vehicle.

[0075] It should be understood that the scene mode can be a mode that describes comprehensive traffic environments such as daytime, nighttime, rainy day, foggy day, sunny day, highway section, city road, parking lot, congestion, smooth flow, etc.

[0076] S400: The decision control module adjusts the data weight of each sensor according to the scene mode, fuses the fourth data of each sensor to generate environmental perception data, and controls the vehicle driving according to the environmental perception data.

[0077] Specifically, the communication optimization module processes the third data based on communication bandwidth requirements to obtain fourth data, and then transmits the fourth data to the decision control module. The decision control module adjusts the data weights of each sensor based on the scenario mode, fuses the fourth data from each sensor to generate environmental perception data, and controls vehicle driving based on the environmental perception data. The scenario mode is the scenario mode of the vehicle's current environment determined based on the fourth data.

[0078] For example, when the vehicle is currently in a scene mode in which the environment undergoes a sudden change, the point cloud data obtained by the lidar is prioritized for controlling the vehicle's travel. That is, the lidar's data weight is adjusted to be greater than the camera's data weight. Based on the lidar's data weight and the camera's data weight, the camera's image data and the lidar's point cloud data are fused to generate environmental perception data. The vehicle's speed, travel direction, etc. are adjusted in real time based on the environmental perception data, thereby achieving control over the vehicle's travel.

[0079] See also Figure 6 , Figure 6 A sub-flow chart of step S400 in the vehicle control method provided by some embodiments of the present invention is schematically shown.

[0080] like Figure 6 As shown, in some embodiments, the decision control module adjusts the data weight of each sensor according to the scene mode, fuses the fourth data of each sensor to generate environmental perception data, and controls the vehicle driving according to the environmental perception data, specifically including but not limited to the following steps S410-S440: S410: The decision control module determines a target scene mode according to the target image data and the target point cloud data.

[0081] In this embodiment, the sensor includes a laser radar and a visual sensor. The laser radar is used to obtain point cloud data, and the visual sensor is used to obtain image data.

[0082] Exemplarily, the decision control module receives target image data (from a visual sensor) and target point cloud data (from a lidar) transmitted by the communication optimization module, and determines the target scene mode corresponding to the vehicle's current driving environment based on the semantic, structural and environmental characteristics reflected in the target image data and the target point cloud data. The target scene mode can be a complex urban intersection, an open highway section, etc.

[0083] For example, in some embodiments, determining the target scene mode according to the target image data and the target point cloud data specifically includes but is not limited to the following steps S411-S412: S411: Determine the current environmental conditions of the vehicle based on the target image data and the target point cloud data.

[0084] Specifically, the structural and environmental features in the target image data and target point cloud data are extracted, and the number of obstacles in front of and around the vehicle, such as pedestrians, vehicles, and traffic cones, is determined based on the structural and environmental features in the target image data and target point cloud data. The traffic density state (i.e., traffic flow) of the current road is estimated by obtaining the target density, target movement speed, and driving trajectory between consecutive data frames. The presence of rain, fog, snow, nighttime weather, etc. is judged by the illumination distribution, color information, and contrast changes of the image data, and the reflection characteristics of the laser point cloud signal are used to assist in determining the weather conditions, thereby comprehensively determining the current environmental conditions of the vehicle, which include the number of obstacles, traffic flow, and weather conditions.

[0085] S412: Determine a target scene mode adapted to the environmental conditions based on the number of obstacles, traffic flow, and weather conditions.

[0086] In this embodiment, the target scene mode includes a first mode or a second mode, and the complexity of the environmental conditions corresponding to the first mode is higher than the complexity of the environmental conditions corresponding to the second mode.

[0087] Specifically, the decision control module determines a target scenario mode that is appropriate for the vehicle's current environmental conditions, i.e., determines whether the vehicle's target scenario mode is the first or second mode. The first mode corresponds to high-complexity environmental conditions, such as a large number of obstacles, dense traffic, intersection interference, inclement weather, or insufficient lighting. The second mode corresponds to low-complexity environmental conditions, such as smooth roads, sparse obstacles, ample ambient lighting, and no significant interference. The complexity of the vehicle's current driving scenario is assessed by detecting the number of obstacles, traffic flow, and weather conditions in the driving environment.

[0088] In some embodiments, a rule table or a machine learning model (such as a decision tree or a lightweight neural network) may be used to classify and determine the target scene mode in which the vehicle is currently located.

[0089] S420: According to the target scene mode, adjust and determine the data weight of the laser radar and the data weight of the visual sensor.

[0090] Specifically, the decision control module adjusts the weights of the lidar data and the visual sensor data based on the target scene mode determined above, increasing or decreasing the weight of the lidar data or increasing or decreasing the weight of the visual sensor data. The data weights can be dynamically output by a predefined weight table, a scene classifier, or a neural network model. For example, in a night scene mode, the visual sensor data weight is decreased, while the lidar data weight is increased.

[0091] In some embodiments, according to the target scene mode, the data weight of the laser radar and the data weight of the visual sensor are adjusted and determined, specifically including but not limited to the following steps S421-S422: S421: If the target scene mode is the first mode, the data weight of the laser radar is adjusted to the first weight, and the data weight of the visual sensor is adjusted to the second weight.

[0092] S422: If the target scene mode is the second mode, the data weight of the laser radar is adjusted to the third weight, and the data weight of the visual sensor is adjusted to the fourth weight.

[0093] For example, if it is determined based on environmental conditions that the vehicle is currently in the first target scene mode, the lidar data weight is adjusted to the first weight, and the visual sensor data weight is adjusted to the second weight, where the first weight is greater than the second weight, so that the lidar target point cloud data and the visual sensor target image data are fused according to the first and second weights. If it is determined based on environmental conditions that the vehicle is currently in the second target scene mode, the lidar data weight is adjusted to the third weight, and the visual sensor data weight is adjusted to the fourth weight, where the third weight is less than the fourth weight, so that the lidar target point cloud data and the visual sensor target image data are fused according to the third and fourth weights.

[0094] S430: According to the data weight of the laser radar and the data weight of the visual sensor, the target point cloud data of the laser radar and the target image data of the visual sensor are fused to generate environmental perception data.

[0095] Specifically, based on the adjusted and updated data weights of the lidar and the data weights of the visual sensor, a weighted fusion algorithm is executed to fuse the target point cloud data of the lidar and the target image data of the visual sensor to generate high-precision environmental perception data.

[0096] S440: Adjust the speed and driving direction of the vehicle based on the environmental perception data.

[0097] Specifically, based on the road boundaries, obstacles, and dynamic pedestrian / vehicle positions identified in the environmental perception data, path planning algorithms (such as A*, DWA) and vehicle dynamics models are used to adjust and control the vehicle's current speed and direction of travel to ensure vehicle safety.

[0098] To summarize, the vehicle control method provided by the embodiment of the present invention includes: a preprocessing module performs data preprocessing on the first data acquired by each sensor to obtain second data; a feature fusion module performs feature alignment and preliminary fusion of multimodal data on the second data to obtain third data; a communication optimization module processes the third data according to the communication bandwidth requirement to obtain fourth data; a decision control module adjusts the data weight of each sensor according to the scene mode, fuses the fourth data of each sensor to generate environmental perception data, and controls vehicle driving according to the environmental perception data.

[0099] The embodiment of the present invention obtains third data by preprocessing, feature aligning and preliminary fusion of the first data acquired by each sensor, and processes the third data according to the communication bandwidth requirement to obtain fourth data that meets the communication bandwidth requirement, thereby improving system performance and data transmission efficiency, ensuring efficient and accurate fusion of the fourth data of each sensor to generate environmental perception data, thereby controlling vehicle driving, reducing the transmission delay of vehicle control information, and reducing the occurrence of traffic accidents.

[0100] An embodiment of the present invention provides a computer-readable storage medium, which stores computer program instructions executable by a processor. When the computer program instructions are executed by the processor, the computer executes any vehicle control method provided by the embodiment of the present invention, or executes the steps in any possible implementation of any vehicle control method provided by the embodiment of the present invention.

[0101] In some embodiments, the storage medium may be a flash memory, a hard disk, an optical disk, a register, a magnetic surface storage, a removable disk, a CD-ROM, a random access memory (RAM), a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM or any other form of storage medium known in the art, or various devices including one or any combination of the above storage media.

[0102] In some embodiments, computer program instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0103] As an example, computer program instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, for example, in one or more scripts within an HTML (HyperText Markup Language) document, or in a single file dedicated to the program in question, or in multiple coordinated files (for example, files storing one or more modules, subroutines, or code portions).

[0104] By way of example, computer program instructions can be deployed to be executed on one computer (e.g., generally intelligent terminal and server), or on multiple computers of one location that are connected through a communication network, or on multiple computers distributed in multiple locations and connected to each other through a communication network. It is easily understood that all or part of the steps of the method described in the above embodiments of the present application can be directly implemented by using electronic hardware or computer program instructions executable by a processor, or a combination of both.

[0105] It is understood by those skilled in the art that the embodiments provided by the present application are only illustrative, and the writing order of each step in the method of the embodiments does not mean a strict execution order and constitutes any limitation on the implementation process, and the order can be adjusted, combined and deleted according to actual needs. The modules or sub-modules, units or sub-units, etc. in the device or system of the embodiments can be combined, divided and deleted according to actual needs. For example, the division of the unit is only a logical function division, and another division mode can also be used in actual implementation. For example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0106] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus a general hardware platform, and of course, can also be implemented by hardware. Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments.

[0107] It should be noted that the above embodiments are intended to illustrate the technical concepts and characteristics of the present application, and the purpose is to enable those skilled in the art to understand the content of the present application and implement it accordingly, and cannot limit the scope of protection of the present application. Those skilled in the art can understand that all or part of the processes of the above-mentioned embodiments can be modified according to the technical solutions described in the embodiments of the present application, or some technical features can be replaced. It can be understood that these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and should be regarded as equivalent changes and modifications based on the embodiments of the present application, and should be regarded as within the scope of the claims of the present application.

Claims

1. A vehicle control method, characterized in that: include: The preprocessing module performs data preprocessing on the first data acquired by each sensor to obtain second data; The feature fusion module performs multimodal data feature alignment and preliminary fusion on the second data to obtain third data; The communication optimization module processes the third data according to the communication bandwidth requirement to obtain fourth data; The decision control module adjusts the data weight of each sensor according to the scene mode, fuses the fourth data of each sensor to generate environmental perception data, and controls the vehicle driving according to the environmental perception data.

2. The method according to claim 1, characterized in that The first data includes initial point cloud data and initial image data. The preprocessing module performs data preprocessing on the first data acquired by each sensor to obtain second data, including: The preprocessing module removes noise points from the initial point cloud data to obtain candidate point cloud data; Determining key feature points in the initial image data, marking the key feature points as feature data in the initial image data, and obtaining candidate image data; The second data includes the candidate point cloud data and the candidate image data.

3. The method according to claim 2, characterized in that The feature fusion module performs multimodal data feature alignment and preliminary fusion on the second data to obtain third data, including: The feature fusion module generates image description text based on the candidate image data; The visual features in the candidate image data are aligned and fused with the text features in the image description text, and the visual features in the candidate image data are aligned and fused with the point cloud features in the candidate point cloud data to obtain standard image data and standard point cloud data after feature alignment and fusion, and the third data includes the standard image data and the standard point cloud data.

4. The method according to claim 3, characterized in that The communication optimization module processes the third data according to the communication bandwidth requirement to obtain fourth data, including: The communication optimization module compresses the standard image data and the standard point cloud data according to data compression requirements, dynamically adjusts the data resolution and update frequency according to the scene mode, and obtains target image data and target point cloud data, wherein the fourth data includes the target image data and the target point cloud data.

5. The method according to claim 4, characterized in that The method further comprises: The third data is interpreted in real time through the AI ​​perception algorithm embedded in the sensor hardware to determine the scene mode that is adapted to the current environmental conditions.

6. The method according to claim 4, characterized in that The sensors include laser radar and visual sensors; The decision control module adjusts the data weight of each sensor according to the scene mode, fuses the fourth data of each sensor to generate environmental perception data, and controls the vehicle driving according to the environmental perception data, including: The decision control module determines a target scene mode according to the target image data and the target point cloud data; Adjusting and determining the data weight of the laser radar and the data weight of the visual sensor according to the target scene mode; fusing the target point cloud data of the laser radar and the target image data of the visual sensor according to the data weight of the laser radar and the data weight of the visual sensor to generate environmental perception data; Adjust the speed and driving direction of the vehicle according to the environmental perception data.

7. The method according to claim 6, characterized in that The determining of the target scene mode according to the target image data and the target point cloud data includes: Determining, based on the target image data and the target point cloud data, the current environmental conditions of the vehicle, the environmental conditions including the number of obstacles, traffic flow, and weather conditions; Based on the number of obstacles, the traffic flow and the weather conditions, a target scene mode adapted to the environmental conditions is determined, the target scene mode including a first mode or a second mode, wherein the complexity of the environmental conditions corresponding to the first mode is higher than the complexity of the environmental conditions corresponding to the second mode.

8. The method according to claim 7, characterized in that The adjusting and determining the data weight of the laser radar and the data weight of the visual sensor according to the target scene mode includes: If the target scene mode is the first mode, adjusting the data weight of the laser radar to the first weight, and adjusting the data weight of the visual sensor to the second weight, wherein the first weight is greater than the second weight; If the target scene mode is the second mode, the data weight of the lidar is adjusted to the third weight, and the data weight of the visual sensor is adjusted to the fourth weight, and the third weight is smaller than the fourth weight.

9. A control device, characterized in that: include: A preprocessing module, a feature fusion module, a communication optimization module, and a decision control module, wherein the feature fusion module is communicatively connected to the preprocessing module and the communication optimization module respectively, and the communication optimization module is also communicatively connected to the decision control module; The preprocessing module, the feature fusion module, the communication optimization module and the decision control module are used to execute the vehicle control method as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions executable by a processor. When the computer program instructions are executed by the processor, the computer executes the vehicle control method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Unmanned aerial vehicle load control method, system and equipment integrating visible light and laser radar

    CN121541662A