Vehicle driving control method and device, vehicle and computer readable storage medium

By introducing a multi-task joint perception mechanism into the autonomous driving system, and utilizing multiple image acquisition modules and a deep multi-task classification model, synchronous understanding of obstacle information and driving scenario information is achieved. This solves the coordination problem between the perception module and decision control, and improves the perception robustness and safety of the system.

CN122126259APending Publication Date: 2026-06-02SHENZHEN QIYANG SPECIAL EQUIP TECH ENG CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN QIYANG SPECIAL EQUIP TECH ENG CO LTD
Filing Date
2026-02-05
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing autonomous driving systems struggle to balance perception accuracy and scene adaptability under varying weather, lighting, or special road conditions. The perception module and decision control lack an efficient coordination mechanism, which affects driving safety and the level of intelligence.

Method used

Multiple image acquisition modules are used to collect image data of the vehicle's surrounding environment. By using a preset obstacle recognition model and a deep multi-task classification model, obstacle information and driving scene information are identified, driving control commands are generated, and a collaborative mechanism for multi-source perception and behavioral decision-making is established.

Benefits of technology

It significantly improves perception robustness under complex dynamic conditions, enhances the safety and intelligence of autonomous driving systems, and bridges the information gap between perception and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122126259A_ABST
    Figure CN122126259A_ABST
Patent Text Reader

Abstract

This application discloses a vehicle driving control method, device, vehicle, and computer-readable storage medium. The method includes: acquiring image data of the vehicle's surrounding environment through multiple image acquisition modules; inputting the image data into a preset obstacle recognition model for identification to determine obstacle information present in the vehicle's surrounding environment; inputting the image data into a preset deep multi-task classification model for semantic understanding of the current driving environment to determine the vehicle's current driving scenario information; generating driving control commands based on the obstacle information and driving scenario information; and controlling the vehicle to operate according to the driving control commands. Thus, through multi-view image data acquisition, dual-model parallel perception (preset obstacle recognition model and preset deep multi-task classification model), and scene perception-driven control generation, highly robust environmental understanding and efficient perception-decision collaboration are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle control technology, and more specifically, to a vehicle driving control method, device, vehicle, and computer-readable storage medium. Background Technology

[0002] In autonomous driving systems, existing technologies often struggle to balance perception accuracy with scene adaptability, especially under varying weather, lighting, or special road conditions, where system robustness significantly decreases. Furthermore, the lack of an efficient coordination mechanism between perception modules and decision-making control impacts overall driving safety and the level of intelligence. Summary of the Invention

[0003] In view of the above problems, this application proposes a vehicle driving control method, device, vehicle, and computer-readable storage medium that can solve the above problems.

[0004] In a first aspect, embodiments of this application provide a vehicle driving control method. The vehicle includes multiple image acquisition modules disposed at the front, rear, and sides of the vehicle. The method includes: acquiring image data of the vehicle's surrounding environment through the multiple image acquisition modules; inputting the image data into a preset obstacle recognition model for recognition to determine obstacle information existing in the vehicle's surrounding environment; inputting the image data into a preset deep multi-task classification model for semantic understanding of the current driving environment to determine the vehicle's current driving scenario information; generating driving control commands based on the obstacle information and driving scenario information; and controlling the vehicle to operate according to the driving control commands.

[0005] Secondly, embodiments of this application also provide a vehicle driving control device. The vehicle includes multiple image acquisition modules disposed at the front, rear, and sides of the vehicle. The device includes: an acquisition module for acquiring image data of the vehicle's surrounding environment through the multiple image acquisition modules; a first recognition module for inputting the image data into a preset obstacle recognition model for recognition to determine obstacle information existing in the vehicle's surrounding environment; a second recognition module for inputting the image data into a preset deep multi-task classification model for semantic understanding of the current driving environment to determine the vehicle's current driving scenario information; an instruction generation module for generating driving control instructions based on obstacle information and driving scenario information; and a control module for controlling the vehicle to operate according to the driving control instructions.

[0006] Thirdly, embodiments of this application also provide a vehicle, including a processor, a memory, and one or more application programs; the one or more application programs are stored in the memory and configured to be executed by the processor to implement the above-described vehicle driving control method.

[0007] Fourthly, embodiments of this application also provide a computer-readable storage medium storing program code, wherein the above-described vehicle driving control method is executed when the program code is run by a processor.

[0008] The technical solution provided in this application includes a vehicle with multiple image acquisition modules installed at the front, rear, and sides of the vehicle. The method includes: acquiring image data of the vehicle's surrounding environment through the multiple image acquisition modules; inputting the image data into a preset obstacle recognition model for identification to determine obstacle information in the vehicle's surrounding environment; inputting the image data into a preset deep multi-task classification model for semantic understanding of the current driving environment to determine the vehicle's current driving scenario information; generating driving control commands based on the obstacle information and driving scenario information; and controlling the vehicle to operate according to the driving control commands. Thus, this application introduces a preset deep multi-task classification model to simultaneously perform multi-dimensional semantic understanding on the same image data, outputting a semantic understanding of the current driving environment. This multi-task joint perception mechanism enables the vehicle to actively identify the contextual features of the current environment, thereby providing a more comprehensive and adaptive environmental cognitive basis for decision-making, significantly improving perception robustness under complex dynamic conditions. Furthermore, this application explicitly uses obstacle information and driving scenario information together as the basis for generating driving control commands. This establishes an explicit and interpretable collaborative mechanism from multi-source perception to behavioral decision-making, bridging the information gap between perception and control, and improving the safety and intelligence level of the autonomous driving system. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0010] Figure 1 A schematic flowchart of a vehicle driving control method provided in an embodiment of this application is shown.

[0011] Figure 2 A schematic diagram of the structure of a vehicle driving control device provided in an embodiment of this application is shown.

[0012] Figure 3 A schematic diagram of the structure of a vehicle provided in an embodiment of this application is shown.

[0013] Figure 4 This illustration shows a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation

[0014] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0015] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0016] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0017] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0018] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0019] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.

[0020] This application provides a vehicle driving control method, device, vehicle, and computer-readable storage medium. The vehicle includes multiple image acquisition modules disposed at the front, rear, and sides of the vehicle. The method includes: acquiring image data of the vehicle's surrounding environment through the multiple image acquisition modules; inputting the image data into a preset obstacle recognition model for recognition to determine obstacle information existing in the vehicle's surrounding environment; inputting the image data into a preset deep multi-task classification model for semantic understanding of the current driving environment to determine the vehicle's current driving scenario information; generating driving control commands based on the obstacle information and driving scenario information; and controlling the vehicle to operate according to the driving control commands.

[0021] Therefore, this application introduces a pre-defined deep multi-task classification model to simultaneously perform multi-dimensional semantic understanding on the same image data, outputting a semantic understanding of the current driving environment. This multi-task joint perception mechanism enables the vehicle to actively identify the contextual features of the current environment, thereby providing a more comprehensive and adaptive environmental cognitive foundation for decision-making and significantly improving perception robustness under complex dynamic conditions. Furthermore, this application explicitly uses obstacle information and driving scenario information together as the basis for generating driving control commands. This establishes an explicit and interpretable collaborative mechanism from multi-source perception to behavioral decision-making, bridging the information gap between perception and control, and improving the safety and intelligence level of the autonomous driving system.

[0022] This invention provides a vehicle driving control method. The executing entity of this method includes, but is not limited to, at least one of the following devices that can be configured to execute the method provided in this application: a server, a terminal, or a vehicle. In other words, the vehicle driving control method can be executed by software or hardware installed on a terminal device, a server device, or a vehicle. The software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0023] Please see Figure 1 , Figure 1 This illustration shows a flowchart of a vehicle driving control method according to an embodiment of this application. The vehicle includes multiple image acquisition modules disposed at the front, rear, and sides of the vehicle, such as... Figure 3 As shown, the method may include steps 110 to 150.

[0024] In step 110, image data of the environment surrounding the vehicle is acquired through multiple image acquisition modules.

[0025] In some implementations, the image acquisition module may include multiple high-resolution cameras, which are respectively installed in key locations such as the front, rear, left, right and inside the vehicle, as well as near the rearview mirror, to cover multi-angle observation areas such as the road in front of the vehicle, blind spots on the sides, traffic behind, and the driver's field of vision.

[0026] The image acquisition module can include different types such as a front-view camera and a wide-angle camera (e.g., a fisheye lens) to meet the needs of both long-distance target recognition and wide field-of-view environmental perception. The image acquisition module preferably features high dynamic range, low-light enhancement (e.g., near-infrared night vision or starlight-level sensitivity), and electronic / optical image stabilization, enabling it to stably output clear and continuous image frame sequences under various complex lighting and weather conditions, including daytime, nighttime, backlighting, rain, and fog.

[0027] Image data acquired by multiple image acquisition modules is transmitted in real time to the vehicle computing unit via a high-speed vehicle communication bus (e.g., GMSL, FPD-Link III, or vehicle Ethernet).

[0028] However, due to differences in image acquisition module models, installation locations, lighting conditions, and weather conditions, the image data acquired by these modules often suffers from problems such as inconsistent resolution, color distortion, geometric distortion, or low-light blurring. Directly inputting this data into subsequent model steps may lead to unstable feature extraction, affecting the accuracy of decision-making. Therefore, in some implementations, this vehicle driving control method also includes preprocessing the acquired image data. Specifically: using a Gaussian filter to remove noise from the image data; and / or performing white balance and contrast adaptive enhancement on the image data to adapt to different lighting conditions; and / or separating foreground objects in the image data based on a thresholding method; and / or performing distortion correction on the image data to restore the true geometric structure; and / or uniformly scaling all image data to a preset resolution (e.g., 512×512).

[0029] The preprocessed image data is converted to RGB format, and the pixel values ​​are normalized to the [0,1] range to meet the input requirements of the model in subsequent steps.

[0030] In some implementations, to reduce the risk of information leakage of relevant personnel, the image preprocessing process also includes automatic detection and desensitization of privacy-sensitive areas. Specifically, the vehicle can use a pre-trained face detection model (e.g., deep learning-based MTCNN, YOLO-Face, or lightweight CNN) to locate pedestrian facial regions in the input image data in real time; desensitization operations such as blurring, pixelation, or occlusion (e.g., Gaussian blur, mosaic, or black rectangle coverage) are applied to the detected facial regions to eliminate identifiable biometric information; the desensitized image data no longer contains identity information that can be associated with a specific individual, but still retains key semantic content such as road structure, traffic signs, vehicles, and non-sensitive objects, ensuring that the recognition task of subsequent models is not affected.

[0031] In some implementations, the aforementioned privacy protection processing can be completed in the onboard computing unit, and the original image containing the face is not uploaded to the cloud, further ensuring user data security.

[0032] After acquiring and processing image data of the vehicle's surrounding environment, the processed image data is input into a preset obstacle recognition model and a preset deep multi-task classification model, respectively. This allows for the parallel extraction of two types of key environmental information from the image data, providing a comprehensive and robust environmental cognitive foundation for path planning, behavioral decision-making, and vehicle driving control. Specifically: In step 120, the image data is input into a preset obstacle recognition model for recognition to determine the obstacle information existing in the vehicle's surrounding environment.

[0033] In some implementations, the preset obstacle recognition model may be a target detection or instance segmentation model that has been pre-trained and deployed in the vehicle computing unit, used to automatically identify and locate various obstacles that may affect vehicle driving safety from input image data.

[0034] In one specific implementation, the pre-defined obstacle recognition model can be implemented based on a deep learning architecture. For example, Faster R-CNN, the YOLO series (e.g., YOLOv5, YOLOv8), DETR, or the lightweight MobileNet-SSD, etc., are trained end-to-end using a large amount of image data labeled with obstacle locations and categories, and have robust detection capabilities for multi-scale targets in complex road scenes.

[0035] Specifically, the datasets used for training the preset obstacle recognition model include publicly available benchmark datasets (e.g., Cityscapes, KITTI) and custom image data collected by test vehicles in real road environments. These training data cover a variety of weather conditions (e.g., sunny, rainy, snowy, foggy) and lighting scenarios (e.g., daytime, nighttime, backlight) to improve the generalization ability of the preset obstacle recognition model under complex real-world conditions.

[0036] To enhance the robustness of the pre-defined obstacle recognition model and alleviate overfitting, various data augmentation techniques are employed during training, including but not limited to: random horizontal flipping, scaling, brightness adjustment, contrast adjustment, and color jitter.

[0037] The pre-defined obstacle recognition model is optimized using an adaptive gradient descent algorithm (e.g., the Adam optimizer), with an initial learning rate set to a small value (e.g., 0.001), dynamically adjusted through a learning rate decay strategy. Training continues for several epochs (e.g., 100 epochs) until the validation loss converges.

[0038] In some implementations, the loss function of the preset obstacle recognition model takes into account both classification accuracy and localization accuracy. For example, it combines cross-entropy loss (for category prediction) and intersection-over-union (IoU) loss (for bounding box regression) to synergistically optimize the performance of obstacle category discrimination and location estimation.

[0039] The training process of the preset obstacle recognition model can be completed on a high-performance computing platform (e.g., a cloud cluster equipped with GPUs such as Tesla V100). The trained preset obstacle recognition model is then compressed and calibrated before being deployed to the vehicle edge device.

[0040] To address new types of obstacles that may appear in the road environment or not covered in the initial training set (e.g., new shared scooters, temporary traffic facilities, special operation vehicles, etc.), the preset obstacle recognition model supports remote incremental updates via over-the-air (OTA) technology.

[0041] In one specific implementation, the vehicle edge device receives model update packages from the cloud server periodically or on demand. The update packages contain expanded weight parameters for newly added obstacle categories or fine-tuned complete models. The vehicle completes model verification, switching, and activation while in a safe parked state (e.g., with the engine off or parked at low speed) to ensure that the update process does not affect driving safety.

[0042] Over-the-air (OTA) updates enable vehicles to continuously improve their ability to recognize unknown or emerging obstacles without replacing hardware, thereby enhancing the long-term adaptability and safety of autonomous driving systems.

[0043] In some implementations, obstacle information may refer to the structured perception results output by a pre-defined obstacle recognition model. Obstacle information includes, but is not limited to: the bounding box coordinates of the obstacle, the obstacle's category, speed, and distance from the vehicle.

[0044] In this context, the bounding box coordinates of an obstacle represent its spatial location and coverage area in the image data. They are typically represented as a two-dimensional rectangle, including center coordinates, width, and height, or defined by the coordinates of its top-left and bottom-right corners. Obstacle categories include, but are not limited to: pedestrians, bicycles, motor vehicles, traffic cones, guardrails, and animals.

[0045] By cooperating between the image acquisition module and the preset obstacle recognition model, a core intelligent visual perception engine is achieved that is oriented towards real-world autonomous driving scenarios, possesses multi-task perception capabilities, exhibits high robustness, and is engineering-ready. Specifically, in some implementations, the step of "inputting image data into the preset obstacle recognition model for recognition to determine the obstacle information existing in the vehicle's surrounding environment" may include the following steps: (1) Input the image data into the preset obstacle recognition model based on YOLOv5 or Faster R-CNN architecture, and output the category and bounding box coordinates of each obstacle; (2) Based on the bounding box coordinates of each obstacle in consecutive frames, cross-frame matching is performed on the same obstacle, and based on the change of bounding box coordinate information of the matched obstacle in adjacent frames, combined with the time interval between adjacent frames, the velocity information of each obstacle is calculated. (3) Calculate the distance information between each obstacle and the vehicle based on the bounding box coordinates of each obstacle and the real-time depth information of each obstacle.

[0046] Image data is input into a pre-defined obstacle recognition model based on YOLOv5 or Faster R-CNN architecture, which outputs the category label, 2D bounding box coordinates, and corresponding confidence score for each obstacle. The 2D bounding box coordinates are typically (…). , , , )or( , , (), () represents the center coordinates of the two-dimensional bounding box. and The length and width of the two-dimensional bounding box.

[0047] Based on the bounding box coordinates of each obstacle in multiple consecutive frames, a target tracking algorithm (e.g., SORT, DeepSORT, or IoU-based association strategy) is used to perform cross-frame matching of the same obstacle, forming a trajectory sequence; the change in the bounding box center position of the same obstacle between adjacent frames is then considered. With known inter-frame time interval (For example, 1 / 30 of a second), calculate the velocity components of the same obstacle within the image plane. , This allows us to determine the speed information of each obstacle.

[0048] In some implementations, if the velocity information of each obstacle is combined with depth information or camera calibration parameters, the velocity information of each obstacle can be further projected onto the vehicle coordinate system to obtain a relative velocity vector.

[0049] By using the bounding box coordinates of each obstacle and their corresponding real-time depth information, the distance between each obstacle and the vehicle is determined. The real-time depth information can be obtained from: monocular vision depth estimation models (deployed in the vehicle system), binocular stereo vision matching, or fusion of LiDAR or millimeter-wave radar point cloud data.

[0050] Specifically, the depth values ​​of all pixels within the obstacle's bounding box can be taken, and the longitudinal distance from the obstacle to the vehicle can be estimated using statistical methods (e.g., median, minimum, or weighted average). This distance information, together with the aforementioned speed information, constitutes a dynamic state description of the obstacle, providing key inputs for driving control functions such as collision warning, adaptive cruise control, and emergency braking.

[0051] However, in real-world road environments, due to factors such as occlusion, sudden changes in illumination, motion blur, or small targets, obstacle detection results in a single frame image may result in false positives or false negatives. This is especially true when the confidence level of the preset obstacle recognition model is low, significantly reducing its reliability. Therefore, to improve the robustness and safety of the vehicle, in some embodiments, the vehicle driving control method further includes the following steps: (1) Input the image data into the preset obstacle recognition model and output the confidence level corresponding to the category of each obstacle; (2) If the confidence level of any obstacle is lower than the preset confidence threshold, cross-frame matching is performed based on the category and bounding box coordinates of any obstacle in consecutive frames of images; (3) If any obstacle is detected as the same category in at least two frames of images, and the bounding box coordinates of any obstacle in adjacent frames of at least two frames of images meet the motion continuity, then any obstacle is confirmed as a valid obstacle.

[0052] Confidence level is a quantitative value of how confident the preset obstacle recognition model is that the current detection result is a real obstacle. Its value range is usually from 0 to 1 (or 0% to 100%). The higher the value, the more confident the preset obstacle recognition model is that the detection result is valid and the category judgment is accurate.

[0053] When the confidence level of the preset obstacle recognition model output is lower than the preset confidence threshold (e.g., 0.8), the vehicle does not immediately adopt the detection result, but triggers a secondary verification process. The secondary verification mechanism verifies or corrects low-confidence targets by fusing the detection results in the current frame with those in multiple consecutive historical frames and utilizing temporal consistency.

[0054] In at least two consecutive frames, obstacles with a confidence level below a preset confidence threshold are identified as belonging to the same category (e.g., both are "pedestrians" or both are "cones"), and the change in the center position of their bounding boxes between adjacent frames does not exceed a preset value (the preset value is dynamically calculated based on the maximum reasonable relative speed of the vehicle, the camera frame rate, and the road geometry. For example, at a frame rate of 30fps, the lateral displacement of a pedestrian does not exceed 20 pixels per frame). If both of the above conditions are met, the obstacle is confirmed as a valid obstacle, and its final confidence level can be increased or used for subsequent decisions; otherwise, it is considered as transient noise, occlusion artifacts, or false detections and is filtered out.

[0055] The secondary verification process effectively balances the recall and accuracy of the perceived vehicle, avoiding the failure to detect real obstacles due to low confidence in a single frame, and suppressing false alarms caused by environmental interference, thus significantly improving the robustness of the vehicle in complex urban scenarios.

[0056] In step 130, the image data is input into a preset deep multi-task classification model to perform semantic understanding of the current driving environment and determine the current driving scenario information of the vehicle.

[0057] In some implementations, the pre-trained deep multi-task classification model can be a model that has been pre-trained and deployed on an in-vehicle edge device. The pre-trained deep multi-task classification model adopts an architecture that combines a shared backbone network with multi-task sub-networks to simultaneously complete multiple semantic understanding tasks from the input image data.

[0058] Driving scenario information can be structured semantic results output by a pre-defined deep multi-task classification model. Its core content includes the specific traffic scenario category in which the vehicle is currently located, the time of occurrence (e.g., daytime or nighttime), the weather type, and the road attributes.

[0059] The special traffic scenario categories include, but are not limited to: highway toll stations, ramps, traffic control, traffic accidents, road construction, traffic checkpoints, and normal road conditions. The occurrence time can be the lighting condition category at the time of image data acquisition, such as "daytime" or "nighttime." The weather type can be the current weather condition category, such as "sunny," "rainy," "snowy," or "foggy." Road attributes can be the road's structure or function category, such as "urban road," "highway," or "rural road."

[0060] By cooperating with image data and a preset deep multi-task classification model, efficient and robust environmental perception is achieved. Specifically, in some implementations, the step of "inputting image data into a preset deep multi-task classification model for semantic understanding of the current driving environment to determine the vehicle's current driving scenario information" may include the following steps: (1) Input the image data into the shared backbone network for feature extraction and generate a shared feature map; (2) Input the shared feature map into the parallel main task sub-network and a preset number of auxiliary task sub-networks respectively to obtain the special traffic scene, occurrence time, weather type and road attributes of the vehicle.

[0061] A shared backbone network can be a pre-trained deep convolutional neural network (e.g., ResNet or EfficientNet) used to extract general visual features from input image data. As the basis for multi-task learning, the parameters of a shared backbone network are shared across all tasks to achieve feature reuse and reduce computational overhead.

[0062] Shared feature maps can be high-dimensional tensors output from a shared backbone network, containing spatial semantic information of the image (e.g., edges, textures, object parts, etc.). These shared feature maps serve as the common input to subsequent main task sub-networks and a predetermined number of auxiliary task sub-networks, acting as a carrier for multi-task knowledge transfer.

[0063] The main task subnetwork is a classification branch specifically designed to identify special traffic scenarios. The output of the main task subnetwork is the special traffic scenario in which the vehicle is currently located.

[0064] In the embodiments of this application, the preset deep multi-task classification model includes three auxiliary task sub-networks. The main task sub-network and the three auxiliary task sub-networks are structurally independent of each other, each receiving the same shared feature map as input, and completing the final classification through their respective task-specific layers (e.g., global average pooling layers and fully connected layers).

[0065] Through a multi-task network architecture that includes a shared backbone network, a main task sub-network, and a preset number of auxiliary task sub-networks, the vehicle can synchronously output four-dimensional environmental semantic information.

[0066] It is worth noting that the aforementioned preset deep multi-task classification model is not a static rule base, but a deep neural network trained with a large amount of labeled data. To ensure that the model can accurately output multi-dimensional semantic information such as specific traffic scenarios, occurrence time, weather type, and road attributes, in some implementations, the vehicle driving control method further includes the following steps: (1) Obtain the training image dataset; Each image in the training image dataset is labeled with a real label corresponding to the main task and multiple auxiliary tasks. The main task is used to identify the special traffic scene where the vehicle is currently located, and the multiple auxiliary tasks are used to identify the time of occurrence, weather type and road attributes, respectively. (2) Construct a multi-task network architecture with a shared backbone network, a main task sub-network, and a preset number of auxiliary task sub-networks; (3) Input the training image dataset into the multi-task network architecture to obtain the prediction results for each task; (4) Based on the prediction results and the true labels, calculate the cross-entropy loss for each task, and determine the joint loss function according to the cross-entropy loss for each task; (5) Perform end-to-end training of the multi-task network architecture based on the joint loss function; (6) Minimize the joint loss function through the backpropagation algorithm and update the parameters of the multi-task network architecture.

[0067] In the embodiments of this application, the training image dataset used to train the preset deep multi-task classification model consists of multiple scene images. Each scene image corresponds to... Each real label is used for supervision. A different perceptual task.

[0068] Among them, the t-th task ( This is used to identify a certain type of driving environment attribute, and its output category space contains... Possible categories (e.g., "special traffic scenario" tasks might include "construction zone", "toll station", etc.) There are several categories; the "Weather Type" task may include "Sunny," "Rainy," etc. (Categories).

[0069] In one specific implementation, for any training image in the training image dataset Its annotation information includes: main task tag This indicates the specific traffic scene category corresponding to the image; auxiliary task label. Indicates the time of occurrence (e.g., day / night); auxiliary task tags Indicates the weather type (e.g., sunny / rainy / snowy / foggy); auxiliary task tags This indicates the road type (e.g., city road / highway / rural road).

[0070] In other words, each image in the training image dataset is assigned a complete set of multi-dimensional semantic labels to simultaneously guide the learning process of the main task and multiple auxiliary tasks, thereby achieving joint semantic understanding of the driving environment.

[0071] After constructing the multi-task network architecture, forward propagation and loss calculation are performed. The training image dataset is input into the multi-task network architecture to obtain the prediction results for each task. The prediction results can be the class probability distribution output by the multi-task network architecture through the main task sub-network and a predetermined number of auxiliary task sub-networks for each image in the image dataset, after features are extracted by the shared backbone network.

[0072] In one specific implementation, the main task subnetwork outputs a probability vector representing the probability that an image belongs to various special traffic scenarios (e.g., construction zone, toll station, etc.); the first auxiliary task subnetwork in a preset number of auxiliary task subnetworks outputs a probability vector representing the confidence that the occurrence time is "daytime" or "nighttime"; the second auxiliary task subnetwork in a preset number of auxiliary task subnetworks outputs a probability vector representing the confidence of the current weather type (e.g., sunny, rainy, snowy, foggy); and the third auxiliary task subnetwork in a preset number of auxiliary task subnetworks outputs a probability vector representing the confidence of the road attribute (e.g., urban road, highway, rural road).

[0073] In one specific implementation, during the training process of the preset deep multi-task classification model, for each image in the training image dataset, the vehicle first obtains its prediction results for each task—that is, the raw scores output by the main task sub-network and each auxiliary task sub-network, and normalizes them into a class probability distribution using the Softmax function. The ground truth labels corresponding to each image in the training image dataset (e.g., special traffic scene category, occurrence time, weather type, or road attribute) are then converted into one-hot encoded form.

[0074] The cross-entropy loss for this task is obtained by taking the negative logarithm of the probability value corresponding to the true label in the model's prediction. For example, if the true specific traffic scene in an image is a "construction zone," and the model predicts a probability of "construction zone" of 0.85, then the cross-entropy loss for the main task is... This process is executed independently for the main task and all auxiliary tasks, thereby obtaining a set of cross-entropy loss values ​​corresponding to the number of tasks.

[0075] Furthermore, in multi-task learning, the model needs to optimize multiple objectives simultaneously. A common approach is to perform a weighted linear combination of the loss functions of each task to form the total loss function. This process can be expressed as: in," "for the first" The weight coefficient of the first task (which can be a manually set hyperparameter used to adjust the weight of the first task). (The contribution ratio of each task to the total loss), "for the first" A single loss function for each task (which measures the model's performance on the 1st task). (The error between the prediction results and the true labels on each task), "This represents the total number of tasks (i.e., the number of semantic understanding tasks processed simultaneously by the model)." However, the method of weighted linear combination of loss functions for each task is only reasonable if a set of parameters exists that is effective across all tasks. That is, minimizing the weighted sum of empirical risks is only effective when there is no competition between tasks, a situation that is extremely rare. In other words, multi-task learning with conflicting objectives requires modeling the balance between tasks, which is beyond the scope of the above method. Furthermore, pre-defined deep multi-task classification models are insufficient for… "The parameters are extremely sensitive. To obtain the optimal..." "Parameter values ​​require manual searching of the hyperparameter space, which is very costly."

[0076] Based on this, the vehicle driving control method adopts an adaptive weighting mechanism based on homoscedastic uncertainty, designing the loss weights of each task as learnable parameters, so that the model can automatically evaluate the reliability of each task during training and dynamically adjust its contribution ratio in the total loss, thereby avoiding manual setting of hyperparameters and improving the stability and efficiency of multi-task collaborative training.

[0077] Specifically: for the first For each task, a pre-defined deep multi-task classification model outputs the original classification score through the main task sub-network and a pre-defined number of auxiliary task sub-networks, and then inputs it into the Softmax function to obtain the class probability distribution.

[0078] In one specific implementation, the Softmax function can be expressed as: in," "To balance the weights among multiple tasks," " represents all trainable parameters of a convolutional neural network," "Input image" The parameters are After processing by the convolutional neural network, the original category score vector is obtained. " is the first of the output vectors One element is 1, and the rest are 0. " is the total number of task categories (e.g., C=4 in weather type tasks, corresponding to dry, rainy, snowy, and foggy days).

[0079] Based on the one-hot coding method and the corresponding cross-entropy function, the final analytical formula for a single task can be: in," "This is the homoscedasticity uncertainty parameter (used to balance the weights of multiple tasks)," " represents all trainable parameters of a convolutional neural network," "Input image" The parameters are After processing by the convolutional neural network, the original category score vector is obtained. "This represents the final loss value for a single task." "This refers to the total number of categories for a single task (i.e., the number of all possible 'category options' under a given task)." "For the first" under this task Each category (which is an identifier of the category number).

[0080] In the above derivation process, the parameter for introducing task uncertainty is defined. In this case, the model corresponds to the true category The predicted probability is a scaled Softmax function, that is: the original output of the network is scaled... Divide by Then input Softmax. Next, following the principle of maximum likelihood estimation, the loss function is defined as the negative logarithm of the predicted probability (i.e., the negative log-likelihood). Thus, the loss expression naturally contains two terms: one related to the score of the true class, and the other is the logarithm of the exponential sum of the scores of all classes.

[0081] In subsequent processing, in order to... The complex loss is linked to the traditional classification loss by introducing a reasonable approximation that eliminates complex interaction terms through the operational properties of logarithms and exponents. This approximation allows the entire loss to be decomposed into two parts: one part is proportional to the loss without... The standard classification loss and a portion of the uncertainty parameters only Related. The standard classification loss is precisely the cross-entropy loss widely used in deep learning, and its mathematical form is: in," "The true label (i.e., the input image)" The corresponding objective truth category is the standard answer for measuring the accuracy of the model's prediction results. " is the set of all trainable parameters of the model," "Input image" Based on model parameters The original category score vector output after processing.

[0082] In some implementations, simplifying assumptions are made regarding the cross-entropy loss, resulting in the expression for the cross-entropy loss: in," "The homoscedasticity uncertainty parameter for each task," "This refers to the total number of categories for a single task, that is, the number of all possible "category options" under a given task." "The set of all trainable parameters for the entire multi-task network," "Input image" The parameters are The original category score vector is obtained after processing by the convolutional neural network.

[0083] As described above, the resulting negative log-likelihood loss comprises two terms: one related to the score of the true class, and the other being the logarithm of the exponential sum of the scores of all classes. In processing this exponential sum term, an approximation is used so that the entire loss expression can be reorganized into the standard cross-entropy loss. of times, plus one that only depends on regular terms After performing the same operation on all tasks, the losses of each task are summed to determine the joint loss function.

[0084] In one specific implementation, the expression for the joint loss function can be: in," "The set of all trainable parameters for the entire multi-task network," "This represents the homoscedasticity uncertainty parameter for the first task." "for the first" The homoscedasticity uncertainty parameter for each task, "This represents the total loss value for multi-task learning." "This represents the total number of tasks."

[0085] During end-to-end training of a multi-task network architecture based on the joint loss function, the model parameters... With all uncertain parameters The parameters of the shared backbone network, main task subnetwork, and each auxiliary task subnetwork are simultaneously updated and optimized through backpropagation to minimize the joint loss function until the model converges. This allows the system to automatically adjust its loss weights based on the learning difficulty of each task: for tasks with high prediction confidence (low uncertainty), The smaller the loss weight Larger tasks are given higher weights and thus dominate the optimization process; conversely, difficult or noisy tasks are appropriately downweighted to avoid interfering with the overall training.

[0086] This training method, through multi-task knowledge sharing and adaptive loss weighting, enables the model to stably output high-confidence driving scenario information in complex and ever-changing real road environments, significantly improving the robustness and generalization ability of the perception system.

[0087] In step 140, driving control commands are generated based on obstacle information and driving scenario information.

[0088] The vehicle performs a comprehensive risk assessment based on obstacle type, bounding box coordinates, speed, distance to the vehicle, the specific traffic scenario, time of occurrence, weather type, and road attributes, and selects an appropriate driving strategy. This could include decelerating, changing lanes, stopping, or maintaining the current position. The vehicle translates the selected strategy into specific driving control commands, such as acceleration / deceleration commands, steering commands, or stopping commands, and transmits them to the vehicle control system for execution, thereby ensuring safe and efficient vehicle operation in complex environments.

[0089] In some implementations, the vehicle performs a comprehensive risk assessment based on the obstacle's category, bounding box coordinates, speed, distance from the vehicle, the specific traffic scenario the vehicle is currently in, the time of occurrence, weather type, and road attributes, using a pre-defined driving strategy rule base.

[0090] The preset driving strategy rule base can be a structured knowledge set consisting of multiple condition-action rules, used to map perceived obstacle information and driving scenario information into specific driving control commands.

[0091] In one specific implementation, each rule in the preset driving strategy rule base is defined in the form of "if...then...". For example, "If we are currently in a construction zone and there is a pedestrian within 30 meters ahead, then reduce speed to 20 km / h and prepare for emergency braking." Another example is, "If, under nighttime rainy conditions, the distance to the vehicle ahead is less than 40 meters and the relative speed is greater than 5 km / h, then trigger moderate-intensity braking."

[0092] The pre-built driving strategy rule base is manually constructed based on traffic regulations, safe driving standards, and a large amount of real-vehicle test data, and can be configured or updated according to the road environment of different countries or regions. During operation, the vehicle will match the perception results with the conditions in the rule base in real time, select the highest priority rule that meets the conditions, and output the corresponding control action, thereby achieving interpretable, reliable decision-making behavior that conforms to human driving habits.

[0093] In some implementations, when generating driving control commands, the vehicle further integrates with the path planning module to calculate a collision-free feasible trajectory in the local map based on the identified obstacle locations and categories by calling a preset obstacle avoidance algorithm (e.g., A* search algorithm or dynamic window method). This trajectory is used to generate specific navigation commands, such as deceleration, lane changing, or detour, and is output to the vehicle's underlying control system in a structured data format (such as JSON) for the autonomous driving execution module to call.

[0094] Meanwhile, the key information from this perception and decision-making process is recorded as an identification log and stored in an in-vehicle distributed database (such as SQLite) to facilitate subsequent accident backtracking or model iteration and optimization.

[0095] In step 150, the vehicle is controlled to operate according to driving control commands.

[0096] The vehicle sends driving control commands to the vehicle's underlying execution unit via the vehicle communication bus. The electronic control unit then parses the driving control commands and drives the corresponding actuators to complete the operation.

[0097] For example, when the driving control command is to decelerate, the electronic braking system adjusts the braking torque to reduce the vehicle speed. As another example, when the driving control command is to change lanes, the electric power steering system controls the steering wheel angle, coordinating with the adaptive cruise control system to adjust the vehicle speed, achieving smooth lateral and longitudinal coordinated control. For yet another example, when the driving control command is to limit speed or maintain lane position, the relevant control signals are synchronously transmitted to the powertrain and steering controllers, ensuring that the vehicle strictly follows the behavioral requirements output by the decision module.

[0098] The entire process is completed in milliseconds, ensuring a closed-loop response from perception and decision-making to execution, enabling the vehicle to operate safely and accurately according to driving control commands.

[0099] In some implementations, to ensure vehicle reliability and continuous optimization capabilities, the vehicle also integrates an operation monitoring and closed-loop feedback mechanism. Specifically, an onboard monitoring platform is deployed in the vehicle. This platform receives obstacle detection results from the multi-task perception module and presents a heat map of obstacle distribution and the recognition confidence level of each task in a visual format on the human-machine interface. Simultaneously, by combining the onboard Global Positioning System (GPS) and Geographic Information System (GIS), the location information of obstacles is mapped onto an electronic map, allowing the driver to monitor the surrounding environment in real time and manually take over vehicle control when necessary.

[0100] In addition, the vehicle is equipped with a data feedback mechanism: during vehicle operation, it automatically records the perception results and human intervention behaviors in actual driving scenarios, especially key events such as false alarms, missed alarms or abnormal control commands, and stores such samples in the vehicle's distributed storage unit after labeling them; this high-quality data can be uploaded to the training server regularly to incrementally fine-tune the multi-task network model, thereby continuously improving the recognition accuracy under complex working conditions such as special traffic scenarios and weather conditions.

[0101] In terms of vehicle maintenance, the in-vehicle edge computing device has a built-in self-test module that can periodically test the imaging quality of the camera, the status of the sensor, and the operational stability of the deep learning inference framework. When it detects a decline in hardware performance or an outdated software version, it can trigger firmware updates or model replacements through remote upgrades or local maintenance interfaces to ensure the long-term stable operation of the entire perception-decision-execution chain.

[0102] Please see Figure 2 , Figure 2This illustration shows a structural schematic diagram of a vehicle driving control device according to an embodiment of this application. The vehicle includes multiple image acquisition modules disposed at the front, rear, and sides of the vehicle. The vehicle driving control device 200 includes: an acquisition module 210, a first recognition module 220, a second recognition module 230, a command generation module 240, and a control module 250. Specifically: The acquisition module 210 is used to acquire image data of the environment surrounding the vehicle through multiple image acquisition modules; The first recognition module 220 is used to input image data into a preset obstacle recognition model for recognition, and determine the obstacle information existing in the environment around the vehicle. The second recognition module 230 is used to input image data into a preset deep multi-task classification model to perform semantic understanding of the current driving environment and determine the current driving scene information of the vehicle. The instruction generation module 240 is used to generate driving control instructions based on obstacle information and driving scenario information; The control module 250 is used to control the vehicle to operate according to driving control commands.

[0103] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0104] In the several embodiments provided in this application, the coupling or direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication connection through some interface, device or module, and may be electrical, mechanical or other forms.

[0105] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0106] Please see Figure 3 , Figure 3 The diagram illustrates the structure of a vehicle according to an embodiment of this application. The vehicle 300 in this application may include one or more of the following components: a processor 310, a memory 320, and one or more application programs. The one or more application programs may be stored in the memory 320 and configured to be executed by one or more processors 310. The one or more programs are configured to execute the vehicle driving control method as described in the foregoing method embodiments.

[0107] Processor 310 may include one or more processing cores. Processor 310 connects to various parts within the vehicle 300 using various interfaces and lines, and performs various functions and processes data of the vehicle 300 by running or executing instructions, programs, code sets, or instruction sets stored in memory 320, and by calling data stored in memory 320. Optionally, processor 310 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 310 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 310 and may be implemented separately using a communication chip.

[0108] The memory 320 may include random access memory (RAM) or read-only memory (ROM). The memory 320 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 320 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created by the vehicle 300 during use.

[0109] Please see Figure 4 , Figure 4 The diagram shows a computer-readable storage medium 400 provided in an embodiment of this application. The computer-readable storage medium 400 stores program code, which can be called by a processor to execute the vehicle driving control method described in the above method embodiment.

[0110] The computer-readable storage medium 400 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has storage space for program code 410 for performing the steps according to the method embodiments of this application. This program code can be read from or written to one or more computer program devices. The program code 410 for performing the steps according to the method embodiments of this application may be compressed, for example, in a suitable form.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A vehicle driving control method, characterized in that, The vehicle includes multiple image acquisition modules disposed at the front, rear, and sides of the vehicle, and the method includes: Image data of the environment surrounding the vehicle are acquired through the multiple image acquisition modules; The image data is input into a preset obstacle recognition model for recognition, thereby determining the obstacle information present in the environment surrounding the vehicle. The image data is input into a preset deep multi-task classification model to perform semantic understanding of the current driving environment and determine the current driving scenario information of the vehicle. Based on the obstacle information and the driving scenario information, driving control commands are generated; Control the vehicle to operate according to the driving control commands.

2. The vehicle driving control method according to claim 1, characterized in that, The obstacle information includes the obstacle's category, bounding box coordinates, speed, and distance from the vehicle; The step of inputting the image data into a preset obstacle recognition model for recognition, and determining the obstacle information existing in the environment surrounding the vehicle, includes: The image data is input into the preset obstacle recognition model based on YOLOv5 or Faster R-CNN architecture, and the category and bounding box coordinates of each obstacle are output. Based on the bounding box coordinates of each obstacle in consecutive frames of images, cross-frame matching is performed on the same obstacle, and the velocity information of each obstacle is calculated based on the change of bounding box coordinate information of the matched obstacle in adjacent frames, combined with the time interval between adjacent frames. Based on the bounding box coordinates of each obstacle and the real-time depth information of each obstacle, the distance information between each obstacle and the vehicle is calculated.

3. The vehicle driving control method according to claim 2, characterized in that, The method includes: The image data is input into the preset obstacle recognition model, and the confidence level corresponding to the category of each obstacle is output. If the confidence level of any obstacle is lower than the preset confidence threshold, cross-frame matching is performed based on the category and bounding box coordinates of the obstacle in multiple consecutive frames of images. If any obstacle is detected as belonging to the same category in at least two frames of images, and the bounding box coordinates of any obstacle in adjacent frames of at least two frames of images conform to motion continuity, then any obstacle is confirmed as a valid obstacle.

4. The vehicle driving control method according to claim 1, characterized in that, The driving scenario information includes the specific traffic scenario in which the vehicle is currently located, the time of occurrence, the weather type, and the road attributes; The step of inputting the image data into a preset deep multi-task classification model to perform semantic understanding of the current driving environment and determine the current driving scenario information of the vehicle includes: The image data is input into a shared backbone network for feature extraction to generate a shared feature map. The shared feature map is input into the parallel main task sub-network and a preset number of auxiliary task sub-networks to obtain the special traffic scene, occurrence time, weather type and road attributes of the vehicle.

5. The vehicle driving control method according to claim 1, characterized in that, The method further includes: Obtain a training image dataset; each image in the training image dataset is labeled with a real label corresponding to the main task and multiple auxiliary tasks. The main task is used to identify the special traffic scene in which the vehicle is currently located, and the multiple auxiliary tasks are used to identify the time of occurrence, weather type and road attributes, respectively. Construct a multi-task network architecture with a shared backbone network, a main task sub-network, and a preset number of auxiliary task sub-networks; The training image dataset is input into the multi-task network architecture to obtain the prediction results for each task; Based on the prediction results and the true labels, calculate the cross-entropy loss for each task, and determine the joint loss function based on the cross-entropy loss for each task; The multi-task network architecture is trained end-to-end based on the joint loss function. The parameters of the multi-task network architecture are updated by minimizing the joint loss function using the backpropagation algorithm.

6. The vehicle driving control method according to claim 5, characterized in that, The cross-entropy loss is: in," "The homoscedasticity uncertainty parameter for each task," "The total number of categories for a single task," "The set of all trainable parameters for the entire multi-task network," "Input image" The parameters are The original category score vector is obtained after processing by the convolutional neural network.

7. The vehicle driving control method according to claim 5, characterized in that, The joint loss function is: in," "The set of all trainable parameters for the entire multi-task network," "This represents the homoscedasticity uncertainty parameter for the first task." "for the first" The homoscedasticity uncertainty parameter for each task, "This represents the total loss value for multi-task learning." "This represents the total number of tasks." 8. A vehicle driving control device, characterized in that, The vehicle includes multiple image acquisition modules disposed at the front, rear, and sides of the vehicle, and the device includes: The acquisition module is used to acquire image data of the environment surrounding the vehicle through the multiple image acquisition modules; The first recognition module is used to input the image data into a preset obstacle recognition model for recognition, and to determine the obstacle information existing in the environment around the vehicle; The second recognition module is used to input the image data into a preset deep multi-task classification model to perform semantic understanding of the current driving environment and determine the current driving scenario information of the vehicle. The instruction generation module is used to generate driving control instructions based on the obstacle information and the driving scenario information; The control module is used to control the vehicle to operate according to the driving control commands.

9. A vehicle, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the vehicle driving control method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code, which can be invoked by a processor to execute the vehicle driving control method as described in any one of claims 1-7.