Vehicle automatic driving method and system based on large model

CN117755336BActive Publication Date: 2026-08-07INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF AUTOMATION CHINESE ACAD OF SCI
Filing Date
2023-11-30
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]本发明提供的基于大模型的车辆自动驾驶方法及系统,用于解决现有技术中存在的自动驾驶车辆在特殊天气及全场景下的安全性低的问题

Benefits of technology

[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the large-model-based vehicle autonomous driving method as described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117755336B_ABST
    Figure CN117755336B_ABST
Patent Text Reader

Abstract

The application provides a large model-based vehicle automatic driving method and system, and the method comprises the following steps: acquiring sensing data; converting audio data into first text data according to an audio conversion model, converting visual data into second text data based on a visual conversion model, processing the text data input by a user, the first text data and the second text data based on a language large model to obtain perception auxiliary information and decision auxiliary information; obtaining a preliminary decision result according to the environmental features of the surroundings of the vehicle, positioning data and driving data, wherein the environmental features are determined by target detection on the information around the vehicle based on the visual data, the positioning data and the perception auxiliary information; obtaining an optimal decision result according to the preliminary decision result and the decision auxiliary information, and generating an optimal planning path according to the optimal decision result to adjust the running state of the vehicle. The application fully considers the environmental features of the surroundings of the vehicle, and improves the safety of the vehicle in special weather and all scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and in particular to a method and system for autonomous driving of vehicles based on a large model. Background Technology

[0002] Autonomous vehicles are a type of equipment that reduces traffic congestion and improves traffic safety.

[0003] Existing autonomous vehicles use LiDAR data, visual data, and millimeter-wave radar data as inputs to achieve autonomous driving through autonomous driving systems. These systems comprise three main modules: perception and localization, decision-making and planning, and control. Therefore, the safety of autonomous vehicles primarily depends on the accuracy and efficiency of each module. Current autonomous driving systems can achieve good results under certain conditions (e.g., normal lighting). However, in real-world scenarios, scenes change rapidly and environments vary greatly (e.g., dusty weather, rain, snow, etc.), under which the safety of autonomous vehicles cannot be guaranteed. Specifically, autonomous vehicles have low safety in extreme weather conditions and in all scenarios. Summary of the Invention

[0004] The present invention provides a vehicle autonomous driving method and system based on a large model, which is used to solve the problem of low safety of autonomous vehicles in special weather and all scenarios in the prior art.

[0005] This invention provides a vehicle autonomous driving method based on a large model, comprising:

[0006] Acquire vehicle sensor data, including user-input text data, audio data, visual data, and positioning data;

[0007] The audio data is converted into text using an audio conversion model to obtain first text data. The visual data is converted into text using a visual conversion model to obtain second text data. The user-input text data, the first text data, and the second text data are processed for information perception based on a language big data model to obtain perception assistance information and decision assistance information. The decision assistance information is used to assist the operation of the vehicle.

[0008] Preliminary decision-making results are obtained based on the environmental features around the vehicle, the positioning data, and the driving data. The environmental features are used to detect and determine targets around the vehicle based on the visual data and the perception assistance information. The driving data includes driving tasks, driving experience, prior driving knowledge, and traffic rules.

[0009] Based on the preliminary decision results and the decision support information, an optimal decision result is obtained, and an optimal planning path is generated based on the optimal decision result. The optimal planning path is used to adjust the motion state of the vehicle.

[0010] According to the present invention, a vehicle autonomous driving method based on a large model is provided, wherein the positioning data includes lidar data, and the environmental features are obtained through the following steps:

[0011] Target detection is performed on the image data in the visual data and the lidar data to obtain a first target detection result;

[0012] Based on the perception assistance information, target detection is performed on the image data and a portion of the image data to obtain a second target detection result. The portion of the image data is obtained by cropping the image data, and the cropped image data includes target information.

[0013] Based on the first target detection result and the second target detection result, the environmental features around the vehicle are obtained.

[0014] According to a vehicle autonomous driving method based on a large model provided by the present invention, the step of performing target detection on the image data and partial image data based on the perception assistance information to obtain a second target detection result includes:

[0015] The perception assistance information, the image data, and a portion of the image data are input into the target detection model to obtain the second target detection result.

[0016] The methods for obtaining the second target detection result include:

[0017] The first image feature of the image data is extracted based on the first image encoder in the target detection model;

[0018] The second image features of the partial image data are extracted based on the second image encoder in the target detection model;

[0019] The text features of the perception assistance information are extracted based on the text encoder in the target detection model.

[0020] Based on the fusion module in the target detection model, the first image features, the second image features, and the text features are fused to obtain fused features;

[0021] Based on the fusion features, the detection result of the second target is obtained.

[0022] According to a large-model-based vehicle autonomous driving method provided by the present invention, obtaining the optimal decision result based on the preliminary decision result and decision assistance information includes:

[0023] Obtain the weight coefficients corresponding to the preliminary decision results and decision support information, respectively;

[0024] The optimal decision result is obtained based on the weighting coefficients, the preliminary decision result, and the decision support information.

[0025] According to the present invention, a vehicle autonomous driving method based on a large model is provided, wherein the methods for acquiring the vehicle's perception assistance information and decision assistance information include:

[0026] Based on the vehicle's audio data, text data, visual data, and auxiliary information, the perception auxiliary information and the decision auxiliary information are obtained. The vehicle's auxiliary information includes descriptive information about the environmental features surrounding the vehicle and decision information for controlling the vehicle's operation.

[0027] According to the vehicle autonomous driving method based on a large model provided by the present invention, the environmental features are further obtained through the following steps:

[0028] Based on the target detection model, the visual data, the positioning data, and the perception assistance information are used to detect targets around the vehicle to obtain the environmental features.

[0029] The target detection model includes multiple image encoders, a text encoder, a region candidate network, an interest region pooling layer, a classification layer, and a regression layer. The multiple image encoders and the text encoder are both composed of an 8-layer Transformer network with multi-head attention.

[0030] The present invention also provides a vehicle autonomous driving system based on a large model, comprising:

[0031] The data acquisition module is used to acquire sensor data, which includes user-input text data, audio data, visual data, and positioning data.

[0032] The data processing module is used to perform text conversion on the audio data according to the audio conversion model to obtain first text data, perform text conversion on the visual data according to the visual conversion model to obtain second text data, and perform information perception processing on the user-input text data, the first text data and the second text data according to the language big data model to obtain perception assistance information and decision assistance information, wherein the decision assistance information is used to assist the operation of the vehicle.

[0033] The auxiliary decision-making module is used to obtain preliminary decision results based on the environmental features around the vehicle, the positioning data, and the driving data. The environmental features are used to perform target detection and determination on the information around the vehicle based on the visual data and the perception assistance information. The driving data includes driving tasks, driving experience, prior driving knowledge, and traffic rules.

[0034] The decision planning module is used to obtain the optimal decision result based on the preliminary decision result and the decision support information, and to generate the optimal planning path based on the optimal decision result. The optimal planning path is used to adjust the motion state of the vehicle.

[0035] The present invention also provides an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the large-model-based vehicle autonomous driving method as described above.

[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the large-model-based vehicle autonomous driving method as described above.

[0037] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the vehicle autonomous driving method based on a large model as described above.

[0038] The vehicle autonomous driving method and system based on a large model provided by this invention fully considers the environmental characteristics around the vehicle and the vehicle's decision-making assistance information when obtaining the optimal decision result for controlling the vehicle's operation. This improves the vehicle's perception and decision-making accuracy in special weather and all scenarios, thereby enhancing the vehicle's safety in special weather and all scenarios. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0040] Figure 1 This is a flowchart illustrating the vehicle autonomous driving method based on a large model provided by the present invention.

[0041] Figure 2 This is a schematic diagram of the structure of the autonomous driving system provided by the present invention;

[0042] Figure 3 This is a schematic diagram of the principle of the sensing module provided by the present invention;

[0043] Figure 4 This is a schematic diagram of the structure of PoVLOB provided by the present invention;

[0044] Figure 5 This is a schematic diagram of the decision-making module provided by the present invention;

[0045] Figure 6 This is a schematic diagram of the structure of the vehicle autonomous driving system based on a large model provided by the present invention;

[0046] Figure 7 This is a schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0048] Figure 1 This is a flowchart illustrating the vehicle autonomous driving method based on a large model provided by the present invention, as shown below. Figure 1 As shown, the method includes:

[0049] Step 110: Acquire vehicle sensor data, including user-input text data, audio data, visual data, and positioning data;

[0050] Step 120: Convert the audio data into text based on the audio conversion model to obtain the first text data; convert the visual data into text based on the visual conversion model to obtain the second text data; perform information perception processing on the user-input text data, the first text data, and the second text data based on the language big model to obtain perception assistance information and decision assistance information. The decision assistance information is used to assist vehicle operation.

[0051] Step 130: Obtain preliminary decision results based on the environmental features around the vehicle, positioning data, and driving data. The environmental features are determined by target detection based on visual data and perception assistance information around the vehicle. The driving data includes driving tasks, driving experience, prior driving knowledge, and traffic rules.

[0052] Step 140: Based on the preliminary decision results and decision support information, obtain the optimal decision result, and generate the optimal planning path based on the optimal decision result. The optimal planning path is used to adjust the vehicle's motion state.

[0053] It should be noted that the above method can be implemented by computer equipment.

[0054] Optionally, the positioning data includes LiDAR data, GNSS (Global Navigation Satellite System) data, and IMU (Inertial Measurement Unit) data.

[0055] Optionally, the aforementioned positioning data can be calculated by the positioning module in the autonomous driving system of a specific vehicle using a multi-sensor fusion positioning algorithm based on an inertial measurement unit (IMU), a global navigation satellite system (GNSS), high-precision maps, and lidar (LiDAR).

[0056] Optionally, the environmental features around the vehicle can be acquired by using the vehicle's visual data, lidar data, and perception assistance information.

[0057] Optionally, target detection can be performed on the vehicle's visual data and lidar data based on perception-assisted information to obtain the environmental features around the vehicle.

[0058] For example, the environmental features around a vehicle can be obtained by inputting the vehicle's visual data, lidar data, and perception assistance information into a target detection model.

[0059] Specifically, the acquired environmental features around the vehicle, location data, and driving data such as driving tasks, driving experience, prior driving knowledge, and traffic rules are input into a behavioral decision model based on a Partially Observable Markov Decision Process (POMDP). Preliminary decision results are obtained based on the output of this behavioral decision model.

[0060] The behavioral decision-making model can be represented by a six-tuple {S,A,Ω,T,O,R}, where S is the finite state space of the vehicle, which can be specifically determined based on the vehicle's environmental characteristics, location data, and driving data; A is the vehicle's behavioral decision space; Ω is the probability distribution of the agent's observation sequence set being O given the action A and the next environmental state s′; T is the state transition function; O is the observation sequence set; R is the activation function; and s′ is the next environmental state (s′∈S). The preliminary decision results (e.g., following another vehicle, changing lanes, turning left / right, etc.) are obtained through the AEMS2 algorithm.

[0061] The vehicle's decision support information is fused with the preliminary decision results to obtain the fused decision result, i.e., the optimal decision result. Based on the optimal decision result, the vehicle's operation is controlled. The decision support information is used to assist the vehicle's operation. For example, the decision support information can specifically be U-turn / left turn / right turn at the next intersection, acceleration / deceleration at the next intersection, etc.

[0062] In this embodiment, after determining the optimal decision result, the vehicle control system generates the corresponding optimal planning path based on the optimal decision structure and user input instructions. The vehicle then travels to the next moving point, at the specified speed or direction based on the optimal planning path, ensuring the vehicle's safety in special weather conditions and all scenarios.

[0063] The vehicle autonomous driving method based on a large model provided by this invention fully considers the environmental characteristics around the vehicle and the vehicle's decision-making assistance information when obtaining the optimal decision result for controlling the vehicle's operation. This improves the vehicle's perception and decision-making accuracy in special weather and all scenarios, thereby enhancing the vehicle's safety in special weather and all scenarios.

[0064] Furthermore, in one embodiment, the methods for acquiring the vehicle's perception assistance information and decision assistance information include:

[0065] Based on the vehicle's audio data, text data, visual data, and vehicle auxiliary information, perception auxiliary information and decision auxiliary information are obtained. The vehicle auxiliary information includes descriptive information about the environmental characteristics around the vehicle and decision information for controlling the vehicle's operation.

[0066] Optionally, specifically, it involves acquiring sensor data collected by multiple sensors deployed on the vehicle, such as LiDAR data, visual data, GNSS data, IMU data, audio data, and text data. Among these, the audio and text data are input from humans into the vehicle, serving as auxiliary information for human interaction. This can include descriptive information about the vehicle's surrounding environment and decision-making information for controlling vehicle operation.

[0067] Figure 2 This is a structural schematic diagram of the autonomous driving system provided by the present invention, as shown below. Figure 2 As shown, it includes a perception and positioning module, a decision planning module, a control module, and an auxiliary module. The perception and positioning module includes a perception module and a positioning module. The auxiliary module includes an audio-text large model, a visual-text large model, and a language large model. The decision planning module is used to obtain the optimal decision result based on the preliminary decision result and decision support information, and then generate the optimal planning path based on the optimal decision result. The auxiliary module is used to control the vehicle to run according to the optimal planning path.

[0068] Audio data is fed into an audio-text model (such as the Whisper speech recognition model) and converted into text information, i.e., the first text information.

[0069] Visual data is fed into a large visual-text model (such as the multimodal Transformer model BLIP) and transformed into text information, i.e., second text information.

[0070] The obtained text information and human-input text data are fed into a large language model (such as the Large Language Series Model LLaMA). The large language model then generates the perception assistance information and decision assistance information required by the vehicle.

[0071] The control module controls vehicle operation based on the optimal decision result. The decision planning module generates the optimal decision result based on the vehicle's positioning data, path planning information, perception assistance information generated by the language model, and decision assistance information generated by the auxiliary module. The perception module performs target detection on the vehicle's visual data, LiDAR data, and perception assistance information generated by the auxiliary module to obtain the environmental features around the vehicle.

[0072] Furthermore, in one embodiment, the positioning data includes lidar data, and environmental features are obtained through the following steps:

[0073] Target detection is performed on image data and LiDAR data in the visual data to obtain the first target detection result;

[0074] Based on perception-assisted information, target detection is performed on image data and partial image data to obtain a second target detection result. The partial image data is obtained by cropping the image data, and the cropped image data includes target information.

[0075] Based on the results of the first target detection and the second target detection, the environmental features around the vehicle are obtained.

[0076] Optionally, based on the environmental features around the vehicle obtained by the perception module, the perception module processes image data, lidar data and perception assistance information in the visual data to achieve target detection of information around the vehicle.

[0077] Specifically, Figure 3 This is a schematic diagram of the principle of the sensing module provided by the present invention, as shown below. Figure 3 As shown, the image data from the LiDAR data and visual data is input into the F-PointNet in the perception module to achieve preliminary detection results of the environment (such as cars, pedestrians, bicycles, electric vehicles, traffic lights, curbs, intersections, lane lines, etc.), and obtain the first target detection bounding box-i&c, which is the first target detection result.

[0078] Based on perception-assisted information, target detection is performed on images and partial images containing targets such as cars, pedestrians, bicycles, electric vehicles, traffic lights, curbs, intersections, and lane lines, yielding a second detection result. The partial image data is obtained by cropping the original image data, and the cropped image data includes the target information to be detected.

[0079] After fusing the first target detection result and the second target detection result, the final target detection result, i.e., the environmental features around the vehicle, is obtained.

[0080] Furthermore, in one embodiment, based on perception assistance information, target detection is performed on image data and a portion of image data to obtain a second target detection result, including:

[0081] The sensory assistance information, image data, and partial image data are input into the target detection model to obtain the second target detection result;

[0082] The methods for obtaining the detection results of the second target include:

[0083] First image features are extracted from image data based on the first image encoder in the object detection model;

[0084] Second image features are extracted from partial image data based on the second image encoder in the object detection model;

[0085] Text features are extracted from perceptual auxiliary information based on the text encoder in the object detection model;

[0086] Based on the fusion module in the object detection model, the first image features, the second image features, and the text features are fused to obtain the fused features;

[0087] Based on the fusion features, the detection result of the second target is obtained.

[0088] Reference Figure 3 The sensory auxiliary information, image data, and partial image data are input into the target detection model PoVLOB in the perception module to obtain the second target detection result.

[0089] Image data, partial image data, and perception assistance information are input into the PoVLOB in the perception module to achieve the target detection result of image data and perception assistance information fusion, and obtain the second target detection boundingbox-i&t, which is the second target detection result.

[0090] The bounding box-i&c and bounding box-i&t are fed into the bounding box regression module in the perception module to obtain the final bounding box of the target detection, which represents the environmental features around the vehicle. The bounding box regression module consists of a three-layer convolutional neural network with a GIOU loss function.

[0091] The structure of PoVLOB is as follows: Figure 4 As shown, its input includes image data, partial image data, and perceptual assistance information (in the form of text information). For example, the text information corresponding to the perceptual assistance information is set to "Attention [MASK1][MASK2]". MASK1 includes three categories: front, left, and right. Based on the classification of MASK1, the area where the target is located in the image data is divided into three parts. Specifically, "left," "front," and "right" correspond to the left 1 / 3, middle 1 / 3, and right 1 / 3 of the image data, respectively. MASK2 represents the target that needs attention, such as: cars, pedestrians, bicycles, electric vehicles, traffic lights, curbs, intersections, lane lines, etc.

[0092] In some embodiments, environmental features are also obtained through the following steps: based on the target detection model, target detection is performed on the information around the vehicle according to visual data, positioning data, and perception assistance information to obtain environmental features; wherein, the target detection model includes multiple image encoders, text encoders, region candidate networks, interest domain pooling layers, classification layers, and regression layers, and the multiple image encoders and text encoders are all composed of 8 layers of Transformer networks with multi-head attention.

[0093] Optionally, PoVLOB comprises six parts: two image encoders (a first image encoder and a second image encoder, respectively), a text encoder, a region proposal network, an interest region pooling module, a classification module, and a regression module. Both the two image encoders and the text encoder consist of eight layers of Transformers with multi-head attention. The architecture of the region proposal network, interest region pooling, classification, and regression modules is the same as the corresponding parts in the Fast R-CNN model.

[0094] Before training the PoVLOB network, the target categories for MASK2 are first defined. This invention designs MASK2 to include cars, pedestrians, bicycles, electric vehicles, traffic lights, curbs, intersections, lane lines, etc. Next, the KITTI object detection dataset is downloaded, and text information corresponding to perceptual auxiliary information is generated based on the object detection image data.

[0095] During training and testing of PoVLOB, the collected image data and perceptual assistance information are input into two image encoders and a text encoder, respectively, to extract two image features and two text features. Specifically, the image data is input into the first image encoder in PoVLOB, and the first image feature is extracted from the image data based on the first image encoder. A portion of the image data is input into the second image encoder in PoVLOB, and the second image feature is extracted from the portion of the image data based on the second image encoder. The text information corresponding to the perceptual assistance information is input into the text encoder, and the text feature corresponding to the text information is extracted from the text encoder.

[0096] Image features and text features are fused together, and the fused features are fed into the region candidate network, interest region pooling, classification and regression modules to obtain the second target detection result.

[0097] In this embodiment, the input to the region candidate network is a feature map, and the output is multiple regions of interest (ROIs). Each ROI output by the region candidate network is specifically represented by a probability value (used to determine whether the anchor is foreground or background) and four coordinate values. The probability value represents the probability that an object exists in the ROI, which is obtained by binary classification of each ROI through the Softmax layer of the classification module. The coordinate values ​​are the predicted position of the object. During training, these coordinates are used to regress with the true coordinates to make the predicted object position more accurate during testing.

[0098] In this embodiment, the network layer corresponding to the region of interest pooling takes the region of interest and the feature map output by the RPN network as input, combines the two to obtain a fixed-size region feature map (Proposal Feature Map), and outputs it to the fully connected network for classification.

[0099] In this embodiment, the classification and regression module takes the ProposalFeature Map obtained from the previous layer as input and outputs the category to which the object belongs in the region of interest and the precise location of the object in the image. The classification and regression module classifies the image through the Softmax layer and corrects the precise location of the object through bounding box regression.

[0100] The vehicle autonomous driving method based on large models provided by this invention obtains perception assistance information by adding audio-text large models, vision-text large models and language large models to the autonomous driving system, and combines the perception assistance information to obtain the optimal decision result for controlling the vehicle operation. It can improve the accuracy of existing perception, localization and decision planning through the results of language large models under special weather and all scenarios, thereby improving vehicle safety and solving the long tail problem of vehicles.

[0101] Furthermore, in one embodiment, obtaining the optimal decision result based on the preliminary decision result and decision support information may include:

[0102] Obtain the weighting coefficients corresponding to the preliminary decision results and decision support information;

[0103] The optimal decision result is obtained based on the weighting coefficients, preliminary decision results, and decision support information.

[0104] Optionally, Figure 5 This is a schematic diagram of the decision-making module provided by the present invention, as shown below. Figure 5 As shown, the decision-making module processes the environmental features (i.e., perception assistance information) output by the vehicle's perception module, path planning information (high-precision map, traffic rules, driving tasks, and driving experience), vehicle positioning data output by the positioning module, and decision assistance information, and outputs a decision-making behavior. During execution, the environmental features (i.e., perception assistance information) output by the vehicle's perception module, path planning information (high-precision map, traffic rules, driving tasks, and driving experience), vehicle positioning data output by the positioning module, and decision assistance information are input into a behavior decision-making model based on a partially Markov chain decision process. Based on the output of this behavior decision-making model, a preliminary decision result for controlling the vehicle's operation is obtained.

[0105] The decision support information is then fused with the preliminary decision results to obtain the optimal fused decision outcome. Specifically, the decision support information and the preliminary decision results are fused based on the following formula:

[0106]

[0107] Among them, D fusion For the optimal decision outcome, D represents the weighting coefficients corresponding to the preliminary decision results. primary This represents the preliminary decision result, where β is the weighting coefficient corresponding to the decision support information, and D... auxiliary This serves as decision support information. If no decision support information is available, then... β and β are 1 and 0, respectively. Conversely, β and β are 0 and 1, respectively.

[0108] The vehicle autonomous driving method based on a large model provided by this invention improves the vehicle's decision-making accuracy and ensures vehicle safety by adding a decision-making module to the autonomous driving system and fusing decision-making assistance information and preliminary decision results based on the decision-making module.

[0109] The following describes the large-model-based autonomous driving system for vehicles provided by this invention. The large-model-based autonomous driving system described below can be referred to in correspondence with the large-model-based autonomous driving method described above.

[0110] Figure 6 This is a schematic diagram of the structure of the vehicle autonomous driving system based on a large model provided by the present invention, such as... Figure 6 As shown, it includes:

[0111] The data acquisition module 610 is used to acquire sensor data, which includes user-input text data, audio data, visual data, and positioning data.

[0112] The data processing module 620 is used to convert audio data into text based on the audio conversion model to obtain first text data, convert visual data into text based on the visual conversion model to obtain second text data, and perform information perception processing on the user-input text data, first text data and second text data based on the language big model to obtain perception assistance information and decision assistance information. The decision assistance information is used to assist vehicle operation.

[0113] The auxiliary decision-making module 630 is used to obtain preliminary decision results based on the environmental features around the vehicle, positioning data, and driving data. The environmental features are determined by target detection based on visual data and perception assistance information around the vehicle. The driving data includes driving tasks, driving experience, prior driving knowledge, and traffic rules.

[0114] The control module 640 is used to obtain the optimal decision result based on the preliminary decision result and decision support information, and to generate the optimal planned path based on the optimal decision result. The optimal planned path is used to adjust the vehicle's motion state.

[0115] The large-model-based autonomous driving system provided by this invention fully considers the environmental characteristics around the vehicle and the vehicle's decision-making assistance information when obtaining the optimal decision result for controlling the vehicle's operation. This improves the vehicle's perception and decision-making accuracy in special weather and all scenarios, thereby enhancing the vehicle's safety in special weather and all scenarios.

[0116] Figure 7 This is a schematic diagram of the physical structure of an electronic device provided by the present invention, such as... Figure 7As shown, the electronic device may include a processor 710, a communication interface 711, a memory 712, and a bus 713. The processor 710, communication interface 711, and memory 712 communicate with each other via the bus 713. The processor 710 can call logical instructions from the memory 712 to execute the following methods:

[0117] Acquire sensor data, including user-input text data, audio data, visual data, and positioning data;

[0118] The audio data is converted into text based on the audio conversion model to obtain the first text data. The visual data is converted into text based on the visual conversion model to obtain the second text data. The user-input text data, the first text data, and the second text data are processed for information perception based on the language big data model to obtain perception assistance information and decision assistance information. The decision assistance information is used to assist vehicle operation.

[0119] Preliminary decision-making results are obtained based on the environmental features around the vehicle, positioning data, and driving data. Environmental features are determined by target detection based on visual data and perception assistance information around the vehicle. Driving data includes driving tasks, driving experience, prior driving knowledge, and traffic rules.

[0120] Based on the preliminary decision results and decision support information, the optimal decision result is obtained, and the optimal planning path is generated based on the optimal decision result. The optimal planning path is used to adjust the vehicle's motion state.

[0121] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer power supply (which may be a personal computer, server, or network power supply, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0122] Furthermore, this invention discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when these instructions are executed by a computer, the computer can execute the large-model-based vehicle autonomous driving method provided in the above-described method embodiments, for example including:

[0123] Acquire sensor data, which includes user-input text data, audio data, visual data, and location data;

[0124] The audio data is converted into text based on the audio conversion model to obtain the first text data. The visual data is converted into text based on the visual conversion model to obtain the second text data. The user-input text data, the first text data, and the second text data are processed for information perception based on the language big data model to obtain perception assistance information and decision assistance information. The decision assistance information is used to assist vehicle operation.

[0125] Preliminary decision-making results are obtained based on the environmental features around the vehicle, positioning data, and driving data. Environmental features are determined by target detection based on visual data and perception assistance information around the vehicle. Driving data includes driving tasks, driving experience, prior driving knowledge, and traffic rules.

[0126] Based on the preliminary decision results and decision support information, the optimal decision result is obtained, and the optimal planning path is generated based on the optimal decision result. The optimal planning path is used to adjust the vehicle's motion state.

[0127] On the other hand, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the large-model-based vehicle autonomous driving method provided in the above embodiments, for example including:

[0128] Acquire sensor data, including user-input text data, audio data, visual data, and positioning data;

[0129] The audio data is converted into text based on the audio conversion model to obtain the first text data. The visual data is converted into text based on the visual conversion model to obtain the second text data. The user-input text data, the first text data, and the second text data are processed for information perception based on the language big data model to obtain perception assistance information and decision assistance information. The decision assistance information is used to assist vehicle operation.

[0130] Preliminary decision-making results are obtained based on the environmental features around the vehicle, positioning data, and driving data. Environmental features are determined by target detection based on visual data and perception assistance information around the vehicle. Driving data includes driving tasks, driving experience, prior driving knowledge, and traffic rules.

[0131] Based on the preliminary decision results and decision support information, the optimal decision result is obtained, and the optimal planning path is generated based on the optimal decision result. The optimal planning path is used to adjust the vehicle's motion state.

[0132] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer power supply (which may be a personal computer, server, or network power supply, etc.) to execute the methods described in various embodiments or some parts of the embodiments.

[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for autonomous driving of vehicles based on a large model, characterized in that, include: Acquire vehicle sensor data, including user-input text data, audio data, visual data collected by sensors, and positioning data calculated by a positioning module; The audio data is converted into text using an audio conversion model to obtain first text data. The visual data is converted into text using a visual conversion model to obtain second text data. The user-input text data, the first text data, and the second text data are processed for information perception based on a language big data model to obtain perception assistance information and decision assistance information. The decision assistance information is used to assist the operation of the vehicle. Preliminary decision-making results are obtained based on the environmental features around the vehicle, the positioning data, and the driving data. The environmental features are used to detect and determine targets around the vehicle based on the visual data and the perception assistance information. The driving data includes driving tasks, driving experience, prior driving knowledge, and traffic rules. Based on the preliminary decision results and the decision support information, an optimal decision result is obtained, and an optimal planning path is generated based on the optimal decision result. The optimal planning path is used to adjust the motion state of the vehicle.

2. The vehicle autonomous driving method based on a large model according to claim 1, characterized in that, The positioning data includes lidar data, and the environmental features are obtained through the following steps: Target detection is performed on the image data in the visual data and the lidar data to obtain a first target detection result; Based on the perception assistance information, target detection is performed on the image data and partial image data to obtain a second target detection result. The partial image data is obtained by cropping the image data, and the cropped image data includes target information. Based on the first target detection result and the second target detection result, the environmental features around the vehicle are obtained.

3. The vehicle autonomous driving method based on a large model according to claim 2, characterized in that, The step of performing target detection on the image data and a portion of the image data based on the perception assistance information to obtain a second target detection result includes: The perception assistance information, the image data, and a portion of the image data are input into the target detection model to obtain the second target detection result. The methods for obtaining the second target detection result include: The first image feature of the image data is extracted based on the first image encoder in the target detection model; The second image features of the partial image data are extracted based on the second image encoder in the target detection model. The text features of the perception assistance information are extracted based on the text encoder in the target detection model. Based on the fusion module in the target detection model, the first image features, the second image features, and the text features are fused to obtain fused features; Based on the fusion features, the detection result of the second target is obtained.

4. The vehicle autonomous driving method based on a large model according to any one of claims 1-3, characterized in that, The step of obtaining the optimal decision result based on the preliminary decision result and decision support information includes: Obtain the weight coefficients corresponding to the preliminary decision results and decision support information, respectively; The optimal decision result is obtained based on the weighting coefficients, the preliminary decision result, and the decision support information.

5. The vehicle autonomous driving method based on a large model according to any one of claims 2-3, characterized in that, The methods for acquiring the vehicle's perception assistance information and decision assistance information include: Based on the vehicle's audio data, text data, visual data, and auxiliary information, the perception assistance information and the decision assistance information are obtained. The vehicle's auxiliary information includes descriptive information about the environmental features surrounding the vehicle and decision information for controlling the vehicle's operation.

6. The vehicle autonomous driving method based on a large model according to claim 1, characterized in that, The environmental characteristics are also obtained through the following steps: Based on the target detection model, the visual data, the positioning data, and the perception assistance information are used to detect targets around the vehicle to obtain the environmental features. The target detection model includes multiple image encoders, a text encoder, a region candidate network, an interest region pooling layer, a classification layer, and a regression layer. The multiple image encoders and the text encoder are both composed of an 8-layer Transformer network with multi-head attention.

7. A vehicle autonomous driving system based on a large model, characterized in that, include: The data acquisition module is used to acquire sensor data, which includes text data and audio data input by the user, visual data collected by the sensor, and positioning data calculated by the positioning module. The data processing module is used to perform text conversion on the audio data according to the audio conversion model to obtain first text data, perform text conversion on the visual data according to the visual conversion model to obtain second text data, and perform information perception processing on the user-input text data, first text data and second text data according to the language big data model to obtain perception assistance information and decision assistance information, wherein the decision assistance information is used to assist the operation of the vehicle. The auxiliary decision-making module is used to obtain preliminary decision results based on the environmental features around the vehicle, the positioning data, and the driving data. The environmental features are used to perform target detection and determination on the information around the vehicle based on the visual data and the perception assistance information. The driving data includes driving tasks, driving experience, prior driving knowledge, and traffic rules. The decision planning module is used to obtain the optimal decision result based on the preliminary decision result and the decision support information, and to generate the optimal planning path based on the optimal decision result. The optimal planning path is used to adjust the motion state of the vehicle.

8. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the vehicle autonomous driving method based on a large model as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the large-model-based vehicle autonomous driving method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the large-model-based vehicle autonomous driving method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Interacting type autonomous instructional car system

    CN105799710A

  • Image processing method and related device

    CN116304146A