Vehicle driving environment detection method and device, electronic equipment and storage medium

By combining deep learning of multi-view point cloud data and multimodal obstacle perception models in vehicle environment detection, the problem that traditional models cannot accurately determine the vehicle environment is solved, and more accurate obstacle detection and environment perception are achieved.

CN120635855APending Publication Date: 2025-09-12VOYAH AUTOMOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510551464.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional single deep learning models cannot accurately determine the vehicle environment, resulting in inaccurate perception of the vehicle's surroundings.

Method used

By acquiring multi-dimensional image data and multi-view point cloud data of the vehicle in the current driving environment, using the trained visual large language model for language fusion learning, combined with the multimodal obstacle perception model for deep learning, the obstacle category and location information are determined, and the detection results are optimized through verification instructions.

Benefits of technology

It improves the perception and detection accuracy of the vehicle's surrounding environment, ensuring the accuracy of obstacle category and location information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635855A_ABST
    Figure CN120635855A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle driving environment detection method and device, electronic equipment and a storage medium. Relates to the technical field of vehicle driving. Acquiring initial demand data of the vehicle in the current driving environment, wherein the initial demand data comprises multi-dimensional image data, multi-view point cloud data and a detection instruction of a user for the current driving environment; inputting the initial demand data into a trained visual big language model for language fusion learning, and determining the types of obstacles existing in the driving environment and the position information of the obstacles; and taking the obstacle category, the position information of the obstacle, the multi-dimensional image data and the original point cloud data of the vehicle in the current driving environment as candidate demand data, inputting the candidate demand data into the trained multi-mode obstacle sensing model for deep learning, and determining the detection information of the obstacle in the current driving environment. Therefore, the detection of the current driving environment is completed. According to the invention, the accuracy and precision of output detection information of the trained multi-mode obstacle sensing model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle driving technology, and in particular to a method, device, electronic device and storage medium for detecting a vehicle driving environment. Background Art

[0002] At present, with the development of society and the advancement of science and technology, more and more users are beginning to use vehicles with intelligent driving or assisted driving as their means of transportation. In the field of autonomous driving, environmental perception is the most important channel for intelligent driving systems to obtain information. In order to obtain more comprehensive and detailed environmental information, traditional intelligent driving systems often introduce multimodal sensor inputs, and then conduct deep learning on the data collected by these sensors to achieve detection and perception of the vehicle's surrounding environment.

[0003] However, traditional single deep learning models have limited extraction and analysis of data features and cannot accurately determine the vehicle environment. Summary of the Invention

[0004] The embodiments of the present application provide a method, device, electronic device and storage medium for detecting a vehicle driving environment. The embodiments provided in the present application solve the technical problem in the prior art that the vehicle environment cannot be accurately determined.

[0005] In a first aspect of the embodiments of the present application, the embodiments of the present application provide a method for detecting a vehicle driving environment, the method for detecting a vehicle driving environment comprising:

[0006] Acquiring initial demand data of the vehicle in a current driving environment, the initial demand data including multi-dimensional image data, multi-view point cloud data, and a user's detection instruction for the current driving environment;

[0007] Inputting the initial demand data into a trained visual language model for language fusion learning to determine the types of obstacles present in the driving environment and the location information of the obstacles;

[0008] The obstacle category, the location information of the obstacle, the multidimensional image data, and the original point cloud data of the vehicle in the current driving environment are used as candidate required data and input into a trained multimodal obstacle perception model for deep learning to determine detection information of the obstacle in the current driving environment so as to complete the detection of the current driving environment, wherein the detection information is used to represent a response statement containing the obstacle category, the location information, and the geometric information of the obstacle.

[0009] In a feasible implementation, the multi-view point cloud data of the vehicle in the current driving environment is obtained in the following manner:

[0010] Obtain the original point cloud data of the vehicle in the current driving environment;

[0011] Based on a preset three-view point cloud division rule and a preset diagonal five-view point cloud division rule, the original point cloud data is divided into multiple view spaces to determine the anchor point position information under different views;

[0012] Based on the anchor point position information in different views, multi-view point cloud data of the vehicle in the current driving environment is constructed.

[0013] In a feasible implementation manner, after inputting the obstacle category, the location information of the obstacle, the multi-dimensional image data, and the original point cloud data of the vehicle in the current driving environment as candidate required data into a trained multimodal obstacle perception model for deep learning and determining the detection information of the obstacle in the current driving environment, the method further includes:

[0014] Obtaining a verification instruction for the current driving environment;

[0015] The initial demand data, the candidate demand data, the detection information and the verification instruction are input into the trained visual large language model for secondary language fusion learning to determine the target detection information after verification of the obstacle in the current driving environment, wherein the target detection information is used to represent the target response sentence including the verified obstacle category, the verified obstacle location information and the verified obstacle geometry information.

[0016] In a feasible implementation, the verification instruction includes a right / wrong verification instruction and / or a supplementary verification instruction, and the inputting of the initial requirement data, the candidate requirement data, the detection information, and the verification instruction into a trained visual large language model for secondary language fusion learning to determine target detection information after verification of the obstacle in the current driving environment includes:

[0017] Inputting the right / wrong verification instruction, the initial requirement data, and the candidate requirement data into a trained visual language model to determine whether the detection information of the obstacle has errors;

[0018] If it does not exist, the supplementary verification instruction, the initial requirement data, the candidate requirement data, and the detection information are input into the trained visual large language model for secondary language fusion learning to determine the target detection information after verification of the obstacle in the current driving environment.

[0019] In a feasible implementation manner, after inputting the right / wrong verification instruction, the initial requirement data, and the candidate requirement data into a trained visual large language model to determine whether the detection information of the obstacle is erroneous, the method further includes:

[0020] If there is an error in the detection information of the obstacle, the detection information is modified, and the modified detection information, the supplementary verification instruction, the initial requirement data and the candidate requirement data are input into the trained visual large language model for secondary language fusion learning to determine the target detection information after verification of the obstacle in the current driving environment.

[0021] In a feasible implementation manner, obtaining a verification instruction for the current driving environment includes:

[0022] Based on the user's verification requirement for the current driving environment, a verification instruction for the current driving environment is obtained.

[0023] In a feasible implementation, the multi-dimensional image data of the vehicle in the current driving environment is obtained in the following manner:

[0024] Acquire multi-dimensional initial image data of the vehicle in the current driving environment;

[0025] De-noising is performed on the multi-dimensional initial image data to determine the vehicle in the multi-dimensional image data.

[0026] According to a second aspect of the embodiments of the present application, an apparatus for detecting a vehicle driving environment is provided. The apparatus for detecting a vehicle driving environment includes:

[0027] A first acquisition module is configured to acquire initial demand data of the vehicle in a current driving environment, wherein the initial demand data includes multi-dimensional image data, multi-view point cloud data, and a user's detection instruction for the current driving environment;

[0028] A first determination module is configured to input the initial demand data into a trained visual language model for language fusion learning, and determine the types of obstacles present in the driving environment and the location information of the obstacles;

[0029] The second determination module is configured to input the obstacle category, the location information of the obstacle, the multidimensional image data, and the original point cloud data of the vehicle in the current driving environment as candidate required data into a trained multimodal obstacle perception model for deep learning, thereby determining detection information of the obstacle in the current driving environment to complete detection of the current driving environment, wherein the detection information is used to represent a response statement including the obstacle category, the location information, and the geometric information of the obstacle.

[0030] In a third aspect of the embodiments of the present application, the embodiments of the present application provide an electronic device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to perform the steps of the above-mentioned vehicle driving environment detection method when the processor is running.

[0031] In a fourth aspect of the embodiments of the present application, the embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the vehicle driving environment detection method as described above are executed.

[0032] The vehicle driving environment detection method, device, electronic device, and storage medium provided in the embodiments of the present application input the initial demand data of the vehicle in the current driving environment into a trained visual large language model for language fusion learning, thereby determining the obstacle category and location information present in the driving environment. The obstacle category, obstacle location information, multi-dimensional image data, and the original point cloud data of the vehicle in the current driving environment are then input as candidate demand data into a trained multimodal obstacle perception model for deep learning to determine the detection information of the obstacles in the current driving environment, thereby completing the detection of the current driving environment. The present application inputs the input data combining multi-view point cloud data and instructions into the trained visual large language model, enabling the trained visual large language model to participate in the analysis of the multimodal input data, thereby providing a basis for feature extraction and deep learning analysis of the trained multimodal obstacle perception model, thereby improving the accuracy and precision of the detection information output by the trained multimodal obstacle perception model, and thereby improving the perception capability and accuracy of the vehicle's surrounding environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A flowchart of a method for detecting a vehicle driving environment provided in an embodiment of the present application is shown;

[0034] Figure 2A schematic diagram of a spatial screenshot of multi-view point cloud data in a method for detecting a vehicle driving environment provided by an embodiment of the present application is shown;

[0035] Figure 3 A structural block diagram of a vehicle driving environment detection device provided in an embodiment of the present application is shown;

[0036] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown.

[0037] Figure 3 and Figure 4 The corresponding relationship between the reference numerals and the names of the drawings is as follows:

[0038] 300 Detection device for vehicle driving environment; 310 First acquisition module; 320 First determination module; 330 Second determination module; 340 Second acquisition module; 350 Third determination module; 4 Electronic device; 401 Processor; 402 Memory; 403 Bus. DETAILED DESCRIPTION

[0039] In order to better understand the technical solutions provided by the embodiments of this specification, the technical solutions of the embodiments of this specification are described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of this specification and the specific features in the embodiments are detailed descriptions of the technical solutions of the embodiments of this specification, rather than limitations on the technical solutions of this specification. In the absence of conflict, the embodiments of this specification and the technical features in the embodiments can be combined with each other.

[0040] In this article, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of further restrictions, the elements defined by the statement "comprising a ..." do not exclude the presence of other identical elements in the process, method, article or equipment comprising the elements. The term "two or more" includes two or more than two cases.

[0041] First, the application scenarios to which this application is applicable are introduced. The embodiments provided in this application are applicable to the field of vehicle driving technology, and in particular, relate to a method, device, electronic device and storage medium for detecting a vehicle driving environment.

[0042] Currently, traditional single deep learning models have limited extraction and analysis of data features and cannot accurately determine the vehicle environment.

[0043] Based on this, the embodiments of the present application provide a method, device, electronic device and storage medium for detecting a vehicle driving environment. The embodiments provided in the present application solve the technical problem in the prior art that the vehicle environment cannot be accurately determined.

[0044] Figure 1 FIG. 1 shows a flow chart of a method for detecting a vehicle driving environment provided by an embodiment of the present application. Figure 1 As shown, the method for detecting the vehicle driving environment includes the following steps:

[0045] S101 , obtaining initial demand data of the vehicle in the current driving environment, where the initial demand data includes multi-dimensional image data, multi-view point cloud data, and a user's detection instruction for the current driving environment.

[0046] In this step, the embodiment provided by the present application first needs to obtain the initial demand data of the vehicle in the current driving environment when the vehicle is in driving or stationary state. In this way, when the user needs to query, detect or perceive the current driving environment, the corresponding initial demand data can be directly obtained.

[0047] It should be noted that the multi-dimensional image data in the embodiments provided by the present application can select different types of input data under multimodality according to different application scenarios, such as weather data, road information data and obstacle data in the current driving environment.

[0048] It can be understood that the initial requirement data in the embodiments provided by the present application can be set as relevant data for identifying obstacles, that is, the initial requirement data in the embodiments provided by the present application can specifically be: multi-dimensional image data, multi-view point cloud data and user detection instructions for the current driving environment.

[0049] Among them, in the embodiments provided in this application, multi-dimensional image data is used to represent image data of multiple dimensions acquired by an image acquisition device. The image acquisition device in the embodiments provided in this application may be a camera or an image sensor, etc.

[0050] In the above, the user's detection instructions for the current driving environment in the embodiment provided by this application can be: when the user wants to detect the current driving environment at any time, the user sends a voice input instruction, gesture input instruction, and touch screen input instruction to the vehicle.

[0051] Exemplarily, multi-view point cloud data of the vehicle in the current driving environment is obtained in the following manner: original point cloud data of the vehicle in the current driving environment is obtained; based on a preset three-perspective view point cloud division (Top-Plan-View, TPV) rule and a preset diagonal five-view point cloud division rule, the original point cloud data is divided into multi-view spaces to determine the anchor point position information under different views; based on the anchor point position information under different views, multi-view point cloud data of the vehicle in the current driving environment is constructed.

[0052] In the above, the multi-view point cloud data in the embodiment provided by the present application is obtained by dividing the original point cloud data in the current driving environment according to the preset three-view point cloud division rule and the preset diagonal five-view point cloud division rule.

[0053] Here, please refer to the specific division method. Figure 2 , Figure 2 FIG. 1 shows a schematic diagram of a spatial screenshot of multi-view point cloud data in a method for detecting a vehicle driving environment provided by an embodiment of the present application. Figure 2 As shown, evenly distributed anchor points are selected from the original point cloud data in the current driving environment, and then the space corresponding to the above anchor points is divided according to the preset three-perspective view (TPV) and the preset diagonal five views to determine the position information of each anchor point under different views. Based on the anchor points and the position information of the anchor points obtained from these space divisions, the multi-view point cloud data of the vehicle in the current driving environment is constructed, so that the subsequent trained visual large language model can more accurately understand the point cloud information in the three-dimensional space, solving the weakness of the trained visual large language model in being unable to perform 3D detection tasks.

[0054] Among them, any one of the TPV three views contains multiple evenly distributed anchor points. The selection of the screenshot interval of TPV and the preset diagonal five views can select different granularities according to different traffic flow scenarios. The denser the scene, the finer the screenshot interval can be selected.

[0055] The screenshot of the TPV+preset diagonal five-view determined in the embodiment provided by this application describes the spatial positions of all anchor points and the semantic prompt words for the spatial positions of the above-mentioned anchor points. The embodiment provided by this application uses TPV+preset diagonal five-view to refine the input original point cloud data, which can solve the problem that traditional large visual language models are difficult to perform 3D detection tasks, and the point cloud data determined by TPV+preset diagonal five-view can fully and carefully understand the three-dimensional spatial data under the current driving environment.

[0056] In the embodiments provided in this application, the device for collecting raw point cloud data can be customized and selected according to different road conditions or different application scenarios. The device for collecting point cloud data selected in this application can be: lidar or millimeter wave radar, etc.

[0057] S102: Input the initial demand data into the trained visual language model for language fusion learning to determine the obstacle categories and location information of the obstacles in the driving environment.

[0058] In this step, after the initial demand data is determined, the embodiment provided by the present application needs to input the multidimensional image data and / or multidimensional initial image, multi-view point cloud data and the user's detection instructions for the current driving environment in the initial demand data into the trained visual large language model for language fusion learning, and judge the obstacle categories (such as obstructions, pedestrians, animals, plants, background, noise, raindrops and snow water, etc.) in the driving environment. After determining the obstacle category, determine the position information (including but not limited to spatial position) of the anchor points corresponding to the above obstacles, so as to complete the refinement of the above data.

[0059] It should be noted that the obstacle categories in the embodiments provided in this application can be customized according to different application scenarios and vehicle and road conditions. The obstacle categories in the embodiments provided in this application can be pedestrians that affect the driving of the vehicle in the current driving environment (obstacle categories include dynamic obstacles and static obstacles).

[0060] It can be understood that the trained visual language model in the embodiment provided in this application can be a llama-7B model, which can specifically realize the fusion of multi-dimensional data.

[0061] S103: Input the obstacle category, obstacle location information, multi-dimensional image data, and original point cloud data of the vehicle in the current driving environment as candidate required data into a trained multimodal obstacle perception model for deep learning. The model then determines obstacle detection information in the current driving environment to complete detection of the current driving environment. The detection information is used to represent a response statement containing the obstacle category, location information, and geometric information of the obstacle.

[0062] In this step, the embodiment provided by the present application utilizes the ultra-long text comprehension capability of the trained visual large language model in the early stage to perform language fusion on the initial demand data, determine the obstacle category and obstacle location information, and input the candidate demand data including the obstacle category, obstacle location information, multi-dimensional image data, and the original point cloud data of the vehicle in the current driving environment into the trained multimodal obstacle perception model for deep learning. This achieves the re-fusion of the fusion output result of the trained visual large language model and the detection result of the trained multimodal obstacle perception model, thereby determining the detection and perception of the vehicle's current driving environment, and outputting a response sentence containing the obstacle category, location information, and obstacle geometric information, to improve the perception accuracy and detection accuracy of the trained multimodal obstacle perception model.

[0063] It is understandable that the trained multimodal obstacle perception model in the embodiments provided in this application may be a multimodal BEV perception model (Bird's Eye View Perception Model, BEV).

[0064] For example, after determining the detection information of obstacles in the current driving environment, the embodiments provided in this application may further:

[0065] Obtain verification instructions for the current driving environment, wherein the verification instructions are used to represent instructions for correctness and / or supplementary verification of the detection information; input the initial demand data, candidate demand data, detection information and verification instructions into the trained visual large language model for secondary language fusion learning to determine the target detection information after obstacle verification in the current driving environment, wherein the target detection information is used to represent a target response sentence containing the verified obstacle category, the verified obstacle location information and the verified obstacle geometry information.

[0066] It should be noted that after determining the preliminary obstacle detection information, the present application verifies the preliminary obstacle detection information output above by specifying verification instructions for correctness and / or supplementary verification of the current driving environment. The specific verification method may be: based on the above verification instructions, the preliminary obstacle detection information, initial demand data and candidate demand data are re-input into the trained visual large language model for secondary language fusion learning, and the target detection information of the verified obstacle is output, thereby realizing the correction and correctness judgment of the preliminary obstacle detection information, so that the trained visual large language model can participate in both pre-detection and post-detection recognition, thereby more accurately reducing the problem of inaccurate data fusion that may exist in the trained visual large language model, and improving the accuracy of multimodal perception.

[0067] Exemplarily, obtaining a verification instruction for the current driving environment includes: obtaining a verification instruction for the current driving environment based on a user's verification requirement for the current driving environment.

[0068] In the above, the verification instruction of the current driving environment in the embodiment provided by this application can be: input and set by the user based on the user's verification requirements; or it can be a preset verification instruction directly triggered after determining the detection information of the obstacle in the current driving environment.

[0069] Exemplarily, the verification instructions include right / wrong verification instructions and / or supplementary verification instructions. Initial requirement data, candidate requirement data, detection information, and verification instructions are input into a trained visual language model for secondary language fusion learning to determine target detection information after obstacle verification in the current driving environment, including:

[0070] The right and wrong verification instructions, initial requirement data and candidate requirement data are input into the trained visual large language model to determine whether there is an error in the obstacle detection information; if not, the supplementary verification instructions, initial requirement data, candidate requirement data and detection information are input into the trained visual large language model for secondary language fusion learning to determine the target detection information after obstacle verification in the current driving environment.

[0071] In the above, when the user wants to verify whether the detection information is correct, the embodiment provided by the present application needs to input the correctness verification instructions, initial requirement data and candidate requirement data into the trained visual large language model, and judge whether there is an error in the detection information of the obstacle. When there is an error in the detection information of the obstacle, the detection information is modified, and the modified detection information, supplementary verification instructions, initial requirement data and candidate requirement data are input into the trained visual large language model for secondary language fusion learning to determine the target detection information after obstacle verification in the current driving environment.

[0072] It should be noted that when it is determined that there are no errors in the obstacle detection information, the user can perform additional verification as needed. The additional verification is used to determine whether there are any omissions in obstacle detection in the current driving environment of the vehicle.

[0073] If the user needs to perform additional verification instructions, or the preset verification instructions require the execution of additional verification instructions in sequence after the execution of the correct or incorrect verification instructions, the additional verification instructions, initial requirement data, candidate requirement data, and detection information will be input into the trained visual language model again for secondary language fusion learning to obtain the target detection information after obstacle correction and verification in the current driving environment.

[0074] Next, a specific embodiment is used to determine the driving environment of a vehicle, wherein the current driving environment of the vehicle is specifically a driving environment favored by a highly reflective road sign with high illumination. The current driving environment to be determined specifically involves determining whether there is a noise point (i.e., an obstacle) under the highly reflective road sign with high illumination.

[0075] First, collect multi-dimensional image data (denoised data) under the vehicle's high-reflective road sign: denoised multi-dimensional image data, multi-view point cloud data, and the position information of spatial anchor points under multi-view: [0m, 0m, 0m]..[0m, 0m, 1m][Lm, Wm, Hm].

[0076] Next, the above data is input into the trained visual language model for language fusion learning, and the category of obstacles under high-reflective road signs with high illumination and the location information of the obstacles are output. For example, if the obstacle category is a bicycle, the location information of the bicycle is: [x1m, y1m, z1m]; if the obstacle category is a noise point cloud, the location information of the noise point cloud is: [x2m, y2m, z2m]; if the obstacle category is a pedestrian, the location information of the pedestrian is: [x3m, y3m, z3m].

[0077] Next, the above detection results, initial demand data, candidate demand data and verification instructions "correct all false detections and detection quality problems in the existing perception results, and supplement all missed perception results" are input into the trained visual large language model for secondary language fusion learning to determine the position information of the verified obstacle and the target response sentence of the verified obstacle geometric information, which can be "there is a pedestrian at the position [43.5m, -22.1m, 0.8m] under the high-reflective road sign with high illumination, the pedestrian's direction is [0.12rad], and the pedestrian's control size is [0.4m, 0.3m, 1.6m].

[0078] The vehicle driving environment detection method provided in an embodiment of the present application inputs initial demand data of the vehicle in the current driving environment into a trained visual large language model for language fusion learning, thereby determining the obstacle categories and location information present in the driving environment. The obstacle categories, obstacle location information, multi-dimensional image data, and the original point cloud data of the vehicle in the current driving environment are then input as candidate demand data into a trained multimodal obstacle perception model for deep learning to determine the detection information of obstacles in the current driving environment, thereby completing the detection of the current driving environment. The present application inputs input data combining multi-view point cloud data and instructions into the trained visual large language model, enabling the trained visual large language model to participate in the analysis of multimodal input data, thereby providing a basis for feature extraction and deep learning analysis of the trained multimodal obstacle perception model, thereby improving the accuracy and precision of the detection information output by the trained multimodal obstacle perception model, and thereby improving the perception capability and accuracy of the vehicle's surrounding environment.

[0079] Figure 3 FIG. 1 shows a structural block diagram of a vehicle driving environment detection device provided by an embodiment of the present application. Figure 3 As shown, the vehicle driving environment detection device 300 includes:

[0080] The first acquisition module 310 is used to acquire initial demand data of the vehicle in the current driving environment. The initial demand data includes multi-dimensional image data, multi-view point cloud data and a user's detection instruction for the current driving environment.

[0081] The first determination module 320 is used to input the initial demand data into the trained visual language model for language fusion learning, and determine the obstacle categories and location information of the obstacles in the driving environment.

[0082] The second determination module 330 is used to input the obstacle category, obstacle location information, multi-dimensional image data, and original point cloud data of the vehicle in the current driving environment as candidate required data into a trained multimodal obstacle perception model for deep learning to determine the detection information of the obstacles in the current driving environment so as to complete the detection of the current driving environment. The detection information is used to represent a response statement containing the obstacle category, location information, and geometric information of the obstacle.

[0083] The second acquisition module 340 is used to obtain a verification instruction for the current driving environment.

[0084] The third determination module 350 is used to input the initial demand data, candidate demand data, detection information and verification instructions into the trained visual large language model for secondary language fusion learning, and determine the target detection information after obstacle verification in the current driving environment, wherein the target detection information is used to represent the target response sentence including the verified obstacle category, the verified obstacle location information and the verified obstacle geometry information.

[0085] Exemplarily, the first acquisition module 310 is specifically configured to:

[0086] Obtain the original point cloud data of the vehicle in the current driving environment.

[0087] Based on the preset three-view point cloud division rules and the preset diagonal five-view point cloud division rules, the original point cloud data is divided into multiple view spaces to determine the anchor point position information under different views.

[0088] Based on the anchor point position information under different views, multi-view point cloud data of the vehicle in the current driving environment is constructed.

[0089] Exemplarily, the verification instruction includes a right / wrong verification instruction and / or a supplementary verification instruction. The third determination module 350 is specifically configured to:

[0090] The right / wrong verification instructions, initial requirement data, and candidate requirement data are input into the trained visual large language model to determine whether there are errors in the obstacle detection information.

[0091] If it does not exist, the supplementary verification instructions, initial requirement data, candidate requirement data, and detection information will be input into the trained visual language model for secondary language fusion learning to determine the target detection information after obstacle verification in the current driving environment.

[0092] For example, if there is an error in the obstacle detection information, the detection information is modified, and the modified detection information, supplementary verification instructions, initial requirement data, and candidate requirement data are input into the trained visual large language model for secondary language fusion learning to determine the target detection information after obstacle verification in the current driving environment.

[0093] Exemplarily, obtaining a verification instruction for the current driving environment includes:

[0094] Based on the user's verification requirement for the current driving environment, a verification instruction for the current driving environment is obtained.

[0095] Exemplarily, the first acquisition module 310 acquires multi-dimensional image data of the vehicle in the current driving environment by:

[0096] Acquire multi-dimensional initial image data of the vehicle in the current driving environment.

[0097] De-noising is performed on the multi-dimensional initial image data to determine the location of the vehicle in the multi-dimensional image data.

[0098] The vehicle driving environment detection device 300 provided in the embodiments of the present application, compared with the prior art, and the vehicle driving environment detection method provided in the embodiments of the present application, compared with the prior art, inputs the initial demand data of the vehicle in the current driving environment into a trained visual large language model for language fusion learning to determine the obstacle category and location information of the obstacle in the driving environment. The obstacle category, obstacle location information, multi-dimensional image data, and the original point cloud data of the vehicle in the current driving environment are used as candidate demand data and input into the trained multimodal obstacle perception model for deep learning to determine the detection information of the obstacles in the current driving environment, thereby completing the detection of the current driving environment. The present application inputs the input data combining multi-view point cloud data and instructions into the trained visual large language model, enabling the trained visual large language model to participate in the analysis of multimodal input data, thereby providing a basis for feature extraction and deep learning analysis of the trained multimodal obstacle perception model, thereby improving the accuracy and precision of the detection information output by the trained multimodal obstacle perception model, and thus improving the perception capability and accuracy of the vehicle's surrounding environment.

[0099] See also Figure 4 , Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the electronic device 4 includes a processor 410 , a memory 420 and a bus 430 .

[0100] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 is running, the processor 410 communicates with the memory 420 via the bus 430. When the machine-readable instructions are executed by the processor 410, the above-mentioned Figure 1 The specific implementation of the steps of the vehicle driving environment detection method in the method embodiment shown can be found in the method embodiment, and will not be repeated here.

[0101] The embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the computer program can execute the above-mentioned Figure 1 The specific implementation of the steps of the vehicle driving environment detection method in the method embodiment shown can be found in the method embodiment, and will not be repeated here.

[0102] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0103] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0104] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-readable program code.

[0105] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0106] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0108] An embodiment of the present application further provides a computer program product, which includes computer software instructions. When the computer software instructions are executed on a processing device, the processing device executes the process of the vehicle driving environment detection method.

[0109] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0110] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0111] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0112] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0113] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0114] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0115] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

[0116] Although the preferred embodiments of this specification have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of this specification.

[0117] Obviously, those skilled in the art may make various changes and modifications to this specification without departing from the spirit and scope of this specification. Thus, if such changes and modifications fall within the scope of the claims of this specification and their equivalents, this specification is intended to include such changes and modifications.

Claims

1. A method for detecting a vehicle driving environment, characterized in that: The vehicle driving environment detection method includes: Acquiring initial demand data of the vehicle in a current driving environment, the initial demand data including multi-dimensional image data, multi-view point cloud data, and a user's detection instruction for the current driving environment; Inputting the initial demand data into a trained visual language model for language fusion learning to determine the types of obstacles present in the driving environment and the location information of the obstacles; The obstacle category, the location information of the obstacle, the multidimensional image data, and the original point cloud data of the vehicle in the current driving environment are used as candidate required data and input into a trained multimodal obstacle perception model for deep learning to determine detection information of the obstacle in the current driving environment so as to complete the detection of the current driving environment, wherein the detection information is used to represent a response statement containing the obstacle category, the location information, and the geometric information of the obstacle.

2. The method for detecting a vehicle driving environment according to claim 1, wherein: Obtain multi-view point cloud data of the vehicle in the current driving environment through the following methods: Obtain the original point cloud data of the vehicle in the current driving environment; Based on a preset three-view point cloud division rule and a preset diagonal five-view point cloud division rule, the original point cloud data is divided into multiple view spaces to determine the anchor point position information under different views; Based on the anchor point position information in different views, multi-view point cloud data of the vehicle in the current driving environment is constructed.

3. The method for detecting a vehicle driving environment according to claim 1, wherein: After inputting the obstacle category, the location information of the obstacle, the multidimensional image data, and the original point cloud data of the vehicle in the current driving environment as candidate required data into a trained multimodal obstacle perception model for deep learning and determining detection information of the obstacle in the current driving environment, the method further includes: Obtaining a verification instruction for the current driving environment; The initial demand data, the candidate demand data, the detection information and the verification instruction are input into the trained visual large language model for secondary language fusion learning to determine the target detection information after verification of the obstacle in the current driving environment, wherein the target detection information is used to represent the target response sentence including the verified obstacle category, the verified obstacle location information and the verified obstacle geometry information.

4. The method for detecting a vehicle driving environment according to claim 3, wherein: The verification instruction includes a right / wrong verification instruction and / or a supplementary verification instruction. The initial requirement data, the candidate requirement data, the detection information, and the verification instruction are input into a trained visual large language model for secondary language fusion learning to determine target detection information after verification of the obstacle in the current driving environment, including: Inputting the right / wrong verification instruction, the initial requirement data, and the candidate requirement data into a trained visual language model to determine whether the detection information of the obstacle has errors; If it does not exist, the supplementary verification instruction, the initial requirement data, the candidate requirement data, and the detection information are input into the trained visual large language model for secondary language fusion learning to determine the target detection information after verification of the obstacle in the current driving environment.

5. The method for detecting a vehicle driving environment according to claim 4, wherein: After inputting the right / wrong verification instruction, the initial requirement data, and the candidate requirement data into a trained visual language model to determine whether the detection information of the obstacle is erroneous, the method further includes: If there is an error in the detection information of the obstacle, the detection information is modified, and the modified detection information, the supplementary verification instruction, the initial requirement data and the candidate requirement data are input into the trained visual large language model for secondary language fusion learning to determine the target detection information after verification of the obstacle in the current driving environment.

6. The method for detecting a vehicle driving environment according to claim 3, wherein: The obtaining of a verification instruction for the current driving environment includes: Based on the user's verification requirement for the current driving environment, a verification instruction for the current driving environment is obtained.

7. The method for detecting a vehicle driving environment according to claim 1, wherein: The multi-dimensional image data of the vehicle in the current driving environment is obtained by the following methods: Acquire multi-dimensional initial image data of the vehicle in the current driving environment; De-noising is performed on the multi-dimensional initial image data to determine multi-dimensional image data of the vehicle in the current driving environment.

8. A vehicle driving environment detection device, characterized in that: The vehicle driving environment detection device includes: A first acquisition module is configured to acquire initial demand data of the vehicle in a current driving environment, wherein the initial demand data includes multi-dimensional image data, multi-view point cloud data, and a user's detection instruction for the current driving environment; A first determination module is configured to input the initial demand data into a trained visual language model for language fusion learning, and determine the types of obstacles present in the driving environment and the location information of the obstacles; The second determination module is configured to input the obstacle category, the location information of the obstacle, the multidimensional image data, and the original point cloud data of the vehicle in the current driving environment as candidate required data into a trained multimodal obstacle perception model for deep learning, thereby determining detection information of the obstacle in the current driving environment to complete detection of the current driving environment, wherein the detection information is used to represent a response statement including the obstacle category, the location information, and the geometric information of the obstacle.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via the bus, and the machine-readable instructions are executed by the processor to perform the steps of the vehicle driving environment detection method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for detecting a vehicle driving environment as described in any one of claims 1 to 7 are executed.