Vehicle surrounding environment analysis method, device, equipment and medium

By using vehicle trajectory points and on-board video to generate geographical distribution data of traffic targets, and using geoecological analysis models, the problem of difficult to understand the panoramic geographical ecology around the vehicle in the prior art is solved, and high-precision environmental analysis and safety supervision are achieved.

CN119938979APending Publication Date: 2025-05-06BEIJING TRANWISEWAY INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411978698.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively understand the panoramic geographical ecology around the vehicle, resulting in poor driving safety.

Method used

Route information is determined through the trajectory points reported by the vehicle, driving videos collected by the on-board equipment are obtained, and geographical distribution data of traffic targets is generated based on the route information, and a large-scale geographic analysis model is used to output the description text of the vehicle's surrounding environment.

Benefits of technology

It realizes high-precision analysis of the panoramic geographical ecology around the vehicle, and outputs environmental description text in natural language form to help users understand the surrounding environment of the vehicle and improve driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938979A_ABST
    Figure CN119938979A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle surrounding environment analysis method and device, equipment and a medium. The method comprises the steps that route information of a to-be-queried vehicle is determined according to track points reported by the to-be-queried vehicle; acquiring a driving video acquired by front vehicle-mounted equipment on the to-be-queried vehicle at the track point; generating traffic target geographical distribution data based on a driving video and the route information; and outputting a vehicle surrounding environment description text based on the traffic target geographic distribution data through a preset geographic ecological analysis large model. The method comprises the following steps: calculating route information of a vehicle by using track points, generating high-precision traffic target geographical distribution data by using a driving video collected at the current track point and combining the route information, and analyzing ecological information around the vehicle based on the traffic target geographical distribution data by using a geographical ecological analysis large model. And outputting a vehicle surrounding environment description text in a natural language form, so that a user can know the panoramic geographic ecology around the vehicle, and the safety supervision of the vehicle is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of monitoring technology, and in particular to a method, device, equipment and medium for analyzing the surrounding environment of a vehicle. Background Art

[0002] As the number of vehicles on the market increases year by year, there are more and more vehicles on the road. Usually, the traffic environment in which vehicles are driving is complicated, resulting in poor driving safety.

[0003] Based on this, an effective means is needed to understand the panoramic geographical and ecological conditions around the vehicle in order to monitor the safety of the vehicle. Summary of the invention

[0004] The purpose of this application is to propose a vehicle surrounding environment analysis method, device, equipment and medium to address the deficiencies of the above-mentioned prior art, and this purpose is achieved through the following technical solutions.

[0005] A first aspect of the present application provides a vehicle surrounding environment analysis method, the method comprising:

[0006] Determine the route information of the vehicle to be queried according to the track points reported by the vehicle to be queried;

[0007] Obtaining a driving video collected by a front-mounted vehicle device on the vehicle to be queried at the track point;

[0008] Based on the driving video and the route information, generating traffic target geographic distribution data for the vehicle to be queried;

[0009] Through the preset geographic ecological analysis model, a vehicle surrounding environment description text is output based on the traffic target geographic distribution data.

[0010] A second aspect of the present application provides a vehicle surrounding environment analysis device, the device comprising:

[0011] A route determination module, used to determine the route information of the vehicle to be queried according to the track points reported by the vehicle to be queried;

[0012] A video acquisition module, used to acquire the driving video collected by the front-mounted device on the vehicle to be queried at the track point;

[0013] A target generation module, used to generate traffic target geographic distribution data for the vehicle to be queried based on the driving video and the route information;

[0014] The analysis module is used to output a description text of the vehicle's surrounding environment based on the traffic target geographical distribution data through a preset geographical ecological analysis model.

[0015] The third aspect of the present application proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect above.

[0016] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the method described in the first aspect above.

[0017] Based on the above-mentioned vehicle surrounding environment analysis method, device, equipment and medium, the present application has at least the following beneficial effects or advantages:

[0018] The route information of the vehicle is calculated by using the trajectory points reported by the vehicle. The driving video collected by the front-mounted equipment on the vehicle at the current trajectory point and combined with the route information are used to generate high-precision traffic target geographic distribution data. Then, by using a large geographic ecological analysis model, the ecological information around the vehicle is analyzed based on the traffic target geographic distribution data, and a description text of the vehicle's surrounding environment is output in natural language, so that users can understand the panoramic geographic ecological conditions around the vehicle and achieve safe supervision of the vehicle.

[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0021] Figure 1 The present invention is a flow chart of an embodiment of a vehicle surrounding environment analysis method according to an exemplary embodiment;

[0022] Figure 2 A process for obtaining an instance segmentation model and a geographic ecological analysis model according to an exemplary embodiment is shown;

[0023] Figure 3 is a schematic structural diagram of a vehicle surrounding environment analysis device according to an exemplary embodiment;

[0024] Figure 4 is a schematic diagram of a hardware structure of an electronic device according to an exemplary embodiment;

[0025] Figure 5The figure is a schematic diagram of the structure of a storage medium according to an exemplary embodiment. DETAILED DESCRIPTION

[0026] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0027] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0028] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0029] In order to achieve panoramic ecological analysis of the vehicle's surroundings, this application uses the vehicle's trajectory points and on-board videos, combined with map matching technology and artificial intelligence technology of visual large models to generate basic geographic data of traffic targets, and uses a large language model to perform a panoramic analysis of the vehicle's surrounding ecology based on the basic geographic data of traffic targets, and outputs a text describing the vehicle's surroundings in natural language, so that users can understand the panoramic geographic ecological situation around the vehicle and achieve safe supervision of the vehicle.

[0030] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0031] Figure 1 The flowchart of an embodiment of a vehicle surrounding environment analysis method according to an exemplary embodiment includes the following steps 101 to 104:

[0032] Step 101: determining the route information of the vehicle to be queried according to the track points reported by the vehicle to be queried;

[0033] Step 102: Obtaining a driving video captured by a front-mounted vehicle device on the vehicle to be queried at the trajectory point;

[0034] Step 103: Generate traffic target geographic distribution data for the vehicle to be queried based on the driving video and route information;

[0035] Step 104: Output a description text of the vehicle's surrounding environment based on the traffic target geographic distribution data through a preset geographic ecological analysis model.

[0036] The track point can be the track point of the current time reported by the vehicle most recently, or it can be a track point of historical time. This application does not limit this. If it is a track point of the current time, it means that the user needs to perform a panoramic analysis of the surrounding environment where the vehicle is currently located. If it is a track point of historical time, it means that the user needs to perform a panoramic analysis of the surrounding environment where the vehicle was located at historical time. Among them, the data representing the track point usually includes information such as geographic location (i.e., longitude and latitude), speed, driving direction, time point, etc.

[0037] The route information may be content related to the road route where the track point is located, such as the road name and the geographical location of the vehicle on the road.

[0038] The route information of the vehicle to be queried is determined based on the trajectory points reported by the vehicle to be queried. The specific method is to use road network matching technology in combination with geographic data (such as maps, road network information, etc.) to map the trajectory points to the route of the road, thereby obtaining the route information of the vehicle. In this way, the trajectory points can be accurately matched to the actual road, avoiding the problem of deviation between the geographical location contained in the trajectory points and the actual road.

[0039] Driving video refers to the on-board video collected by the front-mounted device on the vehicle within a period of time including the time point corresponding to the trajectory point, which records the objects within the vehicle's field of view during the vehicle's driving process. For example, the on-board video within the time range between the trajectory point and the last reported trajectory point is used as the driving video at the trajectory point.

[0040] The traffic target geographic distribution data can be understood as the object distribution data on both sides of the vehicle, such as pedestrians, non-motor vehicles, traffic signs, roads, buildings and other objects.

[0041] The geographic ecological analysis big model is a pre-trained big language model for panoramic analysis, and its output is in the form of natural language. The big language model can adopt a transformer-decoder-based architecture.

[0042] At this point, the above Figure 1 The vehicle surrounding environment analysis process shown in the figure calculates the route information of the vehicle by using the trajectory points reported by the vehicle, and generates high-precision traffic target geographic distribution data by using the driving video collected by the front-mounted equipment on the vehicle at the current trajectory point and combining it with the route information. Then, by using a large geographic ecological analysis model, the ecological information around the vehicle is analyzed based on the geographic distribution data of the traffic target, and a description text of the vehicle's surrounding environment is output in the form of natural language, so that users can understand the panoramic geographic ecological situation around the vehicle and realize safe supervision of the vehicle.

[0043] In some embodiments of the present application, the process of step 103 may include:

[0044] Deduplication is performed on the images in the driving video. The object information contained in each picture frame in the deduplication driving video is identified through a preset instance segmentation model. The geographical location of the object is determined for the object information in each picture frame based on the route information. Based on the object information and geographical location of each picture frame, traffic target geographical distribution data is generated for the vehicle to be queried.

[0045] Deduplication processing refers to deduplication of images with high similarity in video data. For example, the similarity between any two adjacent image frames in a driving video is calculated. If the similarity exceeds a threshold, one of the frames is deleted.

[0046] The instance segmentation model is a large language model obtained through pre-training for object recognition, and the object information it outputs may include the boundary contour of the object in the picture frame, the object type, etc.

[0047] The object's geographic location refers to the object's geographic location in the picture frame, that is, the latitude and longitude position.

[0048] It should be noted here that for the same object appearing in multiple consecutive picture frames, if the object is a static object, then the geographical location of the object in each picture frame is theoretically consistent, but considering the calculation deviation, the geographical location of the object determined by the object information of the same object in each picture frame may not be completely consistent, but this does not affect the purpose of this application.

[0049] In this embodiment, duplicate images in the driving video are removed to reduce the amount of data that needs to be processed, and then the geographic location of the objects in each picture frame in the deduplicated driving video is determined using the route information, thereby generating high-precision traffic target geographic distribution data using the object information contained in each picture frame and the determined object geographic location.

[0050] In some embodiments of the present application, as mentioned above, the object information includes the object type and the boundary contour of the object in the picture frame. Based on this, the process of determining the geographic location of the object for the object information of each picture frame based on the route information may include:

[0051] For any two adjacent picture frames in each picture frame, the same object is searched from the two adjacent picture frames, and the area change rate of each identical object is calculated according to the boundary contours of each identical object found in the two adjacent picture frames. Then, the average value of the area change rate of each identical object is determined as the movement change ratio corresponding to the second picture frame in the two adjacent picture frames, and the geographical location of the object is determined for the object information of each picture frame according to the route information and the movement change ratio corresponding to each picture frame.

[0052] The same object means that the same object appears in two adjacent picture frames. The object has a corresponding boundary contour in the two adjacent picture frames. The pixel area occupied by the object in the picture can be obtained through the boundary contour. Specifically, the number of pixels contained in the rectangular frame of the boundary contour can be counted as the pixel area.

[0053] Regarding the process of searching for the same object from two adjacent picture frames, the specific method is: searching for the same object using a picture matching method.

[0054] The area change rate refers to the pixel area occupied by the object in the second picture frame divided by the pixel area occupied by the object in the first picture frame, and the time of the first picture frame is earlier than the time of the second picture frame.

[0055] The movement change ratio corresponding to the picture frame can represent the movement change of the object in the picture.

[0056] Exemplarily, assuming that the picture frames between two adjacent trajectory points are: T1, T2, ... Tn, then we can obtain n-1 groups of adjacent two picture frames: (T1, T2), (T2, T3) ... (T(n-1), Tn). For each group of adjacent two picture frames, the movement change ratio corresponding to the second picture frame can be obtained. It can be seen that the movement change ratios corresponding to T2, ... Tn are finally obtained. For the movement change ratio of the T1 picture frame, since it belongs to the initial picture frame, its corresponding movement change ratio is regarded as 0.

[0057] Specifically, for the process of determining the geographical location of an object for the object information of each picture frame based on the route information and the movement change ratio corresponding to each picture frame, the length of the route traveled by the vehicle during the process of collecting driving video is obtained according to the route information, and the geographical location of each picture frame on the route between two trajectory points is determined according to the movement change ratio corresponding to each picture frame and the obtained route length, and for each picture frame, the geographical location of the picture frame is determined as the object geographical location of the object information in the picture frame.

[0058] The route length may be determined based on the geographic location of the trajectory point within the time period corresponding to the driving video.

[0059] In this embodiment, by calculating the area change rate of each identical object in two adjacent picture frames, the area change rate can accurately reflect the movement change of the object, so that the movement change ratio of one picture frame between two adjacent picture frames can be accurately obtained by the average value of the area change rate of these identical objects, and the geographical location of the object information of each picture frame can be determined using the obtained movement change ratio of each picture frame.

[0060] It should be noted that the present application also includes the training process of the instance segmentation model and the geographic ecological analysis large model used above. The instance segmentation model and the geographic ecological analysis large model can be trained by collecting vehicle trajectory data and generating sample data through the previous driving video.

[0061] First, the training process for the instance segmentation model includes the following:

[0062] The driving trajectory of the vehicle and the driving video captured by the front-mounted device on the vehicle on the driving trajectory are collected, the picture frames contained in the driving video are deduplicated, a preset number of target pictures are screened from the deduplicated driving video, sample data are generated using the screened target pictures, and a preset visual model is trained using the sample data to obtain an instance segmentation model.

[0063] The target image is labeled to obtain sample data, which includes the target image, the boundary contour of the object in the target image, and the object type.

[0064] When screening a preset number of target images from the deduplicated driving video, the target images can be screened at fixed time intervals, or each screened target image can be compared for similarity with the previously screened target image, and images with a similarity lower than a certain value are retained. This application does not limit the specific screening rules, as long as the diversity of objects in the screened target images is ensured.

[0065] Then, the training process for the large model of geographic ecological analysis includes the following:

[0066] Based on preset geographic data, the geographic location of each trajectory point on the driving trajectory is converted to the geographic location of the route where the driving trajectory is located, and a mapping relationship is established between the driving trajectory and the trajectory points and picture frames with the same timestamp in the deduplicated driving video. The deduplicated driving video is segmented into multiple video segments using the picture frames with a mapping relationship in the deduplicated driving video as segmentation points. For each video segment, traffic target geographic distribution data is generated based on the video segment and the geographic locations of two trajectory points with a mapping relationship corresponding to the picture frames on both sides of the video segment. A description text is generated based on the traffic target geographic distribution data through an initial large language model, and the description text is modified into a standard description text. A sample is generated using the traffic target geographic distribution data and the standard description text, and the initial large language model is fine-tuned using the generated samples to obtain a large model for geographic ecological analysis.

[0067] The geographic data may be data such as maps and road network information, which converts the geographic location of each track point to the actual road.

[0068] It is understandable that the reporting frequency of trajectory points is usually lower than the acquisition frequency of picture frames. For example, a trajectory point is reported every 30 seconds, and a picture frame is acquired every 40 milliseconds. Therefore, by establishing a mapping relationship between trajectory points and picture frames, driving video segmentation can be achieved, and each video segment obtained by segmentation can be regarded as a video segment located between two adjacent trajectory points.

[0069] The description text generated using the initial large language model usually does not meet the actual needs of users. Therefore, the description text needs to be further modified to generate reasonable standard description text, so that the geographical distribution data of traffic targets can be used as the input data of the large language model and the standard description text can be used as label data.

[0070] In the traffic target geographic distribution data, the geographic location of the trajectory point and the object geographic location and object type of each object are recorded in Json format. For example, the traffic target geographic distribution data contains data such as buildings, pedestrians, vehicles, and sky brightness, and the generated standard description text is: "The current lighting is poor, the surrounding buildings are dense, the lanes are narrow, there are many pedestrians and other objects, and there is a greater safety risk."

[0071] Furthermore, for the generation process of the traffic target geographic distribution data of each video segment, the instance segmentation model can be used to identify the object information contained in each picture frame in the video segment, and then the object geographic location can be determined for the object information of each picture frame in the video segment based on the geographic location of two trajectory points corresponding to the picture frames on both sides of the video segment that have a mapping relationship, thereby generating the traffic target geographic distribution data based on the object information and the object geographic location of each picture frame.

[0072] The object information includes the object type and the boundary contour of the object in the picture frame.

[0073] Based on this, based on the geographical locations of two trajectory points corresponding to the picture frames on both sides of the video segment and having a mapping relationship, the geographical location of the object is determined for the object information of each picture frame in the video segment, including: for any two adjacent picture frames in the video segment, the same object is searched from the two adjacent picture frames, and the area change rate of each identical object is calculated according to the boundary contours of each identical object found in the two adjacent picture frames, and then the average value of the area change rate of each identical object is determined as the movement change ratio corresponding to the second picture frame in the two adjacent picture frames, and the geographical location of the object is determined for the object information of each picture frame according to the geographical locations of the two trajectory points corresponding to the picture frames on both sides of the video segment and the movement change ratio corresponding to each picture frame.

[0074] For the detailed description of the implementation process of determining the geographic location of an object for the object information of each picture frame in the video segment given above, please refer to the relevant description in the above model application stage, which will not be repeated here.

[0075] And, based on the object information and the object geographical location of each picture frame, traffic target geographical distribution data is generated, including: selecting target object information corresponding to an object whose area change rate is less than a preset value from the object information of each picture frame, and using the target object information and the object geographical location corresponding to the target object information to generate traffic target geographical distribution data.

[0076] Among them, an object with an area change rate less than a preset value indicates that it is gradually disappearing from the vehicle's field of view, and the preset value may be 1. In other words, if the area change rate is less than 1, it means that the area occupied by the object in the picture frame is gradually getting smaller.

[0077] And, according to the geographical locations of two trajectory points corresponding to the picture frames on both sides of the video segment and having a mapping relationship and the movement change ratio corresponding to each picture frame, the geographical location of the object is determined for the object information of each picture frame, including:

[0078] According to the geographical locations of two trajectory points corresponding to the picture frames on both sides of the video segment and having a mapping relationship, the route length between the two trajectory points is determined; according to the movement change ratio and route length corresponding to each picture frame, the geographical location of each picture frame on the route between the two trajectory points is determined; for each picture frame, the geographical location of the picture frame is determined as the object geographical location of the object information in the picture frame.

[0079] Since the geographical location of a track point is its position on the route, the route length between the track points can be obtained from the geographical locations of two track points.

[0080] Exemplarily, assuming that the movement change ratios of all picture frames between two adjacent trajectory points are r1, r2...rn, and the sum of these movement change ratios is Sr, and the route length between the two trajectory points is set to L, then the geographical location of each picture frame is a point on the route along the route direction with a length of L*r1 / Sr, L*r2 / Sr..., L*rn / Sr, and the longitude and latitude of these points can be used as the geographical location of each picture frame.

[0081] Based on the above description of the training process, Figure 2 The present invention is a flow chart of obtaining an instance segmentation model and a geographic ecological analysis model according to an exemplary embodiment, and includes the following processing steps:

[0082] 1. Obtain driving trajectory and driving video from vehicles equipped with trajectory monitoring equipment and front-mounted vehicle equipment.

[0083] 2. Combined with geographic data, the track points on the driving trajectory are converted to the longitude and latitude of the actual driving route through road network matching technology.

[0084] 3. De-duplicate the image frames in the driving video and establish a mapping relationship between the image frames and trajectory points with the same timestamp.

[0085] 4. Filter a certain number of image frames from the deduplicated driving video, and establish sample data of image-object boundary contour-object type through manual labeling. Based on the large model fine-tuning technology, use the sample data to train and obtain the instance segmentation model.

[0086] 5. Use the instance segmentation model to perform instance segmentation on each picture frame in the deduplicated driving video to determine the object type and boundary contour of the object in each picture frame.

[0087] 6. Using the image frames that have a mapping relationship with the trajectory points as segmentation points, the deduplicated driving video is segmented into multiple video segments.

[0088] 7. For each video segment, according to the boundary contours of the objects in the two adjacent frames, calculate the area change rate of each identical object in the two adjacent frames, and take the average of the area change rates of all identical objects as the movement change ratio of the second frame in the two adjacent frames. Let the movement change ratios of all frames in the video segment be r1, r2…rn, and their sum be Sr. According to the positioning of the trajectory points corresponding to the frames on both sides of the video segment on the route, take the route between the trajectory points, and set its total length as L. Then the position of each frame is the point on the route with a length of L*r1 / Sr, L*r2 / Sr…, L*rn / Sr along the route direction. The longitude and latitude of these points are taken as the longitude and latitude of the position of each frame, that is, the geographical location of the object in the frame is the longitude and latitude of the position of the frame, and then select the target object with an area change rate <1 in each frame (indicates that the object is on both sides of the current position and begins to disappear from the vehicle's field of view). Use the geographical location and object type of the target object to generate traffic target geographical distribution data.

[0089] 8. Analyze the geographical distribution data of traffic targets through the initial large language model to generate description text, modify the description text through manual annotation, and generate reasonable standard description text.

[0090] 9. Fine-tune the initial large language model using standard description text and traffic target geographic distribution data to obtain a large geographic ecological analysis model for panoramic analysis.

[0091] Corresponding to the above-mentioned embodiment of the vehicle surrounding environment analysis method, the present application also provides an embodiment of a vehicle surrounding environment analysis device.

[0092] Figure 3 FIG. 1 is a schematic diagram of a vehicle surrounding environment analysis device according to an exemplary embodiment. The device is used to execute the vehicle surrounding environment analysis method provided in any of the above embodiments. Figure 3 As shown, the vehicle surrounding environment analysis device includes:

[0093] A route determination module 310 is used to determine the route information of the vehicle to be queried according to the track points reported by the vehicle to be queried;

[0094] The video acquisition module 320 is used to acquire the driving video collected by the front-mounted device of the vehicle to be queried at the track point;

[0095] A target generation module 330, for generating traffic target geographic distribution data for the vehicle to be queried based on the driving video and the route information;

[0096] The analysis module 340 is used to output a description text of the vehicle surrounding environment based on the traffic target geographical distribution data through a preset geographical ecological analysis model.

[0097] In an optional implementation, the target generation module 330 is specifically used to deduplicate the images in the driving video; identify the object information contained in each picture frame in the deduplicated driving video through a preset instance segmentation model; determine the geographic location of the object for the object information in each picture frame based on the route information; and generate traffic target geographic distribution data for the vehicle to be queried based on the object information and the object geographic location in each picture frame.

[0098] In an optional implementation, the object information includes the object type and the boundary contour of the object in the picture frame; the target generation module 330 is specifically used to search for the same object from any two adjacent picture frames in each picture frame during the process of determining the geographical location of the object for the object information of each picture frame based on the route information; calculate the area change rate of each identical object based on the boundary contours of each identical object found in the two adjacent picture frames; determine the average value of the area change rate of each identical object as the movement change ratio corresponding to the second picture frame in the two adjacent picture frames; and determine the geographical location of the object for the object information of each picture frame based on the route information and the movement change ratio corresponding to each picture frame.

[0099] In an optional implementation, the device further includes ( Figure 3 Not shown):

[0100] A segmentation model training module is used to collect the driving trajectory of the vehicle and the driving video collected by the front-mounted device on the vehicle on the driving trajectory; and perform deduplication processing on the picture frames contained in the driving video;

[0101] A preset number of target images are screened from the deduplicated driving video; sample data are generated using the screened target images, wherein the sample data includes the target images, boundary contours of objects in the target images, and object types; and a preset visual model is trained using the sample data to obtain the instance segmentation model.

[0102] In an optional implementation, the device further includes ( Figure 3 Not shown):

[0103] The large model training module is used to convert the geographical location of each track point on the driving trajectory to the geographical location of the route where the driving trajectory is located based on preset geographical data; establish a mapping relationship between the driving trajectory and the track points and picture frames with the same timestamp in the deduplicated driving video; use the picture frames in the deduplicated driving video that have the mapping relationship as segmentation points to segment the deduplicated driving video into multiple video segments; for each video segment, generate traffic target geographical distribution data based on the video segment and the geographical locations of two track points that have a mapping relationship corresponding to the picture frames on both sides of the video segment, generate a description text based on the traffic target geographical distribution data through an initial large language model, modify the description text into a standard description text, and generate a sample using the traffic target geographical distribution data and the standard description text; use the generated samples to fine-tune the initial large language model to obtain the geographical ecological analysis large model.

[0104] In an optional implementation, the large model training module is specifically used to identify the object information contained in each picture frame in the video segment through the instance segmentation model during the process of generating the geographical distribution data of the traffic target based on the video segment and the geographical locations of two trajectory points corresponding to the picture frames on both sides of the video segment and having a mapping relationship; determine the object geographical location for the object information of each picture frame in the video segment based on the geographical locations of the two trajectory points corresponding to the picture frames on both sides of the video segment and having a mapping relationship; and generate the geographical distribution data of the traffic target according to the object information and the object geographical location of each picture frame.

[0105] In an optional implementation, the object information includes the object type and the boundary contour of the object in the picture frame; the large model training module is specifically used to search for the same object from any two adjacent picture frames in the video segment in the process of determining the object's geographical location for the object information of each picture frame in the video segment based on the geographical locations of two trajectory points corresponding to the picture frames on both sides of the video segment and having a mapping relationship between the picture frames on both sides of the video segment; calculate the area change rate of each identical object based on the boundary contours of each identical object found in the two adjacent picture frames; determine the average value of the area change rate of each identical object as the movement change ratio corresponding to the second picture frame in the two adjacent picture frames; determine the object's geographical location for the object information of each picture frame based on the geographical locations of the two trajectory points corresponding to the picture frames on both sides of the video segment and the movement change ratio corresponding to each picture frame.

[0106] In an optional implementation, the large model training module is specifically used to select target object information corresponding to an object whose area change rate is less than a preset value from the object information of each picture frame in the process of generating the traffic target geographic distribution data based on the object information and the object geographic location of each picture frame; and generate the traffic target geographic distribution data using the target object information and the object geographic location corresponding to the target object information.

[0107] In an optional implementation, the large model training module is specifically used to determine the geographical location of the object for the object information of each picture frame according to the geographical locations of the two trajectory points corresponding to the picture frames on both sides of the video segment and the movement change ratio corresponding to each picture frame. The length of the route between the two trajectory points is determined according to the geographical locations of the two trajectory points corresponding to the picture frames on both sides of the video segment; the geographical location of each picture frame on the route between the two trajectory points is determined according to the movement change ratio corresponding to each picture frame and the route length; for each picture frame, the geographical location of the picture frame is determined as the object geographical location of the object information in the picture frame.

[0108] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0109] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0110] The embodiment of the present application also provides an electronic device corresponding to the vehicle surrounding environment analysis method provided in the above embodiment, so as to execute the above vehicle surrounding environment analysis method.

[0111] Figure 46 is a hardware structure diagram of an electronic device according to an exemplary embodiment, the electronic device includes: a communication interface 601, a processor 602, a memory 603 and a bus 604; wherein the communication interface 601, the processor 602 and the memory 603 communicate with each other through the bus 604. The processor 602 can execute the vehicle surrounding environment analysis method described above by reading and executing the machine executable instructions corresponding to the control logic of the vehicle surrounding environment analysis method in the memory 603. The specific content of the method is referred to the above embodiment, and will not be repeated here.

[0112] The memory 603 mentioned in this application can be any electronic, magnetic, optical or other physical storage device, and can contain storage information, such as executable instructions, data, etc. Specifically, the memory 603 can be RAM (Random Access Memory), flash memory, storage drive (such as hard disk drive), any type of storage disk (such as optical disk, DVD, etc.), or similar storage medium, or a combination thereof. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 601 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0113] The bus 604 may be an ISA bus, a PCI bus or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 603 is used to store programs, and the processor 602 executes the programs after receiving execution instructions.

[0114] Processor 602 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in processor 602 or an instruction in software form. The above-mentioned processor 602 can be a general-purpose processor, including a network processor (Network Processor, referred to as NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a readily available programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor are combined to execute.

[0115] The electronic device provided in the embodiment of the present application and the vehicle surrounding environment analysis method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented therein.

[0116] The present application also provides a computer-readable storage medium corresponding to the vehicle surrounding environment analysis method provided in the above embodiment. Figure 5 As shown, the computer-readable storage medium is a CD 30 on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, the vehicle surrounding environment analysis method provided by any of the aforementioned embodiments will be executed.

[0117] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.

[0118] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the vehicle surrounding environment analysis method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0119] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0120] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0121] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A vehicle surrounding environment analysis method, characterized in that: The method comprises: Determine the route information of the vehicle to be queried according to the track points reported by the vehicle to be queried; Obtaining a driving video collected by a front-mounted vehicle device on the vehicle to be queried at the track point; Based on the driving video and the route information, generating traffic target geographic distribution data for the vehicle to be queried; Through the preset geographic ecological analysis model, a vehicle surrounding environment description text is output based on the traffic target geographic distribution data.

2. The method according to claim 1, characterized in that The generating of traffic target geographic distribution data for the vehicle to be queried based on the driving video and the route information includes: Deduplication processing is performed on the pictures in the driving video; Using a preset instance segmentation model, identify the object information contained in each frame of the deduplicated driving video; Determining the geographic location of the object for the object information of each picture frame based on the route information; According to the object information and the object geographical location of each picture frame, traffic target geographical distribution data is generated for the vehicle to be queried.

3. The method according to claim 2, characterized in that The object information includes the object type and the boundary contour of the object in the picture frame; The determining the geographic location of the object for the object information of each picture frame based on the route information includes: For any two adjacent picture frames in each picture frame, searching for the same object from the two adjacent picture frames; According to the boundary contours of each identical object found in two adjacent picture frames, the area change rate of each identical object is calculated; Determine the average value of the area change rate of each identical object as the movement change ratio corresponding to the second picture frame of the two adjacent picture frames; The geographical location of the object is determined for the object information of each picture frame according to the route information and the movement change ratio corresponding to each picture frame.

4. The method according to claim 2, characterized in that: The method also includes a training process of the instance segmentation model: Collecting the driving trajectory of the vehicle and the driving video collected by the front-mounted device on the vehicle on the driving trajectory; Deduplication processing is performed on the picture frames contained in the driving video; Filter a preset number of target images from the deduplicated driving video; Generate sample data using the screened target image, the sample data including the target image, the boundary contour and the object type of the object in the target image; The sample data is used to train a preset visual model to obtain the instance segmentation model.

5. The method according to claim 4, characterized in that The method also includes a training process of the geographic ecological analysis large model: Based on preset geographical data, converting the geographical location of each track point on the driving track to the geographical location of the route where the driving track is located; Establishing a mapping relationship between the driving trajectory and the trajectory points and picture frames with the same timestamp in the deduplicated driving video; Using the picture frames in the deduplicated driving video that have the mapping relationship as segmentation points, the deduplicated driving video is segmented into a plurality of video segments; For each video segment, traffic target geographic distribution data is generated based on the video segment and the geographic locations of two trajectory points corresponding to the picture frames on both sides of the video segment and having a mapping relationship, a description text is generated based on the traffic target geographic distribution data through an initial large language model, the description text is modified into a standard description text, and a sample is generated using the traffic target geographic distribution data and the standard description text; The initial large language model is fine-tuned using the generated samples to obtain the large geographic ecological analysis model.

6. The method according to claim 5, characterized in that The generating of the geographical distribution data of the traffic target based on the video segment and the geographical locations of two trajectory points corresponding to the picture frames on both sides of the video segment and having a mapping relationship includes: Using the instance segmentation model, identifying object information contained in each picture frame in the video segment; Determine the geographical location of the object for the object information of each picture frame in the video segment based on the geographical locations of two trajectory points that have a mapping relationship corresponding to the picture frames on both sides of the video segment; The traffic target geographical distribution data is generated according to the object information and the object geographical location of each picture frame.

7. The method according to claim 6, characterized in that The object information includes the object type and the boundary contour of the object in the picture frame; the determining the object geographic location for the object information of each picture frame in the video segment based on the geographic locations of two trajectory points corresponding to the picture frames on both sides of the video segment and having a mapping relationship includes: For any two adjacent picture frames in the video segment, searching for the same object from the two adjacent picture frames; According to the boundary contours of each identical object found in two adjacent picture frames, the area change rate of each identical object is calculated; Determine the average value of the area change rate of each identical object as the movement change ratio corresponding to the second picture frame of the two adjacent picture frames; According to the geographical locations of two trajectory points corresponding to the picture frames at both sides of the video segment and having a mapping relationship and the movement change ratio corresponding to each picture frame, the geographical location of the object is determined for the object information of each picture frame.

8. The method according to claim 7, characterized in that The generating of the traffic target geographical distribution data according to the object information and the object geographical location of each picture frame includes: Selecting target object information corresponding to an object whose area change rate is less than a preset value from the object information of each picture frame; The traffic target geographical distribution data is generated by using the target object information and the object geographical location corresponding to the target object information.

9. The method according to claim 7, characterized in that: The determining the geographical location of the object for the object information of each picture frame according to the geographical locations of the two trajectory points corresponding to the picture frames at both sides of the video segment and the movement change ratio corresponding to each picture frame includes: Determine the route length between the two trajectory points according to the geographical locations of the two trajectory points corresponding to the image frames on both sides of the video segment and having a mapping relationship; Determine the geographical location of each picture frame on the route between the two track points according to the movement change ratio corresponding to each picture frame and the length of the route; For each picture frame, the geographical location of the picture frame is determined as the object geographical location of the object information in the picture frame.

10. A vehicle surrounding environment analysis device, characterized in that: The device comprises: A route determination module, used to determine the route information of the vehicle to be queried according to the track points reported by the vehicle to be queried; A video acquisition module, used to acquire the driving video collected by the front-mounted device on the vehicle to be queried at the track point; A target generation module, used to generate traffic target geographic distribution data for the vehicle to be queried based on the driving video and the route information; The analysis module is used to output a description text of the vehicle's surrounding environment based on the traffic target geographical distribution data through a preset geographical ecological analysis model.