A method and system for prompting the opening of vehicle doors and windows after a vehicle collision

The vehicle's camera collects driving surveillance video, uses target detection and ViT model to determine whether doors and windows can be opened safely after the vehicle crash, solving the problem of judging the risk of secondary collision after the vehicle crash, and ensuring the safety of vehicle personnel.

CN116386005BActive Publication Date: 2025-07-11CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310352379.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-04
Publication Date
2025-07-11
Estimated Expiration
2043-04-04

AI Technical Summary

Technical Problem

After a vehicle collision, rashly opening the doors and windows may lead to a secondary collision and cause further damage. It is difficult for the prior art to accurately determine whether the doors and windows can be safely opened.

Method used

The vehicle's camera collects driving surveillance video, uses the vehicle target detection network and ViT model to extract the vehicle target's area of interest and semantic feature vectors, and combines the context encoder and classifier to determine whether doors and windows can be opened safely.

Benefits of technology

It realizes an accurate judgment on whether a secondary collision occurs when opening doors and windows, and prompts users through buzzers or vibrations to ensure the safety of vehicle personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386005B_ABST
    Figure CN116386005B_ABST
Patent Text Reader

Abstract

Disclosed are a method and system for prompting the opening of vehicle doors and windows after a vehicle collision. First, a plurality of key frames of driving monitoring are extracted from a driving monitoring video and respectively passed through a vehicle target detection network to obtain a plurality of regions of interest of vehicle targets. Then, the plurality of regions of interest of vehicle targets are respectively passed through a ViT model to obtain a plurality of semantic feature vectors of interest of vehicle targets. Next, the plurality of semantic feature vectors of interest of vehicle targets are passed through a context encoder to obtain classification feature vectors. Finally, the classification feature vectors are passed through a classifier to obtain a classification result for indicating whether the doors and windows can be opened. In this way, the risk of secondary collision that may occur when opening the doors and windows can be accurately judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent control, and more specifically, to a method and system for prompting the opening of vehicle doors and windows after a vehicle collision. Background Art

[0002] With the development of the economy, more and more vehicles appear in people's lives. Vehicle collisions not only damage the vehicles but also pose a threat to the lives of the people inside the vehicles. After a vehicle collision, the first reaction of the vehicle owner is to get out of the vehicle. However, after a collision, if the doors and windows are opened rashly, there may be a secondary collision, causing further injuries.

[0003] Therefore, a more intelligent solution for prompting the opening of vehicle doors and windows after a collision is expected. Summary of the Invention

[0004] To solve the above technical problems, the present application is proposed. Embodiments of the present application provide a method and system for prompting the opening of vehicle doors and windows after a vehicle collision. First, a plurality of key frames of the driving monitoring video are extracted from the driving monitoring video of the surrounding vehicles and respectively passed through a vehicle target detection network to obtain a plurality of regions of interest of vehicle targets. Then, the plurality of regions of interest of vehicle targets are respectively passed through a ViT model to obtain a plurality of semantic feature vectors of vehicle targets of interest. Then, the plurality of semantic feature vectors of vehicle targets of interest are passed through a context encoder to obtain a classification feature vector. Finally, the classification feature vector is passed through a classifier to obtain a classification result indicating whether the doors and windows can be opened. In this way, the risk of a secondary collision when opening the doors and windows can be accurately judged.

[0005] According to one aspect of the present application, there is provided a method for prompting the opening of vehicle doors and windows after a vehicle collision, including:

[0006] Obtaining a driving monitoring video of surrounding vehicles collected by a camera deployed on the vehicle;

[0007] Extracting a plurality of key frames of the driving monitoring video from the driving monitoring video;

[0008] Respectively passing the plurality of key frames of the driving monitoring video through a vehicle target detection network to obtain a plurality of regions of interest of vehicle targets;

[0009] Respectively passing the plurality of regions of interest of vehicle targets through a ViT model including an image patch embedding layer to obtain a plurality of semantic feature vectors of vehicle targets of interest;

[0010] Passing the plurality of semantic feature vectors of vehicle targets of interest through a context encoder based on a transformer to obtain a classification feature vector; and

[0011] Pass the classified feature vector through a classifier to obtain a classification result, where the classification result is used to indicate whether the doors and windows can be opened.

[0012] In the above method for prompting the opening of doors and windows after a vehicle collision, extracting a plurality of key frames of driving monitoring from the driving monitoring video includes:

[0013] Sample key frames from the driving monitoring video at a predetermined sampling frequency to extract a plurality of key frames of driving monitoring from the driving monitoring video.

[0014] In the above method for prompting the opening of doors and windows after a vehicle collision, the vehicle target detection network is an anchor window-based target detection network, and the anchor window-based target detection network is Fast R-CNN, Faster R-CNN, or RetinaNet.

[0015] In the above method for prompting the opening of doors and windows after a vehicle collision, passing the plurality of vehicle target regions of interest through a ViT model including an image patch embedding layer to obtain a plurality of vehicle target region of interest semantic feature vectors includes:

[0016] Use the image patch embedding layer to embed each vehicle target region of interest in the plurality of vehicle target regions of interest to obtain a plurality of vehicle target region of interest vectors;

[0017] Based on the plurality of vehicle target regions of interest, enhance the feature distribution in the channel dimension of the plurality of vehicle target region of interest vectors to obtain a plurality of optimized vehicle target region of interest vectors;

[0018] Pass each of the optimized vehicle target region of interest vectors through the ViT model to obtain a plurality of context vehicle target region of interest vectors; and

[0019] Concatenate the plurality of context vehicle target region of interest vectors to obtain the plurality of vehicle target region of interest semantic feature vectors.

[0020] In the above method for prompting the opening of doors and windows after a vehicle collision, passing the plurality of vehicle target region of interest semantic feature vectors through a transformer-based context encoder to obtain a classified feature vector includes:

[0021] Input the plurality of vehicle target region of interest semantic feature vectors into the transformer-based context encoder to obtain a plurality of context vehicle target region of interest semantic feature vectors;

[0022] Calculate the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors, where the mean vector of the Gaussian mixture model is the position-wise mean vector of the multiple context vehicle target interested semantic feature vectors, and the value at each position in the covariance matrix of the Gaussian mixture model is the variance between the eigenvalues of the corresponding two positions in the position-wise mean vector;

[0023] Calculate the Gaussian probability density distribution distance index between each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors and the Gaussian mixture model respectively to obtain multiple Gaussian probability density distribution distance indexes;

[0024] Use the multiple Gaussian probability density distribution distance indexes as weights to weight the multiple context vehicle target interested semantic feature vectors to obtain multiple optimized context vehicle target interested semantic feature vectors; and

[0025] Concatenate the multiple optimized context vehicle target interested semantic feature vectors to obtain the classification feature vector.

[0026] In the above vehicle door and window opening prompt method after a vehicle collision, inputting the multiple vehicle target interested semantic feature vectors into the Transformer-based context encoder to obtain multiple context vehicle target interested semantic feature vectors includes:

[0027] Arrange the multiple vehicle target interested semantic feature vectors in a one-dimensional manner to obtain a first global vehicle target interested feature vector;

[0028] Calculate the product between the first global vehicle target interested feature vector and the transposed vectors of each vehicle target interested semantic feature vector in the multiple vehicle target interested semantic feature vectors to obtain multiple first self-attention correlation matrices;

[0029] Normalize each first self-attention correlation matrix in the multiple first self-attention correlation matrices respectively to obtain multiple first normalized self-attention correlation matrices;

[0030] Pass each first normalized self-attention correlation matrix in the multiple first normalized self-attention correlation matrices through the Softmax classification function to obtain multiple first probability values; and

[0031] Respectively use each first probability value in the multiple first probability values as a weight to weight each vehicle target interested semantic feature vector in the multiple vehicle target interested semantic feature vectors to obtain the multiple context vehicle target interested semantic feature vectors.

[0032] In the above-mentioned method for prompting the opening of vehicle doors and windows after a vehicle collision, calculating the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors includes:

[0033] Calculating the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors according to the following Gaussian mixture formula;

[0034] Among them, the Gaussian mixture formula is:

[0035]

[0036] Among them, μ represents the position-wise mean vector among the multiple context vehicle target interested semantic feature vectors, and the value of each position of σ represents the variance between the feature values of each position in the multiple context vehicle target interested semantic feature vectors.

[0037] In the above-mentioned method for prompting the opening of vehicle doors and windows after a vehicle collision, calculating the Gaussian probability density distribution distance exponent between each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors and the Gaussian mixture model to obtain a plurality of Gaussian probability density distribution distance exponents includes:

[0038] Calculating the Gaussian probability density distribution distance exponent between each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors and the Gaussian mixture model according to the following optimization formula to obtain the plurality of Gaussian probability density distribution distance exponents;

[0039] Among them, the optimization formula is:

[0040]

[0041] Among them, V i is the i-th context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors, (·) T represents the transposed vector of the vector, μ u and Σ u are the mean vector and covariance matrix of the Gaussian mixture model. The multiple context vehicle target interested semantic feature vectors and the mean vector of the Gaussian mixture model are in the form of column vectors. exp(·) represents the exponential operation of the matrix, and the exponential operation of the matrix represents the natural exponential function value with the feature values of each position in the matrix as the power, represents position-wise subtraction, represents matrix multiplication, w i represents the i-th Gaussian probability density distribution distance exponent in the plurality of Gaussian probability density distribution distance exponents.

[0042] In the above method for prompting the opening of vehicle doors and windows after a collision, the classification feature vector is passed through a classifier to obtain a classification result, and the classification result is used to indicate whether the doors and windows can be opened, including:

[0043] Using the fully connected layer of the classifier to perform fully connected encoding on the classification feature vector to obtain an encoded classification feature vector; and

[0044] Inputting the encoded classification feature vector into the Softmax classification function of the classifier to obtain the classification result.

[0045] According to another aspect of the present application, a system for prompting the opening of vehicle doors and windows after a collision is provided, which includes:

[0046] A video acquisition module for acquiring a driving monitoring video of surrounding vehicles collected by a camera deployed on the vehicle;

[0047] A key frame extraction module for extracting a plurality of driving monitoring key frames from the driving monitoring video;

[0048] A vehicle target detection module for passing the plurality of driving monitoring key frames through a vehicle target detection network respectively to obtain a plurality of vehicle target regions of interest;

[0049] An image encoding module for passing the plurality of vehicle target regions of interest through a ViT model including an image patch embedding layer respectively to obtain a plurality of vehicle target region of interest semantic feature vectors;

[0050] A context encoding module for passing the plurality of vehicle target region of interest semantic feature vectors through a context encoder based on a transformer to obtain a classification feature vector; and

[0051] A classification module for passing the classification feature vector through a classifier to obtain a classification result, and the classification result is used to indicate whether the doors and windows can be opened.

[0052] Compared with the prior art, the method and system for prompting the opening of vehicle doors and windows after a collision provided by the present application first pass a plurality of driving monitoring key frames extracted from the driving monitoring video through a vehicle target detection network respectively to obtain a plurality of vehicle target regions of interest, then pass the plurality of vehicle target regions of interest through a ViT model respectively to obtain a plurality of vehicle target region of interest semantic feature vectors, then pass the plurality of vehicle target region of interest semantic feature vectors through a context encoder to obtain a classification feature vector, and finally pass the classification feature vector through a classifier to obtain a classification result for indicating whether the doors and windows can be opened. In this way, the risk of secondary collision when opening the doors and windows can be accurately judged. Description of the Drawings

[0053] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings. The following accompanying drawings are not deliberately drawn to scale in actual size, and the focus is on showing the gist of the present application.

[0054] Figure 1 It is an application scenario diagram of the method for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application.

[0055] Figure 2 It is a flowchart of the method for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application.

[0056] Figure 3 It is a schematic diagram of the architecture of the method for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application.

[0057] Figure 4 It is a flowchart of sub-step S140 of the method for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application.

[0058] Figure 5 It is a flowchart of sub-step S150 of the method for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application.

[0059] Figure 6 It is a flowchart of sub-step S151 of the method for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application.

[0060] Figure 7 It is a flowchart of sub-step S160 of the method for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application.

[0061] Figure 8 It is a block diagram of the system for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application. Detailed implementation manners

[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all of them. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts also fall within the scope of protection of the present application.

[0063] As shown in this application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specific to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0064] Although this application makes various references to certain modules in the system according to the embodiments of this application, however, any number of different modules can be used and run on the user terminal and / or the server. The modules are only illustrative, and different aspects of the system and method can use different modules.

[0065] Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. On the contrary, according to needs, various steps can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or more steps can be removed from these processes.

[0066] Next, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the exemplary embodiments described herein.

[0067] As described above, after a vehicle collision, the first reaction of the vehicle owner is to get out of the vehicle. However, after a collision, if the doors and windows are opened rashly, there may be a secondary collision, causing further injuries. Therefore, a more intelligent scheme for prompting to open the doors and windows after a vehicle collision is expected.

[0068] Accordingly, considering that it is particularly important to analyze the surrounding vehicles to determine whether a secondary collision will occur during the actual process of opening the doors and windows after a vehicle collision, in the technical solution of this application, it is expected to collect the driving surveillance video of the surrounding vehicles through a camera installed on the side of the vehicle, and to judge the risk of a secondary collision when opening the doors and windows from the driving surveillance video of the vehicle. However, since there is a large amount of information in the driving surveillance video of the surrounding vehicles, and the vehicle has different speeds during driving on different sections of the road, and the reactions of the drivers of the surrounding vehicles when seeing a collision accident are also different, it is difficult to capture and extract the driving characteristics of the surrounding vehicles in real time and accurately. Therefore, in this process, the difficulty lies in how to fully and accurately extract the implicit feature distribution information about the vehicle driving in the driving surveillance video of the surrounding vehicles, so as to accurately judge the risk of a secondary collision when opening the doors and windows, and to prompt the user whether the doors and windows can be opened by means of a buzzer or vibration to avoid a secondary collision and ensure the safety of the vehicle occupants.

[0069] The development of deep learning and neural networks provides a new solution idea and plan for extracting the implicit feature distribution information about the vehicle driving in the driving surveillance video of the surrounding vehicles.

[0070] Specifically, in the technical solution of this application, first, the driving surveillance video of the surrounding vehicles is collected through a camera deployed on the vehicle. Then, considering that in the driving surveillance video of the surrounding vehicles, the driving change characteristics of the surrounding vehicles can be represented by the difference between adjacent surveillance frames in the driving surveillance video of the surrounding vehicles, that is, the driving state change of the surrounding vehicles is represented by the image representation of adjacent image frames. However, considering that the difference between adjacent frames in the surveillance video is small and there is a large amount of data redundancy, in order to reduce the computational amount and avoid the adverse effects of data redundancy on detection, the driving surveillance video of the surrounding vehicles is sampled at a predetermined sampling frequency to extract multiple driving surveillance key frames from the driving surveillance video.

[0071] Then, considering that when detecting the driving state of the surrounding vehicles, the hidden features such as the driving speed and distance of the surrounding vehicle targets should be focused on. However, since there is a large amount of information in the driving monitoring key frames, it is difficult to capture and extract the effective information of the vehicle targets. Therefore, considering that if the useless interference feature information can be filtered out when mining the driving state features of the surrounding vehicles, it is obvious that the accuracy of monitoring the driving state of the surrounding vehicles can be improved. Based on this, in the technical solution of this application, the multiple driving monitoring key frames are further respectively passed through a vehicle target detection network to obtain multiple regions of interest (ROIs) of vehicle targets. Specifically, the target anchoring layer of the vehicle target detection network is used to slide with the anchor box B to process the multiple driving monitoring key frames respectively, so as to frame the regions of interest of the vehicle targets, thereby obtaining the multiple regions of interest of vehicle targets. In this way, on the one hand, the interference of useless information in the background can be eliminated, and on the other hand, the vehicle targets can be focused on during subsequent feature extraction. For example, the change of speed features and distance features can be inferred through the change of the size of the vehicle target in time series. In particular, here, the vehicle target detection network is an anchor window-based target detection network, and the anchor window-based target detection network is Fast R-CNN, Faster R-CNN or RetinaNet.

[0072] Furthermore, a convolutional neural network model with excellent performance in implicit feature extraction of images can be used to mine the features of each region of interest of vehicle targets. However, due to the inherent limitations of convolutional operations, it is difficult for the pure CNN method to learn explicit global and long-range semantic information interactions. Therefore, in the technical solution of this application, the multiple regions of interest of vehicle targets are respectively encoded through a ViT model containing an image patch embedding layer to extract the implicit associated semantic features regarding the driving state of each vehicle target in each region of interest of vehicle targets, so as to obtain multiple semantic feature vectors of regions of interest of vehicle targets. In particular, here, the implementation process of embedding can be to first arrange the pixel values of all pixel positions of the image patches regarding the vehicle targets in each region of interest of vehicle targets into a one-dimensional vector, and then use a fully connected layer to perform fully connected encoding on this one-dimensional vector to achieve embedding. It should be understood that ViT can directly process the regions of interest of vehicle targets through the self-attention mechanism like Transformer, so as to respectively extract the implicit associated semantic feature information based on the global driving state of each vehicle target in each region of interest of vehicle targets.

[0073] Then, considering that the implicit correlation features regarding the driving states of the respective vehicle targets in the respective vehicle target regions of interest have temporal dynamic change feature information in the time dimension. Therefore, in order to fully extract the temporal dynamic change features of the driving states of surrounding vehicles and thereby accurately capture the speed features and distance features of surrounding vehicles, in the technical solution of this application, the multiple vehicle target region-of-interest semantic feature vectors are further encoded through a Transformer-based context encoder to extract the form state correlation features regarding the respective vehicle targets in each driving monitoring key frame based on the temporal global dynamic correlation feature distribution information, that is, the temporal dynamic feature information of the driving states of surrounding vehicles, thereby obtaining classification feature vectors.

[0074] Next, the classification feature vectors are further classified through a classifier to obtain a classification result indicating whether the doors and windows can be opened. That is, in the technical solution of this application, the labels of the classifier include that the doors and windows can be opened and that the doors and windows cannot be opened. Among them, the classifier determines which classification label the classification feature vector belongs to through a softmax function. It should be understood that in the technical solution of this application, the classification label of the classifier is the control strategy label for whether the doors and windows can be opened. Therefore, after obtaining the classification result, the risk of secondary collision when opening the doors and windows can be accurately judged based on the classification result, and the user can be prompted whether the doors and windows can be opened by means of a buzzer or vibration to avoid secondary collision and ensure the safety of vehicle occupants.

[0075] Particularly, in the technical solution of this application, here, when the multiple vehicle target region-of-interest semantic feature vectors are passed through a Transformer-based context encoder to obtain classification feature vectors, the multiple context vehicle target region-of-interest semantic feature vectors obtained by passing the multiple vehicle target region-of-interest semantic feature vectors through the Transformer are directly concatenated to obtain the classification feature vectors. In this way, the classification feature vectors will have poor consistency and correlation in the fusion feature dimension of the multiple context vehicle target region-of-interest semantic feature vectors as the target classification dimension, thereby affecting the accuracy of the classification result of the classification feature vectors.

[0076] Therefore, it is desired to converge the differences between the multiple context vehicle target region-of-interest semantic feature vectors at the Gaussian probability density level. Specifically, first, the Gaussian mixture model of the multiple context vehicle target region-of-interest semantic feature vectors is calculated, and then the Gaussian probability density distribution distance index between each context vehicle target region-of-interest semantic feature vector and the Gaussian mixture model is further calculated, expressed as:

[0077]

[0078] Among them, V i is the i-th context vehicle target interested semantic feature vector, μ u and Σ u are the mean vector and covariance matrix of the Gaussian mixture model, that is, μ u represents the weighted mean vector of the multiple context vehicle target interested semantic feature vectors, and Σ u represents the weighted sum mean variance matrix of the variance matrices of the multiple context vehicle target interested semantic feature vectors themselves, where the vector is a column vector.

[0079] Here, by calculating the Gaussian probability density distribution distance index between each context vehicle target interested semantic feature vector and the Gaussian mixture model, the feature distribution distance of the target feature vector relative to the joint Gaussian probability density distribution represented by the Gaussian mixture model can be represented. By weighting each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors respectively, the compatibility of the cascaded classification feature vector with the probability density joint distribution related migration of the Gaussian probability density in the target domain can be improved, thereby enhancing the consistency and correlation of its Gaussian probability density distribution in the fusion feature dimension of the multiple context vehicle target interested semantic feature vectors as the target classification dimension, so as to improve the accuracy of the classification result of the classification feature vector. In this way, the risk of secondary collision when opening the doors and windows can be accurately judged, and the user can be prompted whether the doors and windows can be opened by means of a buzzer or vibration to avoid secondary collision and ensure the safety of vehicle occupants.

[0080] Figure 1 is an application scenario diagram of the vehicle door and window opening prompt method after vehicle collision according to an embodiment of the present application. As Figure 1 shown, in this application scenario, first, obtain the driving monitoring video of surrounding vehicles collected by a camera (such as Figure 1 C shown in Figure 1 ) deployed on a vehicle (such as Figure 1 M shown in Figure 1 D shown in

[0081] After introducing the basic principle of the present application, various non-limiting embodiments of the present application will be specifically introduced with reference to the accompanying drawings.

[0082] Figure 2 The flowchart of the method for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application. As Figure 2 shown, the method for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application includes the steps of: S110, obtaining a driving monitoring video of surrounding vehicles collected by a camera deployed on the vehicle; S120, extracting a plurality of key frames of the driving monitoring from the driving monitoring video; S130, respectively passing the plurality of key frames of the driving monitoring through a vehicle target detection network to obtain a plurality of regions of interest of vehicle targets; S140, respectively passing the plurality of regions of interest of vehicle targets through a ViT model including an image patch embedding layer to obtain a plurality of semantic feature vectors of vehicle targets of interest; S150, passing the plurality of semantic feature vectors of vehicle targets of interest through a context encoder based on a transformer to obtain a classification feature vector; and S160, passing the classification feature vector through a classifier to obtain a classification result, where the classification result is used to indicate whether the doors and windows can be opened.

[0083] Figure 3 The schematic architecture diagram of the method for prompting the opening of vehicle doors and windows after a vehicle collision according to an embodiment of the present application. As Figure 3 shown, in this network architecture, first, obtain a driving monitoring video of surrounding vehicles collected by a camera deployed on the vehicle; then, extract a plurality of key frames of the driving monitoring from the driving monitoring video; then, respectively pass the plurality of key frames of the driving monitoring through a vehicle target detection network to obtain a plurality of regions of interest of vehicle targets; then, respectively pass the plurality of regions of interest of vehicle targets through a ViT model including an image patch embedding layer to obtain a plurality of semantic feature vectors of vehicle targets of interest; then, pass the plurality of semantic feature vectors of vehicle targets of interest through a context encoder based on a transformer to obtain a classification feature vector; finally, pass the classification feature vector through a classifier to obtain a classification result, where the classification result is used to indicate whether the doors and windows can be opened.

[0084] More specifically, in step S110, a driving surveillance video of surrounding vehicles collected by a camera deployed on the vehicle is obtained. During the actual process of prompting the opening of vehicle doors and windows after a vehicle collision, it is particularly important to analyze the surrounding vehicles to determine whether a secondary collision will occur. Therefore, in the technical solution of this application, it is desired to collect the driving surveillance video of surrounding vehicles through a camera installed on the side of the vehicle, and to judge the risk of a secondary collision when opening the doors and windows from the driving surveillance video of the vehicle. However, since there is a large amount of information in the driving surveillance video of surrounding vehicles, and the vehicle has different speeds during driving on different sections of the road, and the reactions of drivers of surrounding vehicles when seeing a collision accident are also different, it is difficult to capture and extract the driving characteristics of surrounding vehicles in real time and accurately. Therefore, in the technical solution of this application, by mining the distribution information of implicit driving characteristics of vehicles in the driving surveillance video of surrounding vehicles, the risk of a secondary collision when opening the doors and windows is accurately judged, and the user is prompted whether the doors and windows can be opened by means of a buzzer or vibration to avoid a secondary collision and ensure the safety of vehicle occupants.

[0085] More specifically, in step S120, multiple driving surveillance key frames are extracted from the driving surveillance video. Considering that in the driving surveillance video of surrounding vehicles, the driving change characteristics of surrounding vehicles can be represented by the difference between adjacent surveillance frames in the driving surveillance video of surrounding vehicles, that is, the driving state change of surrounding vehicles is represented by the image representation of adjacent image frames. However, considering that the difference between adjacent frames in the surveillance video is small and there is a large amount of data redundancy, in order to reduce the computational amount and avoid the adverse effects of data redundancy on detection, the driving surveillance video of surrounding vehicles is sampled for key frames at a predetermined sampling frequency to extract multiple driving surveillance key frames from the driving surveillance video.

[0086] Correspondingly, in a specific example, extracting multiple driving surveillance key frames from the driving surveillance video includes: sampling the driving surveillance video for key frames at a predetermined sampling frequency to extract multiple driving surveillance key frames from the driving surveillance video.

[0087] More specifically, in step S130, the multiple driving monitoring key frames are respectively passed through a vehicle target detection network to obtain multiple regions of interest (ROIs) of vehicle targets. When detecting the driving states of surrounding vehicles, it is necessary to focus on hidden features such as the driving speed and distance of surrounding vehicle targets. However, due to the large amount of information in the driving monitoring key frames, it is difficult to capture and extract the effective information of vehicle targets. Therefore, considering that if the useless interference feature information can be filtered out when mining the driving state features of surrounding vehicles, it is obvious that the accuracy of monitoring the driving states of surrounding vehicles can be improved. Based on this, in the technical solution of this application, the multiple driving monitoring key frames are further respectively passed through a vehicle target detection network to obtain multiple regions of interest (ROIs) of vehicle targets.

[0088] Specifically, the target anchoring layer of the vehicle target detection network is used to slide with the anchor box B to process the multiple driving monitoring key frames respectively, so as to frame the regions of interest of the vehicle targets, thereby obtaining the multiple regions of interest (ROIs) of vehicle targets. In this way, on the one hand, the interference of useless information in the background can be eliminated, and on the other hand, the vehicle targets can be focused on during subsequent feature extraction. For example, the change of speed features and distance features can be inferred by the change of the size of vehicle targets in time series. In particular, here, the vehicle target detection network is an anchor window-based target detection network, and the anchor window-based target detection network is Fast R-CNN, Faster R-CNN or RetinaNet.

[0089] Correspondingly, in a specific example, the vehicle target detection network is an anchor window-based target detection network, and the anchor window-based target detection network is Fast R-CNN, Faster R-CNN or RetinaNet.

[0090] More specifically, in step S140, the multiple regions of interest (ROIs) of vehicle targets are respectively passed through a ViT model including an image patch embedding layer to obtain multiple semantic feature vectors of vehicle targets of interest. A convolutional neural network model with excellent performance in implicit feature extraction of images can be used to mine the features of each region of interest (ROI) of vehicle targets. However, due to the inherent limitations of convolutional operations, it is difficult for a pure CNN method to learn explicit global and long-range semantic information interaction. Therefore, in the technical solution of this application, the multiple regions of interest (ROIs) of vehicle targets are respectively encoded through a ViT model including an image patch embedding layer to extract the implicit associated semantic features regarding the driving states of each vehicle target in each region of interest (ROI) of vehicle targets, so as to obtain multiple semantic feature vectors of vehicle targets of interest.

[0091] In particular, here, the process of embedding can be as follows: first, arrange the pixel values of all pixel positions of the image patches regarding the vehicle target in each of the vehicle target regions of interest into a one-dimensional vector, and then use a fully connected layer to perform fully connected encoding on this one-dimensional vector to achieve embedding. It should be understood that ViT can directly process the vehicle target regions of interest through the self-attention mechanism like Transformer, so as to respectively extract the global implicit correlation semantic feature information regarding the driving states of the respective vehicle targets in each of the vehicle target regions of interest.

[0092] Correspondingly, in a specific example, as Figure 4 shown, pass the multiple vehicle target regions of interest through a ViT model including an image patch embedding layer to obtain multiple vehicle target region of interest semantic feature vectors, including: S141, use the image patch embedding layer to perform embedding on each of the vehicle target regions of interest in the multiple vehicle target regions of interest to obtain multiple vehicle target region of interest vectors; S142, based on the multiple vehicle target regions of interest, perform feature distribution enhancement in the channel dimension on the multiple vehicle target region of interest vectors to obtain multiple optimized vehicle target region of interest vectors; S143, pass each of the optimized vehicle target region of interest vectors through the ViT model to obtain multiple context vehicle target region of interest vectors; and, S144, cascade the multiple context vehicle target region of interest vectors to obtain the multiple vehicle target region of interest semantic feature vectors.

[0093] More specifically, in step S150, pass the multiple vehicle target region of interest semantic feature vectors through a context encoder based on a transformer to obtain classification feature vectors. Since the implicit correlation features regarding the driving states of the respective vehicle targets in each of the vehicle target regions of interest have temporal dynamic change feature information in the time dimension. Therefore, in order to be able to fully extract the temporal dynamic change features of the driving states of surrounding vehicles, so as to accurately capture the speed features and distance features of surrounding vehicles, in the technical solution of this application, further pass the multiple vehicle target region of interest semantic feature vectors through the context encoder based on a transformer for encoding, so as to extract the dynamic correlation feature distribution information based on the temporal global of the form state association features regarding the respective vehicle targets in each of the driving monitoring key frames, that is, the temporal dynamic features of the driving states of surrounding vehicles, thereby obtaining classification feature vectors.

[0094] It should be understood that through the context encoder, the relationship between a certain word segmentation and other word segmentations in the vector representation sequence can be analyzed to obtain the corresponding feature information. The context encoder aims to mine the hidden patterns between the contexts in the word sequence. Optionally, the encoder includes: CNN (Convolutional Neural Network), Recursive NN (Recursive Neural Network), Language Model, etc. The CNN-based method has a better extraction effect on local features, but it is not effective for the long-term dependency problem in sentences. Therefore, the encoder based on Bi-LSTM (Long Short-Term Memory) is widely used. Recursive NN treats sentences as tree structures rather than sequences. In theory, it has stronger representation capabilities, but it has weaknesses such as difficulty in sample annotation, easy gradient disappearance in deep layers, and difficulty in parallel calculation. Therefore, it is rarely used in practical applications. Transformer is a widely used network structure that has the characteristics of both CNN and RNN. It has a good extraction effect on global features and has certain advantages in parallel computing compared to RNN (Recurrent Neural Network).

[0095] Accordingly, in a specific example, Figure 5As shown, passing the multiple vehicle target interested semantic feature vectors through a Transformer-based context encoder to obtain classification feature vectors includes: S151, inputting the multiple vehicle target interested semantic feature vectors into the Transformer-based context encoder to obtain multiple context vehicle target interested semantic feature vectors; S152, calculating the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors, where the mean vector of the Gaussian mixture model is the position-wise mean vector of the multiple context vehicle target interested semantic feature vectors, and the value at each position in the covariance matrix of the Gaussian mixture model is the variance between the eigenvalues of the corresponding two positions in the position-wise mean vector; S153, respectively calculating the Gaussian probability density distribution distance exponent between each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors and the Gaussian mixture model to obtain multiple Gaussian probability density distribution distance exponents; S154, using the multiple Gaussian probability density distribution distance exponents as weights to weight the multiple context vehicle target interested semantic feature vectors to obtain multiple optimized context vehicle target interested semantic feature vectors; and, S155, concatenating the multiple optimized context vehicle target interested semantic feature vectors to obtain the classification feature vector.

[0096] Correspondingly, in a specific example, as Figure 6 shown, inputting the multiple vehicle target interested semantic feature vectors into the Transformer-based context encoder to obtain multiple context vehicle target interested semantic feature vectors includes: S1511, arranging the multiple vehicle target interested semantic feature vectors in a one-dimensional manner to obtain a first global vehicle target interested feature vector; S1512, calculating the product between the first global vehicle target interested feature vector and the transposed vectors of each vehicle target interested semantic feature vector in the multiple vehicle target interested semantic feature vectors to obtain multiple first self-attention correlation matrices; S1513, respectively performing normalization processing on each first self-attention correlation matrix in the multiple first self-attention correlation matrices to obtain multiple first normalized self-attention correlation matrices; S1514, passing each first normalized self-attention correlation matrix in the multiple first normalized self-attention correlation matrices through the Softmax classification function to obtain multiple first probability values; and, S1515, respectively using each first probability value in the multiple first probability values as a weight to weight each vehicle target interested semantic feature vector in the multiple vehicle target interested semantic feature vectors to obtain the multiple context vehicle target interested semantic feature vectors.

[0097] Accordingly, in a specific example, calculating the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors includes: calculating the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors with the following Gaussian mixture formula; wherein, the Gaussian mixture formula is:

[0098]

[0099] Wherein, μ represents the position-wise mean vector among the multiple context vehicle target interested semantic feature vectors, and the value of each position of σ represents the variance between the feature values of each position in the multiple context vehicle target interested semantic feature vectors.

[0100] It should be understood that, as the learning objective of the neural network model, the Gaussian density map can represent the joint distribution of a single eigenvalue of the feature distribution due to its probability density in the case of the overall distribution composed of multiple eigenvalues. That is, taking the feature distribution as the prior distribution, to obtain the probability density under each prior distribution position due to the correlation effect of other prior distribution positions as the posterior distribution, so as to more accurately describe the feature distribution in a higher dimension.

[0101] Particularly, in the technical solution of the present application, here, when the multiple vehicle target interested semantic feature vectors are passed through a transformer-based context encoder to obtain the classification feature vectors, the multiple context vehicle target interested semantic feature vectors obtained by passing the multiple vehicle target interested semantic feature vectors through the transformer are directly concatenated to obtain the classification feature vectors. In this way, the classification feature vectors will have poor consistency and correlation in the fusion feature dimension of the multiple context vehicle target interested semantic feature vectors that are the target classification dimensions, thus affecting the accuracy of the classification result of the classification feature vectors. Therefore, it is desired to converge the differences between the multiple context vehicle target interested semantic feature vectors at the Gaussian probability density level. Specifically, first calculate the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors, and then further calculate the Gaussian probability density distribution distance index between each context vehicle target interested semantic feature vector and the Gaussian mixture model.

[0102] Accordingly, in a specific example, calculating the Gaussian probability density distribution distance exponent between each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors and the Gaussian mixture model respectively to obtain a plurality of Gaussian probability density distribution distance exponents, including: calculating the Gaussian probability density distribution distance exponent between each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors and the Gaussian mixture model respectively with the following optimization formula to obtain the plurality of Gaussian probability density distribution distance exponents; wherein, the optimization formula is:

[0103]

[0104] wherein, V i is the i-th context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors, (·) T represents the transposed vector of the vector, μ u and Σ u are the mean vector and covariance matrix of the Gaussian mixture model respectively. The multiple context vehicle target interested semantic feature vectors and the mean vector of the Gaussian mixture model are both in the form of column vectors. exp(·) represents the exponential operation of the matrix, and the exponential operation of the matrix represents the natural exponential function value with the eigenvalues at each position in the matrix as the power. represents subtraction by position, represents matrix multiplication, w i represents the i-th Gaussian probability density distribution distance exponent in the plurality of Gaussian probability density distribution distance exponents.

[0105] Here, by calculating the Gaussian probability density distribution distance exponent between each context vehicle target interested semantic feature vector and the Gaussian mixture model, the feature distribution distance of the target feature vector relative to the joint Gaussian probability density distribution represented by the Gaussian mixture model can be represented. By weighting each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors with it respectively, the compatibility of the cascaded classification feature vector with the probability density joint distribution related migration of the Gaussian probability density in the target domain can be improved, thereby enhancing the consistency and correlation of its Gaussian probability density distribution on the fusion feature dimension of the multiple context vehicle target interested semantic feature vectors as the target classification dimension, so as to improve the accuracy of the classification result of the classification feature vector. In this way, the risk of secondary collision when opening the doors and windows can be accurately judged, and the user can be prompted whether the doors and windows can be opened by means of a buzzer or vibration to avoid secondary collision and ensure the safety of vehicle occupants.

[0106] More specifically, in step S160, the classification feature vector is passed through a classifier to obtain a classification result, which is used to indicate whether the doors and windows can be opened. That is, in the technical solution of the present application, the labels of the classifier include generating a door and window locking control instruction and not generating a door and window locking control instruction. Among them, the classifier uses the softmax function to determine which classification label the classification feature vector belongs to. It should be understood that in the technical solution of the present application, the classification label of the classifier is the control strategy label for whether the doors and windows can be opened. Therefore, after obtaining the classification result, the risk of secondary collision when opening the doors and windows can be accurately judged based on the classification result, and the user can be prompted whether the doors and windows can be opened by means of a buzzer or vibration to avoid secondary collision and ensure the safety of vehicle occupants.

[0107] It should be understood that the role of the classifier is to use the given categories and known training data to learn classification rules and classifiers, and then classify (or predict) unknown data. Logistic regression, SVM, etc. are often used to solve binary classification problems. For multi-class classification problems, logistic regression or SVM can also be used, but it requires multiple binary classifications to form a multi-class classification, which is prone to errors and low in efficiency. The commonly used multi-class classification method is the Softmax classification function.

[0108] Correspondingly, in a specific example, as Figure 7 shown, passing the classification feature vector through a classifier to obtain a classification result, which is used to indicate whether the doors and windows can be opened, includes: S161, using the fully connected layer of the classifier to perform fully connected encoding on the classification feature vector to obtain an encoded classification feature vector; and S162, inputting the encoded classification feature vector into the Softmax classification function of the classifier to obtain the classification result.

[0109] In summary, based on the method for prompting the opening of doors and windows after a vehicle collision according to the embodiments of the present application, it first extracts multiple key frames of driving monitoring from the driving monitoring video and passes them through the vehicle target detection network respectively to obtain multiple regions of interest of vehicle targets. Then, the multiple regions of interest of vehicle targets are passed through the ViT model respectively to obtain multiple semantic feature vectors of vehicle targets of interest. Then, the multiple semantic feature vectors of vehicle targets of interest are passed through the context encoder to obtain a classification feature vector. Finally, the classification feature vector is passed through a classifier to obtain a classification result indicating whether the doors and windows can be opened. In this way, the risk of secondary collision when opening the doors and windows can be accurately judged.

[0110] Figure 8The block diagram of the vehicle door and window opening prompt system 100 after a vehicle collision according to an embodiment of the present application. As Figure 8 shown, the vehicle door and window opening prompt system 100 after a vehicle collision according to an embodiment of the present application includes: a video acquisition module 110, configured to acquire a driving monitoring video of surrounding vehicles collected by a camera deployed on the vehicle; a key frame extraction module 120, configured to extract a plurality of driving monitoring key frames from the driving monitoring video; a vehicle target detection module 130, configured to respectively pass the plurality of driving monitoring key frames through a vehicle target detection network to obtain a plurality of vehicle target regions of interest; an image encoding module 140, configured to respectively pass the plurality of vehicle target regions of interest through a ViT model including an image patch embedding layer to obtain a plurality of vehicle target region of interest semantic feature vectors; a context encoding module 150, configured to pass the plurality of vehicle target region of interest semantic feature vectors through a transformer-based context encoder to obtain a classification feature vector; and a classification module 160, configured to pass the classification feature vector through a classifier to obtain a classification result, where the classification result is used to indicate whether the door and window can be opened.

[0111] In one example, in the vehicle door and window opening prompt system 100 after a vehicle collision described above, the key frame extraction module 120 is configured to: perform key frame sampling on the driving monitoring video at a predetermined sampling frequency to extract a plurality of driving monitoring key frames from the driving monitoring video.

[0112] In one example, in the vehicle door and window opening prompt system 100 after a vehicle collision described above, the vehicle target detection network is an anchor window-based target detection network, and the anchor window-based target detection network is Fast R-CNN, Faster R-CNN or RetinaNet.

[0113] In one example, in the vehicle door and window opening prompt system 100 after a vehicle collision described above, the image encoding module 140 is configured to: use the image patch embedding layer to perform embedding on each vehicle target region of interest in the plurality of vehicle target regions of interest to obtain a plurality of vehicle target region of interest vectors; based on the plurality of vehicle target regions of interest, perform feature distribution enhancement in the channel dimension on the plurality of vehicle target region of interest vectors to obtain a plurality of optimized vehicle target region of interest vectors; pass each of the optimized vehicle target region of interest vectors through the ViT model to obtain a plurality of context vehicle target region of interest vectors; and cascade the plurality of context vehicle target region of interest vectors to obtain the plurality of vehicle target region of interest semantic feature vectors.

[0114] In one example, in the above-mentioned vehicle door and window opening prompt system 100 after a vehicle collision, the context encoding module 150 is configured to: input the multiple vehicle target interested semantic feature vectors into the transformer-based context encoder to obtain multiple context vehicle target interested semantic feature vectors; calculate the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors, where the mean vector of the Gaussian mixture model is the position-wise mean vector of the multiple context vehicle target interested semantic feature vectors, and the value at each position in the covariance matrix of the Gaussian mixture model is the variance between the corresponding two position eigenvalues in the position-wise mean vector; calculate the Gaussian probability density distribution distance index between each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors and the Gaussian mixture model respectively to obtain multiple Gaussian probability density distribution distance indices; use the multiple Gaussian probability density distribution distance indices as weights to weight the multiple context vehicle target interested semantic feature vectors to obtain multiple optimized context vehicle target interested semantic feature vectors; and, cascade the multiple optimized context vehicle target interested semantic feature vectors to obtain the classification feature vector.

[0115] In one example, in the above-mentioned vehicle door and window opening prompt system 100 after a vehicle collision, inputting the multiple vehicle target interested semantic feature vectors into the transformer-based context encoder to obtain multiple context vehicle target interested semantic feature vectors includes: arranging the multiple vehicle target interested semantic feature vectors in a one-dimensional manner to obtain a first global vehicle target interested feature vector; calculating the product between the first global vehicle target interested feature vector and the transposed vectors of each vehicle target interested semantic feature vector in the multiple vehicle target interested semantic feature vectors to obtain multiple first self-attention correlation matrices; respectively performing normalization processing on each first self-attention correlation matrix in the multiple first self-attention correlation matrices to obtain multiple first normalized self-attention correlation matrices; passing each first normalized self-attention correlation matrix in the multiple first normalized self-attention correlation matrices through the Softmax classification function to obtain multiple first probability values; and, respectively using each first probability value in the multiple first probability values as a weight to weight each vehicle target interested semantic feature vector in the multiple vehicle target interested semantic feature vectors to obtain the multiple context vehicle target interested semantic feature vectors.

[0116] In one example, in the above-mentioned vehicle door and window opening prompt system 100 after a vehicle collision, calculating the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors includes: calculating the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors with the following Gaussian mixture formula; wherein, the Gaussian mixture formula is:

[0117]

[0118] wherein, μ represents the position-wise mean vector among the multiple context vehicle target interested semantic feature vectors, and the value of each position of σ represents the variance between the feature values of each position in the multiple context vehicle target interested semantic feature vectors.

[0119] In one example, in the above-mentioned vehicle door and window opening prompt system 100 after a vehicle collision, calculating the Gaussian probability density distribution distance exponent of each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors respectively to obtain multiple Gaussian probability density distribution distance exponents includes: calculating the Gaussian probability density distribution distance exponent of each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors respectively with the following optimization formula to obtain the multiple Gaussian probability density distribution distance exponents; wherein, the optimization formula is:

[0120]

[0121] wherein, V i is the i-th context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors, (·) T represents the transposed vector of the vector, μ u and Σ u are the mean vector and covariance matrix of the Gaussian mixture model, the multiple context vehicle target interested semantic feature vectors and the mean vector of the Gaussian mixture model are in the form of column vectors, exp(·) represents the exponential operation of the matrix, and the exponential operation of the matrix represents the natural exponential function value with the feature values of each position in the matrix as the power, represents position-wise subtraction, represents matrix multiplication, w i represents the i-th Gaussian probability density distribution distance exponent among the multiple Gaussian probability density distribution distance exponents.

[0122] In one example, in the above-described vehicle post-collision door and window opening prompt system 100, the classification module 160 is configured to: perform fully-connected encoding on the classification feature vector using the fully-connected layer of the classifier to obtain an encoded classification feature vector; and input the encoded classification feature vector into the Softmax classification function of the classifier to obtain the classification result.

[0123] Here, those skilled in the art can understand that the specific functions and operations of each unit and module in the above-described vehicle post-collision door and window opening prompt system 100 have been described in detail above with reference to Figures 1 to 7 the description of the vehicle post-collision door and window opening prompt method, and thus, the repeated description thereof will be omitted.

[0124] As described above, the vehicle post-collision door and window opening prompt system 100 according to the embodiments of the present application can be implemented in various wireless terminals, such as a server having a vehicle post-collision door and window opening prompt algorithm. In one example, the vehicle post-collision door and window opening prompt system 100 according to the embodiments of the present application can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the vehicle post-collision door and window opening prompt system 100 can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the vehicle post-collision door and window opening prompt system 100 can also be one of the many hardware modules of the wireless terminal.

[0125] Alternatively, in another example, the vehicle post-collision door and window opening prompt system 100 and the wireless terminal can also be separate devices, and the vehicle post-collision door and window opening prompt system 100 can be connected to the wireless terminal through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.

[0126] According to another aspect of the present application, there is also provided a non-volatile computer-readable storage medium, on which computer-readable instructions are stored, and when the instructions are executed by a computer, the method described above can be performed.

[0127] The program part in the technology can be regarded as a "product" or "article" in the form of executable code and / or related data, which is participated in or implemented by a computer-readable medium. Tangible, permanent storage media can include any memory or storage used by a computer, a processor, or similar devices or related modules. For example, various semiconductor memories, tape drives, disk drives, or any similar devices capable of providing storage functions for software.

[0128] All software, or portions thereof, may sometimes communicate over a network, such as the Internet or other communication networks. Such communication can load the software from one computer device or processor to another. For example, from a server or host computer of a video object detection device to a hardware platform in a computer environment, or other computer environments implementing the system, or systems with similar functions related to providing information required for object detection. Therefore, another medium capable of transmitting software elements can also be used as a physical connection between local devices, such as light waves, radio waves, electromagnetic waves, etc., which are propagated through cables, optical fibers, or air. Physical media used to carry the wave, such as cables, wireless connections, or optical fibers and similar devices, can also be considered as media carrying software. As used herein, unless restricted to tangible "storage" media, other terms representing "computer or machine-readable media" refer to media involved in the process of a processor executing any instructions.

[0129] This application uses specific terms to describe the embodiments of this application. Such as "first / second embodiment", "one embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that the "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more at different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application can be appropriately combined.

[0130] In addition, those skilled in the art can understand that various aspects of this application can be illustrated and described by several patentable types or situations, including any new and useful process, machine, product, or combination of substances, or any new and useful improvement thereof. Accordingly, various aspects of this application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software can all be referred to as "data blocks", "modules", "engines", "units", "components", or "systems". In addition, various aspects of this application may be embodied as a computer product located in one or more computer-readable media, which includes computer-readable program code.

[0131] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those of ordinary skill in the art to which this invention belongs. It should also be understood that terms such as those defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense, unless explicitly defined as such herein.

[0132] The foregoing is a description of the present invention and should not be construed as limiting thereof. Although several exemplary embodiments of the present invention have been described, those skilled in the art will readily appreciate that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the present invention. Accordingly, all such modifications are intended to be included within the scope of the present invention as defined by the claims. It should be understood that the foregoing is a description of the present invention and should not be considered limited to the specific embodiments disclosed, and modifications to the disclosed embodiments as well as other embodiments are intended to be included within the scope of the appended claims. The present invention is defined by the claims and their equivalents.

Claims

1. A method for prompting the opening of vehicle doors and windows after a vehicle collision, characterized in that, Including: Obtain the driving monitoring video of surrounding vehicles collected by a camera deployed on a vehicle; Extract multiple driving monitoring key frames from the driving monitoring video; Pass the multiple driving monitoring key frames through a vehicle target detection network respectively to obtain multiple vehicle target regions of interest; Pass the multiple vehicle target regions of interest through a ViT model including an image patch embedding layer respectively to obtain multiple vehicle target interested semantic feature vectors; Pass the multiple vehicle target interested semantic feature vectors through a Transformer-based context encoder to obtain classification feature vectors; and pass the classification feature vectors through a classifier to obtain a classification result, where the classification result is used to indicate whether the doors and windows can be opened; Pass the multiple vehicle target interested semantic feature vectors through a Transformer-based context encoder to obtain classification feature vectors, including: Input the multiple vehicle target interested semantic feature vectors into the Transformer-based context encoder to obtain multiple context vehicle target interested semantic feature vectors; Calculate the Gaussian mixture model of the multiple context vehicle target interested semantic feature vectors, where the mean vector of the Gaussian mixture model is the position-wise mean vector of the multiple context vehicle target interested semantic feature vectors, and the value at each position in the covariance matrix of the Gaussian mixture model is the variance between the corresponding two position-wise eigenvalues of the position-wise mean vector; Calculate the Gaussian probability density distribution distance index between each context vehicle target interested semantic feature vector in the multiple context vehicle target interested semantic feature vectors and the Gaussian mixture model respectively to obtain multiple Gaussian probability density distribution distance indexes; Use the multiple Gaussian probability density distribution distance indexes as weights to weight the multiple context vehicle target interested semantic feature vectors to obtain multiple optimized context vehicle target interested semantic feature vectors; and concatenate the multiple optimized context vehicle target interested semantic feature vectors to obtain the classification feature vectors.

2. The method for prompting the opening of vehicle doors and windows after a vehicle collision according to claim 1, wherein Extract multiple driving monitoring key frames from the driving monitoring video, including: Perform key frame sampling on the driving monitoring video at a predetermined sampling frequency to extract multiple driving monitoring key frames from the driving monitoring video.

3. The method for prompting the opening of vehicle doors and windows after a vehicle collision according to claim 2, wherein, The vehicle target detection network is an anchor window-based target detection network, and the anchor window-based target detection network is Fast R-CNN, Faster R-CNN or RetinaNet.

4. The method for prompting the opening of vehicle doors and windows after a vehicle collision according to claim 3, wherein, Pass the multiple vehicle target regions of interest through a ViT model including an image patch embedding layer respectively to obtain multiple vehicle target interested semantic feature vectors, including: Use the image patch embedding layer to embed each vehicle target region of interest in the multiple vehicle target regions of interest to obtain multiple vehicle target interested vectors; Based on the multiple vehicle target regions of interest, perform feature distribution enhancement in the channel dimension on the multiple vehicle target interested vectors to obtain multiple optimized vehicle target interested vectors; Pass the respective optimized vehicle target interest vectors through the ViT model to obtain multiple context vehicle target interest vectors; and Concatenate the multiple context vehicle target interest vectors to obtain the multiple vehicle target interest semantic feature vectors.

5. The method for prompting the opening of vehicle doors and windows after a vehicle collision according to claim 1, characterized in that Input the multiple vehicle target interest semantic feature vectors into the transformer-based context encoder to obtain multiple context vehicle target interest semantic feature vectors, including:[[]] Arrange the multiple vehicle target interest semantic feature vectors in a one-dimensional manner to obtain a first global vehicle target interest feature vector; Calculate the product between the first global vehicle target interest feature vector and the transposed vectors of the respective vehicle target interest semantic feature vectors among the multiple vehicle target interest semantic feature vectors to obtain multiple first self-attention correlation matrices; Respectively perform normalization processing on each of the multiple first self-attention correlation matrices to obtain multiple first normalized self-attention correlation matrices; Pass each of the multiple first normalized self-attention correlation matrices through the Softmax classification function to obtain multiple first probability values; and Respectively use each of the first probability values among the multiple first probability values as weights to weight the respective vehicle target interest semantic feature vectors among the multiple vehicle target interest semantic feature vectors to obtain the multiple context vehicle target interest semantic feature vectors.

6. The method for prompting the opening of vehicle doors and windows after a vehicle collision according to claim 5, wherein, Calculate the Gaussian mixture model of the multiple context vehicle target interest semantic feature vectors, including:[[]] Calculate the Gaussian mixture model of the multiple context vehicle target interest semantic feature vectors using the following Gaussian mixture formula; wherein, the Gaussian mixture formula is: Among them, represents the position-wise mean vector among the multiple context vehicle target interested semantic feature vectors, and the value of each position of represents the variance among the feature values of each position in the multiple context vehicle target interested semantic feature vectors.

7. The method for prompting the opening of vehicle doors and windows after a vehicle collision according to claim 6, wherein Respectively calculate the Gaussian probability density distribution distance exponents between each context vehicle target interest semantic feature vector among the multiple context vehicle target interest semantic feature vectors and the Gaussian mixture model to obtain multiple Gaussian probability density distribution distance exponents, including:[[]] Calculate the Gaussian probability density distribution distance exponents between each context vehicle target interest semantic feature vector among the multiple context vehicle target interest semantic feature vectors and the Gaussian mixture model using the following optimization formula to obtain the multiple Gaussian probability density distribution distance exponents; wherein, the optimization formula is: Among them, is the th context vehicle target interested semantic feature vector among the multiple context vehicle target interested semantic feature vectors, represents the transposed vector of the vector, and are the mean vector and covariance matrix of the Gaussian mixture model. The multiple context vehicle target interested semantic feature vectors and the mean vector of the Gaussian mixture model are both in the form of column vectors, represents the exponential operation of the matrix. The exponential operation of the matrix represents the natural exponential function values with the eigenvalues at each position in the matrix as the exponents, represents subtraction by position, represents matrix multiplication, represents the th Gaussian probability density distribution distance exponent among the multiple Gaussian probability density distribution distance exponents.

8. The method for prompting the opening of vehicle doors and windows after a vehicle collision according to claim 1, wherein Pass the classification feature vector through a classifier to obtain a classification result, and the classification result is used to indicate whether the doors and windows can be opened, including:[[]] Use the fully connected layer of the classifier to perform fully connected encoding on the classification feature vector to obtain an encoded classification feature vector; and Input the encoded classification feature vector into the Softmax classification function of the classifier to obtain the classification result.

9. A door and window opening prompt system after a vehicle collision, characterized in that including:[[]] A video acquisition module, configured to acquire a driving surveillance video of surrounding vehicles collected by a camera deployed on a vehicle; A key frame extraction module, configured to extract multiple driving surveillance key frames from the driving surveillance video; A vehicle target detection module, configured to respectively pass the multiple driving surveillance key frames through a vehicle target detection network to obtain multiple vehicle target interest regions; An image encoding module, configured to respectively obtain multiple vehicle target region-of-interest semantic feature vectors for the multiple vehicle target regions-of-interest through a ViT model including an image patch embedding layer; A context encoding module, configured to obtain classification feature vectors for the multiple vehicle target region-of-interest semantic feature vectors through a transformer-based context encoder; And A classification module, configured to obtain a classification result for the classification feature vectors through a classifier, where the classification result is used to indicate whether the doors and windows can be opened; Obtaining classification feature vectors for the multiple vehicle target region-of-interest semantic feature vectors through a transformer-based context encoder includes: Inputting the multiple vehicle target region-of-interest semantic feature vectors into the transformer-based context encoder to obtain multiple context vehicle target region-of-interest semantic feature vectors; Calculating a Gaussian mixture model for the multiple context vehicle target region-of-interest semantic feature vectors, where the mean vector of the Gaussian mixture model is the position-wise mean vector of the multiple context vehicle target region-of-interest semantic feature vectors, and the value at each position in the covariance matrix of the Gaussian mixture model is the variance between the eigenvalues of the corresponding two positions in the position-wise mean vector; Respectively calculating the Gaussian probability density distribution distance exponent between each context vehicle target region-of-interest semantic feature vector in the multiple context vehicle target region-of-interest semantic feature vectors and the Gaussian mixture model to obtain multiple Gaussian probability density distribution distance exponents; Using the multiple Gaussian probability density distribution distance exponents as weights to weight the multiple context vehicle target region-of-interest semantic feature vectors to obtain multiple optimized context vehicle target region-of-interest semantic feature vectors; and concatenating the multiple optimized context vehicle target region-of-interest semantic feature vectors to obtain the classification feature vectors.

Citation Information

Patent Citations

  • Aerial photography vehicle detection method and detection system based on multi-scale small samples

    CN112949520A

  • Vehicle target tracking method and device, vehicle and storage medium

    CN115469657A