Collision failure detection method and related device
By extracting and transforming the inter-frame differences of collision videos, and utilizing a visual attention layer and detection model, collision failure images are automatically detected. This solves the problems of cumbersome and complex detection and high error rate in existing technologies, and achieves efficient and accurate collision failure detection.
Patent Information
- Application Number
- CN202411090186.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-10
AI Technical Summary
Existing collision failure detection methods are cumbersome, complex, and prone to errors, making it difficult to achieve efficient and accurate collision failure detection.
By extracting multi-frame image features from the collision video to be tested, calculating the inter-frame differences, and using a visual attention layer and collision detection model to perform visual attention-based feature transformation and detection, collision failure images are automatically detected.
It achieves efficient and accurate collision failure detection, reduces errors caused by manual comparison, and improves the convenience and accuracy of detection.
Smart Images

Figure CN121505313A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a collision failure detection method and related apparatus. Background Technology
[0002] Currently, for collision videos in game scenes, there is a need to efficiently and accurately detect collision failure images in the collision videos so that the collision videos can be optimized based on the collision failure images.
[0003] In related technologies, collision failure detection methods refer to: extracting multiple image features corresponding to multiple frames in a collision video, manually comparing multiple image features, and detecting collision failure images in the collision video.
[0004] However, the above methods involve manually comparing multiple image features, which makes collision failure detection cumbersome, complex, and prone to errors, making it difficult to achieve efficient and accurate collision failure detection. Summary of the Invention
[0005] To address the aforementioned technical problems, this application provides a collision failure detection method and related apparatus, which makes collision failure detection more automatic and convenient, reduces detection errors, and thus achieves collision failure detection efficiently and accurately.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] On one hand, embodiments of this application provide a collision failure detection method, the method comprising:
[0008] Feature extraction is performed on each frame of the collision video to be tested to obtain the visual features to be tested for each frame of the collision video.
[0009] Difference calculation is performed on multiple visual features corresponding to the multiple frames of images to be tested to obtain the inter-frame differences of the multiple frames of images to be tested.
[0010] By using the visual attention layer in the collision detection model, multiple visual features to be tested are transformed based on visual attention according to the inter-frame differences to be tested, thereby obtaining multiple attention features to be tested.
[0011] The collision detection layer in the collision detection model is used to perform collision detection on the multiple attention features to be tested, and multiple collision detection results are obtained.
[0012] If the multiple collision detection results represent multiple collision detection categories including a collision failure category, then the collision video to be tested is determined to include a collision failure image.
[0013] On the other hand, embodiments of this application provide a collision failure detection device, the device comprising: an extraction unit, a calculation unit, a transformation unit, a detection unit, and a determination unit;
[0014] The extraction unit is used to extract features from each frame of the test image in the multi-frame test image of the collision video to obtain the test visual features of each frame of the test image.
[0015] The calculation unit is used to perform difference calculation on multiple visual features corresponding to the multiple frames of images to be tested, and to obtain the inter-frame differences of the multiple frames of images to be tested.
[0016] The transformation unit is used to perform visual attention-based feature transformation on the multiple visual features to be tested according to the inter-frame differences in the collision detection model through the visual attention layer, so as to obtain multiple attention features to be tested.
[0017] The detection unit is used to perform collision detection on the multiple attention features to be tested through the collision detection layer in the collision detection model, and obtain multiple collision detection results;
[0018] The determining unit is configured to determine that the collision video to be tested includes a collision failure image if the multiple collision detection categories represented by the multiple collision detection results include a collision failure category.
[0019] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory:
[0020] The memory is used to store computer programs and to transfer the computer programs to the processor;
[0021] The processor is configured to execute the method described in any of the foregoing aspects according to instructions in the computer program.
[0022] On the other hand, embodiments of this application provide a computer-readable storage medium for storing a computer program that, when run on a computer device, causes the computer device to perform the methods described in any of the foregoing aspects.
[0023] On the other hand, embodiments of this application provide a computer program product, including a computer program that, when run on a computer device, causes the computer device to perform the method described in any of the foregoing aspects.
[0024] As can be seen from the above technical solution, firstly, the visual features to be tested are extracted from each frame of the collision video to be tested; the feature differences between multiple visual features to be tested corresponding to multiple frames of the collision video are calculated to obtain the inter-frame differences to be tested corresponding to multiple frames of the collision video; the inter-frame differences to be tested and multiple visual features to be tested are input into the visual attention layer of the collision detection model to perform feature transformation based on visual attention, and multiple attention features to be tested are output. This method, based on the extraction of the visual features to be tested from each frame of the collision video to be tested, obtains the inter-frame differences to be tested by comparing the feature differences between multiple visual features to be tested corresponding to multiple frames of the collision video. It can detect the dynamic changes of collision between multiple frames of the collision video. By considering the inter-frame differences to be tested through the visual attention layer in the collision detection model, and performing feature transformation based on visual attention for multiple visual features to be tested to obtain multiple attention features to be tested, it can efficiently and accurately detect the detailed features related to collision failure in multiple frames of the collision video. Then, multiple attention features to be tested are input into the collision detection layer of the collision detection model for collision detection, and multiple collision detection results are output. When multiple collision detection results represent multiple collision detection categories including the collision failure category, it is determined that the collision video to be tested includes collision failure images. This method uses the collision detection layer in the collision detection model to achieve collision detection for multiple attention features to be tested and obtain multiple collision detection results. It can efficiently and accurately detect whether each frame of the image to be tested belongs to the collision failure category. When multiple collision detection results represent multiple collision detection categories including the collision failure category, it indicates that multiple frames of the image to be tested include the image to be tested that belongs to the collision failure category. It can efficiently and accurately detect collision failure images in the collision video to be tested.
[0025] Based on this, the method eliminates the need for manual comparison of multiple test visual features corresponding to multiple test images in the collision video. It obtains the inter-frame differences by comparing the feature differences between multiple test visual features, and the inter-frame differences can be considered by the visual attention layer. Multiple test attention features are obtained by performing feature transformation based on visual attention on multiple test visual features, and multiple test attention features are obtained by using the collision detection layer to perform collision detection on multiple test attention features to obtain multiple collision detection results. This makes collision failure detection more automatic and convenient, reduces detection errors, and determines that the collision video includes collision failure images when multiple collision detection categories represented by multiple collision detection results include collision failure categories. This achieves efficient and accurate collision failure detection. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a system schematic diagram of a collision failure detection method provided in an embodiment of this application;
[0028] Figure 2 A flowchart of a collision failure detection method provided in an embodiment of this application;
[0029] Figure 3 A schematic diagram illustrating a collision failure detection step provided in an embodiment of this application;
[0030] Figure 4 A flowchart illustrating a method for generating a collision text sequence provided in this application embodiment;
[0031] Figure 5 A schematic diagram illustrating a step for generating a collision text sequence, provided in an embodiment of this application;
[0032] Figure 6 A structural diagram of a collision failure detection device provided in an embodiment of this application;
[0033] Figure 7 A structural diagram of a server provided in an embodiment of this application;
[0034] Figure 8 This is a structural diagram of a terminal provided in an embodiment of this application. Detailed Implementation
[0035] The embodiments of this application will now be described with reference to the accompanying drawings.
[0036] Currently, in game scenarios, multiple image features corresponding to multiple frames of images in the collision video are extracted, and multiple image features are manually compared to detect collision failure images in the collision video, which facilitates subsequent optimization of the collision video based on the collision failure images. However, research has found that manually comparing multiple image features makes collision failure detection cumbersome, complex, and prone to errors, making it difficult to achieve collision failure detection efficiently and accurately.
[0037] This application provides a collision failure detection method that eliminates the need for manual comparison of multiple test visual features corresponding to multiple frames of a collision video. By comparing the feature differences between multiple test visual features, the inter-frame differences are obtained. The inter-frame differences can be considered using a visual attention layer. Multiple test attention features are obtained by performing a visual attention-based feature transformation on the multiple test visual features. The collision detection layer can then perform collision detection on the multiple attention features to obtain multiple collision detection results. This makes collision failure detection more automatic and convenient, reduces detection errors, and determines that the collision video includes a collision failure image when the multiple collision detection categories represented by the multiple collision detection results include the collision failure category. This achieves efficient and accurate collision failure detection.
[0038] Next, the system architecture of the collision failure detection method will be introduced. See [link / reference] Figure 1 , Figure 1 This is a schematic diagram of a collision failure detection method provided in an embodiment of the present application. The system includes a computer device 100, which is used to execute the collision failure detection method.
[0039] The computer device 100 extracts features from each frame of the multi-frame test image of the collision video to obtain the test visual features of each frame of the test image.
[0040] As an example, the collision video to be tested is collision video X in a game scene. This collision video X includes n frames of images to be tested, where n is a positive integer and n≥2. Computer device 100 extracts the visual features to be tested from each of the n frames of images to be tested in the collision video X. Then, the n visual features to be tested corresponding to the n frames of images to be tested are f1, ..., f2. n .
[0041] Computer device 100 performs difference calculations on multiple visual features corresponding to multiple frames of images to be tested, and obtains the inter-frame differences between the multiple frames of images to be tested.
[0042] As an example, based on the above example, computer device 100 calculates f1, ..., f n The feature differences between the n frames are used to obtain the inter-frame differences of the n frames to be tested, which is Δ(f i f i+1 ), where i is a positive integer and i < n.
[0043] The computer device 100 uses the visual attention layer in the collision detection model to perform feature transformation based on visual attention on multiple visual features to be tested according to the differences between the frames to be tested, and obtains multiple attention features to be tested.
[0044] As an example, based on the above example, computer device 100 will Δ(f)i f i+1 ) and f1, ..., f n The visual attention layer in the input collision detection model performs a feature transformation based on visual attention, outputting n attention features to be tested as f1', ..., f2'. n '.
[0045] The computer device 100 performs collision detection on multiple attention features to be tested through the collision detection layer in the collision detection model, and obtains multiple collision detection results.
[0046] As an example, based on the above example, computer device 100 will use f1', ..., f n The collision detection layer in the input collision detection model performs collision detection and outputs n collision detection results as p1, ..., p2. n .
[0047] If multiple collision detection results represent multiple collision detection categories, including collision failure categories, computer device 100 determines that the collision video to be tested includes collision failure images.
[0048] As an example, based on the above example, when p1, ..., p n When the n collision detection categories include a collision failure category, computer device 100 determines that the collision video X includes a collision failure image.
[0049] In other words, based on the extraction of the visual features of each frame of the collision video, the inter-frame differences are obtained by comparing the feature differences between multiple visual features corresponding to multiple frames of the collision video. This enables the detection of dynamic collision changes between multiple frames of the collision video. Furthermore, by considering the inter-frame differences through the visual attention layer in the collision detection model, multiple attention features are obtained by performing visual attention-based feature transformations on multiple visual features. This allows for the efficient and accurate detection of detailed features related to collision failure in multiple frames of the collision video. Through the collision detection layer in the collision detection model, multiple collision detection results are obtained by performing collision detection on multiple attention features. This allows for the efficient and accurate detection of whether each frame of the collision video belongs to the collision failure category. When multiple collision detection results indicate that multiple collision detection categories include the collision failure category, it means that the multiple frames of the collision video include images belonging to the collision failure category. This enables the efficient and accurate detection of collision failure images in the collision video.
[0050] It should be noted that, in the embodiments of this application, the computer device can be a server or a terminal. The method provided in the embodiments of this application can be executed by the terminal or the server alone, or it can be executed by the terminal and the server in cooperation. Specifically, when the method provided in the embodiments of this application is executed by the terminal or the server alone, its execution method is similar to... Figure 1The corresponding embodiments are similar, mainly replacing the computer device with a terminal or server. Furthermore, when the method provided in this application is executed by a terminal and a server, steps that need to be displayed on the front-end interface can be executed by the terminal, while steps that require background calculations and do not need to be displayed on the front-end interface can be executed by the server.
[0051] The terminal can be a smartphone, tablet, laptop, desktop computer, intelligent voice interaction device, vehicle terminal, extended reality device, or aircraft, but is not limited to these. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, but is not limited to these. The terminal and server can be connected directly or indirectly through wired or wireless communication, and this application does not impose any restrictions. For example, the terminal and server can be connected through a network, which can be wired or wireless.
[0052] Next, the collision failure detection method provided in the embodiments of this application will be described in detail using a computer device executing the method provided in the embodiments of this application as an example, in conjunction with the accompanying drawings.
[0053] See Figure 2 , Figure 2 The flowchart of a collision failure detection method provided in this application embodiment includes the following steps S201-S205.
[0054] S201: Extract features from each frame of the test image in the multi-frame test image of the collision video to obtain the test visual features of each frame of the test image.
[0055] S202: Perform difference calculation on multiple visual features corresponding to multiple test images to obtain the inter-frame differences between multiple test images.
[0056] S203: Through the visual attention layer in the collision detection model, multiple visual features to be tested are transformed based on visual attention according to the differences between the frames to be tested, and multiple attention features to be tested are obtained.
[0057] In related technologies, in order to detect collision failure images in collision videos in game scenes so that the collision videos can be optimized based on the collision failure images, it is usually necessary to extract multiple image features corresponding to multiple frames in the collision video and manually compare multiple image features to detect collision failure images in the collision video. However, manually comparing multiple image features makes collision failure detection cumbersome, complex and prone to errors, making it difficult to achieve collision failure detection efficiently and accurately.
[0058] Therefore, in this embodiment, for the collision video to be tested in a game scene, the visual features to be tested are extracted from each frame of the test image in the multi-frame test image of the collision video to be tested. Considering the feature differences between the multiple visual features to be tested corresponding to the multi-frame test images, it is possible to characterize the dynamic changes of collision between the multi-frame test images. Based on the dynamic changes of collision between the multi-frame test images, a visual attention mechanism is introduced to transform the multiple visual features to be tested into multiple attention features to be tested, which can characterize the detailed features related to collision failure in the multi-frame test images, so as to efficiently and accurately detect whether each frame of the test image belongs to the collision failure category, thereby efficiently and accurately detecting the collision failure images in the collision video to be tested. Based on this, the feature differences between the multiple visual features to be tested corresponding to the multi-frame test images are first calculated to obtain the inter-frame differences to be tested corresponding to the multi-frame test images. Then, the inter-frame differences to be tested and the multiple visual features to be tested are input into the visual attention layer in the collision detection model to perform feature transformation based on visual attention, and output multiple attention features to be tested.
[0059] Among them, the collision video to be tested refers to the collision video of the failed collision images to be detected; the image to be tested refers to the image obtained by segmenting the collision video to be tested into frames; feature extraction refers to extracting the high-level, abstract visual feature representation of the image to be tested; the visual feature to be tested refers to the high-level, abstract visual feature representation of the extracted image to be tested; difference calculation refers to calculating the feature difference between two different visual features to be tested; inter-frame difference to be tested refers to the feature difference between multiple visual features to be tested calculated; collision detection model refers to the model that detects the failed collision images in the collision video to be tested; visual attention layer refers to the attention layer that detects the detailed features related to the failed collision in multiple frames of the image to be tested; feature transformation based on visual attention refers to transforming multiple visual features to be tested considering the inter-frame difference to be tested in order to detect the detailed features related to the failed collision in multiple frames of the image to be tested; multiple attention features to be tested refer to the detailed features related to the failed collision in the multiple frames of the image to be tested detected.
[0060] Based on the extraction of the visual features of each frame of the collision video to be tested by S201-S203, the inter-frame differences are obtained by comparing the feature differences between multiple visual features corresponding to multiple frames of the collision video. This enables the detection of dynamic collision changes between multiple frames of the collision video. Furthermore, by considering the inter-frame differences through the visual attention layer in the collision detection model, multiple attention features are obtained by performing feature transformation based on visual attention for multiple visual features to be tested. This allows for the efficient and accurate detection of detailed features related to collision failure in multiple frames of the collision video. This lays the foundation for the efficient and accurate detection of whether each frame of the collision video belongs to the collision failure category.
[0061] As an example of S201-S203, the collision video to be tested is collision video X in a game scene. This collision video X includes n frames of images to be tested, where n is a positive integer and n≥2. The computer device extracts the visual features to be tested from each of the n frames of images to be tested in the game collision video X. Then, the n visual features to be tested corresponding to the n frames of images to be tested are f1, ..., f2. n Computer equipment calculates f1, ..., f n The feature differences between the n frames are used to obtain the inter-frame differences of the n frames to be tested, which is Δ(f i f i+1 ), where i is a positive integer, i < n; the computer device will Δ(f i f i+1 ) and f1, ..., f n The visual attention layer in the input collision detection model performs a feature transformation based on visual attention, outputting n attention features to be tested as f1', ..., f2'. n '.
[0062] S204: Collision detection is performed on multiple attention features to be tested through the collision detection layer in the collision detection model to obtain multiple collision detection results.
[0063] S205: If multiple collision detection results represent multiple collision detection categories including collision failure categories, determine that the collision video to be tested includes collision failure images.
[0064] In this embodiment, after performing S201-S203 to obtain multiple attention features to be tested, based on the multiple attention features to be tested referring to the detailed features related to collision failure in the detected multiple frames of images to be tested, collision detection is performed on the multiple attention features to be tested to detect whether each frame of images to be tested belongs to the collision failure category. If there is an image to be tested that belongs to the collision failure category, the collision failure image in the collision video to be tested can be detected. Based on this, the multiple attention features to be tested are input into the collision detection layer in the collision detection model to perform collision detection and output multiple collision detection results. When the multiple collision detection categories represented by the multiple collision detection results include the collision failure category, it is determined that the collision video to be tested includes the collision failure image.
[0065] Here, the collision detection layer refers to the detection layer that detects whether each frame of the test image belongs to the collision failure category; collision detection refers to transforming multiple attention features to detect whether each frame of the test image belongs to the collision failure category; the collision detection result refers to the collision result of detecting whether the test image belongs to the collision failure category; the collision detection category refers to the collision category of the test image represented by the collision detection result; the collision failure category refers to the pre-defined collision category representing collision failure; and the collision failure image refers to the test image whose collision detection result represents the collision failure category.
[0066] Based on the multiple attention features to be tested, which are the detailed features related to collision failure in the multiple frames of the test images obtained by detection, the above S204-S205, through the collision detection layer in the collision detection model, achieves collision detection for multiple attention features to be tested and obtains multiple collision detection results. It can efficiently and accurately detect whether each frame of the test image belongs to the collision failure category. When the multiple collision detection categories represented by the multiple collision detection results include the collision failure category, it means that the multiple frames of the test images include test images belonging to the collision failure category. It can efficiently and accurately detect collision failure images in the collision video to be tested.
[0067] As an example of S204-S205, based on the examples of S201-S203 above, the computer device will... n The collision detection layer in the input collision detection model performs collision detection and outputs n collision detection results as p1, ..., p2. n When p1, ..., p n When the n collision detection categories represent a collision failure category, the computer device determines that the collision video X includes a collision failure image.
[0068] See Figure 3 , Figure 3 This diagram illustrates a collision failure detection step provided in an embodiment of this application. The collision failure detection step includes: feature extraction, difference calculation, feature transformation, collision detection, and failure determination. Feature extraction refers to extracting the visual features to be tested from each frame of the multi-frame test image of the collision video; difference calculation refers to calculating the feature differences between multiple test visual features corresponding to multiple test images to obtain the inter-frame differences corresponding to the multi-frame test images; feature transformation refers to inputting the inter-frame differences and multiple test visual features into the visual attention layer of the collision detection model for visual attention-based feature transformation, outputting multiple attention features to be tested; collision detection refers to inputting the multiple attention features into the collision detection layer of the collision detection model for collision detection, outputting multiple collision detection results; failure determination refers to determining that the collision video to be tested includes a collision failure image when multiple collision detection categories corresponding to multiple collision detection results include a collision failure category.
[0069] As can be seen from the above technical solution, there is no need for manual comparison of multiple test visual features corresponding to multiple test images of the collision video. The inter-frame differences are obtained by comparing the feature differences between multiple test visual features. The inter-frame differences can be considered by the visual attention layer. Multiple test attention features are obtained by performing feature transformation based on visual attention. The collision detection layer can then perform collision detection on the multiple test attention features to obtain multiple collision detection results. This makes collision failure detection more automatic and convenient, reduces detection errors, and determines that the collision video includes collision failure images when the multiple collision detection categories represented by the multiple collision detection results include collision failure categories. This achieves efficient and accurate collision failure detection.
[0070] In this embodiment, when calculating the feature differences between multiple visual features corresponding to multiple test images in S202 to obtain the inter-frame differences of the test images, considering that collisions are continuously and dynamically changing, the feature differences between the visual features of consecutive test images are calculated for the multiple test images. That is, the feature difference between the visual features of the i-th test image and the visual features of the (i+1)-th test image in the multiple test images is calculated as the inter-frame difference between the (i+1)-th test image and the i-th test image; i is a positive integer, i < n. Based on this, this application provides a possible implementation, S202 may include, for example, S202a (not shown in the figure): calculating the difference between the visual features of the i-th test image and the visual features of the (i+1)-th test image in the multiple test images to obtain the inter-frame difference between the (i+1)-th test image and the i-th test image; i is a positive integer, i < n.
[0071] The above-mentioned S202a calculates the feature differences between the visual features to be tested in consecutive frames of test images for multiple frames of test images, and obtains the inter-frame differences between consecutive frames of test images, which can more accurately detect the continuous dynamic changes of collision between multiple frames of test images, so as to more accurately detect the detailed features related to collision failure in multiple frames of test images in the future.
[0072] As an example of S202a, based on the examples of S201-S203 above, the computer device calculates the visual feature f to be tested in the i-th frame of the n-frame test images. i The visual features f to be tested in the (i+1)th frame of the image to be tested i+1 The feature difference between them is taken as Δ(f) between the (i+1)th frame test image and the i-th frame test image. i f i+1 ); where Δ(f) i f i+1 )=||f i -f i+1 ||2 Where i is a positive integer, i = 1, ..., n-1.
[0073] In this embodiment, when the above-mentioned S203 inputs the inter-frame difference to be tested and multiple visual features to be tested into the visual attention layer of the collision detection model for visual attention-based feature transformation and outputs multiple attention features to be tested, it is considered that the visual attention layer in the collision detection model has a feature weight matrix, which is used to perform visual attention-based feature transformation on multiple visual features to be tested. In addition, when performing visual attention-based feature transformation on multiple visual features to be tested, it is also necessary to pay attention to the feature dimension of each visual feature to be tested. Therefore, by using the feature weight matrix of the visual attention layer in the collision detection model, the feature dimension of each visual feature to be tested, and the inter-frame difference to be tested, multiple attention features to be tested can be obtained by transforming multiple visual features to be tested based on the visual attention mechanism. Based on this, this application provides a possible implementation method, S203 may include, for example, S203a (not shown in the figure): performing visual attention-based feature transformation on multiple visual features to be tested according to the feature weight matrix of the visual attention layer in the collision detection model, the feature dimension of each visual feature to be tested, and the inter-frame difference to be tested, to obtain multiple attention features to be tested.
[0074] Among them, the feature weight matrix of the visual attention layer refers to the weight matrix for performing feature transformation based on visual attention on multiple visual features to be tested; the feature dimension of the visual features to be tested refers to the number of feature components of the visual features to be tested.
[0075] Based on the consideration of inter-frame differences in the collision detection model through the visual attention layer, the above-mentioned S203a further utilizes the feature weight matrix of the visual attention layer in the collision detection model and further considers the feature dimension of each visual feature to be tested. It achieves feature transformation based on visual attention for multiple visual features to be tested to obtain multiple attention features to be tested. This can more accurately detect the detailed features related to collision failure in multiple frames of test images, so as to more accurately detect whether each frame of test image belongs to the collision failure category in the future.
[0076] As an example of S203a, based on the examples of S201-S203 above, the feature weight matrix of the visual attention layer in the collision detection model is FW, and the feature dimension d of each visual feature to be tested is... m Computer equipment via FW, d m and Δ(f) i f i+1 Based on the visual attention mechanism, transform f1, ..., f n We obtain f1', ..., f n '.
[0077] In addition, when performing visual attention-based feature transformation on multiple visual features to be tested, it is also necessary to pay attention to the time series corresponding to the multiple visual features to be tested in order to determine the temporal relationship between the multiple visual features to be tested.
[0078] In this embodiment of the application, when multiple collision detection categories corresponding to multiple collision detection results include a collision failure category, in the specific implementation of determining that the collision video to be tested includes a collision failure image in the above-mentioned S205, in order to reduce false detections, it can be further determined whether the detection confidence of the collision detection result corresponding to the collision failure category is greater than a preset confidence level. If so, it means that the multiple collision detection categories represented by the multiple collision detection results include a collision failure category, and the collision detection result corresponding to the collision failure category has a high degree of confidence, thus more accurately determining that the collision video to be tested includes a collision failure image. Based on this, this application provides a possible implementation, S205 may include, for example, S205a (not shown in the figure): if multiple collision detection categories include a collision failure category, and the detection confidence of the collision detection result corresponding to the collision failure category is greater than a preset confidence level, determine that the collision video to be tested includes a collision failure image.
[0079] Among them, the detection confidence of the collision detection result refers to the degree of credibility of the collision detection result; the pre-set confidence refers to the lower limit of credibility.
[0080] When multiple collision detection results represent multiple collision detection categories including collision failure categories, the above-mentioned S205a further determines that the detection confidence of the collision detection result corresponding to the collision failure category is greater than the preset confidence, so as to avoid the test image corresponding to the collision failure category with low confidence being detected as a collision failure image, thereby more accurately detecting the collision failure image in the test collision video.
[0081] As an example of S205a, based on the examples of S204-S205 above, the confidence level is preset to θ, and the computer device is at p1, ..., p n If the n collision detection categories include the collision failure category, determine whether the detection confidence of the collision detection result corresponding to the collision failure category is greater than θ. If so, determine that the collision video X includes the collision failure image.
[0082] In this embodiment of the application, the collision detection model in S203-S204 above is obtained by training an initial detection model based on the collision marker category of each frame of sample images in the sample collision video and the multi-frame sample images of the sample collision video; wherein, the initial detection model includes a visual attention layer and a collision detection layer. The process of obtaining the collision detection model is as follows: First, sample collision videos are acquired as training input data, and the collision marker category of each sample image in multiple frames of the sample collision videos is acquired as training target data. Second, the visual features of each sample image are extracted, and the feature differences between multiple visual features corresponding to multiple sample images are calculated to obtain the inter-frame differences of the samples corresponding to multiple sample images. Then, the inter-frame differences and multiple sample visual features are input into the visual attention layer of the initial detection model for feature transformation based on visual attention, and multiple sample attention features are output. Then, the multiple sample attention features are input into the collision detection layer of the initial detection model for collision detection, and multiple collision prediction results are output. Finally, for each sample image, the collision prediction results of the sample image are compared with the collision marker category of the sample image. The model parameters of the initial detection model are iteratively trained through the loss function of the initial detection model until the initial detection model converges or the number of iterations reaches a preset number to complete the training. The trained initial detection model is then used as the collision detection model. Based on this, this application provides a possible implementation method, and the steps for obtaining the collision detection model may include, for example, the following S1-S6 (not shown in the figure).
[0083] S1: Obtain the collision label category for each frame of the sample collision video and the multi-frame sample images of the sample collision video.
[0084] S2: Extract features from each frame of sample image to obtain the visual features of each frame of sample image.
[0085] S3: Perform difference calculation on the visual features of multiple samples corresponding to multiple frames of sample images to obtain the inter-frame differences of the samples corresponding to multiple frames of sample images.
[0086] S4: Using the visual attention layer in the initial detection model, perform visual attention-based feature transformation on the visual features of multiple samples according to the differences between sample frames to obtain multiple sample attention features.
[0087] S5: Collision detection is performed on the attention features of multiple samples through the collision detection layer in the initial detection model to obtain multiple collision prediction results.
[0088] S6: For each sample image, train the initial detection model based on the collision prediction result of the sample image, the collision marker category of the sample image, and the loss function of the initial detection model to obtain the collision detection model.
[0089] Among them, sample collision video refers to collision video including collision failure images labeled with collision failure categories; sample image refers to image obtained by segmenting the sample collision video into frames; feature extraction refers to extracting high-level, abstract visual feature representations of sample images; sample visual features refer to the extracted high-level, abstract visual feature representations of sample images; difference calculation refers to calculating the feature difference between two different sample visual features; sample frame difference refers to the feature difference between multiple calculated sample visual features; initial detection model refers to a model that learns to detect collision failure images in sample collision videos; visual attention layer refers to an attention layer that detects detailed features related to collision failure in multiple frames of sample images; based on visual attention... Attention feature transformation refers to transforming the visual features of multiple samples considering the differences between sample frames to detect detailed features related to collision failure in multi-frame sample images; multiple sample attention features refer to the detailed features related to collision failure in the detected multi-frame sample images; collision detection layer refers to the detection layer that detects whether each frame sample image belongs to the collision failure category; collision detection refers to transforming the attention features of multiple samples to detect whether each frame sample image belongs to the collision failure category; collision prediction result refers to the collision result of detecting whether the sample image belongs to the collision failure category; collision label category refers to the collision category of the pre-labeled sample image; model training refers to iteratively training the model parameters of the initial detection model until the initial detection model converges or the number of iterations reaches the preset number.
[0090] The above steps S1-S6, based on acquiring the collision marker category of each frame of the sample collision video and extracting the visual features of each frame of the sample collision video, obtain the inter-frame difference by comparing the feature differences between multiple visual features corresponding to multiple frames of sample images. This enables the detection of dynamic collision changes between multiple frames of sample images. Furthermore, by considering the inter-frame difference through the visual attention layer in the initial detection model, multiple sample attention features are obtained through visual attention-based feature transformation for multiple sample visual features. This enables the efficient and accurate detection of detailed features related to collision failure in multiple frames of sample images. Through the collision detection layer in the initial detection model, multiple collision prediction results are obtained by performing collision detection for multiple sample attention features. This enables the efficient and accurate detection of whether each frame of sample image belongs to the collision failure category. Based on the loss function of the initial detection model, the collision prediction results of the sample images are compared with the collision marker categories of the sample images. The model parameters of the initial detection model are iteratively trained, enabling the initial detection model to learn to detect collision failure images in the sample collision video, thus obtaining a collision detection model capable of detecting collision failure images in the test collision video.
[0091] As an example, the formula for calculating the collision prediction result by passing the visual features of the sample through the visual attention layer and the collision detection layer in the initial detection model is as follows:
[0092]
[0093] Where I represents the visual features of the sample image, FW represents the feature weight matrix of the visual attention layer in the initial detection model, and d m The feature dimension represents the visual features of the sample, softmax() represents the normalization function, VisualAttention() represents the visual attention function, and P(I) represents the collision prediction result.
[0094] In this embodiment, during the above-mentioned S6 step of comparing the collision prediction result of each sample image with the collision marker category of the sample image for each frame of sample image, and iteratively training the model parameters of the initial detection model through the loss function of the initial detection model to obtain the specific implementation of the collision detection model, considering that the collision prediction result of the sample image includes the collision prediction probability corresponding to the collision marker category of the sample image, it is necessary to maximize the collision prediction probability corresponding to the collision marker category of the sample image through the loss function of the initial detection model, so that the collision prediction category represented by the collision prediction result of the sample image is close to the collision marker category of the sample image to complete the training, thereby using the trained initial detection model as the collision detection model. Based on this, this application provides a possible implementation method, where the collision prediction result includes the collision prediction probability corresponding to the collision marker category; in S6, the initial detection model is trained according to the collision prediction result of the sample image, the collision marker category of the sample image, and the loss function of the initial detection model to obtain the collision detection model, for example, it may include S6a (not shown in the figure): maximizing the collision prediction probability according to the loss function of the initial detection model, training the initial detection model to obtain the collision detection model.
[0095] The collision prediction probability corresponding to the collision marker category refers to the predicted probability that the detected sample image belongs to the collision marker category.
[0096] Based on the collision prediction results of the sample image, which include the collision prediction probability corresponding to the collision marker category of the sample image, the above-mentioned S6a maximizes the collision prediction probability corresponding to the collision marker category of the sample image through the loss function of the initial detection model. This enables the collision prediction category represented by the collision prediction results of the sample image to be close to the collision marker category of the sample image, allowing the initial detection model to learn and detect collision failure images in the sample collision video more accurately, so as to obtain a collision detection model that can detect collision failure images in the collision video to be tested more accurately.
[0097] As an example, the loss function of the initial detection model can be the cross-entropy loss function, as shown below:
[0098]
[0099] Where c represents the collision category, M represents the number of collision categories, and P o,c y represents the collision prediction probability corresponding to sample image c. o,c Represents a binary indicator (0 or 1). If the collision marker category of the sample image is c, then y o,c If the collision marker category of the sample image is not c, then y is 1. o,c It is 0.
[0100] Correspondingly, the formula for calculating the model parameters of the initial detection model during iterative training is as follows:
[0101] μ=μ-η·Loss
[0102] η = η·decay_rate
[0103] Where μ represents the model parameters of the initial detection model, η represents the learning rate, and decay_rate represents the decay rate.
[0104] Furthermore, in this embodiment of the application, after determining in S201-S205 that the collision video to be tested includes collision failure images, in order to enhance the interpretability of the collision detection model and facilitate subsequent optimization of the collision video based on the collision failure images, the visual information of multiple frames of the collision video to be tested can be converted into text information. Based on this, multiple collision text features corresponding to multiple preset collision texts are introduced to convert visual information into text information. First, multiple visual features corresponding to multiple test images and multiple collision text features corresponding to multiple preset collision texts are fused to obtain multiple input fusion features. Then, a multimodal attention mechanism is introduced to transform the multiple input fusion features into multiple output attention features, which can characterize the detailed features related to collision texts in multiple test images. These multiple input fusion features are then input into the multimodal attention layer of the text generation model for feature transformation based on multimodal attention, outputting multiple output attention features. Based on this, text generation using the multiple output attention features can convert the visual information of each test image frame into text information. These multiple output attention features are then input into the text generation layer of the text generation model for text generation, outputting multiple collision description texts. Considering that multiple collision description texts are usually a sequence, the multiple collision description texts are input into a sequence generation model for sequence generation, outputting a sequence of collision texts corresponding to the multiple collision description texts. Based on this, this application provides a possible implementation method. After determining in step S205 that the test collision video includes collision failure images, see... Figure 4 , Figure 4 The flowchart of a method for generating a collision text sequence provided in the embodiments of this application includes the following steps S401-S404.
[0105] S401: Perform feature fusion on multiple visual features to be tested and multiple collision text features corresponding to multiple preset collision texts to obtain multiple input fusion features.
[0106] S402: By using the multimodal attention layer in the text generation model, multiple input fusion features are transformed based on multimodal attention to obtain multiple output attention features.
[0107] S403: By using the text generation layer in the text generation model, multiple output attention features are used to generate text, resulting in multiple collision description texts.
[0108] S404: Generate a sequence of multiple collision description texts using a sequence generation model to obtain a sequence of collision texts corresponding to the multiple collision description texts.
[0109] Among them, preset collision text refers to the pre-defined collision text describing the collision; collision text features refer to the high-level, abstract feature representation of the extracted preset collision text on the text; feature fusion refers to the fusion; input fusion features refer to the visual features to be tested and the collision text features; text generation model refers to the model that generates collision description text for the image to be tested; multimodal attention layer refers to the attention layer that detects detailed features related to collision text in multiple frames of the image to be tested; feature transformation based on multimodal attention refers to transforming multiple input fusion features to detect detailed features related to collision text in multiple frames of the image to be tested; multiple output attention features refer to the detailed features related to collision text in the detected multiple frames of the image to be tested; text generation layer refers to the generation layer that generates collision text describing the collision corresponding to multiple frames of the image to be tested; text generation refers to generating collision text describing the collision corresponding to multiple frames of the image to be tested; collision description text refers to the collision text describing the collision in the image to be tested; sequence generation model refers to; sequence generation refers to the model that generates a sequence of collision texts for multiple frames of the image to be tested; the sequence of collision texts corresponding to multiple collision description texts refers to the text sequence formed by multiple collision description texts corresponding to the generated multiple frames of the image to be tested.
[0110] As an example of S401-S404, the computer device integrates f1, ..., f n The features of multiple collision texts corresponding to multiple preset collision texts are used to obtain n input fusion features, namely f1'', ..., f2''. n ''; The computer equipment will f1'', ..., f n The input text generation model uses a multimodal attention layer to perform feature transformation based on multimodal attention, outputting n output attention features f1, f2, ..., f3.n '''; The computer equipment will f1''', ..., f n The input text generation model uses a text generation layer to generate text, and outputs n collision description texts s1, ..., s2. n Computer equipment will use s1, ..., s n Input a sequence generation model to generate sequences, and output s1, ..., s2. n The corresponding collision text sequence S n .
[0111] See Figure 5 , Figure 5 This diagram illustrates a step for generating a collision text sequence according to an embodiment of this application. The generation of the collision text sequence includes: feature fusion, feature transformation, text generation, and sequence generation. Feature fusion refers to fusing multiple visual features corresponding to multiple frames of test images and multiple collision text features corresponding to multiple preset collision texts to obtain multiple input fused features; feature transformation refers to inputting the multiple input fused features into the multimodal attention layer of the text generation model for multimodal attention-based feature transformation, outputting multiple output attention features; text generation refers to inputting the multiple output attention features into the text generation layer of the text generation model for text generation, outputting multiple collision description texts; sequence generation refers to inputting the multiple collision description texts into the sequence generation model for sequence generation, outputting a collision text sequence corresponding to the multiple collision description texts.
[0112] In this embodiment, when multiple input fusion features are input into the multimodal attention layer of the text generation model in S402 for feature transformation based on multimodal attention, and multiple output attention features are output, for each input fusion feature, considering that the multimodal visual attention layer in the text generation model has feature weights, these feature weights are used to transform the input fusion feature into query features, key features, and value features. The query features and key features are used to perform multimodal attention feature transformation on the value features. In addition, when performing multimodal attention feature transformation on the value features, the feature dimension of the key features also needs to be considered. Therefore, firstly, the input fusion feature is transformed into query features, key features, and value features based on the multimodal attention mechanism using the feature weights of the multimodal attention layer in the text generation model. Then, the value features are transformed into output attention features based on the multimodal attention mechanism using the feature dimensions of the query features, key features, and key features. Based on this, this application provides a possible implementation method, and S402 may include, for example, the following S402a-S402b (not shown in the figure).
[0113] S402a: For each input fusion feature, based on the feature weights of the multimodal attention layer in the text generation model, the input fusion feature is transformed using multimodal attention to obtain query features, key features, and value features.
[0114] S402b: Based on the query features, key features, and the feature dimensions of the key features, perform feature transformation on the value features based on multimodal attention to obtain the output attention features.
[0115] Among them, the feature weight of the multimodal attention layer refers to the weight of the feature transformation based on multimodal attention for the input fusion features; the query feature, key feature and value feature refer to the three features of the feature transformation based on multimodal attention for the input fusion features; the feature dimension of the key feature refers to the number of feature components of the key feature.
[0116] The aforementioned S402a-S402b can more accurately detect detailed features related to the collision text in multiple test images, so as to more accurately convert the visual information of each test image into text information in the subsequent process.
[0117] As an example, the formula for calculating the input fusion feature transformation into the output attention feature is as follows:
[0118]
[0119] Where Q represents the query feature, K T V represents the transpose of the key feature, V represents the value feature, and d represents the value feature. K The key feature represents the feature dimension, softmax() represents the normalization function, and Attention represents the output attention feature.
[0120] Correspondingly, the calculation formula for generating collision description text from the output attention features is as follows:
[0121] D = softmax(W) d Attention+b d )
[0122] Among them, W d b represents the weight matrix of the text generation layer in the text generation model. d W d The corresponding biases are: softmax() represents the normalization function, and D represents the collision description text.
[0123] In this embodiment of the application, when the above-mentioned S404 generates a sequence of multiple collision description text input sequences using a sequence generation model and outputs a collision text sequence corresponding to multiple collision description texts, considering that the sequence is continuously and dynamically generated, the collision text sequence corresponding to the (i+1)th collision description text is generated by the collision text sequence corresponding to the ith collision description text and the (i+1)th collision description text. Therefore, by generating a sequence of multiple collision description texts using a sequence generation model, the collision text sequence corresponding to the ith collision description text and the (i+1)th collision description text input sequence can be output as the collision text sequence corresponding to the (i+1)th collision description text. Based on this, this application provides a possible implementation method, S404 may include, for example, S404a (not shown in the figure): by using a sequence generation model, generating a sequence of multiple collision description texts using the collision text sequence corresponding to the ith collision description text and the (i+1)th collision description text, to obtain the collision text sequence corresponding to the (i+1)th collision description text, where i is a positive integer and i < n.
[0124] The above S402a-S402b can monitor the continuous dynamic changes of the collision text sequence corresponding to multiple collision description texts, so as to generate the collision text sequence corresponding to multiple collision description texts of multiple frames of test images more accurately.
[0125] As an example of S402a-S402b, based on the examples of S401-S404 above, the computer device will use s1, ..., s n Chinese i The corresponding collision text sequence S i and s i+1 Input a sequence generation model to generate a sequence, and output s i+1 The corresponding collision text sequence S i+1 Where i is a positive integer, i = 1, ..., n-1.
[0126] As an example, s i The corresponding S i and s i+1 Generate s i+1 The corresponding S i+1 The calculation formula is as follows:
[0127] S i+1 =Decoder(S i s i+1 )
[0128] Here, Decoder() represents the decoding function.
[0129] In this embodiment, the sequence generation model is obtained by training an initial generation model based on the sample text sequence corresponding to the j-th sample text, the (j+1)-th sample text, and the sample text sequence corresponding to the (j+1)-th sample text among multiple sample texts. The training process of this sequence generation model is as follows: First, the sample text sequence corresponding to the j-th sample text and the (j+1)-th sample text are obtained as input and output data for training, and the sample text sequence corresponding to the (j+1)-th sample text is obtained as the target data for training, where j is a positive integer; then, the sample text sequence corresponding to the j-th sample text and the (j+1)-th sample text are input into the initial generation model for sequence generation, and the sequence prediction result corresponding to the (j+1)-th sample text is output; finally, the sequence prediction result corresponding to the (j+1)-th sample text and the sample text sequence corresponding to the (j+1)-th sample text are compared, and the model parameters of the initial generation model are iteratively trained using the loss function of the initial generation model until the initial generation model converges or the number of iterations reaches a preset number to complete the training. The trained initial generation model is then used as the sequence generation model. Based on this, this application provides a possible implementation method, and the steps for obtaining the sequence generation model may include, for example, the following S7-S9 (not shown in the figure).
[0130] S7: Obtain the sample text sequence corresponding to the j-th sample text, the (j+1)-th sample text, and the sample text sequence corresponding to the (j+1)-th sample text from multiple sample texts; j is a positive integer.
[0131] S8: Using the initial generation model, generate the sequence of sample texts corresponding to the j-th sample text and the (j+1)-th sample text to obtain the sequence prediction result corresponding to the (j+1)-th sample text.
[0132] S9: Based on the sequence prediction result corresponding to the (j+1)th sample text, the sample text sequence corresponding to the (j+1)th sample text, and the loss function of the initial generation model, train the initial generation model to obtain the sequence generation model.
[0133] Based on obtaining the sample text sequence corresponding to the j-th sample text, the (j+1)-th sample text, and the sample text sequence corresponding to the (j+1)-th sample text in multiple sample texts, the initial generation model generates the sequence prediction result corresponding to the (j+1)-th sample text based on the sample text sequence corresponding to the j-th sample text and the (j+1)-th sample text. It can predict whether the predicted text sequence corresponding to the (j+1)-th sample text is the same as the sample text sequence corresponding to the (j+1)-th sample text. Based on the loss function of the initial generation model, the sequence prediction result corresponding to the (j+1)-th sample text and the sample text sequence corresponding to the (j+1)-th sample text are compared. The model parameters of the initial generation model are iteratively trained so that the initial generation model learns the correlation between the sample text sequence corresponding to the j-th sample text and the (j+1)-th sample text, in order to obtain a sequence generation model that can generate collision text sequences of multiple frames of test images.
[0134] In this embodiment of the application, when comparing the sequence prediction result corresponding to the (j+1)th sample text and the sample text sequence corresponding to the (j+1)th sample text in S9 above, and iteratively training the model parameters of the initial generation model through the loss function of the initial generation model to obtain the specific implementation of the sequence generation model, considering that the sequence prediction result corresponding to the (j+1)th sample text includes the sample text sequence corresponding to the first j sample texts and the prediction condition probability of generating the sample text sequence corresponding to the (j+1)th sample text under the condition of the (j+1)th sample text, it is necessary to maximize the prediction condition probability through the loss function of the initial generation model, so that the predicted text sequence represented by the sequence prediction result corresponding to the (j+1)th sample text is close to the sample text sequence corresponding to the (j+1)th sample text to complete the training, and thus the trained initial generation model is used as the sequence generation model. Based on this, this application provides a possible implementation method, wherein the sequence prediction result corresponding to the (j+1)th sample text includes the prediction conditional probability of generating the sample text sequence corresponding to the (j+1)th sample text under the condition of the sample text sequence corresponding to the first j sample texts and the (j+1)th sample text; S9 may include, for example, S9a (not shown in the figure): according to the loss function of the initial generation model, maximize the prediction conditional probability, train the initial generation model, and obtain the sequence generation model.
[0135] The sequence prediction result corresponding to the (j+1)th sample text in S9a mentioned above includes the prediction conditional probability of generating the (j+1)th sample text sequence under the conditions of the sample text sequences corresponding to the first j sample texts and the (j+1)th sample text. By maximizing the prediction conditional probability through the loss function of the initial generation model, the predicted text sequence represented by the sequence prediction result corresponding to the (j+1)th sample text can be made close to the sample text sequence corresponding to the (j+1)th sample text. This allows the initial generation model to learn more accurately the sample text sequence corresponding to the jth sample text and the correlation between the (j+1)th sample text and the sample text sequence corresponding to the (j+1)th sample text, so as to obtain a sequence generation model that can more accurately generate collision text sequences of multiple frames of test images.
[0136] As an example, the loss function for the initial generative model is as follows:
[0137]
[0138] Among them, s j Let S represent the (j+1)th sample text. <j+1 S represents the sequence of sample texts corresponding to the first j sample texts. j+1 Let P(S) represent the sequence of sample texts corresponding to the (j+1)th sample text. j+1 |S <j+1 s j ) represents the predicted conditional probability of generating the sample text sequence corresponding to the (j+1)th sample text given the sample text sequence corresponding to the first j sample texts and the (j+1)th sample text, where m represents the number of sample texts.
[0139] It should be noted that, based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods.
[0140] based on Figure 2 In accordance with the collision failure detection method provided in the corresponding embodiments, this application also provides a collision failure detection device, see [link to relevant documentation]. Figure 6 , Figure 6 The present application provides a structural diagram of a collision failure detection device 600, which includes: an extraction unit 601, a calculation unit 602, a transformation unit 603, a detection unit 604, and a determination unit 605.
[0141] The extraction unit 601 is used to extract features from each frame of the test image in the multi-frame test image of the collision video to obtain the test visual features of each frame of the test image.
[0142] The calculation unit 602 is used to perform difference calculation on multiple visual features corresponding to multiple frames of images to be tested, and to obtain the inter-frame differences of the multiple frames of images to be tested.
[0143] Transformation unit 603 is used to perform visual attention-based feature transformation on multiple visual features to be tested according to the inter-frame differences in the collision detection model through the visual attention layer, so as to obtain multiple attention features to be tested.
[0144] The detection unit 604 is used to perform collision detection on multiple attention features to be tested through the collision detection layer in the collision detection model, and obtain multiple collision detection results.
[0145] The determining unit 605 is used to determine that the collision video to be tested includes a collision failure image if the multiple collision detection categories represented by multiple collision detection results include a collision failure category.
[0146] In one possible implementation, the computing unit 602 is used for:
[0147] The difference between the visual features to be tested in the i-th frame and the visual features to be tested in the (i+1)-th frame is calculated to obtain the inter-frame difference between the (i+1)-th frame and the i-th frame; i is a positive integer, i < n.
[0148] In one possible implementation, the transformation unit 603 is used for:
[0149] Based on the feature weight matrix of the visual attention layer in the collision detection model, the feature dimension of each visual feature to be tested, and the differences between frames to be tested, a feature transformation based on visual attention is performed on multiple visual features to be tested to obtain multiple attention features to be tested.
[0150] In one possible implementation, determining unit 605 is used for:
[0151] If multiple collision detection categories include a collision failure category, and the detection confidence of the collision detection result corresponding to the collision failure category is greater than the preset confidence, then the collision video to be tested is determined to include a collision failure image.
[0152] In one possible implementation, the collision failure detection device 600 further includes: a first training unit;
[0153] The first training unit is used for:
[0154] Obtain the collision label category for each frame of the sample collision video and the multi-frame sample images of the sample collision video;
[0155] Feature extraction is performed on each frame of sample image to obtain the sample visual features of each frame of sample image;
[0156] The difference between the visual features of multiple samples corresponding to multiple frames of sample images is calculated to obtain the inter-frame difference of the samples corresponding to multiple frames of sample images.
[0157] By using the visual attention layer in the initial detection model, visual attention-based feature transformation is performed on the visual features of multiple samples according to the differences between sample frames, and attention features of multiple samples are obtained.
[0158] By using the collision detection layer in the initial detection model, collision detection is performed on the attention features of multiple samples to obtain multiple collision prediction results.
[0159] For each sample image, the initial detection model is trained based on the collision prediction results of the sample image, the collision marker category of the sample image, and the loss function of the initial detection model to obtain the collision detection model.
[0160] In one possible implementation, the collision prediction result includes the collision prediction probability corresponding to the collision marker category;
[0161] The first training unit is used for:
[0162] The initial detection model is trained to maximize the collision prediction probability by using the loss function of the initial detection model, thus obtaining the collision detection model.
[0163] In one possible implementation, the collision failure detection device 600 further includes: a fusion unit and a generation unit;
[0164] The fusion unit is used to fuse multiple visual features to be tested and multiple collision text features corresponding to multiple preset collision texts to obtain multiple input fusion features.
[0165] The transformation unit 603 is also used to perform feature transformation based on multimodal attention on multiple input fusion features through the multimodal attention layer in the text generation model to obtain multiple output attention features;
[0166] The generation unit is used to generate text from multiple output attention features through the text generation layer in the text generation model, thereby obtaining multiple collision description texts.
[0167] The generation unit is also used to generate a sequence of multiple collision description texts through a sequence generation model, thereby obtaining a collision text sequence corresponding to the multiple collision description texts.
[0168] In one possible implementation, the transformation unit 603 is used for:
[0169] For each input fusion feature, based on the feature weights of the multimodal attention layer in the text generation model, a feature transformation based on multimodal attention is performed on the input fusion feature to obtain query features, key features, and value features;
[0170] Based on the query features, key features, and the feature dimensions of the key features, the value features are transformed using multimodal attention to obtain the output attention features.
[0171] In one possible implementation, the generating unit is used for:
[0172] Using a sequence generation model, the collision text sequence corresponding to the i-th collision description text and the (i+1)-th collision description text are generated to obtain the collision text sequence corresponding to the (i+1)-th collision description text, where i is a positive integer and i < n.
[0173] In one possible implementation, the collision failure detection device 600 further includes: a second training unit;
[0174] The second training unit is used for:
[0175] Obtain the sample text sequence corresponding to the j-th sample text, the (j+1)-th sample text, and the sample text sequence corresponding to the (j+1)-th sample text from multiple sample texts; j is a positive integer;
[0176] Using the initial generation model, the sequence of sample texts corresponding to the j-th sample text and the (j+1)-th sample text are generated to obtain the sequence prediction result corresponding to the (j+1)-th sample text.
[0177] Based on the sequence prediction result corresponding to the (j+1)th sample text, the sample text sequence corresponding to the (j+1)th sample text, and the loss function of the initial generation model, the initial generation model is trained to obtain the sequence generation model.
[0178] In one possible implementation, the sequence prediction result corresponding to the (j+1)th sample text includes the prediction conditional probability of generating the sample text sequence corresponding to the (j+1)th sample text under the conditions of the sample text sequence corresponding to the first j sample texts and the (j+1)th sample text.
[0179] The second training unit is used for:
[0180] By maximizing the predicted conditional probability using the loss function of the initial generative model, the initial generative model is trained to obtain the sequence generation model.
[0181] As can be seen from the above technical solution, the collision failure detection device does not require manual comparison of multiple test visual features corresponding to multiple test images of the collision video under test. It obtains the inter-frame differences by comparing the feature differences between multiple test visual features. The inter-frame differences can be considered by the visual attention layer. Multiple test attention features are obtained by performing feature transformation based on visual attention on multiple test visual features. The collision detection layer can then perform collision detection on multiple test attention features to obtain multiple collision detection results. This makes collision failure detection more automatic and convenient, reduces detection errors, and determines that the collision video under test includes collision failure images when multiple collision detection categories represented by multiple collision detection results include collision failure categories. This achieves efficient and accurate collision failure detection.
[0182] This application also provides a computer device, which may be a server, see [link to previous document]. Figure 7 , Figure 7 This application provides a structural diagram of a server 700. The server 700 can vary significantly due to different configurations or performance characteristics. It may include one or more processors, such as a central processing unit (CPU) 722, and a memory 732, as well as one or more storage media 730 (e.g., one or more mass storage devices) for storing application programs 742 or data 744. The memory 732 and storage media 730 can be temporary or persistent storage. The program stored in the storage media 730 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 722 may be configured to communicate with the storage media 730 and execute the series of instruction operations stored in the storage media 730 on the server 700.
[0183] Server 700 may also include one or more power supplies 726, one or more wired or wireless network interfaces 750, one or more input / output interfaces 758, and / or one or more operating systems 741, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.
[0184] In this embodiment, the central processing unit 722 in the server 700 can execute the methods provided in the various optional implementations of the above embodiments.
[0185] The computer device provided in this application embodiment can also be a terminal, see [link to relevant documentation]. Figure 8 , Figure 8This is a structural diagram of a terminal provided in an embodiment of this application. Taking a smartphone as an example, the smartphone includes components such as a radio frequency (RF) circuit 810, a memory 820, an input unit 830, a display unit 840, a sensor 850, an audio circuit 860, a wireless Fidelity (WiFi) module 870, a processor 880, and a power supply 890. The input unit 830 may include a touch panel 831 and other input devices 832, the display unit 840 may include a display panel 841, and the audio circuit 860 may include a speaker 861 and a microphone 862. Those skilled in the art will understand that... Figure 8 The smartphone structure shown does not constitute a limitation on smartphones and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0186] The memory 820 can be used to store software programs and modules. The processor 880 executes various functions and data processing of the smartphone by running the software programs and modules stored in the memory 820. The memory 820 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, phonebook, etc.). In addition, the memory 820 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0187] The processor 880 is the control center of the smartphone, connecting various parts of the smartphone via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 820, and by accessing data stored in the memory 820. Optionally, the processor 880 may include one or more processing units; preferably, the processor 880 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 880.
[0188] In this embodiment, the processor 880 in the smartphone can execute the methods provided in the various optional implementations of the above embodiments.
[0189] According to one aspect of this application, a computer-readable storage medium is provided for storing a computer program that, when run on a computer device, causes the computer device to perform the methods provided in various optional implementations of the above embodiments.
[0190] According to one aspect of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.
[0191] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0192] The terms "first," "second," etc., used in this application's specification and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0193] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0194] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0195] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0196] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing computer programs, such as USB flash drives, portable hard drives, read-only memory (ROM), RAM, magnetic disks, or optical disks.
[0197] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0198] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A collision failure detection method, characterized in that, The method includes: Feature extraction is performed on each frame of the collision video to be tested to obtain the visual features to be tested for each frame of the collision video. Difference calculation is performed on multiple visual features corresponding to the multiple frames of images to be tested to obtain the inter-frame differences of the multiple frames of images to be tested. By using the visual attention layer in the collision detection model, multiple visual features to be tested are transformed based on visual attention according to the inter-frame differences to be tested, thereby obtaining multiple attention features to be tested. The collision detection layer in the collision detection model is used to perform collision detection on the multiple attention features to be tested, and multiple collision detection results are obtained. If the multiple collision detection results represent multiple collision detection categories including a collision failure category, then the collision video to be tested is determined to include a collision failure image.
2. The method according to claim 1, characterized in that, The step of calculating the difference between multiple visual features corresponding to the multiple frames of images to be tested, and obtaining the inter-frame differences of the multiple frames of images to be tested, includes: The difference between the visual features to be tested in the i-th frame and the visual features to be tested in the (i+1)-th frame is calculated to obtain the inter-frame difference between the (i+1)-th frame and the i-th frame; i is a positive integer, i < n.
3. The method according to claim 1, characterized in that, The collision detection model uses a visual attention layer to perform a visual attention-based feature transformation on the multiple visual features to be tested based on the inter-frame differences, thereby obtaining multiple attention features to be tested, including: Based on the feature weight matrix of the visual attention layer in the collision detection model, the feature dimension of each visual feature to be tested, and the inter-frame difference to be tested, a feature transformation based on visual attention is performed on the multiple visual features to be tested to obtain the multiple attention features to be tested.
4. The method according to claim 1, characterized in that, If the multiple collision detection results represent multiple collision detection categories, including a collision failure category, determining that the collision video to be tested includes a collision failure image includes: If the multiple collision detection categories include the collision failure category, and the detection confidence of the collision detection result corresponding to the collision failure category is greater than the preset confidence, then the collision video to be tested is determined to include the collision failure image.
5. The method according to any one of claims 1-4, characterized in that, The steps for obtaining the collision detection model include: Obtain the collision marker category of each frame of the sample collision video and the multi-frame sample images of the sample collision video; Feature extraction is performed on each frame of the sample image to obtain the sample visual features of each frame of the sample image; The difference calculation is performed on the visual features of multiple samples corresponding to the multi-frame sample images to obtain the inter-frame differences of the samples corresponding to the multi-frame sample images. By using the visual attention layer in the initial detection model, the visual features of the multiple samples are transformed based on visual attention according to the inter-frame differences of the samples to obtain multiple sample attention features. The collision detection layer in the initial detection model is used to perform collision detection on the attention features of the multiple samples to obtain multiple collision prediction results. For each sample image, the initial detection model is trained based on the collision prediction result of the sample image, the collision marker category of the sample image, and the loss function of the initial detection model to obtain the collision detection model.
6. The method according to claim 5, characterized in that, The collision prediction result includes the collision prediction probability corresponding to the collision marker category; the step of training the initial detection model based on the collision prediction result of the sample image, the collision marker category of the sample image, and the loss function of the initial detection model to obtain the collision detection model includes: The collision detection model is obtained by maximizing the collision prediction probability using the loss function of the initial detection model.
7. The method according to claim 1, characterized in that, After determining that the collision video to be tested includes collision failure images, the method further includes: The multiple visual features to be tested and the multiple collision text features corresponding to the multiple preset collision texts are fused to obtain multiple input fusion features; By using the multimodal attention layer in the text generation model, feature transformation based on multimodal attention is performed on the multiple input fusion features to obtain multiple output attention features; By using the text generation layer in the text generation model, multiple collision description texts are obtained by generating text from the multiple output attention features. The multiple collision description texts are generated into a sequence using a sequence generation model to obtain the collision text sequence corresponding to the multiple collision description texts.
8. The method according to claim 7, characterized in that, The text generation model uses a multimodal attention layer to perform feature transformation based on multimodal attention on the multiple input fusion features to obtain multiple output attention features, including: For each input fusion feature, based on the feature weights of the multimodal attention layer in the text generation model, a feature transformation based on multimodal attention is performed on the input fusion feature to obtain query features, key features, and value features; Based on the query features, the key features, and the feature dimensions of the key features, the value features are transformed using multimodal attention to obtain the output attention features.
9. The method according to claim 7, characterized in that, The step of generating a sequence of collision description texts using a sequence generation model to obtain a collision text sequence corresponding to the multiple collision description texts includes: Using the sequence generation model, the collision text sequence corresponding to the i-th collision description text and the (i+1)-th collision description text are generated to obtain the collision text sequence corresponding to the (i+1)-th collision description text, where i is a positive integer and i < n.
10. The method according to any one of claims 7-9, characterized in that, The steps for obtaining the sequence generation model include: Obtain the sample text sequence corresponding to the j-th sample text, the (j+1)-th sample text, and the sample text sequence corresponding to the (j+1)-th sample text from multiple sample texts; j is a positive integer; Using the initial generation model, sequence generation is performed on the sample text sequence corresponding to the j-th sample text and the (j+1)-th sample text to obtain the sequence prediction result corresponding to the (j+1)-th sample text; Based on the sequence prediction result corresponding to the (j+1)th sample text, the sample text sequence corresponding to the (j+1)th sample text, and the loss function of the initial generation model, the initial generation model is trained to obtain the sequence generation model.
11. The method according to claim 10, characterized in that, The sequence prediction result corresponding to the (j+1)th sample text includes the predicted conditional probability of generating the sample text sequence corresponding to the (j+1)th sample text under the conditions of the sample text sequence corresponding to the first j sample texts and the (j+1)th sample text; the step of training the initial generation model based on the sequence prediction result corresponding to the (j+1)th sample text, the sample text sequence corresponding to the (j+1)th sample text, and the loss function of the initial generation model to obtain the sequence generation model includes: The initial generation model is trained by maximizing the predicted conditional probability using the loss function of the initial generation model to obtain the sequence generation model.
12. A collision failure detection device, characterized in that, The device includes: an extraction unit, a calculation unit, a transformation unit, a detection unit, and a determination unit; The extraction unit is used to extract features from each frame of the test image in the multi-frame test image of the collision video to obtain the test visual features of each frame of the test image. The calculation unit is used to perform difference calculation on multiple visual features corresponding to the multiple frames of images to be tested, and to obtain the inter-frame differences of the multiple frames of images to be tested. The transformation unit is used to perform visual attention-based feature transformation on the multiple visual features to be tested according to the inter-frame differences in the collision detection model through the visual attention layer, so as to obtain multiple attention features to be tested. The detection unit is used to perform collision detection on the multiple attention features to be tested through the collision detection layer in the collision detection model, and obtain multiple collision detection results; The determining unit is configured to determine that the collision video to be tested includes a collision failure image if the multiple collision detection categories represented by the multiple collision detection results include a collision failure category.
13. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to execute the method according to any one of claims 1-11 according to instructions in the computer program.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that, when run on a computer device, causes the computer device to perform the method according to any one of claims 1-11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is run on a computer device, it causes the computer device to perform the method according to any one of claims 1-11.