Scene retrieval method, device, equipment and product for driving video

By extracting the timing fusion features in driving videos and matching them, the problem of high complexity in driving video feature extraction in the prior art is solved, and efficient feature extraction and scene retrieval are achieved.

CN119938984APending Publication Date: 2025-05-06ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510017592.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the preprocessing and feature extraction of driving videos are of high complexity, resulting in low feature extraction efficiency.

Method used

By obtaining driving video data, the timing fusion features are extracted, including the fused timing features and image features, and sorting them into a timing feature sequence, matching the similarity between the scene feature vector and the timing feature sequence, and outputting the scene search results.

Benefits of technology

It reduces the complexity of feature extraction and improves the efficiency of feature extraction, and is suitable for application scenarios of real-time monitoring and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938984A_ABST
    Figure CN119938984A_ABST
Patent Text Reader

Abstract

The invention provides a scene retrieval method and device for a driving video, equipment and a product. The method comprises the following steps: acquiring driving video data, and acquiring a scene feature vector of a target retrieval scene; the driving video data comprises a plurality of video frames; extracting time sequence fusion features from each video frame, wherein the time sequence fusion features comprise fused time sequence features and image features; the image features are used for representing events in the video frames; sorting the time sequence fusion features into a time sequence feature sequence according to a time sequence; wherein each time sequence feature sequence comprises a plurality of time sequence fusion features with the same image features but different time sequence features; obtaining a matching result according to the similarity between the scene feature vector and a time sequence fusion feature in the time sequence feature sequence; the matching result at least comprises a scene category corresponding to the time sequence feature sequence and a probability that the time sequence feature sequence is the scene category; and outputting a scene retrieval result according to a matching result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent transportation technology, and in particular to a scene retrieval method, device, equipment and product for driving video. Background Art

[0002] Scene retrieval and positioning technology in driving videos is an important research direction in the field of intelligent transportation. It aims to accurately retrieve and locate specific scenes by analyzing and processing driving videos. With the continuous development of intelligent driving technology and smart transportation technology, the research and application of scene retrieval and positioning technology in driving videos are also receiving more and more attention.

[0003] In order to achieve scene retrieval in driving videos, researchers usually use technical means such as computer vision and deep learning to process and analyze driving videos. By modeling and learning the information extracted from the video sequences of driving videos, they can more accurately identify and track various target objects on the road and achieve positioning and retrieval of specific scenes.

[0004] In related video data retrieval schemes, video dynamic information is usually decomposed into RGB information and optical flow information, and background information and action information are processed separately. In this scheme, the complexity of data preprocessing and feature extraction is high, resulting in low efficiency of feature extraction. Summary of the invention

[0005] In view of this, an embodiment of the present invention is directed to providing a scene retrieval method for driving videos, so as to solve the problem of high complexity of driving video preprocessing and feature extraction in the prior art.

[0006] According to a first aspect of an embodiment of the present application, a scene retrieval method for a driving video is provided, the method comprising:

[0007] Acquire driving video data and obtain a scene feature vector of a target retrieval scene; the driving video data includes a plurality of video frames;

[0008] Extracting temporal fusion features from each video frame respectively, the temporal fusion features include fused temporal features and image features; the image features are used to characterize events in the video frames;

[0009] Arrange the time series fusion features into a time series feature sequence according to the time series order; wherein each time series feature sequence includes a plurality of time series fusion features with the same image features but different time series features;

[0010] A matching result is obtained according to the similarity between the scene feature vector and the temporal fusion feature in the temporal feature sequence; the matching result at least includes the scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is the scene category;

[0011] Output scene retrieval results based on the matching results.

[0012] In a possible design of the first aspect, obtaining a scene feature vector of a target retrieval scene includes:

[0013] Receiving input of a scene text description of a target retrieval scene;

[0014] According to the text feature extraction algorithm, a scene feature vector is obtained from the scene text description.

[0015] In a possible design of the first aspect, extracting temporal fusion features from each video frame includes:

[0016] For each video frame, a first feature is extracted from the video frame through a spatial attention mechanism, a second feature is extracted from the video frame through a channel attention mechanism, and a temporal feature is extracted from the video frame;

[0017] The first feature and the second feature are concatenated to obtain an image feature;

[0018] The image features are fused with the time series features to obtain the time series fusion features.

[0019] In a possible design of the first aspect, the time series fusion features are sorted into a time series feature sequence according to the time series order, including:

[0020] Start multiple threads to record different events in the video frame in chronological order, wherein the number of threads is the same as the number of events in the video frame;

[0021] The time series fusion features corresponding to the events recorded by each thread are determined as a time series feature sequence.

[0022] In a possible design of the first aspect, obtaining a matching result according to a similarity between a scene feature vector and a temporal fusion feature in a temporal feature sequence includes:

[0023] Calculate the similarity between the scene feature vector and the image features in the temporal fusion feature;

[0024] According to the similarity and weighted voting strategy, the scene category corresponding to the time series feature sequence and the probability that the time series feature sequence is a scene category are obtained;

[0025] The scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is a scene category are determined as the matching result.

[0026] In a possible design of the first aspect, before calculating the similarity between the scene feature vector and the image feature in the temporal fusion feature, the method further includes:

[0027] The time series fusion features in the time series feature sequence are processed to reduce the dimension.

[0028] In a possible design of the first aspect, outputting a scene retrieval result according to the matching result includes:

[0029] When the probability that the time series feature sequence is a scene category is greater than a preset probability threshold, determining the time period corresponding to the scene category;

[0030] Extracting video clips corresponding to the time period from the driving video data;

[0031] Output the scene category, the time period corresponding to the scene category, and the captured video clip.

[0032] According to a second aspect of an embodiment of the present application, a scene retrieval device for a driving video is provided, the device comprising:

[0033] An acquisition module is used to acquire driving video data and obtain a scene feature vector of a target retrieval scene; the driving video data includes a plurality of video frames;

[0034] An extraction module is used to extract time series fusion features from each video frame respectively, where the time series fusion features include fused time series features and image features; the image features are used to characterize events in the video frames;

[0035] A sorting module, used to sort the time series fusion features into a time series feature sequence in time series order; wherein each time series feature sequence includes a plurality of time series fusion features with the same image features but different time series features;

[0036] A matching module, used to obtain a matching result according to the similarity between the scene feature vector and the temporal fusion feature in the temporal feature sequence; the matching result at least includes the scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is a scene category;

[0037] The output module is used to output the scene retrieval results according to the matching results.

[0038] According to a third aspect of an embodiment of the present application, there is provided an electronic device, including a memory and a processor;

[0039] The memory is connected to the processor and is used for storing programs;

[0040] The processor is used to implement a scene retrieval method for driving video as described in any one of the first aspects of the embodiments of the present application by running a program in the memory.

[0041] According to a fourth aspect of an embodiment of the present application, a computer program product is provided, comprising computer program instructions. When the computer program instructions are executed by a processor, the processor executes a scene retrieval method for driving video as described in any one of the first aspects of the embodiment of the present application.

[0042] According to the fifth aspect of the embodiments of the present application, a chip is provided, including a processor and a data interface, and the processor reads and runs the program stored in the memory through the data interface to execute the scene retrieval method for driving video as any one of the first aspects of the embodiments of the present application.

[0043] According to a sixth aspect of an embodiment of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, a scene retrieval method for driving video as described in any one of the first aspects of the embodiment of the present application is implemented.

[0044] According to the technical solution of the embodiment of the present application, by obtaining driving video data, and obtaining the scene feature vector of the target retrieval scene; the driving video data includes multiple video frames; the temporal fusion features are extracted from each video frame respectively, and the temporal fusion features include fused temporal features and image features; the image features are used to characterize the events in the video frames. In this way, by introducing temporal features into image features, the ability to understand dynamic behaviors in the video can be significantly enhanced. The temporal fusion features are sorted into a temporal feature sequence in a temporal order; wherein each temporal feature sequence includes a plurality of temporal fusion features with the same image features but different temporal features; the matching result is obtained according to the similarity between the scene feature vector and the temporal fusion features in the temporal feature sequence. By matching real-time features, it is possible to quickly respond to user needs and can be applied to application scenarios of real-time monitoring and analysis. The matching result includes at least the scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is a scene category; the scene retrieval result is output according to the matching result. It can be seen that the technical solution of the embodiment of the present application does not undergo complex feature extraction preprocessing. Compared with the traditional RGB feature extraction and optical flow feature distribution extraction solutions, it reduces the complexity of feature extraction and improves the efficiency of feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0046] Figure 1A schematic diagram of a flow chart of a scene retrieval method for a driving video provided in an embodiment of the present application;

[0047] Figure 2 A schematic diagram of a scene retrieval method for driving video provided in an embodiment of the present application;

[0048] Figure 3 A schematic diagram of a flow chart of matching a scene feature vector with a Shi Xu fusion feature in a scene retrieval method for a driving video provided in an embodiment of the present application;

[0049] Figure 4 A schematic diagram of the structure of a scene retrieval device for driving video provided in an embodiment of the present application;

[0050] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0052] The technical terms involved in the embodiments of the present application are explained below:

[0053] Spatial attention mechanism: It is an attention mechanism used in deep learning to improve model performance. Its goal is to make the deep learning model focus on important areas in the image. The spatial attention mechanism can dynamically enhance or weaken the features of different areas in the image, thereby helping the deep learning model to better understand the main content of the image, thereby improving the model's feature expression ability.

[0054] Channel attention mechanism: It is an attention mechanism used in deep learning to improve model performance. It focuses on the importance of feature channels and allows deep learning models to adaptively weight the dimensions of feature channels. For example, some feature channels may appear more important in some characters, so the channel attention mechanism can enhance the influence of these feature channels and suppress redundant information.

[0055] Text feature extraction algorithm: It is an important step in natural language processing, which is used to convert text data into numerical form so that machine learning models can understand and process it. Common text feature extraction algorithms include bag of words model, word frequency-inverse document frequency, and word embedding.

[0056] Voting strategy: It is a decision-making mechanism commonly used in ensemble learning to determine the final prediction results. The voting strategy can be majority voting, weighted voting, or soft voting based on the prediction results of the model. Among them, the weighted voting strategy assigns a weight to each base model, and the final prediction result is weighted voting based on the weight.

[0057] Average pooling: It is a downsampling technique commonly used in deep learning. It slides a window (usually called a pooling window) on the feature map, calculates the average value of all pixel values ​​in the window, and then replaces the pixel value at the center of the window with this average value to generate a new feature map.

[0058] Maximum pooling: It is a downsampling technique commonly used in deep learning. It slides a pooling window on the feature map, finds the maximum value of all pixel values ​​in the window, and then replaces the pixel value at the center of the window with this maximum value to generate a new feature map.

[0059] Scene retrieval and positioning technology in driving videos is an important research direction in the field of intelligent transportation. It aims to accurately retrieve and locate specific scenes by analyzing and processing driving videos. With the continuous development of intelligent driving technology and smart transportation technology, the research and application of scene retrieval and positioning technology in driving videos are also receiving more and more attention.

[0060] In order to achieve scene retrieval in driving videos, researchers usually use computer vision, deep learning and other technical means to process and analyze driving videos. By modeling and learning the information extracted from the video sequence of driving videos, various target objects on the road can be more accurately identified and tracked, and the positioning and retrieval of specific scenes can be achieved. These classified specific scenes can support more downstream tasks, such as labeling specific scene data, training models with labeled specific scene data to improve the perception ability of intelligent driving vehicles; reconstructing specific scenes based on specific scene data, such as 3D reconstruction of scenes for editing and generalization, and deriving more long-tail scenes as input for intelligent driving and intelligent transportation system research.

[0061] In addition to its application in intelligent driving and intelligent transportation systems, the specific scene retrieval and positioning technology of driving video also has a wide range of application prospects. In the construction of smart cities, by using driving video for traffic monitoring and management, it can help urban management departments better plan urban traffic, reduce traffic congestion and traffic accidents, and improve the efficiency of urban traffic operation. In the field of intelligent security, the specific scene retrieval and positioning technology of driving video can also be used to monitor and identify criminal behavior and improve the level of social security.

[0062] In related video data retrieval schemes, video dynamic information is usually decomposed into RGB information and optical flow information, and background information and action information are processed separately. In this scheme, the complexity of data preprocessing and feature extraction is high, resulting in low efficiency of feature extraction.

[0063] Based on this, the present application proposes a new scene retrieval method for driving videos, which can extract the temporal features and image features in the video frames according to the acquired driving video data, and extract the scene feature vector of the target retrieval scene. After fusing the image features with the temporal features, they are matched and calculated with the scene feature vector in turn, and finally the corresponding scene category, the probability of the scene category and the time period are output, thereby constructing a method for scene retrieval and automatic capture of driving videos to improve the efficiency of feature extraction and reduce the complexity of feature extraction.

[0064] Specifically, in the present application, by obtaining driving video data and obtaining the scene feature vector of the target retrieval scene; extracting the temporal fusion features that fuse the temporal features and the image features from each video frame. In this way, by introducing the temporal features into the image features, the ability to understand the dynamic behavior in the video can be significantly enhanced. Afterwards, the temporal fusion features are sorted into a temporal feature sequence in a temporal order; wherein each temporal feature sequence includes a plurality of temporal fusion features with the same image features but different temporal features; and the matching result is obtained according to the similarity between the scene feature vector and the temporal fusion features in the temporal feature sequence. Thus, by matching the real-time features, the user's needs can be quickly responded to, so that the scheme of the present application can be applied to the application scenarios of real-time monitoring and analysis. Then, the scene retrieval result is output according to the matching result. It can be seen that the technical scheme of the embodiment of the present application, without complex feature extraction preprocessing, reduces the complexity of feature extraction and improves the efficiency of feature extraction compared with the traditional RGB feature extraction and optical flow feature distribution extraction schemes.

[0065] The following specific embodiments are used to describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0066] For example, Figure 1 A schematic diagram of a flow chart of a scene retrieval method for a driving video provided in an embodiment of the present application, Figure 2 Schematic diagram of a scene retrieval method for driving video provided in an embodiment of the present application. Figure 1 and Figure 2In an exemplary embodiment, a scene retrieval method for a driving video may include the following steps 1100 to 1500:

[0067] 1100. Obtain driving video data and obtain a scene feature vector of a target retrieval scene; the driving video data includes a plurality of video frames.

[0068] 1200. Extracting temporal fusion features from each video frame respectively, wherein the temporal fusion features include fused temporal features and image features; the image features are used to characterize events in the video frames.

[0069] 1300. Arrange the time series fusion features into a time series feature sequence in time series order; wherein each time series feature sequence includes a plurality of time series fusion features having the same image features but different time series features.

[0070] 1400. Obtain a matching result based on the similarity between the scene feature vector and the temporal fusion feature in the temporal feature sequence; the matching result at least includes the scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is the scene category.

[0071] 1500. Output the scene retrieval result according to the matching result.

[0072] Usually, driving videos are taken by on-board cameras, such as driving recorders, while the vehicle is moving. Obtaining driving video data from the driving recorder for retrieval of specific scenes is helpful for studying driver behavior, time backtracking, and evidence support.

[0073] After obtaining the driving video data, it is necessary to extract the temporal fusion features for each video frame in the driving video data. Specifically, for each video frame, the first feature is extracted from the video frame through the spatial attention mechanism, the second feature is extracted from the video frame through the channel attention mechanism, and the temporal feature is extracted from the video frame; the first feature and the second feature are spliced ​​to obtain the image feature; the image feature is fused with the temporal feature to obtain the temporal fusion feature. Figure 2 As shown, the process of extracting the first feature and the second feature can be performed by an image feature extraction module, for example.

[0074] It should be noted that the goal of the spatial attention mechanism is to make the model focus on important areas in the image. In this embodiment, by introducing spatial attention, the features of different areas in the image can be dynamically enhanced or weakened. For example, when processing a video frame containing a person and a background, the use of the spatial attention mechanism will highlight the characteristics of the person and suppress the background noise, thereby helping the model to better understand the main content of the image. In this way, the feature expression ability of the model can be improved, making important location information more prominent.

[0075] The channel attention mechanism focuses on the importance of feature channels, allowing the model to adaptively weight them in the channel dimension. Some feature channels may be more important in some tasks, and the channel attention mechanism can enhance the influence of these important channels and suppress redundant information. In this way, the expressiveness of features is optimized, helping downstream tasks (such as classification, detection, etc.) achieve better results.

[0076] In this embodiment, after using the spatial attention mechanism and the channel attention mechanism to extract features of the video frame respectively, the first feature and the second feature are spliced, so that a richer feature representation can be obtained, which enables the model to make full use of information at different levels in subsequent processing.

[0077] At the same time, in order to capture the dynamic information in the time dimension, this embodiment also introduces time series features based on the spliced ​​image features, that is, the time series features and image features are fused to obtain time series fusion features. The fusion method can be, for example, to combine the time series features (i.e. Figure 2 The temporal coding in the image features is concatenated with the image features. For example, the temporal fusion features can be expressed as Fv1, Fv2, ……………, Fvk, …. In this way, by introducing temporal features into image features, the model's ability to understand dynamic behaviors can be significantly enhanced, especially in tasks such as action recognition and video analysis.

[0078] In this embodiment, in addition to obtaining driving video data, it is also necessary to obtain the scene feature vector of the target retrieval scene. The target retrieval scene is some specific scenes that the user wants to retrieve. The user uses text to describe the specific scenes that the user wants to retrieve, and the model converts the text description into the corresponding scene feature vector. The specific scene can be, for example, a U-turn of the vehicle itself, a sudden brake of the vehicle ahead, etc., which are not listed here one by one.

[0079] In this embodiment, the specific scene description input by the user may be a description of a specific scene or a description of multiple scenes. This embodiment does not specifically limit the number of target retrieval scenes. In one example, for the scene text descriptions of n target retrieval scenes input, the text feature extraction module may receive the scene text descriptions of the n target retrieval scenes input, and obtain corresponding scene feature vectors Sc1, Sc2, ..., Scn from the n scene text descriptions according to the text feature extraction algorithm, such as Figure 2 shown.

[0080] After obtaining the temporal fusion features and scene feature vectors, multiple threads are started to record different events in the video frames in chronological order; the number of threads is the same as the number of events in the video frames; the temporal fusion features corresponding to the events recorded by each thread are determined as a temporal feature sequence.

[0081] For example, assuming that there are k specific scenarios, the number of events is k, and k threads are started accordingly. One thread is used to record one event. In actual applications, these k events are traversed once at every preset time interval. In this way, they are recorded in chronological order as k time series feature sequences Fset1, Fset2, ………Fsetk.

[0082] Afterwards, the n scene feature vectors are matched with the k time series feature sequences for similarity. Figure 3 As shown, in an exemplary embodiment, the process of obtaining a matching result according to the similarity between a scene feature vector and a temporal fusion feature in a temporal feature sequence may include the following steps 3100 to 3400:

[0083] 3100. Perform dimension reduction processing on the time series fusion features in the time series feature sequence.

[0084] 3200. Calculate the similarity between the scene feature vector and the image features in the temporal fusion feature.

[0085] 3300. According to the similarity and the weighted voting strategy, the scene category corresponding to the time series feature sequence and the probability that the time series feature sequence is the scene category are obtained.

[0086] 3400. Determine the scene category corresponding to the time series feature sequence and the probability that the time series feature sequence is a scene category as a matching result.

[0087] Specifically, the temporal fusion features in the temporal feature sequence can be processed by methods such as average pooling and maximum redundancy to reduce the dimension and enhance the representativeness of the features. Afterwards, the similarity between the scene feature vector and the image features in the temporal fusion features can be calculated by methods such as cosine similarity and Euclidean distance.

[0088] Afterwards, the weight of the task is dynamically adjusted according to the feedback of each scene target, so that the scene category SC corresponding to the temporal feature sequence is obtained in real time based on the weighted voting strategy. m , and the probability ProbSC corresponding to each scene category m .

[0089] The probability ProbSC of the scene category in the temporal feature sequence m When the probability is greater than the preset threshold, the time period corresponding to the scene category is determined Extract a video clip corresponding to a time period from the driving video data; output the scene category, the time period corresponding to the scene category, and the extracted video clip. Figure 2 As shown in , two scene retrieval results are output: and

[0090] In this way, according to the task requirements, a preset probability threshold is set, and only when the probability of the scene category is greater than the preset probability threshold, they are considered to be related. At this time, for these scene categories, the corresponding time period is determined and the video clips are automatically intercepted and saved to form new short videos or video clips. It can be understood that for the case where the probability of the scene category is less than the preset probability threshold, no information is output.

[0091] It can be seen that through real-time feature extraction and matching, it is possible to quickly respond to user needs and is suitable for real-time monitoring and analysis scenarios. By setting a preset probability threshold, video clips can be automatically captured, reducing manual intervention and improving efficiency. At the same time, alignment and matching based on image and text features can improve the relevance of video content, ensure that the captured clips are highly consistent with user needs, and flexibly adjust the feature extraction and matching algorithms and threshold settings according to different application scenarios and needs to adapt to various types of video analysis tasks.

[0092] According to the technical solution of this embodiment, by obtaining driving video data and obtaining the scene feature vector of the target retrieval scene; extracting the temporal fusion features that fuse the temporal features and the image features from each video frame. In this way, by introducing the temporal features into the image features, the ability to understand the dynamic behavior in the video can be significantly enhanced. Afterwards, the temporal fusion features are sorted into a temporal feature sequence in a temporal order; wherein each temporal feature sequence includes a plurality of temporal fusion features with the same image features but different temporal features; and the matching result is obtained according to the similarity between the scene feature vector and the temporal fusion features in the temporal feature sequence. Thus, by matching the real-time features, the user's needs can be quickly responded to, so that the solution of this application can be applied to the application scenarios of real-time monitoring and analysis. Then, the scene retrieval result is output according to the matching result. It can be seen that the technical solution of the embodiment of this application, without complex feature extraction preprocessing, reduces the complexity of feature extraction and improves the efficiency of feature extraction compared with the traditional RGB feature extraction and optical flow feature distribution extraction schemes.

[0093] Exemplary Devices

[0094] Accordingly, the present application also provides a device, such as Figure 4As shown, the scene retrieval device 400 for driving video provided in this embodiment may include: an acquisition module 410 , an extraction module 420 , a sorting module 430 , a matching module 440 and an output module 450 .

[0095] The acquisition module 410 is used to acquire driving video data and obtain a scene feature vector of a target retrieval scene; the driving video data includes a plurality of video frames;

[0096] An extraction module 420 is used to extract time series fusion features from each video frame respectively, where the time series fusion features include fused time series features and image features; the image features are used to characterize events in the video frame;

[0097] The sorting module 430 is used to sort the time series fusion features into a time series feature sequence according to the time series order; wherein each time series feature sequence includes a plurality of time series fusion features having the same image features but different time series features;

[0098] A matching module 440 is used to obtain a matching result according to the similarity between the scene feature vector and the temporal fusion feature in the temporal feature sequence; the matching result at least includes the scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is a scene category;

[0099] The output module 450 is used to output the scene retrieval result according to the matching result.

[0100] In a feasible implementation, the acquisition module 410 may be specifically configured to receive an input scene text description of a target retrieval scene; and acquire a scene feature vector from the scene text description according to a text feature extraction algorithm.

[0101] In a feasible implementation, the extraction module 420 can be specifically used to extract a first feature from each video frame through a spatial attention mechanism, extract a second feature from the video frame through a channel attention mechanism, and extract a temporal feature from the video frame; concatenate the first feature and the second feature to obtain an image feature. The image feature is fused with the temporal feature to obtain a temporal fusion feature.

[0102] In a feasible implementation, the sorting module 430 can be specifically used to start multiple threads to record different events in the video frame in a time sequence, wherein the number of threads is the same as the number of events in the video frame, and the time sequence fusion features corresponding to the events recorded by each thread are determined as a time sequence feature sequence.

[0103] In a feasible implementation, the matching module 440 can be specifically used to calculate the similarity between the scene feature vector and the image feature in the temporal fusion feature. According to the similarity and the weighted voting strategy, the scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is a scene category are obtained. The scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is a scene category are determined as the matching result.

[0104] In a feasible implementation, the matching module 440 is used to calculate the similarity between the scene feature vector and the image feature in the temporal fusion feature, and also to reduce the dimension of the temporal fusion feature in the temporal feature sequence.

[0105] In a feasible implementation, the output module 450 can be specifically used to determine the time period corresponding to the scene category when the probability that the temporal feature sequence is a scene category is greater than a preset probability threshold, intercept a video clip corresponding to the time period from the driving video data, and output the scene category, the time period corresponding to the scene category, and the intercepted video clip.

[0106] The scene retrieval device for driving video provided in this embodiment belongs to the same application concept as the scene retrieval method for driving video provided in the above embodiment of this application, and can execute the scene retrieval method for driving video provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects of executing the scene retrieval method for driving video. For technical details not fully described in this embodiment, please refer to the specific processing content of the scene retrieval method for driving video provided in the above embodiment of this application, and will not be repeated here.

[0107] It should be understood that the modules in the above devices can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device, wherein the processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory in the device or a memory outside the device. Alternatively, the unit in the device can be implemented in the form of a hardware circuit, and the functions of some or all units can be realized by designing the hardware circuit. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by PLD, taking FPGA as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of some or all of the above units. All units of the above devices can be implemented in the form of a processor calling software, or in the form of a hardware circuit, or in part by a processor calling software, and the remaining part is implemented in the form of a hardware circuit.

[0108] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and run instructions, such as a CPU, a microprocessor, a GPU, or a DSP; in another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.

[0109] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0110] In addition, all or part of the units in the above device can be integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0111] Exemplary Electronic Devices

[0112] The present application embodiment provides an electronic device, see Figure 5 As shown, the electronic device includes a memory 500 and a processor 510 connected to the memory 500 .

[0113] The memory 500 is used to store programs.

[0114] The processor 510 is used to execute any one of the scene retrieval methods for driving videos of any of the above embodiments, obtain driving video data, and obtain the scene feature vector of the target retrieval scene; extract the time-series fusion features that fuse the time-series features and the image features from each video frame. In this way, by introducing the time-series features into the image features, the ability to understand the dynamic behavior in the video can be significantly enhanced. Afterwards, the time-series fusion features are sorted into a time-series feature sequence in a time-series order; wherein each time-series feature sequence includes a plurality of time-series fusion features with the same image features but different time-series features; and obtain the matching result according to the similarity between the scene feature vector and the time-series fusion features in the time-series feature sequence. Thus, by matching the real-time features, the user's needs can be quickly responded to, so that the solution of the present application can be applied to the application scenarios of real-time monitoring and analysis. Then, the scene retrieval result is output according to the matching result. It can be seen that the technical solution of the embodiment of the present application, without complex feature extraction preprocessing, reduces the complexity of feature extraction and improves the efficiency of feature extraction compared with the traditional RGB feature extraction and optical flow feature distribution extraction solutions.

[0115] The specific processing process of the processor 510 can refer to the introduction of the above method embodiment, and the specific implementation of the processor 510 can also refer to the introduction of the above embodiment.

[0116] Specifically, the electronic device may further include: a bus, a communication interface 520 , an input device 530 and an output device 540 .

[0117] The processor 510, the memory 500, the communication interface 520, the input device 530 and the output device 540 are connected to each other via a bus.

[0118] A bus may include a pathway that transfers information between components of a computer system.

[0119] The processor 510 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the scheme of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0120] The processor 510 may include a main processor, and may also include a baseband chip, a modem, and the like.

[0121] The memory 500 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes a computer operation instruction. More specifically, the memory 500 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.

[0122] The input device 530 may include a device for receiving data and information input by a user, such as an error microphone, a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.

[0123] Output device 540 may include devices that allow information to be output to a user, such as a speaker, display screen, printer, speakers, etc.

[0124] The communication interface 520 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0125] The processor 510 executes the program stored in the memory 500 and calls other devices, which can be used to implement each step of any one of the scene retrieval methods for driving videos provided in the above embodiments of the present application.

[0126] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored in a memory through the data interface to execute the scene retrieval method for driving video introduced in any of the above embodiments. The specific processing process and its beneficial effects can be found in the above-mentioned embodiment introduction of the scene retrieval method for driving video.

[0127] Exemplary computer program products and storage media

[0128] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the scene retrieval method for driving video according to various embodiments of the present application described in any of the above embodiments of this specification.

[0129] The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0130] In addition, the embodiment of the present application may also be a storage medium on which a computer program is stored. The computer program is executed by a processor to perform the steps of the scene retrieval method for driving video according to various embodiments of the present application described in any of the above embodiments of this specification, and specifically the following steps may be implemented:

[0131] 1100. Obtain driving video data and obtain a scene feature vector of a target retrieval scene; the driving video data includes a plurality of video frames.

[0132] 1200. Extracting temporal fusion features from each video frame respectively, wherein the temporal fusion features include fused temporal features and image features; the image features are used to characterize events in the video frames.

[0133] 1300. Arrange the time series fusion features into a time series feature sequence in time series order; wherein each time series feature sequence includes a plurality of time series fusion features having the same image features but different time series features.

[0134] 1400. Obtain a matching result based on the similarity between the scene feature vector and the temporal fusion feature in the temporal feature sequence; the matching result at least includes the scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is the scene category.

[0135] 1500. Output the scene retrieval result according to the matching result.

[0136] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0137] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0138] The steps in the methods of each embodiment of the present application can be adjusted in sequence, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0139] The modules and sub-modules in the devices and terminals of the various embodiments of the present application can be combined, divided and deleted according to actual needs.

[0140] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0141] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0142] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.

[0143] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0144] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0145] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0146] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A scene retrieval method for driving video, characterized in that: The method comprises: Acquire driving video data and obtain a scene feature vector of a target retrieval scene; the driving video data includes a plurality of video frames; Extracting temporal fusion features from each of the video frames respectively, wherein the temporal fusion features include fused temporal features and image features; the image features are used to characterize events in the video frames; Arrange the time series fusion features into a time series feature sequence in time series order; wherein each of the time series feature sequences includes a plurality of time series fusion features having the same image features but different time series features; Obtaining a matching result according to the similarity between the scene feature vector and the temporal fusion feature in the temporal feature sequence; the matching result at least includes a scene category corresponding to the temporal feature sequence and a probability that the temporal feature sequence is the scene category; The scene retrieval result is output according to the matching result.

2. The method according to claim 1, characterized in that: The step of obtaining a scene feature vector of a target retrieval scene includes: Receiving an input scene text description of the target retrieval scene; The scene feature vector is obtained from the scene text description according to a text feature extraction algorithm.

3. The method according to claim 1, characterized in that: The extracting the temporal fusion feature from each of the video frames respectively includes: For each of the video frames, extracting a first feature from the video frame by a spatial attention mechanism, extracting a second feature from the video frame by a channel attention mechanism, and extracting a temporal feature from the video frame; Concatenate the first feature and the second feature to obtain the image feature; The image feature is fused with the time series feature to obtain the time series fusion feature.

4. The method according to claim 1, characterized in that: The step of arranging the time series fusion features into a time series feature sequence according to the time series order includes: Starting a plurality of threads to respectively record different events in the video frame in a time sequence; wherein the number of the threads is the same as the number of events in the video frame; The time series fusion features corresponding to the events recorded by each of the threads are determined as a time series feature sequence.

5. The method according to claim 1, characterized in that The obtaining a matching result according to the similarity between the scene feature vector and the temporal fusion feature in the temporal feature sequence includes: Calculating the similarity between the scene feature vector and the image feature in the temporal fusion feature; According to the similarity and the weighted voting strategy, the scene category corresponding to the time series feature sequence and the probability that the time series feature sequence is the scene category are obtained; The scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is the scene category are determined as the matching result.

6. The method according to claim 5, characterized in that Before calculating the similarity between the scene feature vector and the image feature in the temporal fusion feature, the method further includes: The time series fusion features in the time series feature sequence are processed to reduce the dimension.

7. The method according to claim 1, characterized in that Outputting the scene retrieval result according to the matching result includes: When the probability that the temporal feature sequence is the scene category is greater than a preset probability threshold, determining a time period corresponding to the scene category; Extracting a video clip corresponding to the time period from the driving video data; The scene category, the time period corresponding to the scene category, and the captured video segment are output.

8. A scene retrieval device for driving video, characterized in that: The device comprises: An acquisition module is used to acquire driving video data and obtain a scene feature vector of a target retrieval scene; the driving video data includes a plurality of video frames; An extraction module, used to extract time series fusion features from each of the video frames respectively, wherein the time series fusion features include fused time series features and image features; the image features are used to characterize events in the video frames; A sorting module, used for sorting the time series fusion features into a time series feature sequence in time series order; wherein each of the time series feature sequences includes a plurality of time series fusion features with the same image features but different time series features; A matching module, configured to obtain a matching result according to the similarity between the scene feature vector and the temporal fusion feature in the temporal feature sequence; the matching result at least includes the scene category corresponding to the temporal feature sequence and the probability that the temporal feature sequence is the scene category; An output module is used to output scene retrieval results according to the matching results.

9. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the scene retrieval method for driving video as described in any one of claims 1 to 7 by running the program in the memory.

10. A computer program product, characterized in that The method comprises computer program instructions, which, when executed by a processor, enable the processor to execute the scene retrieval method for driving video according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method and system for converting scene classification into scene time domain detection

    CN120635813A