Video content detection method

By generating text scene maps and video scene maps, and using video content detection models to process multimodal features and semantic features, the problems of user variable detection needs and false detection in the prior art are solved, and efficient and accurate detection of video content is achieved.

CN120356136AActive Publication Date: 2025-07-22ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB)

Patent Information

Application Number
CN202510846373.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The existing video content detection methods are difficult to meet the user's changing detection needs, cannot effectively identify implicit and complex video content, and are prone to false detection and missed detection in multi-object video scenarios, and rely on manual auditing efficiency.

Method used

By generating text scene maps and video scene maps, using the video content detection model to process multimodal features and semantic features, and generate object detection results, including video extraction modules, text extraction modules and detection modules, and generate object detection results based on the similarity of multimodal features and semantic features, reducing manual review and improving detection efficiency.

Benefits of technology

Accurately position the target object and its related actions in complex multi-object video scenarios, reduce false detection and missed detection, improve detection efficiency, meet the requirements of real-time and accuracy, and is suitable for video review and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356136A_ABST
    Figure CN120356136A_ABST
Patent Text Reader

Abstract

The invention provides a video content detection method, which comprises the following steps: in response to a received target text input by an object, generating a text scene graph based on the target text, the text scene graph comprising a to-be-detected action related to a detection intention and a target object; acquiring a video scene graph of a to-be-detected video; wherein the video scene graph is used for representing action time sequence information of a first object in the to-be-detected video; and processing the text scene graph and the video scene graph by using a video content detection model to generate a target detection result. Wherein the target detection result indicates whether the to-be-detected video comprises the to-be-detected action of the target object. According to the video content detection method provided by the invention, in a complex multi-object video scene, the target object and related actions thereof can be accurately positioned, and false detection and missing detection conditions are effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video content recognition, and more particularly, to a video content detection method. Background Art

[0002] In the era of rapid spread of digital information today, video data has become an important carrier for information dissemination and sharing. With the increasing popularity of social media and video platforms, a large number of videos are uploaded, distributed, and shared. However, in the process of video data distribution and sharing, how to effectively detect the content in videos and avoid the leakage of key information has become a key issue in the information society. However, existing video content detection methods are difficult to perform dynamic recognition based on the changing detection requirements of users. Summary of the Invention

[0003] In view of this, the present invention provides a video content detection method, including: in response to receiving a target text input by an object, generating a text scene graph based on the target text, where the text scene graph includes a to-be-detected action and a target object related to the detection intention; obtaining a video scene graph of the to-be-detected video; where the video scene graph is used to represent the action timing information of a first object in the to-be-detected video; using a video content detection model to process the text scene graph and the video scene graph to generate a target detection result; where the target detection result indicates whether the to-be-detected video includes the to-be-detected action of the target object.

[0004] According to an embodiment of the present invention, the video content detection model includes a video extraction module, a text extraction module, and a detection module; using the video content detection model to process the text scene graph and the video scene graph to generate a target detection result includes: using the video extraction module to process the video scene graph and the to-be-detected video to generate multi-modal features for describing the action changes of the first object; using the text extraction module to process the text scene graph to generate semantic features for describing the to-be-detected action; using the detection module to generate a target detection result based on the similarity between the multi-modal features and the semantic features.

[0005] According to an embodiment of the present invention, the multimodal features include frame fusion features and action fusion features; the video extraction module includes a visual extraction sub-module, a semantic extraction sub-module, and a convolutional sub-module; the video extraction module is used to process the video scene graph and the video to be detected to generate multimodal features for describing the action changes of the first object, including: using the semantic extraction sub-module to semantically process the video scene graph to obtain the category features of the first object included in each video frame of the video to be detected and the category features of the action performed by the first object; using the visual extraction sub-module to process the video scene graph and the video to be detected to obtain the visual features of the first object included in each video frame of the video to be detected and the visual features of the action performed by the first object; for each video frame, splicing the category features and visual features of the first object to obtain object fusion features; using the convolutional sub-module to perform spatio-temporal graph convolution operations on the object fusion features to obtain frame fusion features; for each video frame, splicing the category features and visual features of the action performed by the first object to obtain action fusion features.

[0006] According to an embodiment of the present invention, using the visual extraction sub-module to process the first object in the video scene graph and the video to be detected to obtain the visual features of the first object and the visual features of the action performed by the first object, including: using the visual extraction sub-module to extract features from the video frames including the first object in the video to be detected within the range of the area where the first object is located to obtain the visual features of the first object and the visual features of the action performed by the first object. According to an embodiment of the present invention, the semantic features include the category features of the target object and the category features of the action to be detected; using the text extraction module to process the text scene graph to generate semantic features for describing the action to be detected, including: using the text extraction module to semantically process the text scene graph to obtain the category features of the target object and the category features of the action to be detected.

[0007] According to an embodiment of the present invention, using the detection module to generate a target detection result based on the similarity between the multimodal features and the semantic features, including: using the detection module to respectively obtain a node matching score, a frame matching score, and a video matching score based on the multimodal features and the semantic features; comparing the average value of the node matching score, the frame matching score, and the video matching score with a first threshold, and determining the target detection result according to the comparison result.

[0008] According to an embodiment of the present invention, the detection module includes a node matching sub-module, a frame matching sub-module, and a video matching sub-module; the node matching score is determined as follows: for a single video frame, the node matching sub-module is used to calculate the first initial similarity between the category feature of each target object and the frame fusion feature, and the maximum value of the first initial similarity is determined as the object similarity corresponding to each target object; after summing up the object similarities of multiple target objects and taking the average, the first similarity of the video frame is obtained; for a single video frame, the node matching sub-module is used to calculate the second initial similarity between the category feature of each action to be detected and the action fusion feature of each first object on the video frame, and the maximum value of the second initial similarity is determined as the action similarity, after summing up the action similarities of multiple actions to be detected and taking the average, the second similarity of the video frame is obtained; the first similarity and the second similarity of the same video frame are added to obtain the node-level similarity of the corresponding video frame, and the maximum node-level similarity among the node-level similarities of multiple video frames is determined as the node matching score.

[0009] According to an embodiment of the present invention, the frame fusion sub-features corresponding to each first object in a single video frame are weighted and aggregated to obtain the first graph-level feature of the single video frame; the action fusion features corresponding to each action in a single video frame are weighted and aggregated to obtain the second graph-level feature of the single video frame; the first graph-level feature and the second graph-level feature of the same video frame are concatenated to obtain the object graph-level feature; the category features of multiple target objects are weighted and then fused to obtain the third graph-level feature; the category features of multiple actions to be detected are weighted and then fused to obtain the fourth graph-level feature; the third graph-level feature and the fourth graph-level feature are concatenated to obtain the text graph-level feature; the similarity scores between the text graph-level feature and multiple object graph-level features are respectively determined to obtain the frame matching similarity scores of multiple video frames, and the maximum frame matching similarity score is used as the frame matching score.

[0010] According to an embodiment of the present invention, the object graph-level features of multiple video frames in the video to be detected are subjected to average pooling operation by using the video matching sub-module to obtain the object global feature; the similarity between the object global feature and the text graph-level feature is calculated to obtain the video matching score.

[0011] According to an embodiment of the present invention, when the target detection result indicates that there is an action to be detected of a target object in the video to be detected, the frame matching score is used to enhance the feature of the object graph-level feature of the video frame to which the frame matching score belongs to obtain an enhanced feature; the object global feature and the text graph-level feature are concatenated to obtain an action vector; the action vector and the enhanced feature are concatenated to obtain a video frame sequence; the video frame sequence is input into a recurrent neural network, and the start time and end time of the target object in the video to be detected performing the action to be detected are output.

[0012] Limiting the actions of the first object based on the target text can enable the user's retrieval intention to cover more customized actions, making video content detection not only limited to the limited actions specified by the model. The method of the present invention detects whether the video contains the action to be detected through the model, reduces the manual review of video content, and improves the video detection efficiency. In a complex multi-object video scenario, it can accurately locate the target object and its related actions, effectively avoiding misdetection and missed detection. At the same time, by processing the video scene graph of the video to be detected and the text scene graph of the target text through the video content detection model, the detection efficiency is greatly improved, and a large number of videos can be quickly screened in a short time, meeting the high requirements for real-time and accuracy in fields such as video review. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features and advantages of the present invention will become clearer. In the drawings:

[0014] Figure 1 Schematically shows an exemplary system architecture to which the video content detection method according to an embodiment of the present invention can be applied;

[0015] Figure 2 Schematically shows a flowchart of the video content detection method according to an embodiment of the present invention;

[0016] Figure 3 Schematically shows a schematic diagram of the principle of the video content detection method according to an embodiment of the present invention;

[0017] Figure 4 Schematically shows a schematic diagram of the principle of the video extraction module according to an embodiment of the present invention;

[0018] Figure 5 Schematically shows a schematic diagram of the principle of the convolutional sub-module according to an embodiment of the present invention;

[0019] Figure 6 Schematically shows a data processing flowchart of the video content detection method according to an embodiment of the present invention;

[0020] Figure 7 Schematically shows a block diagram of the video content detection device according to an embodiment of the present invention; and

[0021] Figure 8 Schematically shows a block diagram of an electronic device suitable for video content detection according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0023] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0025] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0026] In the embodiments of the present invention, in terms of the collection, update, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the data involved (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and to maintain the security of user personal information and network security.

[0027] In the embodiments of the present invention, before obtaining or collecting user personal information, the authorization or consent of the user is obtained.

[0028] Currently, in the field of video content detection, the demand for detecting content in videos is increasing day by day. However, existing methods for detecting content in videos usually predefine the object to be detected before detection, and then simply determine whether the content to be detected exists in the corresponding video by detecting specific objects in the video. The specific objects can be human faces or license plates, etc. The above methods have many deficiencies, for example:

[0029] Although the above method can identify explicit content in a video to a certain extent, some information may not be directly presented through the visual features of an object, but rather through the relationships between objects or changes in actions. Therefore, there is a lack of effective detection means for implicit and complex content. Among them, implicit and complex content can be behaviors or social relationships, etc. Moreover, the existing technology generally realizes the detection of complex content through manual review, which is inefficient and prone to misdetection and missed detection.

[0030] On the one hand, the above method generally only applies to the detection of a single object, ignoring the content of interactions between objects that may contain multiple dimensions in the video. Video content such as the positional relationship between multiple objects in space and the interactive changes in the time dimension may cause information leakage.

[0031] On the other hand, the above method generally can only detect based on the specified detection requirements in the detection model. However, different users have different requirements for the content to be detected, and it is often impossible to perform flexible and variable detections, making it difficult to meet the dynamic detection needs of users.

[0032] In view of this, an embodiment of the present invention provides a video content detection method, including: in response to receiving a target text input by an object, generating a text scene graph based on the target text, where the text scene graph includes a to-be-detected action and a target object related to the detection intention; obtaining a video scene graph of the to-be-detected video; where the video scene graph is used to represent the action timing information of a first object in the to-be-detected video; using a video content detection model to process the text scene graph and the video scene graph to generate a target detection result; where the target detection result indicates whether the to-be-detected video includes the to-be-detected action of the target object.

[0033] The method of the present invention can define the actions of the first object according to the target text, and can cover more custom actions according to the actual situation; it meets the dynamic detection needs of users. By using the model to detect whether the to-be-detected action is included in the video, it reduces the situation of missed detection and misdetection caused by manual detection and improves the video detection efficiency.

[0034] Figure 1 Schematically shows an exemplary system architecture to which the video content detection method according to an embodiment of the present invention can be applied. It should be noted that Figure 1 The shown is only an example of the system architecture to which the embodiments of the present invention can be applied, to help those skilled in the art understand the technical content of the present invention, but it does not mean that the embodiments of the present invention cannot be used in other devices, systems, environments or scenarios.

[0035] Such as Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0036] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only for example).

[0037] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0038] The server 105 may be a server that provides various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only for example). The background management server may analyze and process data such as received user requests, etc., and feedback the processing results (such as web pages, information, or data, etc. obtained or generated according to user requests) to the terminal device.

[0039] It should be noted that the video content detection method provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the video content detection system provided by the embodiments of the present invention can generally be set in the server 105. The video content detection method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the video content detection device provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Alternatively, the video content detection method provided by the embodiments of the present invention can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or can also be executed by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the video content detection device provided by the embodiments of the present invention can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or set in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.

[0040] For example, the video to be detected can be stored in any one of the first terminal device 101, the second terminal device 102, or the third terminal device 103 (for example, the first terminal device 101, but not limited thereto), or stored on an external storage device and can be imported into the first terminal device 101. Then, the first terminal device 101 can execute the video content detection method and the location positioning method provided by the embodiments of the present invention locally, or send the video to be detected to other terminal devices, servers, or server clusters, and the other terminal devices, servers, or server clusters that receive the video to be detected execute the video content detection method and the location positioning method provided by the embodiments of the present invention.

[0041] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0042] Figure 2 is only illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0043] As Figure 2 shown, the video content detection method 200 includes operations S210~S230.

[0044] In operation S210, in response to receiving the target text input by the object, a text scene graph is generated based on the target text, and the text scene graph includes the action to be detected and the target object related to the detection intention.

[0045] According to an embodiment of the present invention, the above-mentioned object is a user with a need for video content detection. The target text is the text information written by the user based on his / her detection intention for the video to be detected. In this embodiment, the text scene graph model can be used to convert the target text input by the user into a text scene graph, and the action to be detected and the target object that the user wants to detect can be represented in the text scene graph.

[0046] According to an embodiment of the present invention, the detection intention of the user can be: the user wants to detect whether the action of hugging appears in the video to be detected, or the user wants to detect whether the action of a man and a woman hugging appears in the video to be detected, or the user wants to detect whether the action of a man and a woman looking at each other appears in the video to be detected. Both the above-mentioned man and woman can be used as the target object, and hugging and looking at each other can be used as the actions to be detected. That is to say, the detection intention in this application can be the action to be detected and the target object.

[0047] In operation S220, a video scene graph of the video to be detected is obtained; wherein, the video scene graph is used to represent the action timing information of the first object in the video to be detected.

[0048] According to an embodiment of the present invention, a video scene graph generation model can be used to convert the video to be detected into a video scene graph.

[0049] According to an embodiment of the present invention, the action timing information of the first object included in the above-mentioned video scene graph is the action change of the first object or the action change between two or more first objects. For example: the two first objects can be a man and a woman, and the action change between the two first objects can be that they change from looking at each other to dancing or from looking at each other to hugging.

[0050] According to an embodiment of the present invention, before detection, video scene graphs can be generated in advance for multiple videos to be detected, so that the video scene graphs can be directly used for detection after the user inputs the target text. Or a video scene graph can be generated temporarily according to the current video to be detected for video content detection based on the target text.

[0051] In operation S230, a video content detection model is used to process the text scene graph and the video scene graph to generate a target detection result; wherein, the target detection result indicates whether the action to be detected of the target object is included in the video to be detected.

[0052] According to an embodiment of the present invention, the text scene graph obtained in operation S210 and the video scene graph generated in operation S220 can be input into a video content detection model for comparison by the video content detection model, and a target detection result can be obtained. The target detection result can indicate whether the video to be detected includes the action to be detected. For example, if the user inputs an action of detecting a hug in the target text, a frame with a hug action in the video to be detected will be detected. Further, if the user inputs in the target text that the object to be detected performs the action to be detected, it can also be recognized in the video to be detected. For example, if the user inputs to detect a man and a woman hugging, the target detection result can indicate whether there is a frame in the video to be detected where a man and a woman are hugging.

[0053] According to an embodiment of the present invention, the target detection result can also indicate whether the video to be detected includes a target object changing from a first action to be detected to a second action to be detected. If the user inputs in the target text that there is a change in the action to be detected between the objects to be detected, it can also be recognized in the video to be detected. For example, if the user inputs to detect a man and a woman first looking at each other and then hugging, the target detection result can indicate whether there is a frame in the video to be detected where a man and a woman change from looking at each other to hugging.

[0054] According to an embodiment of the present invention, the video content detection model is trained based on sample texts and sample videos used to describe the action information of sample objects in the sample videos. When preparing the training data, videos in the public dataset can be used as video samples, and then sample texts can be written according to the video content. The sample texts can be divided into two types. One is a positive sample that accurately describes the action changes of the objects in the video sample, and the other is a negative sample that has nothing to do with the action changes of the objects in the video sample. The collected sample videos and sample texts are divided into a training set, a validation set, and a test set in a ratio of 8:1:1 and stored for subsequent use when training the video content detection model.

[0055] During the training process, the validation set can be used to evaluate the model in a timely manner, and evaluation metrics such as accuracy, recall rate, and F1 value can be calculated. Analyze the problems existing in the model according to the evaluation results, such as overfitting or underfitting. If overfitting occurs, a regularization term can be added or the Dropout technique can be adopted; if there is underfitting, the architecture of the video content detection model can be adjusted, the number of network layers can be increased, or the hyperparameters can be adjusted. When the performance of the video content detection model on the validation set tends to be stable, the test set is used to comprehensively evaluate the finally trained video content detection model to obtain the generalization performance index of the video content detection model, ensuring that the video content detection model can accurately detect video content based on sample texts and sample videos in actual applications.

[0056] According to an embodiment of the present invention, defining the actions of the first object based on the target text can enable the user's retrieval intention to cover more customized actions, so that video content detection is not limited to the limited actions specified by the model. The method of the present invention detects whether the video contains the action to be detected through the model, reduces the manual review of video content, and improves the video detection efficiency. In a complex multi-object video scenario, it can accurately locate the target object and its related actions, effectively avoiding false detection and missed detection. At the same time, through automated model processing, the detection efficiency is greatly improved, and a large number of videos can be quickly screened in a short time, meeting the high requirements for real-time and accuracy in fields such as security monitoring and video review, providing strong technical support for related industries, and having broad application value and good social benefits.

[0057] According to an embodiment of the present invention, the video content detection model includes a video extraction module, a text extraction module, and a detection module; using the video content detection model to process the text scene graph and the video scene graph to generate a target detection result, including: using the video extraction module to process the video scene graph and the video to be detected to generate multi-modal features for describing the action changes of the first object; using the text extraction module to process the text scene graph to generate semantic features for describing the action to be detected; using the detection module to generate a target detection result based on the similarity between the multi-modal features and the semantic features.

[0058] The following refers to Figure 3 , and further illustrates the Figure 2 shown method in combination with specific embodiments.

[0059] Figure 3 Schematically shows a schematic diagram of the principle of the video content detection method according to an embodiment of the present invention.

[0060] According to an embodiment of the present invention, the above multi-modal features can be fusion features that fuse visual information features and text information features, and the multi-modal features can more comprehensively display the information of the action changes of the first object. The semantic features can represent the specific information expressed by the words in the target text output by the user. Converting the action to be detected into semantic features is to facilitate calculating the similarity between the multi-modal features and the semantic features.

[0061] According to an embodiment of the present invention, when generating a text scene graph (Text Scene Graph, TSG), the target object can also be used as a node, and the relationship between the nodes of the target object can be connected as the relationship edge of the target object. For example: using a text scene graph generation model to parse the target text input by the user to generate a corresponding text scene graph . Among them, represents the node set of the target object in the text, A node that can represent the first target object in the target text, A node that can represent the second target object in the target text, The total number of nodes of the target objects included in the target text. The following occurrences have the same meaning. A set of relationship edges of the nodes representing the target objects. The relationship edge of the node that can represent the first target object in the target text, The relationship edge of the node that can represent the second target object in the target text, The total number of relationship edges of the nodes of the target objects included in the target text. The following occurrences have the same meaning. The above Can represent the target text identifier. The following occurrences in other positions have the same meaning, Used to indicate that the nodes and relationship edges of the target objects are both obtained through the target text.

[0062] In order to more clearly describe all objects in the embodiments of the present invention, hereinafter, the target object / first object is used to replace the nodes of the target object / first object.

[0063] Exemplarily, in the target object Among them, Represents the number of the target object, and the target object corresponds to a word in the target text. The relationship edge of each target object Is represented in the form of a triple, Among them, Can be the target object And the target object The relationship predicate between them. Can represent the number of the target object.

[0064] According to the embodiments of the present invention, when generating a video scene graph (VSG), a static scene graph corresponding to the video frame can be generated first, and then the static scene graph can be associated along the time to obtain the video scene graph. Exemplarily, a scene graph model can be used to generate a static scene graph for each video frame in the video to be detected . In the static scene graph Among them, the node set of the first object is , ; among them, Can represent the In the video frame The node of the first object, Is the number of the video frame. The following occurrences in other positions Have the same meaning, and are both the numbers of the first object, and can respectively represent the nodes of different first objects. can represent the node of the first object 's visual vector, can be the class label of the node of the first object ; can be the bounding box coordinates of the node of the first object in the video frame, can represent the total number of the first objects in the video frame , and the following has the same meaning.

[0065] In the static scene graph , the set of relationship edges of the first object , . Among them, in the video frame , the relationship predicate of the first object represents the interaction relationship between the node of the first object and the node of the first object.

[0066] According to the embodiment of the present invention, after obtaining the static scene graph, the first object in the video frame of the video to be detected and the first object in another video frame form a pair of first objects ( , . A cross-frame co-reference relationship is established through the connection score between the two first objects in the pair of first objects. The calculation method of the connection score between the first objects refers to formula (1):

[0067]

[0068] wherein, represents the cosine similarity, which is used to measure the visual similarity of the first object, respectively represent the visual vectors of the first object in the video frame and the first object . respectively represent the bounding box coordinates of the first object in the video frame and the first object . represents the intersection over union (Intersection-over-Union, ). The spatial position weight parameter is used to control the influence of the spatial positions of two first objects on the connection score. Moreover, when calculating the connection score, the time distance between the two first objects limits the influence of the spatial positions of the two first objects, that is, the connection score of the first objects with a relatively large time distance is mainly determined by the feature similarity of the visual vectors of the two first objects, rather than the degree of spatial overlap.

[0069] According to an embodiment of the present invention, based on the obtained multiple connection scores, for the first object , select another first object whose connection score is greater than the connection score threshold and perform connection. The obtained connection relationships of all cross-frame first objects form a temporal co-reference edge set . Among them, represents the relationship edge between the first object and the first object and the first object .

[0070] According to an embodiment of the present invention, combine the temporal co-reference edges of adjacent frames in the temporal co-reference edge set with the static scene graph to obtain a video scene graph . Among them, is the total number of frames of the video to be detected. The meaning of appearing in other positions below remains unchanged. and are the static scene graphs of the first frame and the second frame respectively. is the temporal co-reference edge from the first frame to the second frame, and is the temporal co-reference edge from the second frame to the third frame. In actual application, the spatial position weight parameter can be 0.5, and the connection score threshold can be 0.7.

[0071] For example Figure 3As shown, the video scene graph generation model is used for the video 301 to be detected to obtain the video scene graph 302; the video extraction module 310 is used to process the video 301 to be detected and the video scene graph 302 to generate multimodal features 303, and the multimodal features 303 are used to describe the action changes of the first object. Since the video scene graph is in the form of text information, the video frames in the video to be detected also need to be combined for feature extraction when performing feature extraction. For example, the video scene graph can show through the text features it contains that the current frame includes a man and a woman hugging, and the video extraction module can extract the visual features of the man and the woman hugging in the corresponding video frame, and then combine the visual features with the text features shown in the video scene graph to form the multimodal features 303. The information contained in the multimodal features 303 is richer than the simple text information in the video scene graph, and for the cross-frame action information shown in the text information contained in the video scene graph, the video extraction module can also perform visual feature extraction.

[0072] According to an embodiment of the present invention, the text scene graph generation model is used for the target text 304 to obtain the text scene graph 305. The text extraction module 320 is used to semantically extract the text scene graph 305 to obtain semantic features 306, and the semantic features 306 may include the action to be detected input by the user, the target object, and the action to be detected of the target object.

[0073] According to an embodiment of the present invention, for the multimodal features 303 and semantic features 306 obtained above, the detection module 330 can perform detection through semantic similarity to obtain the target detection result 307. For example, the cosine similarity algorithm is used to calculate the similarity between the multimodal features 303 and the semantic features 306. Generally, the value range of the cosine similarity is between -1 and 1, and the closer the value is to 1, the higher the similarity; the detection module 330 can also set a similarity threshold for judgment. For example, if the obtained cosine similarity is higher than the pre-set threshold, the possibility that the action to be detected appears in the corresponding video frame is higher. The detection module 330 can also perform optimization processing on the detection results higher than the threshold. The redundant detection results are eliminated through the non-maximum suppression algorithm. For multiple overlapping and highly similar detection results, only the one with the highest score is retained as the final result; at the same time, the confidence of the detection results can be calibrated, and the similarity score can be converted into a confidence score more suitable for the actual application scenario, and finally an accurate target detection result is generated and output.

[0074] According to an embodiment of the present invention, by comparing the multimodal features obtained based on the video to be detected with the semantic features in the target text, the coherent actions in the video to be detected can be detected, so that the detection is not limited to the actions within a single frame image, the detection results are more comprehensive, and the omission of video content can be avoided.

[0075] According to an embodiment of the present invention, the multi-modal features include frame fusion features and action fusion features; the video extraction module includes a visual extraction sub-module, a semantic extraction sub-module, and a convolutional sub-module; the video extraction module is used to process the video scene graph and the video to be detected, and generate multi-modal features for describing the action changes of the first object, including: using the semantic extraction sub-module to semantically process the video scene graph to obtain the category features of the first object included in each video frame of the video to be detected and the category features of the action performed by the first object; using the visual extraction sub-module to process the video scene graph and the video to be detected to obtain the visual features of the first object included in each video frame of the video to be detected and the visual features of the action performed by the first object; for each video frame, concatenating the category features and the visual features of the first object to obtain object fusion features; using the convolutional sub-module to perform spatio-temporal graph convolution operations on the object fusion features to obtain frame fusion features; for each video frame, concatenating the category features and the visual features of the action performed by the first object to obtain action fusion features.

[0076] The following refers to Figure 4 and further describes the feature extraction operation performed by the video extraction module 310 in combination with specific embodiments.

[0077] Figure 4 Schematically shows a schematic diagram of the principle of the video extraction module according to an embodiment of the present invention.

[0078] According to an embodiment of the present invention, for example, the category features of the first object represent what the first object specifically is, such as a man, a woman, or a plant, etc.; the category features of the action performed by the first object refer to what the first object is doing, such as dancing, making eye contact, or hugging, etc. The above category features are all semantic text information. The visual features of the first object and the visual features of the action performed by the first object are the visual information extracted from the video frame images. Therefore, after feature concatenation, for example, the object fusion features of a man include text information and visual information describing the man. The action fusion features of dancing describe text information and visual information of dancing.

[0079] According to an embodiment of the present invention, the semantic extraction sub-module 311 is used to semantically process the video scene graph 302 to obtain the category features 401 of the first object included in each video frame and the category features 402 of the action performed by the first object.

[0080] Exemplarily, for the video frames included in the video scene graph of the static scene graph , using the word embedding matrix to convert the category label of the first object in the video frame into the category features of the first object; using the word embedding matrix to convert the video frame The relational predicate of the first object Converted into the category feature of the action performed by the first object ; where and are respectively the one-hot encodings of the first object and the relational predicate of the video frame , and are the pre-trained video word embedding matrices.

[0081] According to an embodiment of the present invention, the visual extraction sub-module 312 is used to process the video to be detected 301 and the video scene graph 302, and the visual features 403 of the first object included in each video frame in the video to be detected and the visual features 404 of the action performed by the first object are obtained.

[0082] According to an embodiment of the present invention, specifically, the visual extraction sub-module 312 is used to extract features from the video frame including the first object in the video to be detected within the range of the area where the first object is located, and the visual features 403 of the first object and the visual features 404 of the action performed by the first object are obtained.

[0083] Exemplarily, the visual extraction sub-module is used based on the first object bounding box to extract visual features from the corresponding video frame of the video to be detected, and the visual features of the first object are obtained ; the visual extraction sub-module is used to extract visual features based on the joint area where the relational edge of the first object is located to obtain the visual features of the action performed by the first object . Where represents the visual extraction sub-module represents the union area of the bounding boxes of two first objects represents the video frame the first object in bounding box. In practical applications, the dimension of the visual features when extracting visual features can be 2048.

[0084] For each video frame, the category feature 401 of the first object and the visual feature 403 of the first object are concatenated to obtain the object fusion feature 405; the convolutional sub-module 313 is used to perform spatio-temporal graph convolution operations on the object fusion feature 405 to obtain the frame fusion feature 406; for each video frame, the category feature 402 of the action performed by the first object and the visual feature 404 of the action performed by the first object are concatenated to obtain the action fusion feature 407.

[0085] Exemplarily, for the video frame in the video to be detected, the category feature of the first object and the visual feature of the first object Perform splicing to obtain the object fusion feature . It can be a pre-trained splicing parameter matrix, and \(\sigma\) is a non-linear activation function. Use the convolutional sub-module 313 to perform spatio-temporal graph convolution operation on the object fusion feature to obtain the frame fusion feature . Similar to the calculation method of the object fusion feature, for the video frames in the video to be detected , splice the category feature of the action performed by the first object and the visual feature to obtain the action fusion feature . Among them, the non-linear activation function can be the ReLU function or the Sigmoid function.

[0086] According to an embodiment of the present invention, the multi-modal feature 303 may include the frame fusion feature and the action fusion feature obtained in the above steps.

[0087] Next, with reference to Figure 5 , in combination with specific embodiments, the spatio-temporal graph convolution operation performed by the convolutional sub-module 313 will be further described.

[0088] Figure 5 Schematically shows a schematic diagram of the principle of the convolutional sub-module according to an embodiment of the present invention.

[0089] According to an embodiment of the present invention, for the above spatio-temporal graph convolution operation, the object fusion feature 405 can be convolved by the spatial graph convolution unit 3131 and the temporal graph convolution unit 3132 respectively to obtain the spatial feature 501 and the temporal feature 502; then the spatial feature 501 and the temporal feature 502 are summed to obtain the frame fusion feature 406.

[0090] Exemplarily, when using a graph convolutional network to update the features in a video scene graph. To capture the visual relationship between the first objects within a single video frame, an adjacency matrix of the spatial relationship can be defined which can represent the identifier of the spatial relationship, and \(n\) represents the number of the video frame. For the static scene graph of the video frame in the video to be detected, if the first object pair is connected by the predicate relationship of the first object, then the corresponding element of the adjacency matrix is 1, otherwise it is 0. For the adjacency matrix of the above video frame Each line in it is normalized, and the normalization process refers to formula (2):

[0091]

[0092] Among them, represents the element in the adjacency matrix corresponding to the first object pair Corresponding to the adjacency matrix

[0093] Exemplarily, when using a graph convolutional network to update the features in a video scene graph. In order to obtain the dynamic changes of the first object in adjacent frames in the time dimension, multiple adjacency matrices of the first object in the time dimension can be defined , which can represent the identifier of the time relationship

[0094] According to the above-mentioned time co-reference edge set assign values to the adjacency matrix in the time dimension. When the first object in the video frame and its adjacent next video frame in the first object represent the same first object, set the element of the corresponding adjacency matrix to 1, otherwise 0. Similarly, normalize each line in the adjacency matrix , and the normalization process refers to formula (3):

[0095]

[0096] Among them, can represent the total number of first objects included in the frame, represents the element in the adjacency matrix corresponding to the first object in the video frame and its adjacent next video frame in the first object Corresponding to the adjacency matrix

[0097] Exemplarily, the feature matrix composed of the object fusion features of all the first objects in the video frame can be denoted as , among which, is the transpose operator, and the meaning of appearing in other positions below remains unchanged, is the feature dimension of the first object, represents the object fusion feature of the first object . The process of performing spatial graph convolution operation refers to formula (4); the process of performing temporal graph convolution operation refers to formula (5):

[0098]

[0099]

[0100] Among them, can be the pre-trained spatial convolutional graph parameter matrix of layer can be the pre-trained temporal graph convolutional parameter matrix of layer In the temporal graph convolutional operation, the result of the previous layer of temporal graph convolution of the video frame is used as the input for the temporal graph convolutional operation. can be the number of the convolutional layer represents the non-linear activation function is the feature matrix after the -th layer of spatial graph convolution of the video frame is the feature matrix after the +1-th layer of temporal graph convolution of the video frame Exemplarily, in order to fuse the feature representations of spatial and temporal relationships, the result of the

[0101] -th layer of spatial graph convolutional operation of the video frame and the result of the -th layer of temporal graph convolutional operation of the video frame are summed up, and the feature matrix of the -th layer convolutional video frame can be obtained: the graph convolutional feature representation The calculation reference formula (6) is as follows:

[0102]

[0103] According to the embodiments of the present invention, after layers of spatio-temporal graph convolution, the set of feature representations of all first objects in the video frame after spatio-temporal graph convolution is the frame fusion feature where is the feature dimension of the -th layer graph convolutional network. For the object fusion feature of the first object in the video frame after spatio-temporal graph convolution, it can be counted as the frame fusion sub-feature In practical applications, the spatio-temporal feature update of the video scene graph adopts a 2-layer graph convolutional network, namely The feature dimensions of the spatial graph convolution and the temporal graph convolution ​Both can be set to 512.

[0104] According to an embodiment of the present invention, the frame fusion features obtained by using spatio-temporal graph convolution can capture the dynamic changes of the same first object in the video frames, and can effectively represent the complex and implicit actions in the video to be detected. When generating the object fusion features and the action fusion features, both the first object and the actions performed by the first object are described from both semantic and visual aspects, which can make the detection of the video content more accurate based on the target text.

[0105] According to an embodiment of the present invention, the semantic features include the category features of the target object and the category features of the action to be detected; the text extraction module is used to process the text scene graph to generate semantic features for describing the action to be detected, including: the text extraction module processes the text scene graph semantically to obtain the category features of the target object and the category features of the action to be detected.

[0106] According to an embodiment of the present invention, in order to realize the unified cross-modal representation of the video to be detected and the text scene graph, the text extraction module is used to respectively perform feature encoding on the target object and the relationship edges of the target object in the text scene graph, and perform a bidirectional GRU operation on the obtained feature encoding to obtain semantic features, which include the category features of the target object and the category features of the action to be detected.

[0107] Exemplarily, using a pre-trained text word embedding matrix to encode each target object to obtain an initial embedding vector representation , where represents the one-hot encoded vector of the word in the target text corresponding to the target object , and the text word embedding matrix adopts the same initialization method as the video word embedding matrix.

[0108] Exemplarily, a neural network can be used to perform feature encoding on the sequence path formed by the target object in the original word order to obtain the forward and backward hidden state representations of the target object. The neural network can use a Bi-GRU network. The forward hidden state representation of the target object can be ; represents the forward hidden state representation of the target object , represents the forward GRU operation. The backward hidden state representation of the target object can be ; represents the backward hidden state representation of the target object , ​​Represents a backward GRU operation. The category feature of the target object is obtained by calculating the average of the hidden states in the forward and backward directions. of the target object .

[0109] Exemplarily, for the set of relationship edges of the target object in the text scene graph , the above method is used to form the target objects in the form of triples , target object and the relationship predicate . The corresponding word embedding features are respectively represented as . Using the embedding vector sequence of the triple nodes as the input for feature encoding, finally, the hidden state output by the Bi-GRU network is taken, and the average of the hidden states is obtained to get the category feature of the action to be detected . The calculation method of the category feature of the action to be detected refers to Formula (7):

[0110]

[0111] where can represent the sequential concatenation of the word embedding features of the target object. In practical applications, the dimension of the hidden layer of the Bi-GRU network can be uniformly set to 512.

[0112] According to the embodiments of the present invention, the text extraction module disassembles the target text into target objects and actions to be detected, so that video content detection can be carried out from two aspects of target object detection and target action detection, improving the detection accuracy.

[0113] According to the embodiments of the present invention, for a single video frame, the node matching sub-module calculates the first initial similarity between the category feature of each target object and the frame fusion feature, and determines the maximum value of the first initial similarity as the object similarity corresponding to each target object; after summing the object similarities of multiple target objects and taking the average, the first similarity of the video frame is obtained; for a single video frame, the node matching sub-module calculates the second initial similarity between the category feature of each action to be detected and the action fusion feature of each first object on the video frame, and determines the maximum value of the second initial similarity as the action similarity, and after summing the action similarities of multiple actions to be detected and taking the average, the second similarity of the video frame is obtained; the first similarity and the second similarity of the same video frame are added to obtain the node-level similarity of the corresponding video frame, and the maximum node-level similarity among the node-level similarities of multiple video frames is determined as the node matching score.

[0114] According to an embodiment of the present invention, for a single video frame, the first initial similarity may represent the similarity between each target object and multiple first objects, and the largest one may be used as the object similarity. The average value of the object similarities of multiple target objects is taken to obtain the first similarity. The second initial similarity may represent the similarity between each action to be detected and the actions performed by multiple first objects, and the largest one may be used as the action similarity. The average value of the action similarities of multiple actions to be detected is taken to obtain the second similarity. Since the node-level similarity combines the first similarity and the second similarity, the node matching score obtained through the node-level similarity can detect the entire video to be detected from both the object and action aspects. The method for obtaining the node matching score is specifically described below in conjunction with embodiments.

[0115] Exemplarily, the multi-modal features of the video frame in the video to be detected can be obtained according to the above steps. The multi-modal features include frame fusion features and action fusion features . Based on the target text, semantic features can be obtained, and the semantic features include the category features of the target objects and the category features of the actions to be detected . The multi-modal features and semantic features are used for node granularity matching to obtain the node matching score.

[0116] Exemplarily, for each target object, calculate its first initial similarity with the frame fusion features composed of the frame fusion sub-features corresponding to all first objects in the video frame , and determine the maximum value of the first initial similarity as the object similarity corresponding to each target object; sum and average the object similarities of multiple target objects respectively to obtain the first similarity of the video frame , and the first similarity The calculation reference formula (8) is as follows: The first similarity The calculation reference formula (8) is as follows: The calculation reference formula (8) is as follows:

[0117]

[0118] Among them, represents the number of the target object in the target text, can represent the total number of target objects in the target text.

[0119] Exemplarily, for the category feature corresponding to the relationship edge of each target object and the action fusion feature of each first object on the video frame The second initial similarity between them, determine the maximum value of the second initial similarity as the action similarity, sum the action similarities of multiple actions to be detected and take the average to obtain the video frame of the second similarity , the second similarity The calculation reference formula (9):

[0120]

[0121] wherein, can represent the total number of relationship edges of the first object in the video frame , can represent the total number of relationship edges of the nodes of the target object included in the target text

[0122] Exemplarily, add the first similarity of the video frame and the second similarity to obtain the node-level similarity of the video frame . Finally, for all the total video frames in the video to be detected, take the maximum value among the node-level similarities to obtain the node matching score = , represents the node matching identifier

[0123] According to the embodiments of the present invention, the first similarity effectively focuses on the feature information in the video frame that best matches the target object, avoiding the matching error caused by the confusion of multiple object features. For example, in a video frame containing multiple types of vehicles, such as cars, bicycles, motorcycles, etc., this method can accurately find the part with the highest matching degree of the target vehicle category features, highlighting the key object features, so that the first similarity can accurately reflect the overall matching degree of the target object in this video frame. The second similarity can accurately capture the features that best match the action to be detected from the complex video action scene, improving the accuracy of action matching. For example, in a sports event video, for the action to be detected of "shooting a basket", it can accurately identify the moment with the highest matching degree of the "shooting a basket" feature among the various action postures of the player, so that the second similarity can effectively measure the matching situation of the action in the video frame. Selecting the maximum node-level similarity as the score ensures that the selected result is the one that best represents the matching degree of the video frame, providing an accurate data basis for subsequent in-depth analysis and decision-making based on video content

[0124] According to an embodiment of the present invention, the frame fusion feature includes frame fusion sub-features corresponding to each first object, and the frame matching score is determined as follows: performing weighted aggregation on the frame fusion sub-features corresponding to each first object in a single video frame to obtain a first graph-level feature of the single video frame; performing weighted aggregation on the action fusion features corresponding to each action in a single video frame to obtain a second graph-level feature of the single video frame; splicing the first graph-level feature and the second graph-level feature of the same video frame to obtain an object graph-level feature; performing weighted fusion on the category features of multiple target objects to obtain a third graph-level feature; performing weighted fusion on the category features of multiple actions to be detected to obtain a fourth graph-level feature; splicing the third graph-level feature and the fourth graph-level feature to obtain a text graph-level feature; respectively determining the similarity scores between the text graph-level feature and the multiple object graph-level features to obtain the frame matching similarity scores of each of the multiple video frames, and taking the maximum frame matching similarity score as the frame matching score.

[0125] According to an embodiment of the present invention, the first graph-level feature is the feature of each first object after weighted aggregation, which can represent the features exhibited by all first objects in a single video frame; similarly, the second graph-level feature can represent the features exhibited by the actions performed by all first objects in a single video frame. The object graph-level feature obtained by splicing the first graph-level feature and the second graph-level feature can represent the overall features of the first object itself and the actions it performs. Similarly, the text graph-level feature can represent the overall features of the target object and the action to be detected. So that the frame matching score can detect whether the content to be detected appears in the video to be detected as a whole. The following specifically describes the method for obtaining the frame matching score in combination with embodiments.

[0126] According to an embodiment of the present invention, the frame fusion sub-feature is the frame fusion feature corresponding to all first objects included in the video frame corresponding to the video frame .

[0127] According to an embodiment of the present invention, inputting the multi-modal features into the graph embedding layer to obtain an object graph-level feature , inputting the semantic features into the graph embedding layer to obtain a text graph-level feature , and performing frame granularity matching on the object graph-level feature and the text graph-level feature to obtain the frame matching score. The following specifically describes in combination with the embodiments of the present invention.

[0128] According to an embodiment of the present invention, performing weighted aggregation on the frame fusion sub-features corresponding to each first object in a single video frame to obtain a first graph-level feature of the single video frame.

[0129] Exemplarily, to obtain the video frame The first graph-level feature , using multi-scale attention to aggregate the frame fusion sub-features of the first object . The calculation method of the first graph-level feature refers to Formula (10):

[0130]

[0131] wherein, can represent the attention weight of the th first object in the video frame . The calculation method of refers to Formula (11):

[0132]

[0133] wherein, can represent the exponential function, can represent the frame fusion sub-feature of the first object , represents the number of the first object in the video frame , can be a multi-scale attention mapping function.

[0134] According to the embodiments of the present invention, weighted aggregation is performed on the action fusion features corresponding to each action in the video frame to obtain the second graph-level feature of the video frame .

[0135] Exemplarily, similar to the first graph-level feature , the calculation method of the second graph-level feature refers to Formula (12):

[0136]

[0137] wherein, is the total number of the first object relationship edges in the video frame , can represent the attention weight of the th relationship edge of the first object , represents the number of the relationship edge of the first object, The calculation method of

[0138]

[0139] wherein, represents the first object pair in the video frame ( , The action fusion feature of represents the number of the relational edge of the first object.

[0140] Exemplarily, the video frame The first graph-level feature of and the second graph-level feature are concatenated to obtain the object graph-level feature = . Similar to the above calculation of the object graph-level feature , the category features of multiple target objects in the target text can be weighted and fused to obtain the third graph-level feature ; the category features of multiple actions to be detected in the target text are weighted and fused to obtain the fourth graph-level feature ; the third graph-level feature and the fourth graph-level feature are concatenated to obtain the text graph-level feature .

[0141] According to the embodiments of the present invention, the similarity scores between the text graph-level feature and the object graph-level feature corresponding to each video frame are determined respectively, and multiple frame matching similarity scores of each video frame are obtained. The calculation method of the frame matching similarity score refers to formula (14):

[0142]

[0143] wherein, represents the frame matching identifier.

[0144] Exemplarily, finally, the maximum frame matching similarity score is used as the frame matching score , which represents the frame matching identifier.

[0145] According to an embodiment of the present invention, the frame matching score is used to represent the semantic similarity between the text scene graph obtained from the target text and the overall of the first object and the action performed by the first object on each frame of the video to be detected. By assigning weights based on the importance of different first objects in the video scene, key first object features can be highlighted, effectively avoiding the interference of secondary first objects, making the obtained first graph-level features more representative, and being able to accurately reflect the core features of the first object in a single video frame. Similarly, the second graph-level features obtained by weighted aggregation of the action fusion features can also focus on key action information and enhance the expression ability of action features. The text graph-level features and the object graph-level features encode the video frame content and text semantics from different perspectives. By calculating the similarity score between them, the potential association between the video frame and the text description can be deeply mined. Using the maximum frame matching similarity score as the frame matching score ensures that the selected score is the result that best reflects the matching degree between the two, effectively improving the accuracy of the video frame and text description matching. Compared with the traditional frame matching method, this method can more accurately locate the video frame that matches the target text description in a complex video scene, significantly improving the accuracy and reliability of the target detection system in the video analysis scenario, providing more accurate basic data for subsequent intelligent analysis and decision-making based on video content, and enhancing the practicality and adaptability of the video content detection method.

[0146] According to an embodiment of the present invention, the video matching sub-module is used to perform average pooling operation on the object graph-level features of multiple video frames in the video to be detected to obtain object global features; calculate the similarity between the object global features and the text graph-level features to obtain a video matching score.

[0147] According to an embodiment of the present invention, the above-mentioned calculation of similarity can be cosine similarity.

[0148] According to an embodiment of the present invention, the video matching sub-module is used to perform mean pooling operation on the object graph-level features of multiple video frames in the video to be detected to obtain object global features ; perform video granularity matching between the object global features and the text graph-level features to obtain a video matching score , which is a video matching identifier. The calculation method of the video matching score refers to formula (15):

[0149]

[0150] wherein, represents the calculation of the Euclidean norm.

[0151] According to an embodiment of the present invention, when performing average pooling operation on the object graph-level features of multiple video frames, this process can effectively integrate the local differences of the features of the first object within a single video frame, aggregate the first object information scattered in different frames, and eliminate the feature fluctuations caused by factors such as shooting angles and lighting changes, so as to obtain object global features with global representativeness. Compared with traditional video-text matching methods, the embodiments of the present invention globally integrate features through average pooling operation and calculate similarity based on the integrated features, significantly improving the accuracy and stability of video matching scores, being able to quickly and accurately screen out video content that matches the text description from a large amount of video data, greatly improving the efficiency of video retrieval and analysis, providing more reliable technical support for application scenarios such as intelligent monitoring and video recommendation, and enhancing the adaptability and practicality of the system in complex actual applications.

[0152] According to an embodiment of the present invention, the detection module is used to generate a target detection result based on the similarity between multi-modal features and semantic features, including: the detection module respectively obtains a node matching score, a frame matching score, and a video matching score based on the multi-modal features and semantic features; comparing the average value of the node matching score, the frame matching score, and the video matching score with a first threshold, and determining the target detection result according to the comparison result.

[0153] According to an embodiment of the present invention, when training the detection module, the value range of the first threshold can be determined. Generally, if the calculated average value is greater than the first threshold, it indicates that the video to be detected contains the content to be detected.

[0154] Exemplarily, calculate the above-mentioned node matching score , frame matching score and video matching score to obtain an average score . is the average score identifier.

[0155] When detecting the content of the video to be detected, if the average score is greater than the first threshold , it can be determined that the video to be detected contains the content to be detected of the target object in the target text input by the user; otherwise, it is determined that the video does not contain the content to be detected of the target object. In the actual application process, the first threshold can be set to 0.5 according to the performance of the validation set.

[0156] According to an embodiment of the present invention, taking the average of three matching scores can comprehensively balance the influence of scores in different dimensions and reduce the risk of misjudgment caused by score deviation in a certain dimension. For example, when the node matching score is low, but the frame matching score and the video matching score are high, taking the average can avoid ignoring the overall matching situation due to excessive attention to the deficiencies at the node level. This fusion method makes the detection result more robust and can adapt to the diversity and complexity of multimodal data.

[0157] According to an embodiment of the present invention, the video content detection method of the present invention includes: in the case where the target detection result indicates that the to-be-detected video includes a to-be-detected action of a target object, using the frame matching score, the object graph-level feature of the video frame to which the frame matching score belongs is subjected to feature enhancement to obtain an enhanced feature; the object global feature and the text graph-level feature are concatenated to obtain an action vector ; the action vector is concatenated with the enhanced feature to obtain a video frame sequence; the video frame sequence is input into a recurrent neural network, and the start time and end time of the target object performing the to-be-detected action in the to-be-detected video are output.

[0158] Exemplarily, the frame matching similarity score of each video frame can be used as a weight to perform associated frame feature enhancement on the object graph-level feature of each frame to obtain an enhanced feature , and the calculation method of the enhanced feature refers to formula (16):

[0159]

[0160] where , can represent the result of mapping the frame matching similarity score of video frame to , can be used to characterize the correlation degree between video frame and the target text. The function used for mapping can be function.

[0161] Exemplarily, the object global feature and the text graph-level feature are concatenated to obtain an action vector ; the action vector is concatenated to the front end of the enhanced feature to obtain a video frame sequence . represents the feature dimension of the video frame sequence, represents the enhanced feature of the first frame, represents the enhanced feature of the second frame, is the total number of frames of the video to be detected, represents the real number space.

[0162] Among them, the action vector has an index of 0 in the video frame sequence, and the action vector includes the global relationship information between the video to be detected and the target text. If the content to be detected does not exist in the video to be detected, this node can be used as an identifier for both the start and end positions to clearly indicate that there is no segment of the content to be detected.

[0163] According to an embodiment of the present invention, if the target detection result indicates that the video to be detected includes the action to be detected of the target object, the above video frame sequence is input into two one-way recurrent neural networks (Long Short-Term Memory, ) for position prediction. The recurrent neural network is used to capture the temporal dependence relationships of the start and end boundaries of the segment where the action to be detected of the target object appears respectively: The calculation method of the temporal dependence relationship refers to Formula (17) and Formula (18):

[0164]

[0165]

[0166] Among them, can represent the starting hidden state of the video frame , can represent the starting hidden state of the video frame , can represent the ending hidden state of the video frame , can represent the ending hidden state of the video frame , represents the start identifier, represents the end identifier, represents the video frame sequence in the video frame features.

[0167] Exemplarily, the probability distribution of the start of the content segment and the probability distribution of the end position are predicted through two feed-forward neural networks (Feed-Forward Neural Network, FFN), the calculation method of refers to Formula (19),

[0168]

[0169]

[0170] Among them, represents the initial hidden state sequence composed of the initial hidden states of all video frames, represents the end hidden state sequence composed of the end hidden states of all video frames, represents the splicing operation of features, represents the activation function.

[0171] Exemplarily, finally, through and the positions where the content to be detected appears and ends are obtained, realizing the prediction of the position of the video content.

[0172] According to an embodiment of the present invention, the frame matching score can reflect the matching degree between the video frame and the target text. Based on this, feature enhancement is performed, which can highlight the key features related to the target object and the action to be detected, and suppress irrelevant information. For example, when detecting the action of a pedestrian running in the target text, features related to running such as the limb posture and movement speed of the pedestrian can be strengthened, and irrelevant factors such as the clothing color of the pedestrian can be weakened, making the enhanced features more targeted and representative, and facilitating the prediction of the position where the content to be detected appears using a neural network.

[0173] According to an embodiment of the present invention, the following refers to Figure 6 to further explain the data processing process of the present invention.

[0174] Figure 6 Schematically shows the data processing flow chart of the video content detection method according to an embodiment of the present invention;

[0175] As Figure 6 shown, the target text input by the user is: "A woman in a blue shirt is looking at a man". In the generated text scene graph, the user's detection intention can be: the user wants to detect whether a woman in a blue shirt looking at a man appears in the video to be detected. The man, woman, and blue shirt in the text scene graph can all be used as target objects, and the looking at and wearing in the text scene graph can be used as the actions to be detected.

[0176] According to an embodiment of the present invention, a video scene graph is generated according to a video to be detected. The action timing information of the first object included in the video scene graph can be the action change of the first object itself or the action change between two first objects. For example, as Figure 6Shown as follows: The two first objects can be a man and a woman, and the action change between the two first objects can be that they change from looking at each other to hugging. The man, woman, and blue shirt in the video scene graph can be the first objects, and hugging, looking at each other, and wearing can be the actions performed by the first objects.

[0177] According to an embodiment of the present invention, next, the text extraction module is used to respectively perform feature encoding on the target object and the action to be detected in the text scene graph, and perform a bidirectional GRU operation on the obtained feature encodings to obtain semantic features. The convolutional sub-module is used to perform spatio-temporal graph convolution on the object fusion features of the first object to obtain frame fusion features, and the frame fusion features and the action fusion features of the actions performed by the first object are used as multi-modal features.

[0178] According to an embodiment of the present invention, multi-modal features and semantic features can be used for node granularity matching to obtain node matching scores.

[0179] According to an embodiment of the present invention, the multi-modal features are input into the graph embedding layer to obtain object graph-level features , the semantic features are input into the graph embedding layer to obtain text graph-level features , the object graph-level features and the text graph-level features are subjected to frame granularity matching to obtain frame matching scores.

[0180] According to an embodiment of the present invention, the video matching sub-module is used to perform mean pooling operation on the object graph-level features of multiple video frames in the video to be detected to obtain object global features . Video granularity matching is performed between the object global features and the text graph-level features to obtain video matching scores.

[0181] According to an embodiment of the present invention, in the video content detection method, the frame matching similarity score of each video frame can be used as a weight to perform associated frame feature enhancement on the object graph-level features of each frame to obtain enhanced features. The object global features and the text graph-level features are concatenated to obtain an action vector . The action vector is concatenated to the front end of the enhanced features to obtain a video frame sequence. Finally, the video frame is input into a recurrent neural network for position prediction. As can be seen from Figure 6 , in the video to be detected, the start position time and end position time of the picture when the man and woman are hugging are located.

[0182] Generally, the supervision for the training of a model is determined by a loss function. In the video content detection method according to the embodiments of the present invention, a position prediction model can also be used for prediction. When training the position prediction model, the final loss function can be used for the supervision of the training. The final loss function can refer to Formula (21): wherein, is the frame-level loss function,

[0183]

[0184] wherein, is the frame-level loss function, is the video-level loss function, is the cross-entropy loss function. is the first hyperparameter, is the second hyperparameter, and are used to balance the above three loss functions.

[0185] According to the embodiments of the present invention, the frame-level loss function is used to narrow the feature similarity between the text scene graph and the positive sample video frame and increase the difference from the negative sample video frame. To this end, a frame-level loss function with a margin can be constructed , which can refer to Formula (22):

[0186]

[0187] wherein, is the similarity score between the text scene graph and the static scene graph of the positive sample of the video frame , is the similarity score between the text scene graph and the static scene graph of the negative sample of the video frame . The similarity score can be the cosine similarity. is the margin hyperparameter for frame-level matching. By minimizing the similarity between the target text and the positive sample video frame and maximizing the similarity with the negative sample frame at the same time, accurate detection of the action to be detected and the target object is achieved. In the actual training process, the margin hyperparameter in the frame-level loss function can be set to 0.4.

[0188] According to the embodiments of the present invention, the video-level loss function can refer to Formula (23):

[0189]

[0190] Among them, represents the object global feature of the positive sample and the text graph-level feature between the similarity scores, represents the object global feature of the negative sample and the text graph-level feature between the similarity scores. The similarity score can be the cosine similarity.

[0191] According to an embodiment of the present invention, the positive sample represents a video containing the target object and the action to be detected, and the negative sample represents a video not containing the target object and the action to be detected. By minimizing the similarity between the target text and the positive sample frames of the video, while maximizing the similarity with the negative sample frames, accurate detection of the action to be detected and the target object is achieved. During the actual training process, the video granularity loss function interval parameter can be set to 0.3.

[0192] In an embodiment of the present invention, the cross-entropy loss function can refer to formula (24):

[0193]

[0194] Among them, is the probability that the content to be detected starts to appear in the video frame at, is the probability that the content to be detected ends in the video frame at. is the one-hot vector label that the content to be detected starts to appear in the video frame at, is the one-hot vector label that the content to be detected ends in the video frame at.

[0195] Among them, during the actual training process, the first weight parameter can be and the second weight parameter can be . The optimization of the model uses the Adam optimizer, the initial learning rate is 1e-4, the batch size is set to 32, the training can be carried out for 150 epochs, and the model with the best performance on the validation set is selected for the final test.

[0196] Figure 7 Schematically shows a block diagram of a video content detection device according to an embodiment of the present invention.

[0197] As Figure 7As shown, the video content detection device 700 includes a text scene graph generation module 710, a video scene graph acquisition module 720, and a target detection module 730.

[0198] The text scene graph generation module 710 is configured to generate a text scene graph based on the target text in response to receiving the target text input by an object.

[0199] The video scene graph acquisition module 720 is configured to acquire a video scene graph of the video to be detected.

[0200] The target detection module 730 is configured to process the text scene graph and the video scene graph by using a video content detection model to generate a target detection result.

[0201] According to embodiments of the present invention, any multiple of the modules, sub-modules, units, and sub-units, or at least part of the functions of any multiple of them, can be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present invention can be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present invention can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in a suitable combination of any several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present invention can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.

[0202] For example, any combination of the text scene graph generation module 710, the video scene graph acquisition module 720, and the target detection module 730 can be combined and implemented in one module / unit / sub-unit, or any one of these modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units can be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present invention, at least one of the text scene graph generation module 710, the video scene graph acquisition module 720, and the target detection module 730 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the text scene graph generation module 710, the video scene graph acquisition module 720, and the target detection module 730 can be at least partially implemented as a computer program module, which can execute corresponding functions when the computer program module is run.

[0203] It should be noted that the video content detection device in the embodiments of the present invention corresponds to the video content detection method in the embodiments of the present invention. For the description of the video content detection device part, please refer to the video content detection method part specifically, and it will not be elaborated here.

[0204] Figure 8 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present invention is schematically shown. Figure 8 The shown electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0205] As Figure 8 shown, the electronic device 800 according to an embodiment of the present invention includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 88 into a random access memory (RAM) 803. The processor 801 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and so on. The processor 801 can also include on-board memory for caching purposes. The processor 801 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0206] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to the embodiments of the present invention by executing programs in the ROM 802 and / or the RAM 803. It should be noted that the programs can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method flow according to the embodiments of the present invention by executing programs stored in the one or more memories.

[0207] According to an embodiment of the present invention, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read therefrom can be installed into the storage portion 808 as needed.

[0208] According to an embodiment of the present invention, the method flow according to the embodiments of the present invention can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication portion 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system according to the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.

[0209] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the methods according to the embodiments of the present invention are implemented.

[0210] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0211] For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories other than the above-described ROM 802 and / or RAM 803 and / or ROM 802 and RAM 803.

[0212] An embodiment of the present invention further includes a computer program product, which includes a computer program containing program code for executing the method provided by the embodiment of the present invention. When the computer program product runs on an electronic device, the program code is used to enable the electronic device to implement the video content detection method provided by the embodiment of the present invention.

[0213] When the computer program is executed by the processor 801, the above functions defined in the system / apparatus of the embodiment of the present invention are executed. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0214] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium and be downloaded and installed through the communication part 809, and / or be installed from the removable medium 811. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0215] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0216] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0217] The above describes the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A video content detection method, characterized in that, The method includes: In response to receiving the target text input by an object, generating a text scene graph based on the target text, where the text scene graph includes a to-be-detected action and a target object related to the detection intention; Obtaining a video scene graph of the to-be-detected video; wherein, the video scene graph is used to characterize the action timing information of a first object in the to-be-detected video; Using a video content detection model to process the text scene graph and the video scene graph to generate a target detection result; wherein, the target detection result indicates whether the to-be-detected action of the target object is included in the to-be-detected video.

2. The method according to claim 1, wherein The video content detection model includes a video extraction module, a text extraction module, and a detection module; using the video content detection model to process the text scene graph and the video scene graph to generate a target detection result includes: Using the video extraction module to process the video scene graph and the to-be-detected video to generate multimodal features for describing the action change of the first object; Using the text extraction module to process the text scene graph to generate semantic features for describing the to-be-detected action; Using the detection module to generate a target detection result based on the similarity between the multimodal features and the semantic features.

3. The method according to claim 2, wherein The multimodal features include frame fusion features and action fusion features; the video extraction module includes a visual extraction sub-module, a semantic extraction sub-module, and a convolutional sub-module; Using the video extraction module to process the video scene graph and the to-be-detected video to generate multimodal features for describing the action change of the first object includes: Using the semantic extraction sub-module to semantically process the video scene graph to obtain the category features of the first object and the category features of the action performed by the first object included in each video frame of the to-be-detected video; Using the visual extraction sub-module to process the video scene graph and the to-be-detected video to obtain the visual features of the first object and the visual features of the action performed by the first object included in each video frame of the to-be-detected video; For each video frame, splicing the category features and visual features of the first object to obtain object fusion features; Using the convolutional sub-module to perform spatio-temporal graph convolution operations on the object fusion features to obtain frame fusion features; For each video frame, splicing the category features and visual features of the action performed by the first object to obtain action fusion features.

4. The method according to claim 3, characterized in that The Using the visual extraction sub-module to process the first object in the video scene graph and the to-be-detected video to obtain the visual features of the first object and the visual features of the action performed by the first object includes: Using the visual extraction sub-module to perform feature extraction on the video frames including the first object in the to-be-detected video within the range of the area where the first object is located to obtain the visual features of the first object and the visual features of the action performed by the first object.

5. The method according to claim 3, characterized in that, The semantic features include the category features of the target object and the category features of the to-be-detected action; using the text extraction module to process the text scene graph to generate semantic features for describing the to-be-detected action includes: Using the text extraction module to semantically process the text scenario graph to obtain the category features of the target object and the category features of the action to be detected.

6. The method according to claim 5, characterized in that, The generating of the target detection result by using the detection module based on the similarity between the multi-modal features and the semantic features includes: Using the detection module to respectively obtain a node matching score, a frame matching score, and a video matching score based on the multi-modal features and the semantic features; Comparing the average value of the node matching score, the frame matching score, and the video matching score with a first threshold, and determining the target detection result according to the comparison result.

7. The method according to claim 6, characterized in that: The detection module includes a node matching sub-module, a frame matching sub-module, and a video matching sub-module; the node matching score is determined by the following method: For a single video frame, using the node matching sub-module to calculate the first initial similarity between the category features of each target object and the frame fusion feature, and determining the maximum value of the first initial similarity as the object similarity corresponding to each target object; after summing the object similarities of multiple target objects and taking the average, the first similarity of the video frame is obtained; For a single video frame, using the node matching sub-module to calculate the second initial similarity between the category features of each action to be detected and the action fusion feature of each first object on the video frame, and determining the maximum value of the second initial similarity as the action similarity, after summing the action similarities of multiple actions to be detected and taking the average, the second similarity of the video frame is obtained; Adding the first similarity and the second similarity of the same video frame to obtain the node-level similarity of the corresponding video frame, and determining the maximum node-level similarity among the node-level similarities of multiple video frames as the node matching score.

8. The method according to claim 7, wherein: The frame fusion feature includes frame fusion sub-features corresponding to each first object, and the frame matching score is determined by the following method: Performing weighted aggregation on the frame fusion sub-features corresponding to each first object in a single video frame to obtain the first graph-level feature of the single video frame; Performing weighted aggregation on the action fusion features corresponding to each action in a single video frame to obtain the second graph-level feature of the single video frame; Concatenating the first graph-level feature and the second graph-level feature of the same video frame to obtain the object graph-level feature; Performing weighted fusion on the category features of multiple target objects to obtain a third graph-level feature; Performing weighted fusion on the category features of multiple actions to be detected to obtain a fourth graph-level feature; Concatenating the third graph-level feature and the fourth graph-level feature to obtain the text graph-level feature; Respectively determining the similarity scores between the text graph-level feature and multiple object graph-level features to obtain the frame matching similarity scores of multiple video frames, and taking the maximum frame matching similarity score as the frame matching score.

9. The method according to claim 8, wherein: The video matching score is determined by the following method: Using the video matching sub-module to perform average pooling operation on the object graph-level features of multiple video frames in the video to be detected to obtain the object global feature; Calculate the similarity between the global object feature and the text graph-level feature to obtain the video matching score.

10. The method according to claim 9, wherein The method further includes: When the target detection result indicates that the video to be detected includes the target action of the target object, use the frame matching score to enhance the object graph-level feature of the video frame to which the frame matching score belongs to obtain an enhanced feature; Concatenate the global object feature and the text graph-level feature to obtain an action vector; Concatenate the action vector and the enhanced feature to obtain a video frame sequence; Input the video frame sequence into a recurrent neural network, and output the start time and end time of the target object in the video to be detected performing the target action.

Citation Information

Patent Citations

  • Action object recognition in cluttered video scene using text

    CN116324906A

  • Video object positioning method and device, storage medium and program product

    CN117351382A

  • Image and text semantic similarity calculation method and system applied to social media

    CN117421609A

  • Method and device for determining similarity between text and video

    CN117556276A

  • Image text retrieval method and system based on scene graph

    CN118673166A

Cited By

  • Disaster situation video description generation method for emergency disaster scene

    CN121121593A

  • Video and text similarity evaluation method

    CN121167332A