A video understanding method and device integrating multimodal knowledge graph

By extracting multimodal features in video and combining knowledge graphs for video understanding, the problem of simple coarse granularity and multimodal fusion method in the prior art is solved, and a more accurate video understanding and a more stable retrieval process are achieved.

CN119672606BActive Publication Date: 2025-06-06BEIJING YOUREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411749413.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-06-06
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

The existing video understanding technology has problems such as coarse video understanding, lack of ability to extract key details of video information, and simple multimodal fusion method and not in-depth modeling of semantic relationships between various modes, which leads to inaccurate question-and-answer responses and lack of detection of search status, which cannot guarantee the healthy operation of search.

Method used

The multimodal features (visual features, audio features, and text features) are extracted from the video, converted into feature vectors, and analyzed and fused to obtain the fusion features. Then, the fusion features are combined with the knowledge graph for video understanding, and the semantic relationships between various modes are deeply modeled to improve the accuracy of video understanding. At the same time, by conducting a comprehensive analysis of the search data and search related content, search status information is obtained, and search self-inspection and early warning are carried out to ensure the stability of the search.

Benefits of technology

Improve the accuracy of video comprehension, enables more in-depth extraction of key details of videos, and enhances semantic understanding through multimodal fusion. At the same time, by monitoring the search status, the search process is optimized, and the stability and response speed of the search system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672606B_ABST
    Figure CN119672606B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing technology, and specifically to a video understanding method and device that integrates a multimodal knowledge graph. The present invention extracts multimodal features from a video and converts the multimodal features into corresponding feature vectors, and analyzes and fuses the feature vectors corresponding to the multimodal features of each video to obtain fused features, and then obtains the fused features and combines the fused features with the knowledge graph to understand the video to obtain video understanding information, and improves the depth of video understanding and the accuracy of response by analyzing the multimodal features; the present invention obtains retrieval data and comprehensively analyzes the retrieval data and retrieval-related content to obtain retrieval status information, and performs retrieval self-checking based on the retrieval status information to obtain retrieval optimization information, and issues an early warning for abnormal retrieval status, so as to facilitate monitoring of the retrieval status and improve the stability of the retrieval system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and specifically to a video understanding method and device integrating a multimodal knowledge graph. Background Art

[0002] In the process of understanding the world, humans obtain and process information through multiple senses. The construction of multimodal knowledge graphs is closer to human cognitive patterns. It can simulate humans' comprehensive processing and understanding capabilities of information in different modalities, making machines more natural and accurate in processing and understanding complex real-world problems, laying the foundation for smarter human-computer interaction and broader artificial intelligence applications. Video understanding is one of the core research issues in the multimedia field, and its research results are of great significance for promoting the development of multimedia technology. It involves the cross-integration of multiple disciplines such as computer vision, natural language processing, and machine learning, providing new research ideas and methods for the development of these disciplines. At the same time, multimodal knowledge graphs, as a powerful knowledge representation and reasoning tool, provide richer semantic information and knowledge support for video understanding, and video understanding is one of the important application scenarios of multimodal knowledge graphs. The combination of the two is of great significance for promoting the development of artificial intelligence and multimedia technology.

[0003] Existing video understanding technologies mainly rely on visual encoders or large visual models to extract video features, which are combined with text features obtained by speech-to-text or voiceprint features obtained by audio models, as well as user query text features, and are transmitted to the large model for information combination reasoning and giving the final question response. However, existing technologies have problems such as coarse video understanding granularity, lack of ability to extract key video detail information, simple multimodal fusion methods, and lack of in-depth modeling of semantic relationships between modalities, resulting in inaccurate question and answer responses, and lack of detection of retrieval status, which cannot ensure the healthy operation of retrieval. Summary of the invention

[0004] The present invention provides a video understanding method and device integrating multimodal knowledge graphs to solve the above-mentioned technical problems.

[0005] The first aspect of the present invention provides a video understanding method integrating a multimodal knowledge graph, comprising the following steps:

[0006] Step 1: Video content information extraction, extracting multimodal features from the video and converting the multimodal features into corresponding feature vectors.

[0007] As a further improvement of the present invention, the multimodal features include visual features, audio features and text features; the video is divided into multiple video segments using video processing technology, and the visual features of each video segment are extracted using visual recognition technology, and the corresponding visual feature vector is obtained based on the visual features; the audio information in the video is extracted, and the audio information is converted into a spectrogram based on audio processing technology, and the audio features are obtained through the spectrogram, and the corresponding audio feature vector is obtained based on the audio features; the video title, the text obtained by speech recognition and the video text recognition information are collected to obtain text data, and the text features corresponding to the text data are extracted based on natural language processing technology, and the corresponding text feature vector is obtained based on the text features.

[0008] Step 2: Analyze and fuse the feature vectors corresponding to the multimodal features of each video to obtain the fusion feature; identify the multimodal features corresponding to each video to obtain the visual feature vector, audio feature vector and text feature vector, set the corresponding dimension direction of the visual feature vector, audio feature vector and text feature vector respectively, and record the dimension direction corresponding to the visual feature vector as m 1 , then the corresponding visual feature vector is [s 1 ,s 2 , …, ], the dimension direction corresponding to the audio feature vector is recorded as m 2 , then the corresponding audio feature vector is [y 1 ,y 2 , …, ], the dimension direction corresponding to the text feature vector is recorded as m 3 , then the corresponding text feature vector is [z 1 , z 2 , …, ], the visual feature vector, audio feature vector and text feature vector are integrated to obtain the fusion feature, and the corresponding fusion feature vector is [s 1 ,s 2 , …, ,y 1 ,y 2 , …, , z 1 , z 2 , …, ].

[0009] Step 3: Obtain fusion features and combine the fusion features with the knowledge graph to perform video understanding to obtain video understanding information.

[0010] As a further improvement of the present invention, the fusion features are combined with the knowledge graph to perform video understanding, and the specific analysis steps are as follows:

[0011] A pre-understood video is obtained and recognized to obtain a text to be analyzed, which includes but is not limited to the video title, speech recognition text and video text recognition information; entities in the text to be analyzed are recognized based on a pre-trained language model to obtain corresponding text entities, which include but are not limited to persons, place names, objects and events; the text entities are matched with the knowledge graph to obtain the corresponding related entities in the knowledge graph, the text entities are linked with the corresponding related entities, and the corresponding rich semantic information in the knowledge graph corresponding to the related entities is obtained, and the rich semantic information corresponding to the text entities is marked as connection semantic information.

[0012] Based on the semantic information of each connection corresponding to the video, semantic reasoning is performed to obtain the reasoning semantics corresponding to the text to be analyzed in the video, the existing semantic knowledge in the database is obtained, the current reasoning semantics are matched with the existing semantic knowledge to obtain the semantic label corresponding to the video, and the semantic labels of the video are aggregated to obtain the video label information.

[0013] Obtain the corresponding semantic tags in the video tag information, identify each semantic tag to obtain the corresponding semantic features, and match the semantic features corresponding to each semantic tag with the fusion vector to obtain the corresponding vector information [ , , ], overlap and match the vector information corresponding to each semantic label to obtain the semantic repetition of each semantic label, obtain a pre-set semantic repetition threshold, and compare the semantic repetition of each semantic label with the semantic repetition threshold, mark the semantic labels corresponding to those greater than the semantic repetition threshold as high-proportion semantic labels, and mark the vector information corresponding to the high-proportion semantic labels as center vectors, and obtain a pre-set similarity vector space range, count the number of semantic labels corresponding to each center vector within the similarity vector space range, and calculate the ratio of the number of semantic labels corresponding to each center vector to the total number of semantic labels to obtain the semantic proportion of each center vector, obtain the final label according to the semantic proportion, and use the final label as the corresponding video understanding information.

[0014] Step 4: User query response analysis: obtain the user's search content information by identifying the user's query and retrieval operations, obtain the corresponding search features based on the search content information, and match the search features with the knowledge graph to obtain the corresponding search-related content.

[0015] Step 5: Obtain search data and conduct comprehensive analysis on the search data and search-related content to obtain search status information.

[0016] As a further improvement of the present invention, a comprehensive analysis is performed on the search data and the search-related content, and the specific analysis method is as follows:

[0017] The retrieval response time, video understanding processing time and retrieval text are obtained by identifying the retrieval data; the retrieval time of each retrieval operation is obtained, the pre-set retrieval upper limit time is obtained, the retrieval time corresponding to the current retrieval operation is compared with the retrieval upper limit time, the part of the retrieval time that exceeds the retrieval upper limit time is marked as a retrieval excess value, the retrieval excess value is divided into multiple retrieval excess value intervals, each retrieval excess value interval corresponds to a retrieval time shadow value, the retrieval excess value corresponding to the current retrieval time is matched with multiple retrieval excess value intervals to obtain the corresponding retrieval time shadow value.

[0018] Get the retrieval time corresponding to the video understanding processing time of each video, calculate the ratio of the retrieval time to the video understanding processing time to get the retrieval time ratio, get the pre-set retrieval time ratio threshold, and when the retrieval time ratio is greater than the retrieval time ratio threshold, mark the corresponding retrieval time ratio as an abnormal ratio.

[0019] By identifying the retrieval-related content, the ranking position of the retrieval content and the semantics of the related content are obtained, and the retrieval text corresponding to the retrieval data is obtained. The retrieval text semantics are identified by the semantic recognition module, and the semantic similarity of the retrieval text semantics and the semantics of the related content are compared to obtain the corresponding semantic similarity, and the pre-set semantic standard similarity is obtained. If the current semantic similarity is less than the semantic standard similarity, the difference between the semantic similarity and the semantic standard similarity is calculated to obtain the similarity difference. For example, a video ranks 7th in one retrieval and 48th in the next retrieval. This large position change indicates that the retrieval state is unstable, and it is necessary to further optimize the retrieval algorithm or adjust the integration method of video understanding and retrieval.

[0020] According to the ranking position of the search content, the search content corresponding to multiple ranking positions is obtained. Based on the same search content, the search content of the ranking positions corresponding to multiple searches are obtained. The search content corresponding to the same ranking positions of multiple searches are compared to obtain the ranking change value corresponding to the search content. When the ranking change value is greater than the preset ranking change threshold, the ranking change value is marked as a ranking instability value.

[0021] The retrieval time shadow value, anomaly ratio, similarity difference and sorting instability value are normalized and their values ​​are taken, and the formula is used The retrieval status value JS is calculated; jsy, ycb, and xsc represent the retrieval time shadow value, anomaly ratio, and similarity difference, respectively. It is expressed as the sum of n ranking instability values, where n is a positive integer; m1, m2, m3, and m4 are all preset weight factors, with values ​​of 2.425, 2.722, 1.702, and 1.145, respectively; when the retrieval status value is greater than the preset threshold, the corresponding retrieval status information generated is the retrieval status abnormality.

[0022] Step 6: Perform a search self-check based on the search status information to obtain search optimization information, and issue an early warning for abnormal search status.

[0023] The second aspect of the present invention provides a video understanding device that integrates multimodal knowledge graphs, including: a video information extraction module, a feature fusion module, a knowledge graph fusion construction module, a user query response module, and a data storage module.

[0024] The video information extraction module is used to extract and process the video information through a data processor to obtain the multimodal features corresponding to each piece of video information, process and extract features of the text information in the video through a text processing module; process and extract features of the audio information in the video through an audio processing module; and process and extract features of the visual information in the video through a visual processing module.

[0025] The feature fusion module is used to fuse various feature information corresponding to the multimodal features to obtain fused features.

[0026] The knowledge graph fusion construction module is responsible for associating and fusing the fused multimodal features with the pre-built multimodal knowledge graph, including entity recognition and linking units, semantic reasoning and label generation units.

[0027] The user query response module is used to obtain corresponding search key content by processing the user's search information, and to make content response according to the search key content.

[0028] The data storage module is used to store the data information generated by the video information extraction module, the feature fusion module, the knowledge graph fusion construction module and the user query response module.

[0029] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects:

[0030] 1. The present invention extracts multimodal features from videos and converts the multimodal features into corresponding feature vectors, analyzes and fuses the feature vectors corresponding to the multimodal features of each video to obtain fused features, then obtains the fused features and combines the fused features with the knowledge graph to perform video understanding to obtain video understanding information, and improves the depth of video understanding and the accuracy of response by analyzing multimodal features.

[0031] 2. The present invention obtains retrieval data and performs a comprehensive analysis on the retrieval data and retrieval-related content to obtain retrieval status information, performs retrieval self-check based on the retrieval status information to obtain retrieval optimization information, and issues an early warning for abnormal retrieval status, thereby facilitating monitoring of the retrieval status and improving the stability of the retrieval system. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. The following drawings are not intentionally scaled to the actual size, and the focus is on illustrating the main purpose of the present application.

[0033] Figure 1 is a flow chart of the method of the present invention;

[0034] Figure 2 This is a principle block diagram of a video understanding device that integrates a multimodal knowledge graph according to the present invention. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0036] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1-2 In an embodiment of the present invention, an embodiment of a video understanding method integrating a multimodal knowledge graph includes the following steps:

[0037] Step 1: Extract video content information, extract multimodal features from the video and convert the multimodal features into corresponding feature vectors, specifically: multimodal features include visual features, audio features and text features; use video processing technology to divide the video into multiple video segments, and use visual recognition technology to extract the visual features of each video segment, and obtain the corresponding visual feature vector based on the visual features; extract audio information from the video, and convert the audio information into a spectrogram based on audio processing technology, and obtain audio features through the spectrogram, and obtain the corresponding audio feature vector based on the audio features; collect the video title, the text obtained by speech recognition, and the video text recognition information to obtain text data, and extract the text features corresponding to the text data based on natural language processing technology, and obtain the corresponding text feature vector based on the text features.

[0038] Step 2: Analyze and fuse the feature vectors corresponding to the multimodal features of each video to obtain the fusion feature; identify the multimodal features corresponding to each video to obtain the visual feature vector, audio feature vector and text feature vector, set the corresponding dimension direction of the visual feature vector, audio feature vector and text feature vector respectively, and record the dimension direction corresponding to the visual feature vector as m 1 , then the corresponding visual feature vector is [s 1 ,s2 , …, ], the dimension direction corresponding to the audio feature vector is recorded as m 2 , then the corresponding audio feature vector is [y 1 ,y 2 , …, ], the dimension direction corresponding to the text feature vector is recorded as m 3 , then the corresponding text feature vector is [z 1 , z 2 , …, ], the visual feature vector, audio feature vector and text feature vector are integrated to obtain the fusion feature, and the corresponding fusion feature vector is [s 1 ,s 2 , …, ,y 1 ,y 2 , …, , z 1 , z 2 , …, ].

[0039] Step 3: Obtain fusion features and combine the fusion features with the knowledge graph to perform video understanding to obtain video understanding information;

[0040] The fusion features are combined with the knowledge graph to perform video understanding, which is specifically as follows: a pre-understood video is obtained and the pre-understood video is identified to obtain the text to be analyzed, the text to be analyzed includes but is not limited to the video title, speech recognition text and video text recognition information; based on the pre-trained language model, the entities in the text to be analyzed are identified to obtain the corresponding text entities, the text entities include but are not limited to people, place names, objects, events; the text entities are matched with the knowledge graph to obtain the corresponding related entities in the knowledge graph, the text entities are linked with the corresponding related entities, and the corresponding rich semantic information in the knowledge graph corresponding to the related entities is obtained, and the rich semantic information corresponding to the text entities is marked as connection semantic information.

[0041] Based on the semantic information of each connection corresponding to the video, semantic reasoning is performed to obtain the reasoning semantics corresponding to the text to be analyzed in the video, the existing semantic knowledge in the database is obtained, the current reasoning semantics are matched with the existing semantic knowledge to obtain the semantic label corresponding to the video, and the semantic labels of the video are aggregated to obtain the video label information.

[0042] Obtain the corresponding semantic tags in the video tag information, identify each semantic tag to obtain the corresponding semantic features, and match the semantic features corresponding to each semantic tag with the fusion vector to obtain the corresponding vector information [ , , ], overlap and match the vector information corresponding to each semantic label to obtain the semantic repetition of each semantic label, obtain a pre-set semantic repetition threshold, and compare the semantic repetition of each semantic label with the semantic repetition threshold, mark the semantic labels corresponding to those greater than the semantic repetition threshold as high-proportion semantic labels, and mark the vector information corresponding to the high-proportion semantic labels as center vectors, and obtain a pre-set similarity vector space range, count the number of semantic labels corresponding to each center vector within the similarity vector space range, and calculate the ratio of the number of semantic labels corresponding to each center vector to the total number of semantic labels to obtain the semantic proportion of each center vector, obtain the final label according to the semantic proportion, and use the final label as the corresponding video understanding information.

[0043] Step 4: User query response analysis: obtain the user's search content information by identifying the user's query and retrieval operations, obtain the corresponding search features based on the search content information, and match the search features with the knowledge graph to obtain the corresponding search-related content.

[0044] Step 5: Obtain search data and conduct comprehensive analysis on the search data and search-related content to obtain search status information;

[0045] Comprehensive analysis of search data and search-related content is carried out in the following specific ways:

[0046] The retrieval response time, video understanding processing time and retrieval text are obtained by identifying the retrieval data; the retrieval time of each retrieval operation is obtained, the pre-set retrieval upper limit time is obtained, the retrieval time corresponding to the current retrieval operation is compared with the retrieval upper limit time, the part of the retrieval time that exceeds the retrieval upper limit time is marked as a retrieval excess value, the retrieval excess value is divided into multiple retrieval excess value intervals, each retrieval excess value interval corresponds to a retrieval time shadow value, the retrieval excess value corresponding to the current retrieval time is matched with multiple retrieval excess value intervals to obtain the corresponding retrieval time shadow value.

[0047] Get the retrieval time corresponding to the video understanding processing time of each video, calculate the ratio of the retrieval time to the video understanding processing time to get the retrieval time ratio, get the pre-set retrieval time ratio threshold, and when the retrieval time ratio is greater than the retrieval time ratio threshold, mark the corresponding retrieval time ratio as an abnormal ratio.

[0048] By identifying the retrieval-related content, the ranking position of the retrieval content and the semantics of the related content are obtained, and the retrieval text corresponding to the retrieval data is obtained. The retrieval text semantics are identified by the semantic recognition module, and the semantic similarity of the retrieval text semantics and the semantics of the related content are compared to obtain the corresponding semantic similarity, and the pre-set semantic standard similarity is obtained. If the current semantic similarity is less than the semantic standard similarity, the difference between the semantic similarity and the semantic standard similarity is calculated to obtain the similarity difference. For example, a video ranks 7th in one retrieval and 48th in the next retrieval. This large position change indicates that the retrieval state is unstable, and it is necessary to further optimize the retrieval algorithm or adjust the integration method of video understanding and retrieval.

[0049] According to the ranking position of the search content, the search content corresponding to multiple ranking positions is obtained. Based on the same search content, the search content of the ranking positions corresponding to multiple searches are obtained. The search content corresponding to the same ranking positions of multiple searches are compared to obtain the ranking change value corresponding to the search content. When the ranking change value is greater than the preset ranking change threshold, the ranking change value is marked as a ranking instability value.

[0050] The retrieval time shadow value, anomaly ratio, similarity difference and sorting instability value are normalized and their values ​​are taken, and the formula is used The retrieval status value JS is calculated; jsy, ycb, and xsc represent the retrieval time shadow value, anomaly ratio, and similarity difference, respectively. It is expressed as the sum of n ranking instability values, where n is a positive integer; m1, m2, m3, and m4 are all preset weight factors, with values ​​of 2.425, 2.722, 1.702, and 1.145, respectively; when the retrieval status value is greater than the preset threshold, the corresponding retrieval status information generated is the retrieval status abnormality.

[0051] Step 6: Perform a search self-check based on the search status information to obtain search optimization information, and issue an early warning for abnormal search status.

[0052] The present invention also provides a video understanding device integrating multimodal knowledge graphs, including a video information extraction module, a feature fusion module, a knowledge graph fusion construction module, a user query response module, and a data storage module.

[0053] The video information extraction module extracts and processes the video information through a data processor to obtain the multimodal features corresponding to each piece of video information, processes and extracts features from the text information in the video through a text processing module; processes and extracts features from the audio information in the video through an audio processing module; and processes and extracts features from the visual information in the video through a visual processing module.

[0054] The feature fusion module fuses various feature information corresponding to the multimodal features to obtain fused features.

[0055] The knowledge graph fusion construction module is responsible for associating and fusing the fused multimodal features with the pre-built multimodal knowledge graph, including entity recognition and linking units, semantic reasoning and label generation units.

[0056] The user query response module obtains the corresponding search key content by processing the user's search information, and responds with content according to the search key content.

[0057] The data storage module stores the data information generated by the video information extraction module, the feature fusion module, the knowledge graph fusion construction module and the user query response module.

[0058] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video understanding method integrating multimodal knowledge graphs, characterized in that: The steps include: Step 1: Extract multimodal features from the video and convert the multimodal features into corresponding feature vectors; Step 2: Analyze and fuse the feature vectors corresponding to the multimodal features of each video to obtain fusion features; Step 3: Obtain fusion features and combine the fusion features with the knowledge graph to perform video understanding to obtain video understanding information; The specific analysis method of combining fusion features with knowledge graph for video understanding is as follows: Obtain a pre-understood video and identify the pre-understood video to obtain a text to be analyzed; identify entities in the text to be analyzed based on a pre-trained language model to obtain corresponding text entities; match the text entity with the knowledge graph to obtain the corresponding related entity in the knowledge graph, link the text entity with the corresponding related entity, and obtain the corresponding rich semantic information in the knowledge graph corresponding to the related entity, and mark the rich semantic information corresponding to the text entity as connection semantic information; Based on the semantic information of each connection corresponding to the video, semantic reasoning is performed to obtain the reasoning semantics corresponding to the text to be analyzed corresponding to the video, the existing semantic knowledge in the database is obtained, the current reasoning semantics are matched with the existing semantic knowledge to obtain the semantic label corresponding to the video, and the semantic labels of the video are aggregated to obtain the video label information; Obtain the corresponding semantic tags in the video tag information, identify each semantic tag to obtain the corresponding semantic features, and match the semantic features corresponding to each semantic tag with the fusion vector to obtain the corresponding vector information [ , , ], overlap and match the vector information corresponding to each semantic label to obtain the semantic repetition of each semantic label, obtain a pre-set semantic repetition threshold, and compare the semantic repetition of each semantic label with the semantic repetition threshold, mark the semantic label corresponding to a value greater than the semantic repetition threshold as a high-weight semantic label, and mark the vector information corresponding to the high-weight semantic label as a center vector, and obtain a pre-set similarity vector space range, count the number of semantic labels corresponding to each center vector within the similarity vector space range, and calculate the ratio of the number of semantic labels corresponding to each center vector to the total number of semantic labels to obtain the semantic proportion of each center vector, obtain the final label according to the semantic proportion, and use the final label as the corresponding video understanding information; Step 4: User query response analysis; Step 5: Obtain search data and conduct comprehensive analysis on the search data and search-related content to obtain search status information; Step 6: Perform a search self-check based on the search status information to obtain search optimization information, and issue an early warning for abnormal search status.

2. The video understanding method integrating multimodal knowledge graph according to claim 1, characterized in that: The method of extracting multimodal features from a video and converting the multimodal features into corresponding feature vectors is specifically as follows: the multimodal features include visual features, audio features and text features; the video is divided into a plurality of video segments by using video processing technology, and the visual features of each video segment are extracted by using visual recognition technology, and the corresponding visual feature vector is obtained according to the visual features; the audio information in the video is extracted, and the audio information is converted into a spectrogram based on the audio processing technology, and the audio features are obtained by using the spectrogram, and the corresponding audio feature vector is obtained according to the audio features; The video title, the text obtained by speech recognition, and the video text recognition information are combined to obtain text data, and the text features corresponding to the text data are extracted based on natural language processing technology, and the corresponding text feature vector is obtained based on the text features.

3. The video understanding method integrating multimodal knowledge graph according to claim 1, characterized in that: The feature vectors corresponding to the multimodal features of each video are analyzed and fused to obtain fusion features, and the specific analysis method is as follows: The multimodal features corresponding to each video are identified to obtain visual feature vectors, audio feature vectors and text feature vectors. The visual feature vectors, audio feature vectors and text feature vectors are respectively set to corresponding dimensional directions. The dimensional direction corresponding to the visual feature vector is recorded as m1, and the corresponding visual feature vector is [s1, s2, …, ], and the dimension direction corresponding to the audio feature vector is recorded as m2, then the corresponding audio feature vector is [y1, y2, …, ], and the dimension direction corresponding to the text feature vector is recorded as m3, then the corresponding text feature vector is [z1, z2,…, ], the visual feature vector, audio feature vector and text feature vector are integrated to obtain the fusion feature, and the corresponding fusion feature vector is [s1, s2, …, , y1, y2, …, , z1, z2, …, ].

4. The video understanding method integrating multimodal knowledge graph according to claim 1, characterized in that: The user query response analysis obtains the user's search content information by identifying the user's query retrieval operation, obtains corresponding search features based on the search content information, and matches the search features with the knowledge graph to obtain corresponding search-related content.

5. The video understanding method integrating multimodal knowledge graph according to claim 1, characterized in that: The comprehensive analysis of the search data and the search-related contents is specifically conducted in the following manner: The search response time, video understanding processing time and search text are obtained by identifying the search data; the search time of each search operation is obtained, the search upper limit time which is set in advance is obtained, the search time corresponding to the current search operation is compared with the search upper limit time, the part of the search time which exceeds the search upper limit time is marked as the search excess value, the search excess value is divided into multiple search excess value intervals, each search excess value interval corresponds to a search time shadow value, the search excess value corresponding to the current search time is matched with multiple search excess value intervals to obtain the corresponding search time shadow value; Get the retrieval time corresponding to the video understanding processing time of each video, calculate the ratio of the retrieval time to the video understanding processing time to get the retrieval time ratio, get the pre-set retrieval time ratio threshold, and when the retrieval time ratio is greater than the retrieval time ratio threshold, mark the corresponding retrieval time ratio as an abnormal ratio.

6. The video understanding method integrating multimodal knowledge graph according to claim 5, characterized in that: The retrieval of related content includes the ranking position of the retrieval content and the semantics of the related content, obtaining a retrieval text corresponding to the retrieval data, identifying the retrieval text semantics through a semantic recognition module, performing a semantic similarity evaluation on the retrieval text semantics and the related content semantics to obtain a corresponding semantic similarity, obtaining a pre-set semantic standard similarity, and if the current semantic similarity is less than the semantic standard similarity, performing a difference calculation on the semantic similarity and the semantic standard similarity to obtain a similarity difference; According to the ranking position of the search content, the search content corresponding to multiple ranking positions is obtained. Based on the same search content, the search content of the ranking positions corresponding to multiple searches are obtained. The search content corresponding to the same ranking positions of multiple searches are compared to obtain the ranking change value corresponding to the search content. When the ranking change value is greater than the preset ranking change threshold, the ranking change value is marked as a ranking instability value.

7. The video understanding method integrating multimodal knowledge graph according to claim 6, characterized in that: The retrieval time shadow value, anomaly ratio, similarity difference and sorting instability value are normalized and their values ​​are taken, and the formula is used The retrieval status value JS is calculated; jsy, ycb, and xsc represent the retrieval time shadow value, anomaly ratio, and similarity difference, respectively. It is expressed as the sum of n ranking instability values, where n is a positive integer; m1, m2, m3, and m4 are all preset weight factors; when the retrieval state value is greater than the preset threshold, the corresponding retrieval state information generated is the retrieval state abnormality.

8. A video understanding device integrating multimodal knowledge graph, characterized in that A video understanding method integrating a multimodal knowledge graph as described in any one of claims 1 to 7 comprises a video information extraction module, a feature fusion module, a knowledge graph fusion construction module, a user query response module, and a data storage module.

9. The video understanding device integrating multimodal knowledge graph according to claim 8, characterized in that: The video information extraction module is used to extract and process the video information through a data processor to obtain multimodal features corresponding to each video information, process and extract features of text information in the video through a text processing module; process and extract features of audio information in the video through an audio processing module; and process and extract features of visual information in the video through a visual processing module; The feature fusion module is used to fuse various feature information corresponding to the multimodal features to obtain fused features; The knowledge graph fusion construction module is responsible for associating and fusing the fused multimodal features with the pre-built multimodal knowledge graph, including entity recognition and linking units, semantic reasoning and label generation units; The user query response module is used to obtain the corresponding search key content by processing the user's search information, and to respond to the content according to the search key content; The data storage module is used to store the data information generated by the video information extraction module, the feature fusion module, the knowledge graph fusion construction module and the user query response module.

Citation Information

Patent Citations

  • Information retrieval method and device

    CN117033657A

  • Data storage method and system based on network teaching resources

    CN117235187A