Video fact and viewpoint alignment traceability method

Through multimodal information fusion and systematic verification framework, and using computer vision, speech recognition and natural language processing technologies, we solved the problems of semantic association and temporal consistency between facts and opinions in videos, and achieved automated and transparent verification of the authenticity of video content.

CN120747809AInactive Publication Date: 2025-10-03CHONGQING QINGZHI NET EAGLE TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510823171.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies lack multimodal information fusion and systematic verification frameworks, making it difficult to accurately distinguish facts and opinions in videos and establish their semantic associations and temporal consistency, resulting in difficulties in assessing the authenticity of video content.

Method used

Through computer vision, speech recognition and natural language processing technologies, the scenes, objects, dynamic features and background information of the video are extracted. Combined with event extraction technology, the semantic similarity between opinions and event facts is calculated, and time series data alignment technology is used for verification to generate a transparent verification report.

Benefits of technology

It achieves semantic association modeling of subjective opinions and objective facts in video content, ensures the verifiability of events in time and content, improves the accuracy and credibility of information traceability, and realizes the automation and transparency of video content authenticity judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747809A_ABST
    Figure CN120747809A_ABST
Patent Text Reader

Abstract

The invention discloses a video fact and viewpoint alignment traceability method, which relates to the technical field of information retrieval and verification, and comprises the following steps of: performing frame analysis on a video by using a computer vision technology, extracting scenes, objects, dynamic characteristics and background information in the video, extracting audios from the video by using a voice recognition technology, converting the audios into text information, and storing the text information in a database; utilizing an event extraction technology to extract event elements from the video; the event elements are identified through scenes, objects, dynamic features and background information, and event facts are obtained; and based on the text information, using a natural language processing technology to calculate semantic similarity between the viewpoints and the event facts in the video, and automatically aligning the viewpoints and the event facts in the video according to the semantic similarity. According to the method, high-dimensional semantic modeling is carried out on the text, the similarity is calculated, the internal relation between viewpoints and facts can be accurately recognized, and therefore automatic matching of the viewpoints and the facts is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information retrieval and verification, and in particular to a method for aligning and tracing the origin of video facts and opinions. Background Art

[0002] With the rapid development of social media and short video platforms, the amount of user-generated content has grown exponentially, and the speed and breadth of information dissemination far exceed any previous period. How can we quickly and accurately verify whether the opinions expressed in videos are consistent with the facts stated to ensure the authenticity and reliability of the information? Research on automatic understanding and authenticity verification of video content has gradually deepened. Some technologies have attempted to combine multimodal analysis methods such as computer vision, speech recognition, and natural language processing to perform preliminary identification and judgment of events in videos. Existing systems extract image features from video frames to identify visual information such as scenes and objects, and combine speech recognition technology to convert audio into text, thereby achieving a structured representation of video content. However, most of these methods remain at the level of event element recognition and lack in-depth modeling and consistency verification mechanisms for the semantic relationship between opinions and facts, making it difficult to meet the needs of accurate assessment of information authenticity in practical applications.

[0003] Although existing technologies have made certain progress in video content analysis, they still have obvious limitations when dealing with complex and changeable user-generated content. Most methods focus on the processing of unimodal information, such as relying solely on text or images for event extraction, ignoring the synergy between multimodal information. When faced with highly subjective opinions in videos, they are often unable to effectively distinguish between factual statements and subjective judgments, and it is even more difficult to establish the correlation and consistency between opinions and historical events. Existing technologies generally lack a systematic framework that can organically integrate multiple links such as event element identification, timeline verification, semantic similarity calculation, and opinion tracing, so as to achieve a comprehensive assessment of the authenticity of video content. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a method for aligning and tracing video facts and opinions to solve the technical problems in the prior art that are difficult to accurately distinguish facts and opinions and establish their semantic association and temporal consistency due to the lack of multimodal information fusion and systematic verification framework.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a method for aligning and tracing the origins of facts and opinions in a video, which includes using computer vision technology to perform frame analysis on the video to extract scenes, objects, dynamic features, and background information in the video; using speech recognition technology to extract audio from the video, converting the audio into text information; and using event extraction technology to extract event elements from the video.

[0008] Identify event elements through scenes, objects, dynamic features and background information to obtain event facts;

[0009] Based on text information, natural language processing technology is used to calculate the semantic similarity between the opinions in the video and the event facts, and the opinions and event facts in the video are automatically aligned according to the semantic similarity;

[0010] Compare event elements with the historical event database, verify event elements, and use time series data alignment technology to verify the accuracy of the timeline to obtain event extraction results;

[0011] Based on the event extraction results and the semantic similarity matching results, the source of the opinions in the video is obtained, the opinions in the video are verified, the authenticity of the video content is automatically marked, and transparent verification results are provided to users.

[0012] As a preferred solution of the method for tracing the alignment of video facts and opinions described in the present invention, the method uses computer vision technology to perform frame analysis on the video to extract scenes, objects, dynamic features and background information in the video, including the following steps:

[0013] Extract image data frame by frame from the video, generate static image frames, and classify the scene type to which each static image frame belongs;

[0014] Apply object detection algorithms to identify objects in each frame, and use object tracking technology to track the position changes of objects across frames to form dynamic trajectories;

[0015] According to the position change of the object between multiple frames, dynamic features are obtained;

[0016] Combine scene types, dynamic trajectories, and dynamic features to analyze background information in the video.

[0017] As a preferred solution of the method for tracing the alignment of video facts and opinions described in the present invention, wherein: using speech recognition technology to extract audio from the video and converting the audio into text information, the following steps are included:

[0018] Separate and extract the audio stream from the video file to generate independent audio, and perform noise reduction and volume standardization on the independent audio;

[0019] The independent audio after noise reduction and volume standardization is segmented, and the segmented audio is converted into text information using speech recognition technology.

[0020] As a preferred solution of the method for tracing the alignment of video facts and opinions of the present invention, event elements are extracted from the video using event extraction technology, including the following steps:

[0021] Use event extraction technology to analyze text information, mark time segments, and identify and classify them;

[0022] The entities identified and classified are analyzed to obtain event elements.

[0023] As a preferred solution of the method for aligning and tracing the source of video facts and opinions described in the present invention, the event elements are identified through scenes, objects, dynamic features and background information to obtain event facts, including the following steps:

[0024] Analyze scenes, objects, dynamic features, and background information and identify core event features;

[0025] The core event features are compared and verified with the historical event database to obtain the event facts.

[0026] As a preferred solution of the method for aligning and tracing the source of video facts and opinions described in the present invention, the method includes: calculating the semantic similarity between opinions and event facts in the video based on text information using natural language processing technology, and automatically aligning opinions and event facts in the video according to the semantic similarity, including the following steps:

[0027] Clean and standardize text information, and use a pre-trained language model to convert the cleaned and standardized text into a high-dimensional vector to capture the deep semantic features of the high-dimensional vector.

[0028] Calculate the semantic similarity between the opinions expressed in the video and the identified event facts based on semantic features;

[0029] Sort and pair the opinions and event facts in the video according to semantic similarity and perform automatic alignment.

[0030] As a preferred solution of the method for aligning and tracing the source of video facts and opinions described in the present invention, the event elements are compared with the historical event database, the event elements are verified, and the accuracy of the timeline is verified by combining the time series data alignment technology to obtain the event extraction result, including the following steps:

[0031] Generate a data structure of event elements based on text information, scenes, objects, dynamic features, background information and event elements extracted from the video;

[0032] The data structure of event elements is accurately matched with the historical event database, the Jaccard similarity score is calculated, and matching events are screened out.

[0033] The DTW algorithm is used to compare the time series of matching events with the time series of historical events, adjust and calibrate the time sequence of events, and generate a time accuracy score.

[0034] As a preferred solution of the method for aligning and tracing the source of video facts and opinions described in the present invention, the following steps are included: according to the event extraction results and the results of semantic similarity matching, the source of the opinions in the video is obtained, the opinions in the video are verified, the authenticity of the video content is automatically marked, and transparent verification results are provided to users.

[0035] Combine the event extraction results with the semantic similarity matching results to form a comprehensive dataset;

[0036] Based on the comprehensive dataset, natural language processing technology is used to identify the opinions and sources mentioned in the video, and the background information corresponding to the video opinions is marked;

[0037] Use the historical event database and the background information corresponding to the video opinions for comparison and verification, and generate a score for opinion verification.

[0038] Based on the opinion verification score, the authenticity of the video content is automatically marked and compiled into a detailed transparent verification report.

[0039] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the method for aligning and tracing the origin of video facts and opinions as described in the first aspect of the present invention.

[0040] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the method for aligning and tracing the origin of video facts and opinions as described in the first aspect of the present invention.

[0041] The beneficial effects of the present invention are as follows: by using natural language processing technology based on text information to calculate the semantic similarity between opinions and event facts in the video, automatically aligning opinions and event facts in the video according to the semantic similarity, and comparing event elements with the historical event database, verifying the event elements, and combining with the time series data alignment technology to verify the accuracy of the timeline, and obtaining event extraction results, it realizes the modeling of the semantic relationship between subjective opinions and objective facts in the video content, and effectively verifies the rationality of the event in the time dimension. By performing high-dimensional semantic modeling on the text and calculating the similarity, it is possible to accurately identify the intrinsic connection between opinions and facts, thereby realizing automatic matching between the two. The DTW time series alignment algorithm is introduced, combined with structured event element data, to ensure that events are verifiable in both time and content. A multi-dimensional verification mechanism is constructed from the semantic level and the time level respectively, which realizes the automation and transparency of the authenticity judgment of the video content and improves the accuracy and credibility of information traceability. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 Flowchart of the method for aligning and tracing the provenance of video facts and opinions.

[0044] Figure 2 A schematic diagram of the analysis background information.

[0045] Figure 3 Flowchart for converting audio segments into text information.

[0046] Figure 4 Flowchart showing the authenticity of event elements. DETAILED DESCRIPTION

[0047] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0048] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0049] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0050] Reference Figures 1 to 4 , is an embodiment of the present invention, which provides a method for tracing the alignment of video facts and opinions, including the following steps:

[0051] S1. Use computer vision technology to analyze video frames and extract scenes, objects, dynamic features and background information from the video.

[0052] S1.1. Extract image data from the video frame by frame to generate static image frames, and classify the scene type to which each static image frame belongs.

[0053] Furthermore, the original video is decomposed according to a set time interval or frame rate, and a continuous image sequence is obtained frame by frame. Each frame is saved as an independent static image in a common image format. A deep learning-based image classification method is applied to each static image frame, and a pre-trained convolutional neural network is used to extract features of the image content. The probability distribution corresponding to the predefined scene category is output through the fully connected layer; wherein the predefined scene categories include but are not limited to common scene types such as indoor, outdoor, city streets, and natural landscapes.

[0054] S1.2. Apply object detection algorithms to identify objects in each frame, and use object tracking technology to track the position changes of objects across frames to form dynamic trajectories.

[0055] Furthermore, after extracting image data frame by frame from the video and classifying the scene type of each static image frame, a region-based convolutional neural network is used to detect objects in the image for each static image frame, and the bounding box coordinates and category labels of each object are output. A multi-target tracking method based on the Hungarian algorithm is applied between consecutive frames. By calculating the similarity and position overlap of the appearance features of the object detected in the current frame and the object in the previous frame, the matching of objects between different frames is achieved. In this process, a Kalman filter is used to predict and update the motion state of the object to improve the stability of cross-frame tracking, and the dynamic trajectory of the object is generated based on the position information of the object in each frame and its matching results in the time series.

[0056] S1.3. According to the position change of the object between multiple frames, dynamic features are obtained.

[0057] Furthermore, after identifying the object in each frame through the object detection algorithm and tracking the position change of the object across frames using the object tracking technology to form a dynamic trajectory, the center coordinates of the bounding box of the object in each frame and its width and height are extracted to construct the spatial position sequence of the object in consecutive frames; then, a differential operation is performed on the position sequence, and the horizontal and vertical displacements and size change rates of the object between adjacent frames are combined with the time interval to derive the motion speed and acceleration information of the object. The above displacement, size change, speed and acceleration information are combined into a set of numerical values ​​that characterize the motion state of the object, which serves as the dynamic characteristics of the object.

[0058] S1.4. Combine scene type, dynamic trajectory, and dynamic features to analyze background information in the video.

[0059] Furthermore, on the basis of completing the scene type classification, object dynamic trajectory generation and dynamic feature extraction of static image frames, the scene type of each frame is first used as environmental context information, and spatially aligned with the dynamic trajectories and dynamic features of all objects in the frame. Based on the alignment results, the typical motion patterns of objects under the same scene type are counted, including but not limited to motion direction, speed distribution, size change trend and other features. Combining the spatial attributes of the scene type with the temporal characteristics of the object movement, the behavioral semantics closely related to the scene are identified. By integrating the multi-dimensional information of scene type, object dynamic trajectory and dynamic features, background information for describing the overall environmental characteristics of the video content is constructed.

[0060] S2. Use speech recognition technology to extract audio from the video and convert the audio into text information

[0061] S2.1. Separate and extract the audio stream from the video file, generate independent audio, and perform noise reduction and volume standardization on the independent audio.

[0062] Furthermore, an audio and video encapsulation format parsing tool is used to decapsulate the original video file, extract the original audio data encapsulated in the video container, and obtain the original audio stream that has not been processed in any way. The original audio stream is converted into an audio file in a unified format, and an audio noise reduction method based on spectral subtraction is applied to the audio signal for noise reduction to eliminate the possible impact of background noise on speech recognition. The peak normalization method is used to standardize the volume of the noise-reduced audio, and the peak amplitude of the audio is adjusted to within a preset range. The independent audio file after noise reduction and standardization is output for subsequent speech recognition and text information extraction.

[0063] S2.2. Segment the independent audio after noise reduction and volume normalization processing, and use speech recognition technology to convert the segmented audio into text information.

[0064] Furthermore, an endpoint detection method based on energy threshold detection is used to analyze the audio after noise reduction and volume standardization processing, identify the voice activity area, and accordingly divide the continuous audio into multiple voice segments, each segment containing a continuous voice content. For each voice segment, automatic speech recognition combining a pre-trained hidden Markov model and a deep neural network is applied to extract the Mel-frequency cepstral coefficient features of the audio, and input them into the speech recognition model to generate the corresponding text sequence. Furthermore, the generated text sequence is reordered by the language model to improve the language fluency and semantic accuracy of the recognition results, and output the text information converted from each voice segment.

[0065] S3. Use event extraction technology to extract event elements from the video.

[0066] S3.1. Use event extraction technology to analyze text information, mark time segments, and identify and classify them.

[0067] Furthermore, the text information obtained by speech recognition conversion is preprocessed to remove irrelevant symbols and stop words, and is segmented according to sentences or semantic units. A rule-based time expression recognition method is applied, combined with regular matching and time dictionaries, to identify paragraphs containing time information in the text and mark their positions in the text. The marked time paragraphs are input into a sequence labeling model based on conditional random fields, the time expression is converted into a standardized format, and classified according to a predefined time category system. The time categories include absolute time, relative time, and fuzzy time, and structured time information with time labels and classification results is output.

[0068] S3.2. Analyze the entities identified and classified to obtain event elements.

[0069] Furthermore, a named entity recognition method based on conditional random fields is adopted, combined with a predefined entity dictionary, to identify various entities in the text and their contextual semantic relationships. Through dependency syntax analysis technology, the action association information between entities is extracted, and the basic semantic framework of the event is constructed. Combined with the identified time category information and entity semantic relationships, the key attributes of the event, such as the subject, object, time and place, are determined, and classified and integrated according to the preset event element structured template, and structured data containing complete event elements is output.

[0070] S4. Identify event elements through scenes, objects, dynamic features and background information to obtain event facts.

[0071] S4.1. Analyze scenes, objects, dynamic features, and background information and identify core event features.

[0072] Furthermore, the obtained scene type results are spatially matched with the object category information obtained from object detection to determine the semantic role of each object in a specific scene. Combined with the dynamic characteristics of the object, including displacement, movement direction, speed change and other information, the behavior pattern of the object and its interaction with other objects are analyzed. The above analysis results are integrated with the environmental context described in the background information. The background information includes the spatial attributes, typical behavior patterns and temporal characteristics of the scene. Based on the predefined core event feature template, features that can represent the key events of the video content are extracted from the multi-dimensional information.

[0073] S4.2. Compare and verify the core event features with the historical event database to obtain the event facts.

[0074] Furthermore, historical event records that have potential relevance to the current video content in terms of time, location, and event type are retrieved from the historical event database. The extracted core event features, including the event subject, behavioral actions, objects of action, occurrence environment, and time clues, are matched item by item with the corresponding features in the retrieved historical event records. Historical events with a higher matching degree are screened out based on the similarity threshold and output as event facts as the real events described in the current video content.

[0075] S5. Based on text information, natural language processing technology is used to calculate the semantic similarity between the opinions and event facts in the video, and the opinions and event facts in the video are automatically aligned according to the semantic similarity.

[0076] S5.1. Clean and standardize text information, and use a pre-trained language model to convert the cleaned and standardized text into a high-dimensional vector to capture the deep semantic features of the high-dimensional vector.

[0077] Furthermore, the original text information extracted from the video audio is segmented and tagged with parts of speech, and then meaningless symbols, stop words and repeated content are removed, and synonyms and abbreviations are uniformly replaced to achieve formal standardization of the text content and generate a tag sequence. Through a multi-layer self-attention mechanism, the context-related representation of the text is extracted step by step, and the deep semantic features of the high-dimensional vector are output.

[0078] S5.2. Calculate the semantic similarity between the opinions expressed in the video and the identified event facts based on semantic features.

[0079] Furthermore, the opinion text information extracted from the video and the verified event fact description text were cleaned and standardized respectively to remove irrelevant symbols, stop words and redundant expressions; then, the processed text was encoded using a pre-trained natural language processing model, and each piece of text was mapped into a high-dimensional vector representation of a fixed dimension. The high-dimensional vector can capture the deep semantic features of the text, and the cosine similarity calculation method was used to evaluate the similarity of the high-dimensional vectors corresponding to the opinion text and the event fact text to generate a numerical semantic similarity.

[0080] Specifically, the expression is,

[0081]

[0082] Among them, S(V,E) is the semantic similarity, E vec For the facts of the event, V vec For video views,

[0083] S5.3. Sort and pair the opinions and event facts in the video according to semantic similarity and perform automatic alignment.

[0084] Furthermore, the semantic similarity scores between multiple opinion texts extracted from the video and multiple verified event fact description texts are combined to form a one-to-many matching matrix, in which each row corresponds to the similarity calculation result between an opinion text and all event fact texts; then, a normalization processing method is applied to the semantic similarity scores in each row to make the scores comparable under a unified scale; further, the event fact candidate set corresponding to each opinion text is arranged in descending order according to the normalized semantic similarity score, and a threshold is set to screen out event facts with a similarity higher than a preset standard as potential matching objects, and a one-to-one correspondence between opinion and event facts is established based on the matching results, thereby realizing automatic pairing and semantic alignment of opinion and event facts in the video.

[0085] S6. Compare the event elements with the historical event database, verify the event elements, and combine the time series data alignment technology to verify the accuracy of the timeline to obtain the event extraction results.

[0086] S6.1. Generate a data structure of event elements based on text information, scenes, objects, dynamic features, background information, and event elements extracted from the video.

[0087] Furthermore, the text information extracted from the video is cleaned and standardized, and then semantically aligned with the identified event element content to ensure accurate correspondence between key information such as the event subject, time, location, and behavior in the text. The object category and its spatiotemporal position information obtained through object detection and tracking, the dynamic features and scene classification results obtained through multi-frame analysis are integrated as the visual context information of the event, and the visual information is associated with the environmental attributes described in the background information. The background information comes from a comprehensive analysis of the scene type, dynamic trajectory, and dynamic features. Finally, according to the predefined structured field format, the event's timestamp, subject, object, behavior, location, and contextual features are organized into a unified format of event element data structure.

[0088] S6.2 accurately matches the data structure of event elements with the historical event database, calculates the Jaccard similarity score, and filters out matching events.

[0089] Furthermore, all stored historical event records are selected from the historical event database, and each historical event record is formatted according to the same data structure as the event element generated in the current video. The key fields between each historical event record and the current video event element are compared item by item. The key fields include time range, event subject, behavior, object and place of occurrence. The Jaccard similarity coefficient is calculated based on the intersection and union of the above fields, and a numerical Jaccard similarity score is generated to measure the degree of matching between the current video event and the historical event at the structured information level. The historical events are filtered according to the set similarity threshold, and the historical events with a Jaccard similarity score higher than the threshold are retained as candidate matching events.

[0090] Specifically, the expression is,

[0091]

[0092] J(A,B) is a matching event, A is a set of event elements, and B is a set of records in the historical event database.

[0093] S6.3 uses the DTW algorithm to compare the time series of matching events with the time series of historical events, adjust and calibrate the time sequence of events, and generate a time accuracy score.

[0094] Furthermore, the time series is composed of the distribution of event elements identified in the video in the time dimension; at the same time, the time series data of the corresponding historical events in the historical event database are obtained. The time series contains the occurrence time points of each sub-event in the historical event. The two time series are passed as input to the dynamic time warping (DTW) algorithm. By constructing a nonlinear alignment path between the time points, the cumulative distance difference between the two sequences is minimized. The normalized DTW distance value is calculated based on the optimal alignment path and converted into a time accuracy score. The score reflects the degree of consistency between the video event and the historical event in the time sequence. The alignment result with the time accuracy score is output for the comprehensive evaluation of the authenticity of the subsequent event elements.

[0095] Specifically, the expression is,

[0096]

[0097] Where T(X,Y) is the time accuracy score, n is the length of the time series, X i is the value of the i-th time point of the time series composed of event elements, Y i is the value of the i-th time point in the time series of the corresponding event in the historical event database, and i is the time series index.

[0098] S6.4 scores the time accuracy and matches the events to evaluate the authenticity of the event elements and obtain the event extraction results.

[0099] Furthermore, a linear combination method is adopted through weighted fusion, in which the weight coefficient is set according to the relative importance of time consistency and structural similarity in event verification; then, the matching events are sorted based on the comprehensive score after fusion, and a threshold is set to filter out events with scores higher than the preset standard as high-confidence matching results. Combined with the key fields such as the subject, behavior, and location contained in the data structure of the event elements, consistency verification is performed with the corresponding fields in the historical event database to identify information units with deviations or conflicts. According to the comprehensive score and field consistency analysis results, the structured event extraction results are output.

[0100] S7. Based on the event extraction results and the semantic similarity matching results, the source of the opinions in the video is obtained, the opinions in the video are verified, the authenticity of the video content is automatically marked, and transparent verification results are provided to users.

[0101] S7.1. Combine the event extraction results with the semantic similarity matching results to form a comprehensive dataset.

[0102] Furthermore, structured event element information is extracted from the event extraction results. The event element information includes fields such as event subject, behavior, occurrence time, object and location, and the semantic similarity matching results between the opinions and event facts in the video are obtained. The semantic similarity matching results include the similarity scores and alignment relationships between the opinion text and the event facts. With the event subject and occurrence time as the association keys, the event element information and the semantic similarity matching results are mapped one-to-one and merged into a unified data entry under the same event identifier. The fields of the merged data entries are expanded, and source identifiers, matching confidence levels and consistency status labels are added to generate a comprehensive data set with multi-dimensional attributes.

[0103] S7.2. Based on the comprehensive dataset, use natural language processing technology to identify the opinions and sources mentioned in the video and mark the background information corresponding to the video opinions.

[0104] Furthermore, the video text information contained in the comprehensive dataset is subjected to syntactic structure analysis and dependency parsing to identify sentence fragments expressing subjective judgments or position statements, and marker words, sentiment words and position verbs are matched according to the predefined opinion expression pattern library, so as to locate the text units in the video that clearly express opinions. Combined with the event subject and time clues in the event extraction results, a rule-based reference resolution method is used to identify the opinion holders and the sources of their claims. The sources of opinions include direct statements, quotations from others or retelling of historical events. The identified opinion text is semantically bound to its associated background information. The background information comes from multimodal analysis results such as scene types, object dynamic trajectories and event occurrence environments extracted from the video. In the comprehensive dataset, source identification and background information labels are added to each opinion information to mark the background information corresponding to the video opinion.

[0105] S7.3. Use the historical event database and the background information corresponding to the video viewpoints to perform comparison and verification, and generate a viewpoint verification score.

[0106] Furthermore, historical event records that have potential relevance to the video viewpoint in terms of time, place and event subject are retrieved from the historical event database, and key descriptive information in the historical event records is extracted. The key descriptive information includes the background of the event, the subject's behavior pattern and the environmental context characteristics; then, the background information corresponding to the video viewpoint is compared item by item with the retrieved historical event background information. The background information corresponding to the video viewpoint comes from the multimodal analysis results such as the scene type, object dynamic trajectory and event environment extracted from the video. A matching evaluation method based on the combination of word overlap rate and semantic vector similarity is adopted to calculate the structural consistency score and semantic matching score between the video viewpoint background and the historical event background, and the above two scores are weighted fused to generate a viewpoint verification score.

[0107] Specifically, the expression is,

[0108] P(V,H)=w1×S(V,E)+w2×T(X,Y);

[0109] Among them, P(V,H) is the score of opinion verification, w1 is the weight of semantic similarity, and w2 is the weight of time accuracy score.

[0110] S7.4. Based on the opinion verification score, the authenticity of the video content is automatically marked and compiled into a detailed transparent verification report.

[0111] Furthermore, a grading threshold is set based on the score of opinion verification, and the opinions in the video are divided into three levels: high credibility, medium credibility and low credibility. The confidence status corresponding to each opinion is marked on the timeline of the original video. Combined with the time accuracy score, structured event elements and semantic similarity matching results in the event extraction results, multi-dimensional verification information related to each opinion is generated. The multi-dimensional verification information covers the source of the opinion, background consistency, historical event matching and time logic rationality. The above verification information is structured according to a predefined report template to generate a transparent verification report entry containing opinion description, verification basis, scoring results and conclusion explanation, and the verification results of all opinions are integrated into a complete transparent verification report.

[0112] This embodiment also provides a computer device, which is suitable for the method of tracing the alignment of video facts and opinions, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the method of tracing the alignment of video facts and opinions proposed in the above embodiment.

[0113] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.

[0114] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for aligning and tracing video facts and opinions as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk.

[0115] In summary, the present invention calculates the semantic similarity between opinions and event facts in the video based on text information using natural language processing technology, automatically aligns opinions and event facts in the video according to the semantic similarity, compares event elements with the historical event database, verifies event elements, and combines time series data alignment technology to verify the accuracy of the timeline to obtain event extraction results, thereby realizing the modeling of semantic associations between subjective opinions and objective facts in video content, and effectively verifying the rationality of events in the time dimension. By performing high-dimensional semantic modeling on the text and calculating the similarity, the intrinsic connection between opinions and facts can be accurately identified, thereby realizing automatic matching between the two. The DTW time series alignment algorithm is introduced and combined with structured event element data to ensure that events are verifiable in both time and content. A multi-dimensional verification mechanism is constructed from the semantic level and the time level, respectively, which realizes the automation and transparency of video content authenticity judgment and improves the accuracy and credibility of information traceability.

[0116] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for aligning and tracing the source of video facts and opinions, characterized by: include, Use computer vision technology to analyze video frames, extract scenes, objects, dynamic features, and background information from the video, use speech recognition technology to extract audio from the video, convert the audio into text information, and use event extraction technology to extract event elements from the video; Identify event elements through scenes, objects, dynamic features and background information to obtain event facts; Based on text information, natural language processing technology is used to calculate the semantic similarity between the opinions in the video and the event facts, and the opinions and event facts in the video are automatically aligned according to the semantic similarity; Compare event elements with the historical event database, verify event elements, and use time series data alignment technology to verify the accuracy of the timeline to obtain event extraction results; Based on the event extraction results and the semantic similarity matching results, the source of the opinions in the video is obtained, the opinions in the video are verified, the authenticity of the video content is automatically marked, and transparent verification results are provided to users.

2. The method for tracing the alignment of video facts and opinions according to claim 1, characterized in that: Use computer vision technology to analyze video frames and extract scenes, objects, dynamic features and background information from the video. The following steps are included: Extract image data frame by frame from the video, generate static image frames, and classify the scene type to which each static image frame belongs; Apply object detection algorithms to identify objects in each frame, and use object tracking technology to track the position changes of objects across frames to form dynamic trajectories; According to the position change of the object between multiple frames, dynamic features are obtained; Combine scene types, dynamic trajectories, and dynamic features to analyze background information in the video.

3. The method for aligning video facts and opinions according to claim 2, wherein: Using speech recognition technology to extract audio from video and convert the audio into text information includes the following steps: Separate and extract the audio stream from the video file to generate independent audio, and perform noise reduction and volume standardization on the independent audio; The independent audio after noise reduction and volume standardization is segmented, and the segmented audio is converted into text information using speech recognition technology.

4. The method for aligning and tracing video facts and opinions according to claim 3, wherein: Using event extraction technology, event elements are extracted from the video. The following steps are included: Use event extraction technology to analyze text information, mark time segments, and identify and classify them; The entities identified and classified are analyzed to obtain event elements.

5. The method for tracing the alignment of video facts and opinions according to claim 4, characterized in that: Identify event elements through scenes, objects, dynamic features and background information to obtain event facts, including the following steps: Analyze scenes, objects, dynamic features, and background information and identify core event features; The core event features are compared and verified with the historical event database to obtain the event facts.

6. The method for tracing the alignment of video facts and opinions according to claim 5, characterized in that: Based on text information, natural language processing technology is used to calculate the semantic similarity between the opinions in the video and the event facts, and the opinions and event facts in the video are automatically aligned according to the semantic similarity, including the following steps: Clean and standardize text information, and use a pre-trained language model to convert the cleaned and standardized text into a high-dimensional vector to capture the deep semantic features of the high-dimensional vector. Calculate the semantic similarity between the opinions expressed in the video and the identified event facts based on semantic features; Sort and pair the opinions and event facts in the video according to semantic similarity and perform automatic alignment.

7. The method for tracing the alignment of video facts and opinions according to claim 6, characterized in that: Compare the event elements with the historical event database, verify the event elements, and combine the time series data alignment technology to verify the accuracy of the timeline to obtain the event extraction results. The following steps are included: Generate a data structure of event elements based on text information, scenes, objects, dynamic features, background information and event elements extracted from the video; Accurately match the data structure of event elements with the historical event database, calculate the Jaccard similarity score, and filter out matching events; Use the DTW algorithm to compare the time series of matching events with the time series of historical events, adjust and calibrate the time sequence of events, and generate a time accuracy score; The authenticity of event elements is evaluated by scoring the time accuracy and matching events to obtain the event extraction results.

8. The method for tracing the alignment of video facts and opinions according to claim 7, wherein: According to the event extraction results and the semantic similarity matching results, the source of the opinions in the video is obtained, the opinions in the video are verified, the authenticity of the video content is automatically marked, and transparent verification results are provided to users, including the following steps. Combine the event extraction results with the semantic similarity matching results to form a comprehensive dataset; Based on the comprehensive dataset, natural language processing technology is used to identify the opinions and sources mentioned in the video, and the background information corresponding to the video opinions is marked; Use the historical event database and the background information corresponding to the video opinions to compare and verify, and generate a score for opinion verification; Based on the opinion verification score, the authenticity of the video content is automatically marked and compiled into a detailed transparent verification report.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for aligning and tracing the origin of video facts and opinions described in any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for aligning and tracing the origin of video facts and opinions described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Smoke identification and classification early warning method, system and device and storage medium

    CN121302295A

  • Project whole-process data tracing method and system based on video semantic analysis

    CN121681870A

  • A project whole-process data tracing method and system based on video semantic analysis

    CN121681870B