A method for evaluating video and text similarity

By performing semantic hierarchical analysis and narrative element comparison on videos and texts, a multidimensional semantic dataset is generated and weighted and fused, which solves the problems of low efficiency and low accuracy in video and text similarity assessment in existing technologies, and realizes efficient assessment of video and text in terms of narrative logic and content presentation.

CN121167332BActive Publication Date: 2026-03-24BEIJING BANGCLE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing video analysis methods mainly rely on visual feature detection or object recognition, which cannot accurately reflect the semantic coherence of video and text in terms of narrative structure, resulting in low efficiency and low accuracy in video-text similarity assessment.

Method used

By performing semantic level analysis on the novel text and video to be tested, a multidimensional semantic dataset is generated. Narrative elements are compared, and narrative similarity calculation indicators are used to weightedly integrate with content coverage to achieve a comprehensive evaluation of the video and text in terms of narrative logic and content presentation.

Benefits of technology

It achieves a unified semantic representation of video and text, improves the objectivity and reproducibility of evaluation results, and can flexibly adjust weights according to different application needs to dynamically evaluate similarity in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167332B_ABST
    Figure CN121167332B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of video similarity evaluation, and particularly relates to a video and text similarity evaluation method.The method comprises the following steps: performing semantic hierarchical analysis on a to-be-tested novel text to generate a text semantic dataset containing multi-dimensional semantic elements; performing frame-by-frame processing on a to-be-tested video in sequence to construct a video semantic dataset corresponding to the text; comparing the text semantic dataset and the video semantic dataset in terms of narrative elements, arranging the narrative elements of the text and the video in chronological order, and evaluating the similarity values of the narrative elements; and comparing the scene and action descriptions in the text of the text semantic dataset with the video picture semantic descriptions sentence by sentence, and performing counting when the similarity exceeds a preset threshold value.The present application evaluates the similarity between videos and texts to realize the evaluation of the semantic matching degree between videos and texts, and improves the accuracy of cross-modal content correspondence.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video similarity evaluation, and particularly relates to a video and text similarity evaluation method. BACKGROUND

[0002] The similarity evaluation is usually only based on keywords or plot summaries for text-level comparison, lacking understanding of narrative logic, scene transition and action semantics and other deeper elements, and it is difficult to accurately reflect the consistency of videos and texts in narrative structure. The existing video analysis methods are mainly based on visual feature detection or picture object recognition, for example, simple labeling is performed by detecting characters, scenes or object categories. Although this kind of method can provide basic visual information, it lacks semantic coherence modeling at the narrative level and cannot perform semantic disassembly and comparison equivalent to texts. When it is necessary to judge whether a film and television segment accurately reflects the scene or action description in the original work, the existing method can only rely on retrieval and subjective judgment, which is low in efficiency and accuracy. SUMMARY

[0003] Therefore, it is necessary to provide a video and text similarity evaluation method to solve at least one of the above technical problems.

[0004] To achieve the above-mentioned purpose, a video and text similarity evaluation method comprises the following steps:

[0005] Step S1: performing semantic hierarchical analysis on the to-be-tested novel text to generate a text semantic dataset containing multi-dimensional semantic elements, and sequentially performing frame processing on the to-be-tested video to construct a video semantic dataset corresponding to the text;

[0006] Step S2: comparing the narrative elements of the text semantic dataset and the video semantic dataset, arranging the narrative elements of the text and the video in chronological order, and evaluating the similarity value of the narrative elements;

[0007] Step S3: comparing the scene and action descriptions in the text of the text semantic dataset with the picture semantic descriptions of the video sentence by sentence, performing counting when the similarity exceeds a preset threshold, and determining the content coverage rate by using the ratio of the count to the number of text sentences;

[0008] Step S4: performing weighted fusion on the similarity calculation index of the narrative elements and the content coverage rate to obtain a comprehensive similarity value, and outputting a similarity evaluation result when the comprehensive similarity value reaches a judgment threshold.

[0009] The present application has the following beneficial effects:

[0010] By performing semantic level analysis and frame semantic construction on the novel text and video content respectively, the unified expression of different modal contents in the semantic dimension is realized, and the limitation that the existing text and video can only match at the keyword or image label level is overcome.

[0011] By establishing the time sequence comparison relationship of narrative elements, the text and video content can be one-to-one corresponding in the narrative logic, thereby avoiding the incomplete matching problem caused by the traditional method based on static pictures or partial scenes.

[0012] By introducing the fine-grained comparison mechanism of scene and action, the quantitative measurement of content restoration degree is realized by comparing the text description and video picture semantics sentence by sentence, and the objectivity and reproducibility of the evaluation result are effectively improved.

[0013] By weighting and fusing the narrative similarity index and content coverage, not only the matching degree of two modalities in the narrative logic and content presentation is comprehensively reflected, but also the weight can be flexibly adjusted according to different application requirements, so that the dynamic evaluation of similarity in multiple scenes is realized. BRIEF DESCRIPTION OF DRAWINGS

[0014] Fig. 1 It is a step flowchart of a video and text similarity evaluation method;

[0015] Fig. 2 It is a text and video content similarity evaluation flowchart;

[0016] Fig. 3 It is a text and video semantic analysis flowchart;

[0017] The implementation, functional characteristics and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0018] The technical method of the present application will be described clearly and completely below in combination with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor belong to the protection scope of the present application.

[0019] Moreover, the attached drawings are only schematic and are non-limiting. Each drawing figure is a functional description of the various embodiments although the drawings can not be to scale. The same reference signs are used in different drawings to refer to the same or like parts, where appropriate. Some of the drawings can be schematic or exaggerated representations of concepts not drawn to scale. The functional description of the various embodiments are better understood when the drawings are considered in conjunction with the following detailed description and with the claims.

[0020] It should be understood that, although the terms "first", "second" and the like can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of example embodiments. The term "and / or" as used herein encompasses any and all combinations of one or more of the associated associated items.

[0021] To achieve the above object, there is provided Figs. 1 to 3 A video and text similarity evaluation method, comprising the following steps:

[0022] Step S1: performing semantic level analysis on the novel text to be tested to generate a text semantic dataset containing multi-dimensional semantic elements; and performing frame-by-frame processing on the video to be tested to construct a video semantic dataset corresponding to the text;

[0023] Step S2: comparing the narrative elements of the text semantic dataset and the video semantic dataset, arranging the narrative elements of the text and the video in chronological order, and evaluating the similarity value of the narrative elements;

[0024] Step S3: comparing the scene and action descriptions in the text of the text semantic dataset with the video picture semantic descriptions sentence by sentence, performing counting when the similarity exceeds a preset threshold, and determining the content coverage rate by using the ratio of the count to the number of text sentences;

[0025] Step S4: performing weighted fusion on the similarity calculation index of the narrative elements and the content coverage rate to obtain a comprehensive similarity value, and outputting a similarity evaluation result when the comprehensive similarity value reaches a judgment threshold.

[0026] In an embodiment, when performing semantic hierarchical parsing on the novel text to be tested, the original text is split into sentence units (about 500 sentences) with period, exclamation mark, and question mark as the sentence identifiers, and part-of-speech tagging and dependency syntax analysis are performed on each sentence. The BERT-Base Chinese pre-training model is used to extract sentence-level semantic vectors (768 dimensions), and a self-built narrative element dictionary (containing 1200 words in 5 categories of characters, scenes, events, emotions, and time) is used to filter out semantic element items through keyword matching and syntactic dependency relationship, forming a text semantic dataset. The dataset is indexed by sentence order and contains semantic element labels, sentence-level vectors, element categories, and occurrence positions. At the same time, the video to be tested is processed by frame extraction, and 3000 frames of images are extracted from the original video (total length 10 minutes) at a sampling rate of 5 frames per second. The pre-trained CLIP model is used to extract image semantic features (512 dimensions) and scene description labels for each frame, and a video semantic dataset is established with frame sequence as the index. To ensure consistency with the text time narrative, the video frame sequence is time-normalized, and each 100 frames is considered as a narrative segment, forming 30 semantic segments, each storing data structures in three dimensions of scene category, main action, and environmental features.

[0027] The above two types of semantic datasets are input into the narrative comparison module, and the three types of narrative elements extracted from the text dataset, namely characters, events, and backgrounds, are represented by vectors. The character element uses Word2Vec embedding (300 dimensions), the event element uses RoBERTa encoding (768 dimensions), and the background element uses scene word embedding (256 dimensions). The corresponding elements on the video side extract event vectors through the action recognition network (I3D structure), extract character label features through the ResNet-50 classifier, and extract background features through the Places365 network.

[0028] The same type of narrative elements in the text and video are matched in time sequence, the semantic cosine similarity is calculated, and a narrative comparison matrix is formed. The matrix element value ranges from 0 to 1, and when the similarity of a corresponding element exceeds 0.7, the pair is marked as an effective matching item. Count all matching items, and the ratio of the number of matches to the total number of narrative elements is the narrative similarity index. In this embodiment, the calculated narrative similarity is about 0.68.

[0029] 280 sentences with scene and action description in the text semantic dataset were extracted independently. TF-IDF keywords were extracted for each sentence, and sentence-level semantic feature vectors (512 dimensions) were generated using the CLIP text encoder. On the video side, frame segments corresponding to the sentence sequence time period were selected from the video semantic dataset, and picture semantic feature vectors (also 512 dimensions) were extracted, and the cosine distance between the two was calculated. The similarity threshold was set to 0.75, and when the similarity was higher than the threshold, the sentence was recorded as an effective match. Finally, 200 sentences of text were matched with the corresponding pictures, and the content coverage rate was calculated .

[0030] In an embodiment, the narrative similarity index and the content coverage rate are weighted and fused, and the fusion weights are set to 0.6 and 0.4. The comprehensive similarity calculation formula is: , wherein = 0.68S, = 0.714, the comprehensive similarity is 0.68. With 0.65 as the judgment threshold, the result of this embodiment exceeds the threshold, and it is determined that the video content and the novel text have high semantic consistency.

[0031] Please refer to Fig. 2 , the complete workflow of the automated analysis of the similarity between the text and the video content is shown. The original text content and the video source file to be tested are obtained, the original text content is subjected to natural language processing to analyze and extract the multi-dimensional semantic information it contains, and a text semantic dataset is generated. At the same time, the video source file to be tested is subjected to image processing and subtitle extraction, and the video picture and text information are analyzed to generate a video semantic dataset. The text semantic dataset and the video semantic dataset are subjected to multi-dimensional information similarity comparison, including narrative structure comparison and detail content coverage analysis of the extracted semantic elements in the two types of datasets, similarity calculation based on the comparison results, and output of the quantitative similarity evaluation result.

[0032] Especially important is that step S1 includes:

[0033] The semantic hierarchy of the novel text to be tested is analyzed, and the character feature elements, scene information elements and event narrative elements are extracted in turn;

[0034] The extracted element data is indexed according to the sentence, paragraph and chapter structure, and the element type, semantic position and context dependency relationship are recorded in the index;

[0035] According to the semantic hierarchy index, the data of characters, scenes and events are converted into multi-dimensional semantic feature values in a vectorized manner to form a text semantic dataset containing multi-dimensional semantic elements;

[0036] Frame graph segmentation and key frame sampling are sequentially performed on the video to be tested to extract the main character region, background environment information and action behavior characteristics of each frame of picture;

[0037] Speech signal extraction in the video is performed to generate text data.

[0038] Frame features and speech text data are sequentially established according to frame sequence time to establish a video semantic index table, which records time identifiers, object types and semantic contents.

[0039] According to the video semantic index table, a video semantic data set corresponding to the novel text semantic structure is generated.

[0040] In an embodiment, the novel text is about 48,000 words in total, and 3126 sentences are obtained after sentence processing. HanLP is used to perform dependency syntax analysis and semantic role labeling, and the following elements are sequentially extracted: 128 character feature elements (including first appearance description, title, relationship chain), 214 scene information elements (including location, weather, light, prop), and 176 event narrative elements (including trigger words, arguments, cause and effect). The semantic index is established according to the level: sentence level ID, paragraph level ID, chapter level ID, each record contains element type (such as "character-appearance"), starting sentence number, context dependency pointer (pointing to the previous sentence ID). Vectorization uses Sentence-Transformer (all-mpnet-base-v2) to encode the first description sentence to generate a 768-dimensional vector, and additional attribute labels form a text semantic data set. The video is 62 minutes long, and FFmpeg is used to uniformly sample 3720 frames at 1fps. YOLOv8 is used to segment the character region (confidence ≥0.6), Places365 is used to classify the background, and SlowFast is used to identify the action behavior. A total of about 4100 visual elements are extracted. The speech is transcribed into subtitles by Whisper-large-v3 (average CER <3%), and is aligned with the frame time. A video semantic index table is constructed: each row records the frame range (such as 1200-1210), timestamp (40.0-40.3s), object type (character / scene / event), and semantic content (vector+label). Finally, the text chapter average duration (62min / 12 chapters 5.17min / chapter) is mapped to the video interval to generate a dual-modal aligned semantic data set, and the index table supports bidirectional retrieval of sentence ID time period.

[0041] In another embodiment, the novel text is about 52,000 words, and after division, 3418 sentences are obtained. The StanfordCoreNLP is used to complete the dependency analysis, NER and SRL, and 142 character features, 198 scene information and 189 event narrative elements are extracted. The hierarchical index is constructed according to the sentence, paragraph and chapter, and each element record type, location interval and dependency chain. The CLIP text branch is used for vectorization to encode the description sentence, and the output is a 512-dimensional feature with an additional role activity weight. The video length is 58 minutes, and the key frame is sampled based on the content change (scene switching detection) for a total of 3180, combined with InsightFace to recognize the face trajectory, Segment-Anything to segment the main body, and TimeSformer to encode the action sequence. The SeamlessM4T model is used for speech transcription to generate a subtitle stream that is accurately aligned with the frame. The video semantic index table is aggregated in a 10-second sliding window: each record contains the start time (such as 1320s), the duration, the fusion vector (visual 0.6+ speech 0.4 weighted), and the semantic label. The dynamic time warping (DTW) is used for text and video alignment to stretch and match the chapter event density curve, and a fine-grained correspondence relationship is generated. Finally, the video semantic dataset and the text hierarchical structure realize the three-level alignment of sentence-frame, section-fragment, and chapter-period.

[0042] Please refer to FIG. 1, FIG. 2 and FIG. 3. Fig. 3 For the text and video semantic analysis process diagram, the speech extracted from the video is transcribed into text, or the text information is directly received, and the Stanford CoreNLP tool is used for word segmentation, syntax analysis and other natural language processing of the text. Then through the "vectorization" technology, the text content is converted into digital text embedding. Representative pictures are extracted from the video stream. The TimeSformer model is used to analyze the temporal coherence information of the video frames, understand the action and scene changes. The InsightFace model is used to identify and extract features of the faces in the pictures. The Segment Anything model is used to identify and segment each object instance in the pictures. All features obtained from video, image, face and text branch processing are gathered together for unified "vectorization" and "hierarchical index construction".

[0043] Preferably, step S2 comprises:

[0044] In the text semantic dataset and the video semantic dataset, narrative element information is extracted respectively, wherein the narrative element information includes character semantic features, event semantic units and story background elements;

[0045] The text narrative elements and the video narrative elements are executed for time sequence correspondence to form a narrative comparison sequence;

[0046] The semantic similarity of corresponding elements in the narrative contrast sequence is calculated, and the corresponding elements are marked as matching items;

[0047] The number of matching items is counted, and the matching proportion of narrative elements is determined as the similarity calculation index of narrative elements.

[0048] In an embodiment, after dependency parsing, the input text obtains 385 sentences, 56 character entities, 72 event trigger words, and 84 background description keywords. The video input time is 50 minutes, the frame rate is 30 fps, 3000 key frames are extracted at 1 fps using ffmpeg, 212 captions are obtained by combining speech transcription, and the picture description is generated by the image model. The text is divided into 8 time blocks according to the paragraph length (about 480 words per block), the video is divided into 8 segments according to the time length (375 seconds per segment), and the time sequence mapping is established. The semantic features of the characters are converted into 768-dimensional vectors by the sentence encoding model, the trajectories of the characters in the video are combined with the corresponding frame picture description to encode the same-dimensional vectors, the cosine similarity threshold is 0.70, and 18 pairs of characters are marked as matching items. The event semantic unit extracts the trigger words from the text to construct the event vector, the action sequence in the video is encoded by the time sequence model, the vector similarity is calculated after dynamic time warping alignment, the threshold is 0.65, and 22 pairs of events are marked as matching. The background elements extract the scene keywords from the text, the video uses the scene classification model to output the category label, and the weighted consistency score threshold is 0.73, marking 19 pairs of background matching. The total number of narrative elements is 212 (56+72+84), the total number of matching items is 59, the initial proportion is 0.278, after position weight adjustment (the first 1 / 4 block weight is 1.0, the next 1 / 4 weight is 0.9, the next 1 / 4 weight is 0.7, and the last 1 / 4 weight is 0.5), the narrative element similarity index is 0.291.

[0049] In another embodiment, after the input text is segmented, 412 sentences are obtained, 49 character entities, 68 event trigger words, and 76 background keywords are extracted. The video has a duration of 38 minutes and a frame rate of 25 fps, and 2280 key frames are obtained by frame extraction. There are 178 speech transcriptions, and picture descriptions are generated synchronously. The text is divided into 7 time blocks (about 590 words per block), and the video is equally divided into 7 segments (326 seconds per segment). The timing alignment is completed. After the semantic feature coding of the character, it is compared with the video character track vector one by one, the cosine similarity threshold is 0.71, and 15 pairs of matches are marked. The event unit is encoded through the causal chain structure, and the transition node similarity is calculated after the video action sequence alignment, the threshold is 0.67, and 20 pairs of matches are marked. In the background element comparison, the weighted score of the text keyword and the video scene label is calculated, the threshold is 0.74, and 17 pairs of matches are marked. The total number of narrative elements is 193 (49+68+76), the total number of matching items is 52, the initial proportion is 0.269, and after adjusting the position decreasing weight (the first 30% block weight is 1.0, the middle 40% weight is 0.8, and the last 30% weight is 0.6), the narrative element similarity index is 0.277.

[0050] Preferably, the semantic similarity of the semantic features of the characters in the narrative comparison sequence is calculated, and the corresponding elements are marked as matching items, including:

[0051] In the semantic data set of the novel text, the semantic features of the characters are extracted, and the semantic data of the characters are established according to the order of appearance;

[0052] In the video frame sequence, the image and speech data are analyzed, the character recognition features are extracted, and the posture changes are analyzed, and the video character data are established according to the frame time sequence;

[0053] Taking the time identifier index as the alignment reference, the corresponding comparison between the semantic data of the characters and the video character data is performed, and the semantic similarity of the character features is calculated;

[0054] When the similarity of any feature dimension reaches the preset threshold, the corresponding character information is determined as a matching data item, and the matching information is recorded;

[0055] Based on the matching information, the mapping data of the text character information and the video character information is generated, and the mapping sequence is formed by arranging in time identifier order.

[0056] In one embodiment, semantic features of characters are extracted from a text semantic dataset and character data is constructed according to their order of appearance; image and audio data are parsed from a video frame sequence to extract character recognition features and posture changes, and video character data is constructed according to time sequence; the two types of data are aligned based on time markers, and the semantic similarity of character features is calculated; when the similarity in any dimension reaches a threshold, it is marked as a match and recorded; mapping data is generated based on the matching information and arranged into a mapping sequence according to time sequence. After dependency parsing, the text yields 428 sentences, and NER identifies 61 character entities. Character data is constructed by sorting by the chapter in which they first appear. Each record contains a character description sentence vector (Sentence-Transformer 768-dimensional) and an appearance timestamp (normalized chapter position 0.12, 0.27, etc.). The video is 45 minutes long, with a frame rate of 30fps. Using ffmpeg, 2700 frames are extracted per second (1 frame per second). InsightFace recognizes facial trajectories and generates 38 character trajectories. Each trajectory aggregates a 10-second scene description (generated using BLIP2) and a sequence of pose key points (OpenPose 18×2D), encoded as a 768-dimensional vector. The timestamp is the video second (13s, 27s, etc.). The text timestamp is linearly mapped to the video time (text 0.12 → video 648s), and nearest neighbor alignment is used to form 61×38 candidate pairs. Cosine similarity is calculated: appearance description vector similarity, appearance-pose joint vector similarity, and voice emotion vector similarity (extracted using Wav2Vec2), with weights of 0.5, 0.3, and 0.2 respectively. A threshold of 0.73 is set; a match is marked when any dimension of any candidate pair exceeds the threshold. A total of 19 pairs are marked. The matching pair index is recorded (e.g., text character 3 → video trajectory 7), and a mapping sequence is generated and sorted by time: ( ), ( )wait.

[0057] In another embodiment, the text is segmented into 392 sentences, from which 53 human entities are extracted. Human data is constructed according to paragraph order, and the description vector is generated as a 512-dimensional vector by a multimodal encoder (CLIP text branch). The timestamp is the percentage of cumulative words in each paragraph (0.08, 0.19, etc.). The video is 52 minutes long, with 3120 frames extracted. ArcFace recognizes faces and clusters them into 41 trajectories. Every 15 seconds, each trajectory fuses the image description, clothing color histogram (HSV36-dimensional), and motion energy spectrum (SlowFast output), encoding it into a 512-dimensional vector. The timestamp is the number of seconds at the midpoint of the frame (11s, 26s, etc.). Alignment is performed by multiplying the text timestamp by the total video duration to obtain the target seconds, allowing for searching the nearest trajectory within a ±8-second window. Similarity calculation includes: cosine similarity of the description vector, Bach distance of clothing color, and L2 distance of motion energy, with weights of 0.4, 0.4, and 0.2, respectively. The threshold is 0.71; a match is achieved if any dimension meets the threshold, resulting in 16 pairs of tags. The mapping sequence is sorted by text time: ( ), (, etc. The intermediate non-matching role markers are empty. Preferably, calculating semantic similarity of event semantic units in the narrative alignment sequence, marking corresponding elements as matching items includes:

[0058] Extracting plot element information in the novel text semantic dataset, collecting novel text plot node information in the order of plot advancement, establishing plot semantic records, decomposing the plot development process into causal chain sequences, and converting into plot vector features;

[0059] Collecting video event semantic units corresponding to the time period, establishing video plot data, decomposing the plot development process into causal chain sequences, and converting into plot vector features;

[0060] Aligning plot semantic records with video plot records, performing corresponding comparison, and calculating similarity indexes of plots in causal chain structures and turning node features;

[0061] When the similarity indexes of the causal chain structure and the turning node feature reach the preset threshold, the corresponding plot is determined as a matching object, and plot matching information is recorded;

[0062] Based on the plot matching information, mapping data of text plots and video plots is generated, and a plot mapping sequence is formed in the order of time identification.

[0063] In an embodiment, after event extraction, 74 plot nodes are identified in the text, arranged in the order of paragraphs, each node containing trigger words, arguments, and causal relationships before and after, forming a causal chain sequence, and each chain segment is encoded as a 512-dimensional vector (trigger word CLIP embedding 0.6 weight + argument average embedding 0.4 weight). The video is 48 minutes long, divided into 12 segments (240 seconds per segment), and SlowFast detects action sequences within each segment, and Whisper transcribes dialogue trigger words, merging into 62 event units, also constructing causal chains (such as "running → hiding → counterattack"), and vector encoding is the same as the text. Time alignment maps the normalized position of the text node to the second level (position 0.23 → 662s) with the total duration of the video, allowing

[0064] 18-second window searches for the nearest video event. Structural similarity is calculated using chain segment order cosine, and turning node similarity is calculated using key action vector L2 distance, with weights of 0.55 and 0.45. The threshold values are 0.69 and 0.62, respectively. Either one meets the requirements, and a total of 21 pairs are marked as matching. The matching index is recorded (such as text node 12 Video event 17), and the mapping sequence is sorted by text time: ( ), (, etc. The non-matching node markers are empty.

[0065] ​​In another embodiment, 68 narrative nodes are extracted from the text, ordered by chapter progression, and each segment vector is 768-dimensional after causal chain decomposition (0.5 weight for trigger word Sentence-Transformer embedding ). The video is 55 minutes long and is divided into 10 segments (330 seconds each), and 59 units are obtained by combining action recognition and speech event detection. After causal chain construction, the vector is encoded with the same dimension. The time sequence mapping is obtained by multiplying the proportion of cumulative word count of text nodes by video duration (proportion ), and the window is aligned within 25 seconds. The structural similarity is calculated by weighted Jaccard overlap of segments, and the turning node is calculated by the cosine of the action energy peak vector, with weights of 0.6 and 0.4. The threshold values are 0.67 and 0.64, respectively. Any one above the threshold is marked as a match, and there are 18 pairs in total. The mapping sequence is arranged in ascending order of video time: (561s Text node 9), (892s Text node 14), etc., and placeholders are inserted at the intermediate broken links.

[0066] Preferably, the semantic similarity of the story background elements in the narrative alignment sequence is calculated, and the corresponding elements are marked as matching items, including:

[0067] The story background elements of the novel text are extracted from the narrative alignment sequence, the background semantic record data is established according to the appearance order of the background description, and the background features are extracted;

[0068] The story background elements of the corresponding time period are extracted from the video semantic data set, the video background record data is established according to the frame sequence order, and the background features in the video are extracted;

[0069] The background semantic record data and the video background record data are compared, and the semantic consistency score of the background features is calculated;

[0070] The consistency score is adjusted by weighting to adjust the score of the highly consistent area with a preset weight;

[0071] The score is used as a story background similarity index, and when the index reaches a preset threshold, the corresponding elements are marked as matching items.

[0072] In an embodiment, the text is recognized after dependency parsing, 89 background description sentences are identified in the order of paragraphs, the place nouns and modifying phrases of each sentence are extracted, and are encoded into 512-dimensional vectors (0.5 weight of place entity embedding + 0.5 weight of average embedding of modifiers), and the timestamp is the normalized paragraph position (0.11, 0.24, etc.). The video duration is 46 minutes, which is divided into 15 segments (184 seconds per segment), the top 3 scene labels of Places365 classification and the main objects detected by YOLO are used in each segment, which are combined into background description text, and are encoded into 512-dimensional vectors by the CLIP text branch, and the timestamp is the segment midpoint second (507s, 920s, etc.). The alignment is mapped by the text position x the total duration of the video, which allows 12-second window searches the nearest video segment. The semantic consistency score is calculated as: place category cosine similarity (0.6 weight), object set Jaccard (0.3 weight), and illumination vector L2 distance (0.1 weight). The highly consistent area (place + object overlap ≥ 2 items) is weighted up by 1.3 times. The threshold is set to 0.70, and 26 pairs are marked as matching items.

[0073] In another embodiment, the text extracts 76 background description sentences, which are constructed into records in the order of chapters, and the background features are 768-dimensional vectors (scene keywords Sentence-Transformer embedding 0.4 weight + weather / time modification embedding 0.6 weight), and the timestamp is the cumulative word proportion (0.09, 0.22, etc.). The video duration is 53 minutes, which is divided into 13 segments (245 seconds per segment), and the scene classification, sky segmentation (weather inference), and timing illumination curve are fused in each segment, which are encoded into 768-dimensional vectors, and the timestamp is the segment start second (441s, 858s, etc.). The aligned window second. The score calculation includes: scene label cosine (0.5 weight), object co-occurrence Bhattacharyya distance (0.35 weight), and weather state matching (0.15 weight). The highly consistent area (scene + weather full match) is weighted up by 1.4 times. The threshold is 0.68, and 23 pairs are marked as matching items.

[0074] Preferably, the number of matching items is counted, and the narrative element matching proportion is determined as the ratio of the number of matching items to the total number of narrative elements, which is included as the similarity value of the narrative element, including:

[0075] Initialize the matching item counter to zero, and traverse the corresponding elements of the narrative contrast sequence pair by pair, check the marked state of the corresponding elements, and accumulate the number of elements marked as matching items;

[0076] After the traversal is completed, the number of matching items is obtained, and the total number of narrative elements in the narrative contrast sequence is obtained;

[0077] Calculate the initial ratio by dividing the number of matching items by the total number of narrative elements, perform weight adjustment, and assign decreasing weights according to the position of the narrative elements in the sequence. Multiply the initial ratio by the sum of the weights to obtain the adjusted ratio.

[0078] The adjusted ratio is used as the narrative element matching ratio and as an indicator for calculating narrative element similarity.

[0079] In one embodiment, the narrative comparison sequence consists of three types of elements: characters, events, and background. The total number of elements is the union of the text extraction results and the video extraction results. The counter is initialized to zero, and the sequence is iterated pair by pair starting from the beginning, checking the matching status of each pair of elements. Once a matching item is found, the count is incremented. The weight of the first segment (first third of the sequence) is set to 1.0, the weight of the middle segment (middle third) to 0.85, and the weight of the last segment (last third) to 0.6, reflecting the importance of the early plot in similarity judgment. After the iteration is complete, the total number of matching items is obtained, and the initial matching ratio is calculated based on the total number of elements. Subsequently, each matching item is weighted and summed according to the above segment weights. The first segment matching items contribute the most due to their highest weight, followed by the middle segment, and the last segment contributes the least. Finally, the weighted sum is divided by the sum of the weights of the three segments to obtain the normalized adjustment ratio, which serves as the narrative element similarity index in this example.

[0080] In another embodiment, the narrative comparison sequence also contains three types of elements, and the total number of elements is determined by merging and deduplicating the text and video extraction results. The counter is initialized to zero, and the traversal process starts from the beginning of the sequence, checking the labeling status of each element pair and accumulating the matching items. To further refine the positional influence, the sequence is divided into four segments: the first... Weight 1.0, next Weight 0.9, then next Weight 0.75, last The weight is 0.5. Matches in the first part of the story have the highest weight because they occur at the beginning; matches in the later parts, which are mostly concluding events, have lower weights. After traversal, the total number of matches is counted and the initial ratio is calculated. Then, each of the four parts is weighted separately, with the first part contributing significantly and the contribution gradually decreasing in the later parts. The weighted sum is divided by the sum of the weights of the four parts to obtain a more conservative adjustment ratio, which serves as the final narrative element similarity index.

[0081] Preferably, step S3 includes:

[0082] The text semantic dataset is split into a sequence of scene description sentences, and scene keywords are extracted sentence by sentence and converted into scene feature vectors.

[0083] The semantic description of the video frame is broken down into a sequence of frames corresponding to the time period, and keywords are extracted from each frame and converted into frame feature vectors.

[0084] The cosine distance between the scene feature vector and the picture feature vector is calculated, and the distance score is converted;

[0085] A weight coefficient is assigned to the scene keyword according to the frequency of occurrence in the text semantic data set, and the distance score is multiplied by the weight coefficient to obtain a modified score;

[0086] The modified score is used as the scene similarity, and when the scene similarity exceeds the preset threshold, the score is counted, and the content coverage rate is determined by the ratio of the counted score to the total number of scene description sentence sequences.

[0087] In an embodiment, the text semantic data set is processed by sentence to obtain a total of 156 scene description sentence sequences, each sentence extracts keywords such as location, environment and object through dependency analysis, and a pre-trained sentence encoding model is used to generate a 512-dimensional scene feature vector. The video is 42 minutes long, and the time is evenly divided according to the number of text sentences, obtaining 156 corresponding time periods, and the scene mentioned in the key frame picture description (generated by the image description model) and the speech transcription is fused in each period. Similarly, the keywords are extracted and encoded into 512-dimensional picture feature vectors. The two sequences are corresponded one by one based on the text sentence number, the cosine distance between the vectors is calculated and converted into a similarity score (1-distance). The frequency of occurrence of keywords in the full text is counted, and high-frequency words are assigned a weight of 1.4, medium-frequency words are assigned a weight of 1.0, and low-frequency words are assigned a weight of 0.7. Each sentence is modified by the average weight of its keywords to obtain the final scene similarity. The threshold is set to 0.68, and the cumulative number of sentences exceeding the threshold is 98 in the traversal, and the content coverage rate is calculated as .

[0088] In another embodiment, the text scene description sentence sequence contains 142 sentences, and a multi-modal encoder is used to generate a 768-dimensional feature vector after extracting keywords sentence by sentence. The video is 50 minutes long, divided into 142 segments, and the semantic description of multiple frames of pictures and the ASR transcription text are aggregated in each segment. After keyword extraction, the same dimension vector is encoded. After alignment, the cosine distance is calculated and converted into a score. The keyword frequency statistics are divided into four grades according to the global distribution: very high frequency (> 30 times) weight 1.6, high frequency (15-30 times) weight 1.3, medium frequency (5-14 times) weight 1.0, and low frequency (< 5 times) weight 0.8. The modified score is obtained by taking the average weight of the keywords of each sentence. The threshold is set to 0.70, and the number of sentences exceeding the threshold is 86 in the traversal, and the content coverage rate is .

[0089] Preferably, the modified score obtained by multiplying the distance score by the weight coefficient assigned to the scene keyword according to the frequency of occurrence in the text semantic data set includes:

[0090] The number of occurrences of each scene keyword in the text semantic data set is counted to form a frequency statistics table;

[0091] The keywords are sorted by frequency, and the high-frequency keywords are given a weight coefficient greater than 1, and the low-frequency keywords are given a weight coefficient less than 1.

[0092] For each scene description sentence, the corresponding weight coefficient of the keyword is searched, and the average value of the weight of the keyword in the sentence is calculated as the sentence-level weight.

[0093] The sentence-level weight is multiplied by the distance score to obtain the weighted correction score.

[0094] In an embodiment, all sentences in the text semantic data set are processed for word segmentation and part-of-speech tagging, and the nouns and adjectives related to the scene are extracted as the scene keyword set. The frequency statistics of each scene keyword in the text corpus are counted using the word frequency statistical algorithm to obtain a frequency statistics table. In the statistical results, the top 10% of keywords with the highest frequency are selected as high-frequency keywords, and the weight coefficient is set to 1.2; the keywords with frequency in the middle interval are set to 1.0; and the keywords with the lowest frequency are set to 0.8. When processing each scene description sentence, the scene keywords contained in the sentence are searched one by one, the corresponding weight coefficient is searched, and the arithmetic mean of the weights of all keywords in the sentence is calculated to obtain the sentence-level weight of the sentence. For example, the sentence "The street at night is illuminated by streetlights" contains two keywords "night" and "street", and their weights are 1.2 and 1.0 respectively, so the sentence-level weight of the sentence is (1.2+1.0) / 2=1.1.

[0095] After the sentence-level weight calculation is completed, the original distance score D obtained by comparing the text and video semantics and the sentence-level weight w are multiplied to obtain the correction score , which reflects the promotion effect of high-frequency scene keywords on the overall similarity. The weighting process is repeated for all sentences to finally form a sequence of distance scores corrected by weight, which provides input data for subsequent narrative consistency evaluation.

[0096] In another embodiment, when performing keyword statistics on the text semantic data set, a weight calculation strategy based on TF-IDF (Term Frequency-Inverse Document Frequency) is adopted to improve the differentiation of keywords in different scenarios between corpora. The text corpus is divided into several chapter units, and the occurrence frequency TF of each keyword in the chapter and the inverse document frequency IDF=log(N / n) of all chapters are calculated, where N is the total number of chapters and n is the number of chapters containing the word. The calculated weight value TFxIDF is used to measure the importance of the keyword. After obtaining the weight, normalization processing is performed on all keywords to limit the weight value within the interval [0.5, 1.5]. Subsequently, each text description sentence is traversed, the corresponding weight of the keyword in the sentence is found, and the average weight in the sentence is calculated. For example, in a text "at dusk, the street corner coffee shop light is soft", the TF-IDF weights of the keywords "dusk", "street corner" and "coffee shop" are 1.4, 1.1 and 0.9 respectively, and the average value (1.4+1.1+0.9) / 3=1.13 is taken as the sentence-level weight of the sentence. Finally, the sentence-level weight is multiplied by the semantic distance score D obtained in the video text similarity calculation stage to obtain the modified score . The modified score sequence calculated in this way can be input into the subsequent similarity summary model for matching accuracy statistics at the narrative scene level.

[0097] Preferably, multiplying the sentence-level weight and the distance score to obtain the weighted modified score includes:

[0098] The sentence-level weight of each sentence and the distance score of the corresponding picture sequence are paired one by one;

[0099] The sentence-level weight is used as a multiplier and the distance score is used as a multiplicand to generate a preliminary product result;

[0100] The preliminary product result is summed to form a total weighted sum value, and the valid pairing items are recorded synchronously;

[0101] The total weighted sum value is divided by the number of valid pairings to obtain a normalized average modified score as the modified score.

[0102] In an embodiment, the scene description sentence sequence has 156 sentences, and the sentence-level weight of each sentence has been calculated (such as the first sentence weight 0.967, the second sentence 1.12, the third sentence 0.83, etc.), and the distance score of the corresponding picture sequence has been converted into a similarity score (such as the first sentence 0.71, the second sentence 0.65, the third sentence 0.79, etc.). Pairing one by one according to the sentence sequence: the first sentence is paired with the weight 0.967 score , the second sentence , and the third sentence , and so on, generating 156 preliminary product results. The number of valid pairs recorded in the traversal is 156 (all sentences have corresponding pictures). The sum of the 156 products is 103.42, and the final normalized average modified score value is , which is the average of the modified score sequence.

[0103] In another embodiment, the sequence of scene description sentences has 142 sentences, the sequence of sentence-level weights (e.g., 1.233 for the first sentence, 0.95 for the second sentence, 1.41 for the third sentence, etc.), and the corresponding picture similarity scores (e.g., 0.68 for the first sentence, 0.73 for the second sentence, 0.61 for the third sentence, etc.). The product is calculated for each sentence pair: the first sentence , the second sentence , and the third sentence , generating 142 preliminary products. The number of valid pairs is recorded as 142 (no missing alignments). The total weighted sum is 96.18, and the normalized average modified score value is , which is the final modified score.

[0104] Preferably, the modified score is taken as the scene similarity, and when the scene similarity exceeds a preset threshold, the count is performed, and the ratio of the count to the total number of scene description sentences determines the content coverage rate, including:

[0105] Set the scene similarity as the modified score, and initialize the matching counter to zero;

[0106] Check the modified score of each sentence in sequence, and if the modified score is higher than the preset threshold, the matching counter is incremented by one;

[0107] Check the content of all sentences, and read the cumulative value of the matching counter as the number of valid matching sentences;

[0108] Get the complete number of sentences in the sequence of scene description sentences, divide the number of matching sentences by the complete number of sentences to calculate the proportion, and determine the proportion as the content coverage rate.

[0109] In an embodiment, the sequence of scene description sentences has 156 sentences, and the sequence of modified scores for each sentence has been obtained through weighted processing (e.g., 0.687 for the first sentence, 0.728 for the second sentence, 0.656 for the third sentence, …, 0.592 for the 156th sentence), and the normalized average modified score 0.663 is calculated as a reference benchmark. Set the scene similarity threshold to 0.68 (slightly higher than the average to filter significant matches), and initialize the matching counter to 0. Start checking each sentence from the first sentence in the sequence: the first sentence , the counter is incremented by one; the second sentence , the counter is incremented by one; and the third sentence .

[0110] In another embodiment, the sequence of scene description sentences is 142 sentences long, and the revised score of each sentence has been calculated (e.g., 1st sentence 0.838, 2nd sentence 0.694, 3rd sentence 0.860, …, 142nd sentence 0.701), and the normalized average revised score is 0.677. The threshold value is set to 0.70 (for distinguishing strong correlation scenes), and the matching counter is initialized to 0. From the beginning to the end of the sentence sequence, each sentence is judged in turn: the 1st sentence 0.838 > 0.70, the counter +1; the 2nd sentence 0.694 < 0.70, the counter +1; … until the 142nd sentence 0.701 > 0.70, the counter +1. At the end of the traversal, the counter accumulates a value of 86. Taking the total number of sentences 142, the content coverage rate is calculated as 86 / 142 = 0.6059, which is rounded to 0.61 as the final content coverage rate output of this example.

[0111] Especially important is that step S4 includes:

[0112] presetting a narrative element similarity weight coefficient and a content coverage weight coefficient;

[0113] multiplying the similarity value of the narrative element by the narrative element similarity weight coefficient to obtain a narrative weighted score;

[0114] multiplying the content coverage rate by the content coverage weight coefficient to obtain a coverage weighted score;

[0115] performing a sum operation on the narrative weighted score and the coverage weighted score to generate a comprehensive similarity value;

[0116] comparing the comprehensive similarity value with a preset determination threshold value, and when the comprehensive similarity value reaches or exceeds the determination threshold value, outputting a similarity evaluation result.

[0117] In an embodiment, the narrative element similarity index has been calculated to be 0.263, and the content coverage rate has been calculated to be 0.628. When presetting the weight coefficients, considering the core importance of the narrative structure to the adaptation judgment, the narrative element similarity weight coefficient is set to 0.65, and the content coverage weight coefficient is set to 0.35 (the sum of the two coefficients is 1.0 to ensure normalization). The narrative weighted score is calculated as: (rounding to three decimal places); then the coverage weighted score is calculated as: (rounding to three decimal places). The sum of the two weighted scores is: , obtaining a comprehensive similarity value of 0.391. The determination threshold value is set to 0.40 (an empirical threshold value based on historical labeling data statistics), and the comparison result is , which is determined to not meet the similarity standard, and the evaluation result is output as “low similarity, no obvious adaptation relationship”.

[0118] In another embodiment, the narrative element similarity indicator is 0.277, and the content coverage is 0.606. The weight coefficient is adjusted to a more balanced allocation strategy: the narrative element similarity weight coefficient is 0.55, and the content coverage weight coefficient is 0.45 (still satisfying the normalization condition) to balance the structural consistency and the degree of detail restoration. The narrative weighted score is calculated as: The coverage weighted score is calculated as: The comprehensive similarity value is obtained by adding the two scores: The determination threshold is set to 0.42 (calibrated according to the test set in this batch), and the comparison result is , which is determined to meet the similarity standard, and the evaluation result is output as “moderately similar, suspected partial adaptation”. To enhance the interpretability, the itemized contribution details are also output: the narrative structure contribution is 0.152 (35.8%), and the content coverage contribution is 0.273 (64.2%).

[0119] The above description is only a specific implementation of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest range consistent with the principles and novel features disclosed herein.

Claims

1. A method for evaluating the similarity between video and text, characterized in that, Includes the following steps: Step S1: Perform semantic hierarchical parsing on the novel text to be tested to generate a text semantic dataset containing multi-dimensional semantic elements; perform frame-by-frame processing on the video to be tested in sequence to construct a video semantic dataset corresponding to the text; Step S2: Compare the narrative elements of the text semantic dataset and the video semantic dataset, arrange the narrative elements of the text and video in chronological order, and evaluate the similarity value of the narrative elements; wherein, step S2 includes: Narrative element information is extracted from text semantic datasets and video semantic datasets, including character semantic features, event semantic units, and story background elements. The text narrative elements and video narrative elements are matched in a temporal sequence to form a narrative comparison sequence; Semantic similarity is calculated for corresponding elements in the narrative comparison sequence, and corresponding elements are marked as matches. Semantic similarity is also calculated for event semantic units in the narrative comparison sequence. Marking corresponding elements as matches includes: Extract plot element information from the novel text semantic dataset, collect plot node information of the novel text in the order of plot progression, establish plot semantic record, decompose the plot development process into causal chain sequence, and convert it into plot vector features; Collect semantic units of video events in the corresponding time period, establish video plot data, decompose the plot development process into a causal chain sequence, and convert it into plot vector features; Align the semantic records of the plot with the video plot records, perform a corresponding comparison, and calculate the similarity index of the plot in terms of causal chain structure and turning point features; When the similarity index between the causal chain structure and the turning point features reaches a preset threshold, the corresponding plot is identified as the matching object, and the plot matching information is recorded. Based on plot matching information, text plot and video plot mapping data are generated and arranged in chronological order to form a plot mapping sequence; The number of matching items is counted, and the ratio of the number of matching items to the total number of narrative elements is used to determine the narrative element matching ratio, which serves as the index for calculating narrative element similarity. Step S3: Compare the scene and action descriptions in the text of the text semantic dataset with the semantic descriptions of the video footage sentence by sentence. When the similarity exceeds a preset threshold, a count is performed, and the content coverage rate is determined by the ratio of the count to the number of text sentences. Step S3 includes: The text semantic dataset is split into a sequence of scene description sentences, and scene keywords are extracted sentence by sentence and converted into scene feature vectors. The semantic description of the video frame is broken down into a sequence of frames corresponding to the time period, and keywords are extracted from each frame and converted into frame feature vectors. Align the scene description sentence sequence with the image sequence sentence by sentence, calculate the cosine distance between the scene feature vector and the image feature vector, and convert it into a distance score; Based on the frequency of occurrence of scene keywords in the text semantic dataset, an increasing weight coefficient is assigned, and the distance score is multiplied by the weight coefficient to obtain the corrected score; The corrected score is used as the scene similarity. When the scene similarity exceeds the preset threshold, counting is performed, and the ratio of the count to the total number of sentences in the scene description sentence sequence is used to determine the content coverage. Step S4: Perform a weighted fusion of the similarity calculation index of the narrative elements and the content coverage rate to obtain a comprehensive similarity value. When the comprehensive similarity value reaches the judgment threshold, output the similarity evaluation result.

2. The method for evaluating video and text similarity according to claim 1, characterized in that, Calculate semantic similarity for the semantic features of characters in the narrative comparison sequence, and mark the corresponding elements as matching items, including: Extract semantic features of characters from a novel text semantic dataset and build character semantic data according to their order of appearance; The image and audio data in the video frame sequence are parsed to extract human recognition features and analyze posture changes, and video human data is established in frame time sequence. Using the time stamp index as the alignment benchmark, a correspondence comparison is performed between the semantic data of the person and the video person data to calculate the semantic similarity of the person's features; When the similarity of any feature dimension reaches a preset threshold, the corresponding person information is determined as a matching data item, and the matching information is recorded. Based on the matching information, mapping data between textual and video character information is generated and arranged in chronological order to form a mapping sequence.

3. The method for evaluating video and text similarity according to claim 1, characterized in that, Calculate semantic similarity for story background elements in the narrative comparison sequence, and mark the corresponding elements as matches, including: Extract story background elements from the novel text from the narrative comparison sequence, establish background semantic record data according to the order of appearance of background descriptions, and extract background features; Extract story background elements for the corresponding time period from the video semantic dataset, establish video background recording data in frame sequence order, and extract background features from the video; The background semantic recording data and the video background recording data are compared, and the semantic consistency score of the background features is calculated. A weighted adjustment is performed on the consistency scores, adjusting the scores of highly consistent regions with preset weights; The score is used as a similarity indicator for the story background. When the score reaches a preset threshold, the corresponding element is marked as a match.

4. The method for evaluating video and text similarity according to claim 1, characterized in that, The number of matches is counted, and the ratio of the number of matches to the total number of narrative elements is used to determine the narrative element matching ratio. This ratio serves as the similarity value for the narrative elements. Initialize the matching item counter to zero, iterate through the corresponding elements in the narrative comparison sequence one by one, check the marking status of the corresponding elements, and accumulate the number of elements marked as matching items; The system iterates through the completed sequence to obtain the number of matching items and the total number of narrative elements in the narrative comparison sequence. Calculate the initial ratio by dividing the number of matching items by the total number of narrative elements, perform weight adjustment, and assign decreasing weights according to the position of the narrative elements in the sequence. Multiply the initial ratio by the sum of the weights to obtain the adjusted ratio. The adjusted ratio is used as the narrative element matching ratio and as an indicator for calculating narrative element similarity.

5. The method for evaluating video and text similarity according to claim 1, characterized in that, Based on the frequency of occurrence of scene keywords in the text semantic dataset, incremental weight coefficients are assigned. The corrected score is obtained by multiplying the distance score by the weight coefficients, including: Count the number of times each scene keyword appears in the text semantic dataset to form a frequency statistics table; Keywords are sorted by frequency, with high-frequency keywords assigned a weight coefficient greater than 1 and low-frequency keywords assigned a weight coefficient less than 1. Find the corresponding weight coefficient for the keywords in each scene description sentence, and calculate the average weight of the keywords in the sentence as the sentence-level weight; Multiply the sentence-level weight by the distance score to obtain the weighted corrected score.

6. The method for evaluating video and text similarity according to claim 1, characterized in that, Multiplying the sentence-level weights by the distance scores yields a weighted adjusted score, which includes: Each sentence's sentence-level weight is paired one-to-one with the distance score of the corresponding image sequence; Using sentence-level weights as multipliers and distance scores as multiplicands, a preliminary product result is generated. The initial product results are summed to form a total weighted sum, and valid paired terms are recorded simultaneously. Divide the total weighted sum by the number of valid pairs to obtain the normalized average corrected score, which is used as the corrected score.

7. The method for evaluating video and text similarity according to claim 1, characterized in that, The corrected score is used as the scene similarity score. When the scene similarity score exceeds a preset threshold, a count is performed, and the ratio of the count to the total number of sentences in the scene description sentence sequence determines the content coverage, including: Set the scene similarity score as the correction score and initialize the matching counter to zero; The correction score for each sentence is checked sequentially. If the correction score is higher than the preset threshold, the matching counter is incremented by one. Examine the entire sentence content and read the cumulative value of the matching counter as the number of valid matches; Obtain the complete sentence count of the scene description sentence sequence, divide the number of matched sentences by the complete sentence count to calculate the ratio, and determine the ratio as the content coverage rate.

Citation Information

Patent Citations

  • Video text similarity measurement method and system

    CN114092703A

  • Video retrieval method and device, electronic equipment and readable storage medium

    CN119046499A