Automatic video editing method based on semantic analysis
By constructing a knowledge graph and credibility scoring strategy, the automatic video editing method based on semantic analysis solves the problem of logical incoherence in the existing technology, and realizes efficient and logically coherent video editing, improving editing quality and efficiency.
Patent Information
- Application Number
- CN202510904405.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-07-01
AI Technical Summary
The existing video editing technology has problems such as subjective deviations in manual editing and lack of semantic understanding of automatic editing, resulting in logical incoherence and missing key information, which is difficult to meet the massive video processing needs.
The automatic video editing method based on semantic analysis is adopted to extract anchor information by constructing a knowledge graph, establish support relationships, set credibility scoring strategies, dynamically adjust weights, and generate logically coherent video editing results.
It improves the logical coherence and reliability of video clips, ensures the semantic consistency and fluency of editing clips, and improves editing efficiency and quality.
Smart Images

Figure CN120583283A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of video editing and relates to an automatic video editing method based on semantic analysis. Background Art
[0002] With the rapid development of multimedia technology and the explosive growth of video content, the demand for automatic video editing technology in the fields of short video creation, film and television production, etc. is becoming increasingly urgent. Existing video editing methods are mainly divided into two categories: manual editing and automatic editing. Traditional manual editing relies on the experience and subjective judgment of the editor. Due to the deviations in the understanding of video content among different editors, problems such as logical incoherence and omission of key information are prone to occur when splitting video materials. In addition, the editing efficiency is low, making it difficult to meet the needs of massive video processing. Existing automatic editing methods are mostly based on shallow visual or auditory features to screen and splice clips. They lack an in-depth understanding of the semantic level of video content and may rigidly splice content of different themes, resulting in logical discontinuities in the editing results and an inability to accurately convey the core information. Therefore, whether it is the subjective deviation of manual editing or the insufficient semantic understanding of automatic editing, it makes it difficult for existing editing technology to produce logically coherent high-quality videos. A new solution that can deeply analyze video semantics and improve editing logic is urgently needed. Summary of the Invention
[0003] The present invention aims to provide a method for automatic video editing based on semantic analysis. By extracting anchor point information and establishing supporting relationships to form a sequence set of anchor point information, this method addresses the existing problems of logical discontinuities caused by subjective understanding biases in manual editing and the lack of logical construction capabilities in automatic editing, thereby achieving automatic video editing with semantic and logical coherence. To achieve this objective, the present invention employs the following technical solutions.
[0004] On the one hand, the present invention provides a method for automatic video editing based on semantic analysis, including: obtaining a streaming media file, extracting text content in the streaming media file; constructing a knowledge graph, extracting anchor information in the text content based on the knowledge graph, establishing a support relationship, and determining a main anchor, a sub-anchor, and an argument anchor in combination with the support relationship and the anchor information; constructing a dynamic trigger mechanism, the dynamic trigger mechanism including forward prediction and backward verification, determining the semantic extension direction of the anchor information according to the forward prediction, and backward verification judging whether there is a prerequisite chain supporting the anchor information; setting a credibility scoring strategy, the credibility scoring strategy including three-dimensional verification logic, determining the credibility of the anchor information according to logic complexity, language confidence, and text consistency, forming an anchor information sequence set, and sequentially splicing the anchor information sequence set to form a text result after editing the streaming media file.
[0005] Furthermore, the method extracts anchor information from the text content based on the knowledge graph, establishes support relationships, and determines the main anchor, sub-anchor and argument anchor based on the support relationships and anchor information, including: dividing the anchor information into main anchor, sub-anchor and argument anchor according to the hierarchical structure of the knowledge graph, wherein the main anchor corresponds to the top-level core concept in the knowledge graph, the sub-anchor corresponds to the middle-level branch concept, and the argument anchor corresponds to the bottom-level data instance; establishing support relationships between the anchor information, each main anchor supports at least two sub-anchors, each sub-anchor is associated with at least one argument anchor, and the support relationships between the anchor information are verified by the knowledge graph rule engine through the relationship confidence between entities in the knowledge graph.
[0006] Furthermore, the semantic extension direction of the anchor information is determined according to the forward prediction, including: based on the entity association relationship in the knowledge graph, mapping the subject words and corresponding opinion directions in the text content to the entity association relationship in the knowledge graph, converting the subject words and opinion directions into semantic vectors through the BERT model, calculating the cosine similarity between the two, and confirming that the mapping is successful when the semantic similarity exceeds a preset threshold; generating a semantic extension direction based on the entity association relationship in the knowledge graph, and generating a path extending from the current anchor information to the associated entity after the semantic extension direction is successfully matched by the vector similarity between the subject words and the opinion direction.
[0007] Furthermore, the backward verification determines whether there is a prerequisite chain supporting the anchor point information, including: when it is detected that the keyword in the text content generates candidate anchor point information, setting a backtracking threshold, and backtracking the previous sentences of the candidate anchor point information step by step according to the backtracking threshold to verify whether there is a prerequisite chain supporting the viewpoint.
[0008] Furthermore, the three-dimensional verification logic determines the credibility of the anchor information based on logic complexity, language confidence and text consistency, including: the first dimension of logic complexity, which identifies the length of the prerequisite chain of the anchor information by analyzing the logic and calculating the logic depth value; the second dimension of language confidence, which evaluates the degree of certainty of the expression based on text sentiment analysis and calculates the professional score based on the frequency of occurrence of domain terms; the third dimension of text consistency, which calculates the semantic similarity score between the anchor information and the previous sentence.
[0009] Furthermore, forming a sequence set of the anchor information includes: using a dynamic weight allocation algorithm to assign a weight coefficient to each anchor information, the weight coefficient is dynamically determined based on the importance score of the entity in the knowledge graph and the position of the anchor information in the supporting relationship, the logical complexity weight of the main anchor is higher than that of the argument anchor, and the text consistency weight of the argument anchor is higher than that of the main anchor.
[0010] Furthermore, the credibility scoring strategy also includes dividing the anchor credibility level according to the credibility score of the anchor information and the evidence support of the conflicting anchor: calculating the credibility score by analyzing the logical complexity, language confidence and text consistency of the anchor information, and calculating the difference between the supporting evidence and the refuting evidence as the conflict support based on the knowledge graph rule engine; setting a first credibility threshold and a first conflict support threshold, when the anchor information simultaneously meets the credibility score exceeding the first credibility threshold and the conflict support exceeding the first conflict support threshold, and there are at least two independent premise chains, it is divided into a high credibility level; setting a second credibility threshold and a second conflict support threshold lower than the first threshold, when the anchor information meets the credibility score exceeding the second credibility threshold or the conflict support exceeding the second conflict support threshold, and there is a complete premise chain or multiple partially supported evidence chains, it is divided into a medium credibility level; when the anchor information does not meet the above-mentioned high credibility or medium credibility level division conditions, it is divided into a low credibility level.
[0011] Furthermore, the evaluation of the evidence support of the conflict anchor points includes: determining logical contradiction type conflicts through the knowledge graph attribute assertion comparison algorithm, determining insufficient evidence type conflicts by verifying the integrity of the premise conditions using the knowledge graph rule engine, and determining cross-modal inconsistency type conflicts by calculating the semantic vector similarity of the video image and audio text content using the CLIP model; constructing a three-layer analysis framework based on argumentation game theory, in which the advocacy layer extracts the core views of the conflicting parties through the BERT model, the rebuttal layer uses dependency syntax analysis to identify the causal negation or factual opposition relationship between views, and the defense layer adopts evidence theory to integrate the confidence of multi-source arguments, calculate the difference in evidence support between support and refutation, and trigger the retention of anchor information when the difference exceeds the preset threshold.
[0012] Furthermore, the text result of the streaming media file clipping is formed by sequentially splicing the anchor information sequence set, including: setting a video segment retention strategy according to the anchor credibility level division result, the segment corresponding to the high credibility anchor is completely retained as the core content, the segment corresponding to the medium credibility anchor adopts the TextRank algorithm to compress the key information, and the segment corresponding to the low credibility anchor is replaced with a conflict mark or deleted; using the knowledge graph rule engine to parse the causal, temporal and hierarchical relationship between the anchor points, and determining the timeline arrangement order of the video segments according to the support relationship between the anchor points.
[0013] Furthermore, the construction of the knowledge graph includes: collecting data information, identifying entities in the data information through the BERT model, and applying the remote supervision relationship extraction algorithm to extract the semantic relationship between entities, wherein the semantic relationship includes causal relationship, hierarchical relationship, temporal relationship and attribute relationship; establishing an entity alignment mechanism to solve the entity ambiguity problem in multi-source data by calculating the semantic similarity between entities; mapping the identified entities into graph nodes, mapping the extracted relationships into graph edges, and mapping the entity attributes into node or edge attributes.
[0014] Compared with the prior art, the present invention has the following beneficial effects:
[0015] (1) Extract anchor information through a dynamic trigger mechanism, establish support relationships, identify the main anchor, sub-anchor, and argument anchor of the text content, improve the depth of semantic understanding, and arrange the clips according to the anchor sequence set;
[0016] (2) Credibility scoring is performed based on logic complexity, language confidence, and text consistency. Combined with a dynamic weight allocation algorithm, the weights of each dimension are adjusted according to the position of the anchor information in the support relationship, thereby improving the accuracy of anchor credibility assessment;
[0017] (3) Based on argumentation game theory, the credibility of conflicting anchor points is dynamically adjusted, video frame numerical evidence and audio emotional evidence are integrated, and anchor point correction or deletion is triggered by calculating the difference in evidence support, thus avoiding semantic contradictions in the clips and improving the consistency and reliability of the clip content;
[0018] (4) Based on the dynamic time warping algorithm to achieve audio and video synchronization, the clip priority and arrangement order are determined according to the anchor point credibility and evidence support, and the transition effect is automatically generated, which improves the smoothness of the edited clips and the logical coherence of the editing results. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a workflow diagram of the automatic video editing method based on semantic analysis;
[0020] Figure 2 Anchor information extraction flow chart;
[0021] Figure 3 is a schematic diagram of anchor point information;
[0022] Figure 4 Schematic diagram of splicing according to the anchor sequence set. DETAILED DESCRIPTION
[0023] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0024] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0025] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0026] The term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " generally indicates an "or" relationship between the related objects.
[0027] This embodiment introduces a video automatic editing method based on semantic analysis. Figure 1 Shown, including:
[0028] S101, obtaining a streaming media file, and extracting text content from the streaming media file;
[0029] S102: Build a knowledge graph, extract anchor information from the text content based on the knowledge graph, establish a support relationship, and determine the main anchor, sub-anchor, and argument anchor based on the support relationship and anchor information;
[0030] S103: Construct a dynamic trigger mechanism, wherein the dynamic trigger mechanism includes forward prediction and backward verification, wherein the semantic extension direction of the anchor information is determined based on the forward prediction, and the backward verification determines whether there is a prerequisite chain supporting the anchor information;
[0031] S104: Setting a credibility scoring strategy, wherein the credibility scoring strategy includes three-dimensional verification logic. The three-dimensional verification logic determines the credibility of the anchor information based on logic complexity, language confidence, and text consistency, forms a sequence set of anchor information, and sequentially splices the sequence set to form a text result after the streaming media file is clipped.
[0032] This embodiment obtains a streaming media file and extracts its textual content as a semantic carrier for analysis. It then extracts anchor information from the textual content, categorizing it into primary anchors, sub-anchors, and argument anchors. A credibility scoring strategy is then established to generate a sequence of anchor information, ensuring that subsequent clips possess logical relationships. A dynamic triggering mechanism is constructed when extracting anchor information. This mechanism, through a two-way verification method involving forward prediction and backward validation, ensures the rationality of the semantic extension direction of the anchor information and the sufficiency of the prerequisites, thus avoiding semantic deviations. Furthermore, the reliability of the anchor information is quantified along three dimensions: logical complexity, language confidence, and textual consistency. The order in which the audio and video segments of the streaming media file are spliced is determined based on the anchor credibility level.
[0033] In this embodiment, the streaming media file is first obtained and the text content is extracted to complete the preprocessing. FFmpeg is used to separate the audio and video tracks. In the audio track, the WebRTC algorithm is used to reduce noise. The audio is converted to text with the help of the Whisper model. The speech segments are segmented by combining the difference of Mel spectrum features. The audio and video timestamps are aligned by the dynamic time warping algorithm, and the term consistency is corrected based on the knowledge graph. Furthermore, a knowledge graph is constructed, multi-source data such as academic papers and industry reports are collected, entities are identified by the BERT-BiLSTM-CRF model, causal relationships and temporal relationships are extracted using a graph neural network, and after aligning the multi-source entities by cosine similarity, a knowledge graph containing entities and relationships between entities is formed. Furthermore, forward prediction is used to convert the subject words of the text content into BERT vectors, calculate the cosine similarity with the entity vectors in the knowledge graph, and determine whether the subject words in the anchor information are successfully mapped to the entities in the knowledge graph; backward verification generates candidate anchor information by detecting keywords in the anchor information, backtracks to build a premise chain, and verifies the evidence support in the premise chain through the knowledge graph rule engine. This embodiment provides an implementable rule engine, Drools. Drools supports rule definition, reasoning, and execution, and can be combined with knowledge graphs to perform complex logical reasoning, such as calculating support for conflicting evidence. Furthermore, a credibility scoring strategy is established, using a three-dimensional assessment of logical complexity, language confidence, and text consistency. Anchor credibility levels are assigned based on the scoring results and evidence support. Clips are arranged in chronological order according to their support relationships, resulting in editing results and automatic insertion of transitions.
[0034] Furthermore, the extracting of anchor information from the text content based on the knowledge graph, establishing a support relationship, and determining a main anchor, a sub-anchor, and an argument anchor in combination with the support relationship and the anchor information include:
[0035] S201. According to the hierarchical structure of the knowledge graph, the anchor information is divided into a main anchor, a sub-anchor, and an argument anchor, wherein the main anchor corresponds to the top-level core concept in the knowledge graph, the sub-anchor corresponds to the middle-level branch concept, and the argument anchor corresponds to the bottom-level data instance;
[0036] S202. Establish a support relationship between the anchor point information, where each main anchor point supports at least two sub-anchor points, and each sub-anchor point is associated with at least one argument anchor point. The support relationship between the anchor point information is verified by the knowledge graph rule engine through the relationship confidence between entities in the knowledge graph.
[0037] In the process of constructing the knowledge graph, this embodiment first collects multi-source domain data and pre-processes it, and uses the BERT-BiLSTM-CRF model to identify entities in the text content and map them to graph nodes, where the top-level core concept entity serves as the main anchor corresponding to the main node, the middle-level branch concept entity serves as the sub-anchor corresponding to the branch node, and the bottom-level data instance entity serves as the argument anchor corresponding to the leaf node. The top-level core represents the highest-level entity node in the knowledge graph, representing the core theme or key category of domain knowledge, and is the logical starting point and overarching concept of the entire semantic system; the middle-level branch concept represents the intermediate-level entity node between the top-level core concept and the bottom-level data in the knowledge graph, and is a subdivision of the core concept or a branch argument; the bottom-level data instance represents the entity node at the bottom of the knowledge graph, representing specific experimental data, cases, phenomena and other verifiable factual information. As Figure 3 As shown. Furthermore, graph neural networks are used to extract causal relationships, temporal relationships, and hierarchical relationships between entities, and these relationships are set as edges. The cosine similarity of entity vectors is calculated to unify synonyms in multi-source data and achieve node alignment. For example, the semantic similarity between "fast charging" and "fast charging" is calculated through BERT vectors, and entities with the same name and different meanings are merged. Furthermore, the identified entity attributes are mapped to node or edge attributes, and a knowledge graph containing entity nodes and relationship edges is constructed. The rationality of the edges is verified by the knowledge graph rule engine, and the entity importance score is evaluated using the PageRank algorithm. The graph hierarchy is dynamically adjusted to ensure that core concepts are recognized first.
[0038] During the anchor information extraction phase, the BERT-BiLSTM-CRF model is first used to classify entity types based on the knowledge graph and calculate the PageRank score for each entity. The higher the score, the more core it is. Therefore, the top-level core concept entities with the top 10% scores are selected as primary anchors, the mid-level branch concept entities with scores between 10% and 30% are selected as sub-anchors, and the bottom-level data instance entities are selected as argument anchors. The importance threshold is dynamically adjusted to 1.2 times the average score. Furthermore, based on semantic similarity calculations, at least two associated sub-anchors of the primary anchor are found to determine the sufficiency of the logical branching. Using the BM25 algorithm and semantic retrieval algorithms, each sub-anchor is matched with at least one argument anchor to ensure the necessary supporting evidence.
[0039] Furthermore, the anchor support relationship is verified. The knowledge graph rule engine calculates the evidence support of the premise chain by propagating the relationship confidence. If the evidence support is greater than or equal to the set evidence support threshold, it means that the reasoning conclusion of the premise chain in the current knowledge graph is sufficiently reliable and can be determined as a support relationship that meets the reasoning requirements. The formula for relationship confidence is:
[0040]
[0041] Among them, Confidence(h,r,t) represents the relationship confidence, h represents the head entity, r represents the relationship between entities, t represents the tail entity, count(h,r,t) represents the number of occurrences of (h,r,t), and count(h,r,*) represents the total number of times the head entity h is connected to any tail entity through the relationship r.
[0042] The formula for evidence support is:
[0043]
[0044] Among them, Support(h1,r n ,t n ) represents the evidence support, n represents the length of the premise chain, i represents the level depth, counting from 1 to n, h i 、r i , t i They correspond to the head entity, relationship, and tail entity of the i-th link respectively.
[0045] Since the mean support for premise chains in a large number of historical data experiments is 0.65 and the standard deviation is 0.05, based on the statistical characteristics of similar premise chains in the knowledge graph, a minimum standard for premise chain credibility of 0.7 is set, denoted as the evidence support threshold. Evidence support is a measure of the overall reliability of the inference conclusion, derived by the knowledge graph rule engine through reasoning and calculation based on the confidence of the relationships in the premise chain. If the calculated evidence support is ≥ 0.7, the chain relationship is considered to have basic reliability within the knowledge graph logic and can serve as valid supporting evidence for the anchor information. If the evidence support is < 0.7, the premise chain is considered to have a loophole, such as insufficient relationship confidence, and additional evidence or correction of the anchor is required. For example, when verifying "sulfide electrolyte → improved energy density," if there is a chain relationship with a confidence of 0.8 for "sulfide electrolyte → improved ionic conductivity" and a confidence of 0.9 for "improved ionic conductivity → improved energy density," then the evidence support is 0.8 × 0.9 = 0.72 ≥ 0.7.
[0046] Furthermore, determining the semantic extension direction of the anchor point information according to the forward prediction includes:
[0047] S301. Based on the entity association relationships in the knowledge graph, the subject words and corresponding opinion directions in the text content are mapped to the entity association relationships in the knowledge graph, the subject words and opinion directions are converted into semantic vectors using the BERT model, and the semantic similarity between the two is calculated. When the semantic similarity exceeds a preset threshold, the mapping is confirmed to be successful.
[0048] S302. Generate a semantic extension direction based on the entity association relationship in the knowledge graph. After the vector similarity between the subject word and the viewpoint direction is successfully matched, the semantic extension direction generates a path extending from the current anchor point information to the associated entity.
[0049] In this embodiment, the forward prediction is represented as a semantic navigator based on the knowledge graph. Through the predefined entity association relationship and BERT semantic vector calculation, the semantic extension direction of the anchor information in the knowledge graph is set. The purpose of determining the semantic extension direction based on the forward prediction is to organize the illogical text content in an orderly manner according to the logical system of domain knowledge and avoid semantic confusion in the clips. Figure 2As shown in the figure, a viewpoint direction vector library is constructed. Its core purpose is to provide clear, domain-knowledge-logical guidance for semantic extension of anchor information, addressing the issues of disorganized semantic organization and lack of coherence in fragmented text. First, the knowledge graph is decomposed into core directions, such as technical principles, application scenarios, and development trends, based on domain knowledge logic. Each direction corresponds to a set of associated entities, such as "sulfide electrolytes" and "ionic conductivity" associated with the technical principles direction. The TransE algorithm is used to convert these entities into low-dimensional semantic vectors to quantify entity relationships. The associated entity vectors for each direction are weighted and averaged according to PageRank scores to form a viewpoint direction vector. Core entities with high PageRank scores are given higher weights, ensuring that the vectors contain both entity semantics and knowledge importance logic. Furthermore, entity mapping is performed on the subject words. The subject words in the audio text are converted into vectors using BERT. The knowledge graph retrieves the entities with the closest semantics and calculates their cosine similarity. Furthermore, direction prediction and path generation are performed. The similarity between the subject word vector and each viewpoint direction vector is calculated, and the direction with the highest similarity is selected. When the similarity exceeds a preset threshold, the semantic mapping is confirmed to be successful and guidance for semantic extension of the anchor information is provided. Based on the entity association types in the knowledge graph, an extension path is generated from the matching entities. Furthermore, dynamic threshold adjustment is performed. The similarity threshold is automatically adjusted based on the complexity of the knowledge to ensure mapping accuracy. For example, based on the predefined entity associations in the knowledge graph, a library of opinion direction vectors, such as technical principles, application scenarios, and development trends, is constructed. Each direction vector is generated by taking a weighted average of the associated entity vectors according to their PageRank scores. When a keyword such as "cryogenic performance" is detected in the text content, it is converted into a BERT vector, and its cosine similarity is calculated with each opinion direction vector, and then compared with a set average similarity threshold. If the cosine similarity between "cryogenic performance" and "technical challenges" exceeds the average similarity threshold, the mapping is confirmed to be successful, and the path "cryogenic performance → influencing factors → electrolyte properties → material improvement" is generated through the causal edges in the knowledge graph. This mechanism uses the knowledge graph to forward-predict the semantic extension direction of text content, reducing the semantic deviation rate, effectively ensuring the semantic coherence of video content, and providing guidance for the logical extension of anchor information.
[0050] Furthermore, the backward verification determines whether there is a prerequisite chain supporting the anchor point information, including:
[0051] When candidate anchor point information is generated by a keyword in the text content, a backtracking threshold is set, and the preceding sentences of the candidate anchor point information are backtracked step by step according to the backtracking threshold to verify whether there is a prerequisite chain supporting the viewpoint.
[0052] This embodiment implements a deep verification of the logical rationality of anchor information by constructing a backward verification module of a dynamic trigger mechanism. Figure 2As shown in the figure. When keywords, including conclusion words, sequence words, and explanation words, are detected in the text, candidate anchor point generation is automatically triggered. The N-gram sliding window is used to backtrack the preceding sentences of the candidate anchor point information step by step according to the backtracking threshold to extract the premise chain. The preceding text is searched for the existence of premises, and the causal relationship between these premises is verified using the knowledge graph. When the evidence support exceeds the evidence support threshold, the logic is judged to be reasonable. During this process, BERT is used to calculate the semantic similarity between the candidate anchor point and the premise, and the knowledge graph rule engine is used to verify the integrity of the premise chain, ensuring the semantic relevance between the premise and the conclusion.
[0053] Furthermore, the three-dimensional verification logic determines the credibility of the anchor information based on logic complexity, language confidence, and text consistency, including:
[0054] S401: First dimension logic complexity, by analyzing the length of the prerequisite chain of the logic recognition anchor information and calculating the logic depth value;
[0055] S402, the second dimension, language confidence, evaluates the degree of certainty of the expression based on text sentiment analysis and calculates the professionalism score based on the frequency of occurrence of domain terms;
[0056] S403, the third dimension is text consistency, which calculates the semantic similarity score between the anchor information and the previous sentence.
[0057] This embodiment implements a multi-dimensional quantitative evaluation of anchor information by setting a credibility scoring strategy and constructing a three-dimensional verification logic. The first dimension is the logic complexity. The premise chain of the anchor information is parsed by the knowledge graph rule engine, and the length of the premise chain and the average number of associated levels are calculated. The average number of associated levels is used to measure the hierarchical depth of the relationship between entities in the knowledge graph. If the chain length exceeds the average number of associated levels, the exponential decay algorithm is triggered to adjust the score to avoid the difficulty of semantic understanding caused by overly complex premise chains. Decay coefficient = e -0.2×(链长-平均层级数) For example, for the anchor point "sulfide electrolyte improves the energy density of solid-state batteries", extract "sulfide electrolyte → ion conductivity improvement → energy density improvement" and calculate the chain length to be 2 layers. When the chain length is 3 layers, the attenuation coefficient is substituted into the formula to get e -0.2×1.2 ≈0.78. At this point, the logical complexity score is adjusted from the base score of 80 to 80 × 0.78 = 62.4 points to avoid overly complex premise chains that make semantic understanding difficult.
[0058] Furthermore, the second dimension is language confidence, which integrates sentiment analysis and term recognition through natural language processing techniques. The implementation logic for this dimension is as follows: First, sentiment analysis models such as VADER are used to parse the text for deterministic and ambiguous terms, assigning different initial confidence weights to each term. Simultaneously, the BERT-BiLSTM-CRF model is used to identify terms, and a threshold for term frequency is set: at least two occurrences per 100 words, with a set increment. Each occurrence above that threshold increases the professionalism score by one increment. The language confidence score is calculated using the language confidence score formula: language confidence score = initial weight × (1 + amplitude). For example, the statement "sulfide electrolytes inevitably increase ionic conductivity" has an initial confidence weight of 0.9. The terms "sulfide electrolyte" and "ionic conductivity" appear twice, with an amplitude of 10%. The final score is 99 points, representing 0.9 × 1.1 = 0.99.
[0059] Furthermore, the third dimension is text consistency. A sliding window mechanism is used to extract the previous sentence. The anchor information and the previous sentence are converted into semantic vectors through the BERT model, and the cosine similarity between the two is calculated. Since a similarity threshold that is too high will miss content that is weakly semantically related but reasonably extended, and a similarity threshold that is too low will tolerate too many semantic deviations, this embodiment sets the similarity threshold to 0.7. If the calculated cosine similarity is higher than the threshold, it indicates a strong causal relationship; if the calculated cosine similarity is lower than the threshold, a semantic conflict warning is triggered and the anchor credibility is reduced, indicating a weak causal relationship. Taking the anchor information "Solid-state battery thermal runaway risk is low" as an example, by looking back at the previous text "Solid-state batteries have no liquid electrolyte, and sulfide electrolytes have high thermal stability", the cosine similarity of the two vectors calculated using the BERT model is 0.83. Since 0.83>0.7, the text consistency score is rated as 83 points. If the semantic similarity between the anchor information and the previous text is lower than 0.7, it indicates that there is a deviation between the two topics, and the semantic conflict warning mechanism is automatically triggered, and the credibility of the anchor information is adjusted accordingly.
[0060] Furthermore, a weighted fusion is performed based on the scores of the three-dimensional verification logic to determine the total credibility score of the anchor information. In this embodiment, the weights of the preferred logic complexity, language confidence, and text consistency are set to 40%, 30%, and 30%, respectively. This ratio meets the priority requirements of video editing for content "logical rigor > accurate expression = semantic coherence", ensuring that the screened anchor information can be used as a reliable editing basis. Taking the above scores of each dimension as an example, the total credibility score is calculated as 62.4×40%+99×30%+83×30%=79.56 points. When the total score is ≥75 points, the anchor information is identified as highly credible and can be used as the core basis for video editing; when 60 points ≤ total score <75 points, it is necessary to further verify the prerequisites or optimize the expression; when the total score is <60 points, it is determined to be of low credibility.
[0061] Furthermore, forming a sequence set of the anchor point information includes:
[0062] A dynamic weight assignment algorithm is used to assign a weight coefficient to each anchor information. The weight coefficient is dynamically determined based on the importance score of the entity in the knowledge graph and the position of the anchor information in the supporting relationship. The logical complexity weight of the main anchor is higher than that of the argument anchor, and the text consistency weight of the argument anchor is higher than that of the main anchor.
[0063] This embodiment forms a sequence set of anchor information using a dynamic weighting algorithm. First, the importance score of the anchor information is determined based on the PageRank scores of the entities in the knowledge graph. Since the main anchor corresponds to the top-level core concept, its PageRank score is typically in the top 10%, while the argument anchor corresponds to the mid-level branch concept, with a score in the top 10%-30%. Based on this, the weight coefficients are dynamically adjusted based on the anchor information's position in the supporting relationship: the main anchor, as the core of the premise chain, has its logical complexity weight increased to 45%, higher than the 35% of the argument anchor, to ensure the logical rigor of the core idea. The argument anchor, as the evidence supporting the main anchor, has its textual consistency weight increased to 35%, higher than the 30% of the main anchor, to ensure semantic coherence between the evidence and the preceding argument. For example, the main anchor "solid-state battery energy density advantage" has a logical complexity weight of 45%, a language confidence of 30%, and a textual consistency of 25%. Its argument anchor "sulfide electrolyte ion conductivity improvement" has a logical complexity weight of 35%, a language confidence of 30%, and a textual consistency of 35%. Through this dynamic weight distribution mechanism, a hierarchical anchor sequence set is formed, which enables differentiated optimization of core ideas and supporting arguments in terms of logical depth and semantic coherence, ultimately improving the logical hierarchy of video content.
[0064] Furthermore, the credibility scoring strategy further includes dividing the anchor credibility levels according to the credibility scores of the anchor information and the evidence support of the conflicting anchors, including:
[0065] S501. Calculate the credibility score by analyzing the logical complexity, language confidence, and text consistency of the anchor information, and calculate the difference between the supporting evidence and the refuting evidence as the conflict support based on the knowledge graph rule engine;
[0066] S502: Set a first credibility threshold and a first conflict support threshold. When the anchor point information simultaneously satisfies the following conditions: the credibility score exceeds the first credibility threshold, the conflict support exceeds the first conflict support threshold, and there are at least two independent premise condition chains, the information is classified as high credibility.
[0067] S503: Set a second credibility threshold and a second conflict support threshold that are lower than the first threshold. When the anchor information satisfies the credibility score exceeding the second credibility threshold or the conflict support exceeds the second conflict support threshold, and there is a complete premise chain or multiple partially supported evidence chains, classify it as a medium credibility level.
[0068] S504: When the anchor point information does not meet the above-mentioned high credibility or medium credibility classification conditions, it is classified into a low credibility level.
[0069] This embodiment constructs an anchor sequence set by integrating the anchor credibility score and the conflicting evidence support. The knowledge graph rule engine is used to evaluate the evidence support difference of the conflicting anchor, that is, the difference in confidence between the supporting evidence and the refuting evidence. The knowledge graph can organize the evidence supporting and refuting the anchor information into structured information. After scoring these evidences with the knowledge graph rule engine, the difference can be calculated to reflect the degree of supporting evidence and refuting evidence. When calculating the evidence support of the conflicting anchor, based on the knowledge graph rule engine and N-gram sliding window technology, all evidence related to the anchor information is extracted from the text content, and the semantic correlation between the evidence and the anchor information is calculated through the BERT model, and invalid evidence with a correlation below the preset similarity threshold is deleted. At the same time, the evidence confidence is evaluated for valid evidence: on the one hand, the initial confidence is assigned based on the entity relationship in the knowledge graph, such as 0.8 for strong causal relationships and 0.4 for weak correlation relationships; on the other hand, the evidence confidence is weighted and adjusted in combination with the language confidence score of the anchor information. The formula for the confidence level of a single piece of evidence is: confidence level of a single piece of evidence = initial confidence level × language confidence level. Furthermore, when there are multiple pieces of similar evidence, the Dempster-Shafer evidence theory is used to perform a fusion calculation to improve the accuracy of the confidence level of the evidence. The fused confidence level is the result of integrating and calculating the confidence levels of multiple pieces of similar evidence using the Dempster-Shafer evidence theory. When there are two or more pieces of similar evidence, the correlation and conflict between the propositions of each piece of evidence are comprehensively considered, and the confidence levels of multiple pieces of evidence are merged to weaken contradictions and strengthen consistent parts. This allows the fused confidence level to more accurately reflect the overall support or refutation strength of this type of evidence for the anchor point, improving the accuracy of the assessment and providing a more reliable basis for the subsequent calculation of the support level of conflicting evidence and the judgment of the logical relationship between anchor points.
[0070] Furthermore, anchor information with a total credibility score of not less than 75 points and a conflict support of 0.5 or above is determined to be a high-credibility anchor. This type of anchor information is used to construct a complete premise chain, including the main viewpoint, sub-viewpoints and corresponding arguments. The criteria for determining a medium-credible anchor are that the total credibility score is between 60 and 75 points, or the conflict support is between 0.2 and 0.5. A content compression mechanism is activated for such anchors, and the TextRank algorithm is used to extract the core elements of the viewpoint and key evidence fragments. For anchors with a total credibility score below 60 points or a conflict support of less than 0.2, they are classified as low-credibility anchors and replaced with conflict annotation prompt boxes. At the same time, multimodal evidence such as experimental data charts and literature comparison texts are automatically retrieved to generate a visual comparison diagram. For example, for the contradictory views on the high cost of solid-state batteries and the improvement of energy density, the evidence comparison between the two is displayed in the form of a curve to help users quickly identify logical conflict points. This mechanism ensures the logical integrity of core ideas through a credibility level classification strategy, provides visual conflict warnings for low-credibility content, and realizes intelligent stratification and logical optimization of video content.
[0071] Furthermore, the evidence support for evaluating the conflict anchor point includes:
[0072] S601. Determine logical contradiction conflicts using a knowledge graph attribute assertion comparison algorithm, determine insufficient evidence conflicts using a knowledge graph rule engine to verify the integrity of prerequisites, and determine cross-modal inconsistency conflicts using a CLIP model to calculate the semantic vector similarity between the video image and the audio text content.
[0073] S602. A three-layer analysis framework is constructed based on argumentation game theory. The advocacy layer uses the BERT model to extract the core views of the conflicting parties. The rebuttal layer uses dependency syntax analysis to identify causal negation or factual opposition between views. The defense layer uses evidence theory to integrate the confidence of multi-source arguments and calculate the difference in the support of the evidence between support and refutation. When the difference exceeds the preset threshold, the anchor information is retained.
[0074] This embodiment defines the types of conflict anchors and combines them with argumentation game theory to achieve a systematic evaluation of the evidence support for conflict anchors. First, the conflict anchors are identified based on the knowledge graph and multimodal semantic analysis. For logical contradiction conflicts, the knowledge graph attribute assertion comparison algorithm is used to detect the mutual exclusion of attributes of the same entity. For example, "solid-state battery energy density 300Wh / kg" and "less than 200Wh / kg" are asserted at the same time. For insufficient evidence conflicts, the integrity of the premise is verified by the knowledge graph rule engine. For example, the anchor "sulfide electrolyte improves battery performance" lacks the preset premise of "ion conductivity improvement". Cross-modal inconsistency conflicts use the CLIP model to calculate the cosine similarity between the video picture features and the audio text semantic vector, set a multimodal threshold, and determine it as a modal conflict when the similarity is lower than the multimodal threshold.
[0075] Furthermore, a three-layer analysis framework was constructed based on argumentation game theory. At the assertion layer, BERTClassifier was used to extract the core viewpoints of both conflicting parties, such as "solid-state batteries are expensive" and "large-scale production can reduce costs." At the rebuttal layer, dependency syntax analysis was used to identify contradictory relationships between viewpoints, including causal negation and factual opposition. At the defense layer, Dempster Shafer evidence theory was employed to integrate the confidence of multiple sources of evidence. The evidence support difference was calculated using the formula: evidence support difference = total confidence of supporting evidence - total confidence of refuting evidence. When the evidence support difference exceeded the evidence support difference threshold, the anchor information was retained. When the evidence support difference fell below the evidence support difference threshold, the anchor information was corrected or marked for deletion, effectively filtering out low-quality anchor information.
[0076] Furthermore, the step of sequentially splicing the anchor point information sequence set to form a text result after the streaming media file is edited includes:
[0077] S701, setting a video segment retention strategy based on the anchor point credibility level classification result, wherein the segments corresponding to high-credibility anchor points are retained as core content, the segments corresponding to medium-credibility anchor points use the TextRank algorithm to compress key information, and the segments corresponding to low-credibility anchor points are replaced with conflict annotations or deleted;
[0078] S702. Use the knowledge graph rule engine to analyze the causal, temporal, and hierarchical relationships between anchor points, and determine the timeline arrangement order of the video clips based on the support relationship between the anchor points.
[0079] This embodiment realizes intelligent editing and text result generation of streaming media files through anchor point credibility level and support relationship analysis, and completes the timeline arrangement of video clips based on the support relationship between anchor points. Figure 4As shown. First, a video clip retention priority system is constructed based on the anchor credibility level. For high-credible anchors, that is, anchors with a total credibility score of not less than 75 points and a conflict support of 0.5 or above, the corresponding video clips are given the highest retention priority and are marked as core content. For example, in the processing of technical principle explanation clips about the commercial prospects of solid-state batteries, such anchor clips will be fully retained. For medium-credible anchors, if the total credibility score is between 60 and 75 points, or the conflict support is between 0.2 and 0.5, the corresponding clips are set as secondary content, and the compression mechanism is activated to retain key information, such as streamlining the experimental data display clips on the improvement of the low-temperature performance of solid-state batteries. For low-credible anchors, that is, anchors with a total credibility score of less than 60 points or a conflict support of less than 0.2, the corresponding clips are marked as content to be verified, and the system replaces them with conflict annotation prompt boxes, or directly deletes them, such as the processing of clips involving conflicting views on the cost and energy density of solid-state batteries.
[0080] Furthermore, the construction of the knowledge graph includes:
[0081] S801. Collect data information, identify entities in the data information using the BERT model, and apply a remotely supervised relationship extraction algorithm to extract semantic relationships between entities. The semantic relationships include causal relationships, hierarchical relationships, temporal relationships, and attribute relationships.
[0082] S802: Establish an entity alignment mechanism to resolve entity ambiguity in multi-source data by calculating semantic similarity between entities.
[0083] S803: Map the identified entities to graph nodes, map the extracted relationships to graph edges, and map the entity attributes to the attributes of the nodes or edges.
[0084] This embodiment implements semantic parsing based on the principles of deep learning and weakly supervised learning. First, the bidirectional Transformer architecture of the BERT model is used to learn the deep semantic representation of the text through the masked language model and the next sentence prediction task, and the input data is mapped into a semantic vector. The entity boundary is then decoded through the CRF layer to achieve end-to-end recognition of the entity. When extracting relationships, a multi-instance learning framework of a remotely supervised relationship extraction algorithm is adopted. Entity pairs in knowledge bases such as Freebase are used as seeds, such as "sulfide electrolytes-improvement-energy density". Text content containing the entity pair is matched through heuristic rules; the CNN convolutional layer is used to extract local features of the sentence, and the attention mechanism is combined to weight the sentences in the multi-instance package, filtering out irrelevant statements such as "sulfide electrolytes are discussed in a laboratory environment", retaining valid evidence such as "sulfide electrolytes improve energy density by reducing interface impedance", and finally classifying the relationship type through the softmax layer. Furthermore, an entity alignment mechanism is established based on knowledge graph embedding technology to solve the entity ambiguity problem in multi-source data. The TransE algorithm is used to map entities in different data sources into a unified vector space, so that entity pairs that satisfy relational constraints, such as "solid-state battery" and "all-solid-state battery", are close in distance in the vector space. The semantic distance of entity vectors is calculated by cosine similarity, and a threshold is set as the alignment standard.
[0085] In summary, the present invention extracts anchor point information through forward prediction and backward verification, establishes a supporting relationship between the anchor point information, performs credibility scoring on the anchor point information, and solves the problem of mismatch between audio and video content caused by conflicting anchor points, thereby improving the logical coherence and natural fluency of video editing, and providing efficient and intelligent automated editing solutions for fields such as short video creation.
[0086] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which are all protected by the present invention.
Claims
1. A video automatic editing method based on semantic analysis, characterized in that: include: Obtaining a streaming media file, and extracting text content from the streaming media file; Constructing a knowledge graph, extracting anchor information from the text content based on the knowledge graph, establishing a support relationship, and determining a main anchor, sub-anchors, and argument anchors based on the support relationship and anchor information; Constructing a dynamic triggering mechanism, the dynamic triggering mechanism including forward prediction and backward verification, determining the semantic extension direction of the anchor information based on the forward prediction, and backward verification determining whether there is a precondition chain supporting the anchor information; A credibility scoring strategy is set, wherein the credibility scoring strategy includes three-dimensional verification logic. The three-dimensional verification logic determines the credibility of the anchor information based on logic complexity, language confidence, and text consistency, forms an anchor information sequence set, and sequentially splices the anchor information sequence set to form a text result after the streaming media file is clipped.
2. The automatic video editing method based on semantic analysis according to claim 1, characterized in that: The method of extracting anchor information from the text content based on the knowledge graph, establishing a support relationship, and determining a main anchor, a sub-anchor, and an argument anchor by combining the support relationship and the anchor information includes: According to the hierarchical structure of the knowledge graph, the anchor information is divided into main anchors, sub-anchors and argument anchors, where the main anchor corresponds to the top-level core concept in the knowledge graph, the sub-anchor corresponds to the middle-level branch concept, and the argument anchor corresponds to the bottom-level data instance; A support relationship between the anchor information is established, each main anchor supports at least two sub-anchors, each sub-anchor is associated with at least one argument anchor, and the support relationship between the anchor information is verified by the knowledge graph rule engine through the relationship confidence between entities in the knowledge graph.
3. The automatic video editing method based on semantic analysis according to claim 2 is characterized in that: Determining a semantic extension direction of the anchor point information according to the forward prediction includes: Based on the entity association relationship in the knowledge graph, the subject words and corresponding opinion directions in the text content are mapped to the entity association relationship in the knowledge graph, the subject words and opinion directions are converted into semantic vectors through the BERT model, and the semantic similarity between the two is calculated. When the semantic similarity exceeds a preset threshold, the mapping is confirmed to be successful; The semantic extension direction is generated based on the entity association relationship in the knowledge graph. After the vector similarity between the subject word and the viewpoint direction is successfully matched, the semantic extension direction generates a path extending from the current anchor point information to the associated entity.
4. The automatic video editing method based on semantic analysis according to claim 3 is characterized in that: The backward verification determines whether there is a prerequisite chain that supports the anchor point information, including: When candidate anchor point information is generated by a keyword in the text content, a backtracking threshold is set, and the preceding sentences of the candidate anchor point information are backtracked step by step according to the backtracking threshold to verify whether there is a prerequisite chain supporting the viewpoint.
5. The automatic video editing method based on semantic analysis according to claim 4 is characterized in that: The three-dimensional verification logic determines the credibility of the anchor information based on logic complexity, language confidence, and text consistency, including: The first dimension is logic complexity, which calculates the logic depth value by analyzing the length of the prerequisite chain of the logic recognition anchor information; The second dimension, language confidence, evaluates the degree of certainty of the expression based on text sentiment analysis and calculates the professionalism score based on the frequency of occurrence of domain terms. The third dimension is text consistency, which calculates the semantic similarity score between the anchor information and the previous sentence.
6. The automatic video editing method based on semantic analysis according to claim 5 is characterized in that: The sequence set forming the anchor point information includes: A dynamic weight assignment algorithm is used to assign a weight coefficient to each anchor information. The weight coefficient is dynamically determined based on the importance score of the entity in the knowledge graph and the position of the anchor information in the supporting relationship. The logical complexity weight of the main anchor is higher than that of the argument anchor, and the text consistency weight of the argument anchor is higher than that of the main anchor.
7. The automatic video editing method based on semantic analysis according to claim 6 is characterized in that: The credibility scoring strategy also includes dividing the anchor credibility levels according to the credibility scores of the anchor information and the evidence support of the conflicting anchors: The credibility score is calculated by analyzing the logical complexity, language confidence, and text consistency of the anchor information, and the difference between the supporting evidence and the refuting evidence is calculated as the conflict support based on the knowledge graph rule engine; A first credibility threshold and a first conflict support threshold are set, and when the anchor point information simultaneously satisfies the conditions that the credibility score exceeds the first credibility threshold and the conflict support exceeds the first conflict support threshold, and there are at least two independent premise condition chains, it is classified as a high credibility level; A second credibility threshold and a second conflict support threshold are set, which are lower than the first threshold. When the anchor information satisfies the credibility score exceeding the second credibility threshold or the conflict support exceeds the second conflict support threshold, and there is a complete premise chain or multiple partially supported evidence chains, it is classified as a medium credibility level. When the anchor point information does not meet the above-mentioned classification conditions of high credibility or medium credibility level, it is classified as low credibility level.
8. The automatic video editing method based on semantic analysis according to claim 7 is characterized in that: Evaluating the evidential support for conflict anchors includes: The knowledge graph attribute assertion comparison algorithm is used to determine logical contradiction conflicts. The knowledge graph rule engine is used to verify the completeness of the premise to determine insufficient evidence conflicts. The CLIP model is used to calculate the semantic vector similarity between the video image and the audio text content to determine cross-modal inconsistency conflicts. A three-layer analysis framework is constructed based on argumentation game theory. The advocacy layer uses the BERT model to extract the core views of the conflicting parties. The rebuttal layer uses dependency syntax analysis to identify causal negation or factual opposition between views. The defense layer uses evidence theory to integrate the confidence of multi-source arguments and calculate the difference in evidence support between support and refutation. When the difference exceeds the preset threshold, the anchor information retention is triggered.
9. The automatic video editing method based on semantic analysis according to claim 8, characterized in that: The step of sequentially splicing the anchor point information sequence set to form a text result after the streaming media file is clipped includes: The video segment retention strategy is set based on the anchor credibility level classification results. The segments corresponding to high-credibility anchors are retained as core content. The segments corresponding to medium-credibility anchors use the TextRank algorithm to compress key information. The segments corresponding to low-credibility anchors are replaced with conflicting annotations or deleted. The knowledge graph rule engine is used to analyze the causal, temporal and hierarchical relationships between anchor points, and the timeline arrangement order of video clips is determined based on the supporting relationship between anchor points.
10. The automatic video editing method based on semantic analysis according to claim 1, characterized in that: The construction of the knowledge graph includes: Collect data information, identify entities in the data information through the BERT model, and apply the remote supervision relationship extraction algorithm to extract the semantic relationships between entities. The semantic relationships include causal relationships, hierarchical relationships, temporal relationships, and attribute relationships; Establish an entity alignment mechanism to resolve entity ambiguity in multi-source data by calculating the semantic similarity between entities; The identified entities are mapped to graph nodes, the extracted relationships are mapped to graph edges, and the entity attributes are mapped to the attributes of the nodes or edges.
Citation Information
Patent Citations
Automatic video editing method based on semantic recognition
CN112784078A
Video post-editing and video synthesis optimization method
CN116847123A
Video synthesis method and system based on semantic analysis and storage medium
CN118200665A
Cross-modal video clip retrieval method based on dual sequence diagrams
CN119597967A
High-quality video content automatic generation method and related equipment
CN120050487A
Cited By
Video understanding method and system based on multi-mode evidence chain
CN121147827A
A video understanding method and system based on multi-modal evidence chain
CN121147827B
Video intelligent editing quality test method and system based on multi-modal dynamic evaluation
CN121585861A
Event causal reasoning method for intelligence analysis and electronic equipment
CN121960796A
Event causal reasoning method and electronic device for intelligence analysis
CN121960796B