Text generation method and device based on semantic anchor points, equipment and medium
By extracting and causally constraining semantic anchors in video frame sequences, we generate contextually coherent text descriptions, which solves the problem of lack of coherence in text descriptions in traditional methods.
Patent Information
- Application Number
- CN202510856317.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-03
AI Technical Summary
The attention mechanism in traditional methods has difficulty capturing long-distance semantic continuity in videos, resulting in a lack of contextual coherence in the generated text descriptions.
By obtaining the original video frame sequence, extracting semantic anchors and performing causal constraints, text information is generated.
Ensure that the generated text information has strong contextual coherence and can accurately describe the video content.
Smart Images

Figure CN120751210A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text generation, and in particular to a text generation method, device, equipment and medium based on semantic anchor points. Background Art
[0002] Extracting semantic information from dynamic visual content and generating natural language descriptions is a core goal in cross-modal AI. For example, in the fintech sector, it's necessary to extract semantic information from surveillance videos of over-the-counter transactions and generate natural language descriptions of the transaction process to facilitate analysis of abnormal behavior. Alternatively, in healthcare and elderly care, it's necessary to extract semantic information from surveillance videos of patients undergoing rehabilitation training and generate natural language descriptions of the training process to facilitate analysis of the standardization of movements. Traditional approaches rely primarily on end-to-end deep learning frameworks, using convolutional neural networks to extract video features and then combining them with recurrent neural networks to generate text.
[0003] The inventors realized that the attention mechanism in traditional methods has difficulty capturing long-distance semantic continuity in videos, resulting in a lack of contextual coherence in the generated descriptions. Summary of the Invention
[0004] The present invention provides a text generation method, apparatus, computer equipment and medium based on semantic anchor points to solve the technical problem that text information generated by existing text generation methods lacks contextual coherence.
[0005] In a first aspect, a text generation method based on semantic anchors is provided, comprising:
[0006] Acquire an original video and parse the original video to obtain an original video frame sequence;
[0007] Performing semantic anchor point extraction on the original video frame sequence to obtain a plurality of semantic anchor points, wherein one of the semantic anchor points corresponds to a core event in the original video;
[0008] Performing causal constraints on the plurality of semantic anchor points to obtain a semantic anchor point sequence;
[0009] Text information for describing the original video is generated based on the original video frame sequence and the semantic anchor point sequence.
[0010] In a second aspect, a text generation device based on semantic anchors is provided, comprising:
[0011] An acquisition module, configured to acquire an original video and parse the original video to obtain an original video frame sequence;
[0012] an extraction module, configured to perform semantic anchor point extraction on the original video frame sequence to obtain a plurality of semantic anchor points, wherein one of the semantic anchor points corresponds to a core event in the original video;
[0013] a constraint module, configured to perform causal constraints on the plurality of semantic anchor points to obtain a semantic anchor point sequence;
[0014] A generating module is used to generate text information for describing the original video based on the original video frame sequence and the semantic anchor point sequence.
[0015] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for generating text based on semantic anchors are implemented.
[0016] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned text generation method based on semantic anchors are implemented.
[0017] In the scheme implemented by the above-mentioned semantic anchor-based text generation method, device, computer equipment and storage medium, the original video is obtained and parsed to obtain an original video frame sequence; semantic anchor extraction is performed on the original video frame sequence to obtain multiple semantic anchors, wherein one semantic anchor corresponds to a core event in the original video; causal constraints are performed on the multiple semantic anchors to obtain a semantic anchor sequence; text information for describing the original video is generated based on the original video frame sequence and the semantic anchor sequence. In the present invention, semantic anchor extraction can be performed on the original video frame sequence to obtain multiple semantic anchors, and one semantic anchor corresponds to a core event, and then causal constraints are performed on the multiple semantic anchors to obtain a semantic anchor sequence, thereby ensuring contextual coherence. Finally, text information is generated based on the original video sequence and the semantic anchor sequence, and the generated text information has strong contextual coherence. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 This is a flow chart of a method for generating text based on semantic anchors in one embodiment of the present invention;
[0020] Figure 2 yes Figure 1 A schematic flow chart of a specific implementation of step S20;
[0021] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S22;
[0022] Figure 4 yes Figure 1 A schematic flow chart of a specific implementation of step S30;
[0023] Figure 5 yes Figure 1 A schematic flow chart of a specific implementation of step S40;
[0024] Figure 6 1 is a structural diagram of a text generation device based on semantic anchors in one embodiment of the present invention;
[0025] Figure 7 is a structural diagram of a computer device in one embodiment of the present invention;
[0026] Figure 8 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0028] The text generation method based on semantic anchors provided in an embodiment of the present invention can be applied to computer devices, including but not limited to desktop computers and notebooks. Among them, the original video is obtained and parsed to obtain an original video frame sequence; semantic anchors are extracted from the original video frame sequence to obtain multiple semantic anchors, wherein one semantic anchor corresponds to a core event in the original video; causal constraints are applied to the multiple semantic anchors to obtain a semantic anchor sequence; and text information for describing the original video is generated based on the original video frame sequence and the semantic anchor sequence, which can improve the contextual coherence of the generated text information. The present invention is described in detail below through specific embodiments.
[0029] See also Figure 1 As shown, Figure 1 A flowchart of a method for generating text based on semantic anchors provided in an embodiment of the present invention includes the following steps:
[0030] S10: Acquire an original video and parse the original video to obtain an original video frame sequence.
[0031] The semantic anchor-based text generation method provided by the present invention can be applied to various business fields, such as fintech and healthcare and elderly care. In fintech, text information can be generated based on a certain segment of surveillance video, making it easier to understand the events that occurred in that segment through text information. For example, in a banking system, a customer conducts a transaction with a teller at a counter. The transaction process lasts about 30 seconds. Then, text information can be generated based on the 30-second video. The text information is used to describe the transaction process between the teller and the customer, making it easier to identify any abnormal behavior.
[0032] In the medical, health and elderly care business, if a patient is undergoing rehabilitation training, a video clip of a certain part of the rehabilitation training can be extracted from the surveillance video, and then text information can be generated based on the video clip. The text information is used to describe the rehabilitation training process, making it easier for doctors to understand whether the user's movements are standard.
[0033] The original video can be a clip or a complete video. Generally speaking, based on timeliness, the original video is usually a short video, such as a video clip of ten or dozens of seconds, to facilitate the rapid generation of text information. For example, the original video can be a transaction clip during a transaction or a surgical procedure in a surgical video.
[0034] After obtaining the original video, the original video can be parsed to obtain key frames therein. Multiple key frames can then be combined into an original video frame sequence. In other words, the original video frame sequence contains multiple key frames. When extracting key frames, video frames from the original video can be extracted at fixed intervals, such as extracting one video frame every 1 second as a key frame. After obtaining multiple key frames, they can be standardized to uniformly size.
[0035] For example, for a 60-second surveillance video clip of an over-the-counter transaction, the video decoding engine is used to interpret the compressed video stream into an uncompressed image frame sequence, outputting continuous frames. Key frames are then extracted from the continuous frames at intervals of 1 second to obtain 60 key frames. The 60 key frames are standardized to adjust the size of all key frames to a fixed resolution (such as 224×224), and then the pixel values of multiple key frames are normalized.
[0036] S20: Perform semantic anchor point extraction on the original video frame sequence to obtain a plurality of semantic anchor points, wherein one semantic anchor point corresponds to a core event in the original video.
[0037] A semantic anchor corresponds to a core event. For example, a video shows a user inserting a bank card and entering a password, then entering the withdrawal amount, waiting for a while, taking money out of the ATM, and then taking out the card to exit the system. The semantic anchors that can be extracted from this video content are a1-inserting a bank card, a2-entering a password, a3-selecting the withdrawal amount, a4-taking out cash, and a5-exiting the system. That is, this video content contains a total of 5 core events, and 5 semantic anchors can be extracted accordingly.
[0038] Semantic anchors correspond to core events and are used to summarize the content of consecutive video frames. For example, if frames 1-5 in the original video frame sequence are password input, then a1 corresponds to video frames 1 to 5. If frames 6-15 are selection of withdrawal amount, then a2 corresponds to frames 6-15, and so on.
[0039] For example, if a video shows an elderly person walking normally, but suddenly their gait becomes abnormal, causing them to fall. After falling, the elderly person tries to get up on their own, but fails and then chooses to ask for help from others, then the semantic anchor points of this video content are a1-elderly person walking, a2-abnormal gait, a3-elderly person falling, a4-trying to get up, a5-failed to get up, and a6-asking for help from others. That is, there are a total of 6 core events, and correspondingly there are 6 semantic anchor points.
[0040] In some embodiments, as Figure 2 As shown, in step S20, semantic anchor point extraction is performed on the original video frame sequence to obtain multiple semantic anchor points, which specifically includes the following steps:
[0041] S21: performing feature extraction on each frame of the original video frame sequence to obtain multiple frame features;
[0042] S22: Perform cluster analysis on the plurality of frame features to obtain the plurality of semantic anchor points.
[0043] A 3D convolutional neural network model can be used to extract spatiotemporal features from raw video frames (such as financial transaction surveillance videos and elderly care videos) to obtain frame features. For example, in a financial scenario, spatial features such as ATM user gestures and card insertion can be extracted, as well as temporal features of the operation sequence.
[0044] Then, the feature similarity is optimized through the contrast loss function, and semantically similar frames are clustered into one cluster. For example, in the elderly care scene, the consecutive frames of "elderly people walking" are clustered into the same cluster due to the similar motion features, and the "falling" frames are clustered into another cluster. The number of anchor points is automatically determined by the clustering algorithm, and the temporal continuity constraint is added to ensure that each cluster corresponds to a continuous event segment in the video (such as "insert card → enter password → withdraw money" in financial transactions are each an independent continuous segment). For each cluster, the semantic anchor point is calculated, and each anchor point corresponds to a core event. For example, in the financial scene, anchor points such as "insert bank card", "enter password", and "withdraw cash" are generated.
[0045] In some embodiments, as Figure 3 As shown, in step S22, cluster analysis is performed on the plurality of frame features to obtain the plurality of semantic anchor points, which specifically includes the following steps:
[0046] S221: Optimizing the plurality of frame features using a preset contrast loss function, wherein the preset contrast loss function is:
[0047]
[0048] L contrast is the total loss value, s(f i ,f j ) is the cosine similarity between frame feature i and frame feature j, s(f i ,f k ) is the cosine similarity between frame feature i and frame feature k, and τ is the temperature hyperparameter.
[0049] S222: performing cluster analysis on the optimized frame features using a preset clustering algorithm to obtain a plurality of clusters, wherein each cluster includes a plurality of temporally continuous frame features;
[0050] S223: Calculate the intra-cluster feature mean of each cluster to obtain the plurality of semantic anchor points.
[0051] The preset contrast loss function is used to encourage temporally adjacent and semantically similar video frame features (f i ,f j ) aggregates while pushing away irrelevant frame features (f i ,f k ). For frames of the same semantic event (such as the “withdraw money” segment in the transaction video), the high similarity score s(f i ,f k ) is close to 1 to reduce the loss and force feature clustering. For frames of different semantic events (such as "taking out money" and "entering password"), the low similarity score s(f i ,f k) close to -1, increasing the loss and avoiding incorrect clustering. Then, combined with a pre-defined clustering algorithm (such as the dynamic K-means algorithm), the frames in each cluster are required to be temporally continuous (for example, frames from 0 to 10 seconds can only belong to one anchor point). This ensures that the anchor points correspond to complete event segments rather than fragmented features. For example, in an ATM transaction video, the features of consecutive frames of the "card insertion" action have high similarity, resulting in a lower calculated loss value, while the loss values for the "card insertion" and "withdrawal" frames are higher.
[0052] For example, we collect ATM operation videos (such as user card insertion, password entry, withdrawal, and card return), extract key frames at fixed intervals (such as 1 second / frame), and use a 3D convolutional neural network to extract the frame features of each frame. For example, we extract the 5th frame feature f5 (including spatial features such as card insertion angle, hand movement trajectory, and temporal features of the action sequence) and the 8th frame feature f8 of the "card insertion" action. The similarity between the consecutive frames f5 and f8 of "card insertion" is s(f5,f8) = 0.85 (indicating high feature similarity); the similarity between the "card insertion" frame f5 and the "password entry" frame f 15 The similarity s(f5,f 15 ) = 0.3 (large semantic differences in actions), then exp(s(f5,f8) / τ) ≈ 5.47. Then, the sum of the similarity indices of f5 and all other frame features is calculated. Assuming the sum of the similarity indices of f5 and all other frame features is 10, then -log(5.47 / 10) ≈ 0.6, which means that the loss value of (f5,f8) is 0.6. The loss value of any two frame features can be calculated using the same method, and then the total loss value is obtained by summing all the loss values. If the temperature hyperparameter τ is reduced (for example, to 0.3), the similarity difference is amplified, forcing the model to strictly distinguish different actions such as "inserting a card" and "entering a password" (suitable for the security-sensitive requirements of financial scenarios); if the temperature hyperparameter τ is increased (for example, to 1.0), the loss function is less sensitive to similarity differences and allows a certain degree of overlap in cross-action features (suitable for processing noise such as camera angle changes). The contrast loss function optimizes the frame features of the same action (such as "inserting a card") to aggregate in the feature space, reducing the loss value; the frame features of different actions are dispersed, and the loss value increases.
[0053] After optimizing the preset contrast loss function, cluster analysis can be performed using a preset clustering algorithm, such as the dynamic K-means algorithm. The algorithm automatically determines the optimal number of clusters, K, and adds temporal continuity constraints to ensure that the frames within each cluster are temporally continuous. For example, the action segment "insert card → enter password → withdraw cash" in a financial video can be segmented into independent, continuous clusters.
[0054] After obtaining the cluster, the semantic anchor point can be calculated using the following formula:
[0055]
[0056] Among them, a k is the kth semantic anchor point, representing the high-level semantic representation of the kth core event in the video (such as "preparing ingredients" and "cutting vegetables", etc.), which is generated by clustering the semantically consistent and temporally continuous frame features in the video frame sequence. k is the set of all frame features of the kth cluster, where each frame feature f i Belong to the same semantic event and are continuous in time (such as the continuous frames of the “cutting vegetables” action), |C k | is the number of frame features in the kth cluster.
[0057] For example, through cluster analysis, four continuous clusters of "inserting card", "entering password", "withdrawing money", and "returning card" are generated. The timing constraint ensures that each cluster corresponds to a complete action segment. The corresponding semantic anchor points are a1: the feature mean of the card insertion action (such as the spatiotemporal features of the card insertion), a2: the feature mean of the keyboard pressing when entering the password, a3: the feature mean of the withdrawal action, and a4: the feature mean of the card return action.
[0058] S30: Perform causal constraints on the multiple semantic anchor points to obtain a semantic anchor point sequence.
[0059] Causal constraints are used to establish dependencies between semantic anchor sequences, ensuring logical correctness between different semantic anchors. For example, in a transaction, the order of semantic anchors should be "insert card → enter password → withdraw money" rather than "insert card → withdraw money → enter password." In a patient fall, the order of semantic anchors should be "patient stands up → unsteady gait → fall" rather than "patient stands up → fall → unsteady gait."
[0060] In some embodiments, as Figure 4 As shown, in step S30, causal constraints are applied to the plurality of semantic anchor points to obtain a semantic anchor point sequence, which specifically includes the following steps:
[0061] S31: constructing a directed acyclic graph based on the plurality of semantic anchor points to generate an initial semantic anchor point sequence representing a causal relationship between the semantic anchor points;
[0062] S32: Modeling a generating function for each of the semantic anchor points in the initial semantic anchor point sequence, wherein the generating function includes exogenous noise and a previous semantic anchor point;
[0063] S33: Calculate the loss value of the exogenous noise and the previous semantic anchor point through a preset regularizer, and train the preset regularizer through adversarial training to minimize the loss value to construct the semantic anchor point sequence, wherein the preset regularizer is:
[0064]
[0065] L causalis the loss value, HSIC is the preset regularizer, ε k is the exogenous noise, a k` is the previous semantic anchor, a k is the current semantic anchor.
[0066] Using semantic anchors as nodes, directed edges are established based on event timing and logic. For example, in a financial transaction, the "insert card" anchor a1 points to the "enter password" anchor a2, indicating that the former is the prerequisite for the latter. The graph structure ensures that the causal chain is loop-free, such as in a medical scenario, "patient stands up" a1 → "unsteady gait" a2 → "fall" a3, avoiding logical loops. The causal modeling generated by the anchor points is generated by each anchor point a k The generation depends on the previous anchor point a k` and independent noise ε k , modeled as:
[0067] a k =g(a k` ,ε k )(k' <k)
[0068] Where g is a two-layer MLP that learns causal mapping (e.g., the semantics of "entering a password" depends on "successful card insertion" in finance). The default regularizer can be the HSIC (Hilbert-Schmidt Independence Criterion) regularizer, which is used to constrain the independence requirement noise ε k Independent of the previous semantic anchors, the HSIC loss function is used to quantify dependencies. For example, in healthcare, the noise ε3 associated with a fall should be independent of the noise associated with a patient standing up, containing only the randomness of the falling action itself (e.g., the body's tilt angle). Finally, by generating independent noise and distinguishing dependencies, a game theory is used to ensure that the causal structure of the semantic anchor sequence is independent of surface features (e.g., card color in finance, or patient clothing in healthcare).
[0069] For example, the initial semantic anchor sequence is [a1 (insert card), a2 (withdraw), a3 (enter password), a4 (return card)]. It can be seen from the initial semantic anchor sequence that there is a temporal disorder, that is, a2 (withdraw) appears before a3 (enter password), violating the business logic that "entering password is a prerequisite for withdrawal"; there is also a causal gap, that is, there is only a temporal association between the anchor points, and no explicit causal dependency (for example, the necessary relationship of "insert card → enter password" is not modeled); noise interference: the generation of anchor point a2 may contain redundant information related to "insert card" (such as card type), leading to semantic confusion.
[0070] Construct a directed acyclic causal graph, that is, define (a1-a3-a2-a4), and enforce the causal chain of "insert card → enter password → withdraw money → return card". Then generate function modeling and use two-layer MLP to define a k =g(ak' ,ε k ), such as a3=g(a1,ε3). Optimize ε by presetting the regularizer k By eliminating irrelevant noise such as card type and the independence of previous anchor points, we obtain the optimized anchor point sequence [a1 (insert card), a3 (enter password), a2 (withdraw), a4 (return card)]. Causal constraints can ensure the temporal correctness of the generated text information.
[0071] S40: Generate text information for describing the original video based on the original video frame sequence and the semantic anchor point sequence.
[0072] After obtaining the semantic anchor sequence, text information can be generated based on the original video frame sequence and the semantic anchor sequence. For example, by calculating the attention weight between the original video frame sequence and the semantic anchor sequence, then calculating contextual features based on the attention weights, and then generating text information based on the contextual features, the generated text information has strong contextual coherence.
[0073] In some embodiments, as Figure 5 As shown, in step S40, that is, generating text information for describing the original video based on the original video frame sequence and the semantic anchor point sequence, specifically includes the following steps:
[0074] S41: Calculating the attention weight between each frame feature in the original video frame sequence and each semantic anchor point in the semantic anchor point sequence;
[0075] S42: Substituting the attention weight into a preset context feature calculation formula to obtain a context feature;
[0076] S43: Input the context feature into a preset decoder to obtain the text information.
[0077] The attention weight formula can be used to calculate the attention weight between the frame features in the original feature sequence and the semantic anchor points in the semantic anchor point sequence, where the attention weight formula is:
[0078]
[0079] Among them, f t is the frame feature in the original video frame sequence, which is used as the query object to match the semantic anchor point and locate the key event semantics corresponding to the current frame. k As the key in the key-value pair, it provides high-level semantic information and guides attention to focus on the event anchor related to the current frame. t ) maps the frame features to the query space and defines what semantics to look for in the current frame. K(a k) maps semantic anchors to key spaces and defines what semantic identifiers the anchors can provide. Q(f t ) and K(a k ) is multiplied to calculate the semantic similarity between the frame feature and the anchor point. The softmax function is a normalization function that converts the dot product result into a probability distribution to generate the attention weight a t,k , which indicates the degree of attention of the t-th frame to the k-th semantic anchor (the sum of the weights is 1).
[0080] The context feature calculation formula is:
[0081] F t =∑ k a t,k V(a k )
[0082] Among them, F t is the context feature corresponding to the t-th frame, which is a high-level representation that integrates visual details and semantic anchor information, and is used to influence the semantic accuracy and context coherence of text generation. t,k is the attention weight, V(a k ) is the projection feature of the semantic anchor, and V is a learnable linear transformation matrix used to map the high-level semantics of the semantic anchor to a space compatible with the frame features to facilitate feature aggregation. For example, in financial scenarios, V(a k ) can convert the semantics of the “enter password” anchor into a feature vector containing the “password input action”.
[0083] The preset decoder can be a Transformer decoder, which inputs context features into the preset decoder to generate text. The preset decoder generation process is guided by the semantic anchor causal chain to ensure that the text timing logic is correct. Then the text can be refined, time sequence words (such as "first" and "then") can be added, and natural language descriptions can be generated to obtain text information.
[0084] For example, data preprocessing: extract the elderly activity video frames, capture the joint motion features through a three-dimensional convolutional neural network, and generate a semantic anchor sequence A = [a1 (standing up), a2 (unstable gait, a3 (falling), a4 (calling)], with a causal chain of a1-a2-a3-a4. Then, the attention weight is calculated. The Q(f t ) projects the falling trajectory of the body, K(a3) projects the semantics of “body imbalance”, and we get a t,3 =0.85, a t,2 =0.1 (associated prodromal symptoms). Contextual feature F t= 0.85V(a3) + 0.1V(a2), integrating the details of the "falling to the left" action with the historical features of "abnormal stride length." This contextual feature is input into the preset decoder to generate preliminary text information, such as "The elderly person experienced an unstable gait and fell to the left." This preliminary text information is then refined to form the text information, namely, "The elderly person stood up at 10:20, had an abnormal gait at 10:22, and fell and called for care at 10:23."
[0085] It can be seen that in the above scheme, semantic anchor extraction can be performed on the original video frame sequence to obtain multiple semantic anchors, and one semantic anchor corresponds to a core event. Then, causal constraints are performed on the multiple semantic anchors to obtain a semantic anchor sequence, thereby ensuring the coherence of the context. Finally, text information is generated based on the original video sequence and the semantic anchor sequence, and the generated text information has strong context coherence.
[0086] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0087] In one embodiment, a text generation device based on semantic anchors is provided, and the text generation device based on semantic anchors corresponds to the text generation method based on semantic anchors in the above embodiment. Figure 6 As shown, the text generation device based on semantic anchors includes an acquisition module 101, an extraction module 102, a constraint module 103 and a generation module 104. The functional modules are described in detail as follows:
[0088] An acquisition module 101 is configured to acquire an original video and parse the original video to obtain an original video frame sequence;
[0089] An extraction module 102 is configured to extract semantic anchor points from the original video frame sequence to obtain a plurality of semantic anchor points, wherein one semantic anchor point corresponds to a core event in the original video;
[0090] A constraint module 103 is configured to perform causal constraints on the plurality of semantic anchor points to obtain a semantic anchor point sequence;
[0091] The generating module 104 is configured to generate text information for describing the original video based on the original video frame sequence and the semantic anchor point sequence.
[0092] In one embodiment, the extraction module 102 is specifically configured to:
[0093] Performing feature extraction on each frame of the original video frame sequence to obtain multiple frame features;
[0094] Cluster analysis is performed on the plurality of frame features to obtain a plurality of semantic anchor points.
[0095] In one embodiment, the extraction module 102 is further configured to:
[0096] Optimizing the plurality of frame features by using a preset contrast loss function;
[0097] Performing cluster analysis on the optimized frame features by a preset clustering algorithm to obtain a plurality of clusters, wherein each cluster includes a plurality of temporally continuous frame features;
[0098] The mean of the intra-cluster features of each cluster is calculated to obtain a plurality of the semantic anchor points.
[0099] In one embodiment, the constraint module 103 is specifically configured to:
[0100] Constructing a directed acyclic graph based on the plurality of semantic anchor points to generate an initial semantic anchor point sequence representing a causal relationship between the semantic anchor points;
[0101] Modeling a generating function of each semantic anchor point in the initial semantic anchor point sequence, wherein the generating function includes exogenous noise and previous semantic anchor points;
[0102] A loss value between the exogenous noise and the previous semantic anchor point is calculated by a preset regularizer, and the preset regularizer is trained through adversarial training to minimize the loss value to construct the semantic anchor point sequence.
[0103] In one embodiment, the generating module 104 is specifically configured to:
[0104] Calculating an attention weight between each frame feature in the original video frame sequence and each semantic anchor point in the semantic anchor point sequence;
[0105] Substituting the attention weight into a preset context feature calculation formula to obtain a context feature;
[0106] The context features are input into a preset decoder to obtain the text information.
[0107] The present invention provides a text generation device based on semantic anchors, which can extract semantic anchors from an original video frame sequence to obtain multiple semantic anchors, where one semantic anchor corresponds to a core event. Then, causal constraints are applied to the multiple semantic anchors to obtain a semantic anchor sequence, thereby ensuring contextual coherence. Finally, text information is generated based on the original video sequence and the semantic anchor sequence, and the generated text information has strong contextual coherence.
[0108] For the specific definition of the text generation device based on semantic anchors, please refer to the definition of the text generation method based on semantic anchors above, which will not be repeated here. The various modules in the above-mentioned text generation device based on semantic anchors can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0109] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a text generation method based on semantic anchors.
[0110] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the client side of a text generation method based on semantic anchors.
[0111] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0112] Acquire an original video and parse the original video to obtain an original video frame sequence;
[0113] Performing semantic anchor point extraction on the original video frame sequence to obtain a plurality of semantic anchor points, wherein one of the semantic anchor points corresponds to a core event in the original video;
[0114] Performing causal constraints on the plurality of semantic anchor points to obtain a semantic anchor point sequence;
[0115] Text information for describing the original video is generated based on the original video frame sequence and the semantic anchor point sequence.
[0116] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0117] Acquire an original video and parse the original video to obtain an original video frame sequence;
[0118] Performing semantic anchor point extraction on the original video frame sequence to obtain a plurality of semantic anchor points, wherein one of the semantic anchor points corresponds to a core event in the original video;
[0119] Performing causal constraints on the plurality of semantic anchor points to obtain a semantic anchor point sequence;
[0120] Text information for describing the original video is generated based on the original video frame sequence and the semantic anchor point sequence.
[0121] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0122] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0123] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0124] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A text generation method based on semantic anchors, characterized in that: include: Acquire an original video and parse the original video to obtain an original video frame sequence; Performing semantic anchor point extraction on the original video frame sequence to obtain a plurality of semantic anchor points, wherein one of the semantic anchor points corresponds to a core event in the original video; Performing causal constraints on the plurality of semantic anchor points to obtain a semantic anchor point sequence; Text information for describing the original video is generated based on the original video frame sequence and the semantic anchor point sequence.
2. The method according to claim 1, characterized in that The extracting semantic anchor points from the original video frame sequence to obtain a plurality of semantic anchor points includes: Performing feature extraction on each frame of the original video frame sequence to obtain multiple frame features; Cluster analysis is performed on the plurality of frame features to obtain a plurality of semantic anchor points.
3. The method according to claim 2, characterized in that The performing cluster analysis on the plurality of frame features to obtain the plurality of semantic anchor points comprises: Optimizing the plurality of frame features by using a preset contrast loss function; Performing cluster analysis on the optimized frame features by a preset clustering algorithm to obtain a plurality of clusters, wherein each cluster includes a plurality of temporally continuous frame features; The mean of the intra-cluster features of each cluster is calculated to obtain a plurality of the semantic anchor points.
4. The method according to claim 3, characterized in that The preset contrast loss function is: Among them, L contrast is the total loss value, s(f i ,f j ) is the cosine similarity between frame feature i and frame feature j, s(f i ,f k ) is the cosine similarity between frame feature i and frame feature k, and τ is the temperature hyperparameter.
5. The method according to claim 1, wherein The step of performing causal constraints on the plurality of semantic anchor points to obtain a semantic anchor point sequence includes: Constructing a directed acyclic graph based on the plurality of semantic anchor points to generate an initial semantic anchor point sequence representing a causal relationship between the semantic anchor points; Modeling a generating function of each semantic anchor point in the initial semantic anchor point sequence, wherein the generating function includes exogenous noise and previous semantic anchor points; A loss value between the exogenous noise and the previous semantic anchor point is calculated by a preset regularizer, and the preset regularizer is trained through adversarial training to minimize the loss value to construct the semantic anchor point sequence.
6. The method according to claim 5, characterized in that The preset regularizer is: Among them, L causal is the loss value, HSIC is the preset regularizer, ε k is the exogenous noise, a k` is the previous semantic anchor, a k is the current semantic anchor.
7. The method according to claim 1, characterized in that Generating text information for describing the original video based on the original video frame sequence and the semantic anchor point sequence includes: Calculating an attention weight between each frame feature in the original video frame sequence and each semantic anchor point in the semantic anchor point sequence; Substituting the attention weight into a preset context feature calculation formula to obtain a context feature; The context features are input into a preset decoder to obtain the text information.
8. A text generation device based on semantic anchors, characterized in that: include: An acquisition module, configured to acquire an original video and parse the original video to obtain an original video frame sequence; an extraction module, configured to perform semantic anchor point extraction on the original video frame sequence to obtain a plurality of semantic anchor points, wherein one of the semantic anchor points corresponds to a core event in the original video; a constraint module, configured to perform causal constraints on the plurality of semantic anchor points to obtain a semantic anchor point sequence; A generating module is used to generate text information for describing the original video based on the original video frame sequence and the semantic anchor point sequence.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the text generation method based on semantic anchors according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the text generation method based on semantic anchors according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Network API recommendation method and device based on hypergraph contrast learning
CN116821508A
Judicial field semantic abstracting method, system and equipment based on large model and storage medium
CN119559559A
Video generation method and device based on multi-modal information fusion, equipment and medium
CN119906872A
Anchor shot detection method for a news video browsing system
US20020146168A1
Video anchors
US20210110163A1
Cited By
Automatic video clip editing and splicing method and system applied to digital multimedia
CN121531210A
AIGC coherent video generation method and system based on time sequence attention mechanism
CN121865063A