Text generation method and device, equipment and medium

By encoding video frames into discrete feature vectors and mapping them to motion, object, and scene subspaces, and combining dynamic gating networks and decoding models, the problem of insufficient interpretability and semantic depth in video content analysis in traditional methods is solved, achieving highly accurate and rich text generation.

CN120805861APending Publication Date: 2025-10-17PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510872753.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional video content analysis and description generation methods have difficulty accurately distinguishing dynamic actions, key entities and background environments, resulting in a lack of interpretability and semantic depth in the text generated in professional scenarios.

Method used

The encoder encodes video frames into discrete feature vectors, and uses a preset projection matrix to map them to three subspaces: motion, objects, and scene. Combined with a dynamic gating network and a decoding model, multi-dimensional semantic parsing and text generation are performed.

Benefits of technology

It achieves multi-dimensional semantic deep analysis of video content, improves the accuracy and richness of descriptive text, enhances the model's ability to focus on key elements, and improves the semantic consistency and expression accuracy of video-to-text conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805861A_ABST
    Figure CN120805861A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a text generation method and device, equipment and a medium, and the method comprises the steps: obtaining a target video, carrying out the encoding of a video frame of the target video through an encoder, and obtaining a discrete feature vector; respectively mapping the discrete feature vectors to a target subspace through a preset projection matrix to obtain three types of subfeatures; wherein the target subspace comprises a motion subspace, an object subspace and a scene subspace; analyzing the three types of sub-features through a dynamic gating network, and outputting fusion weight vectors corresponding to the three types of sub-spaces; and analyzing the fusion weight vector through a decoding model, and outputting a description text corresponding to the target video. The method can be applied to business program systems such as financial science and technology, medical health care and the like, and text generation of time-space decoupling and semantic refinement can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a text generation method, device, equipment and medium. Background Art

[0002] With the popularization of smart devices and social media platforms, short videos have become an important carrier of information dissemination, especially in key areas such as financial technology, medical health and elderly care. Short videos are gradually becoming a digital medium for risk warning, health education, service promotion and emotional care.

[0003] However, traditional methods for video content analysis and description generation mostly use end-to-end neural network structures, usually using long short-term memory networks or standard Transformer models to encode the video as a whole and then generate descriptive text. This makes it difficult to accurately distinguish dynamic actions, key entities, and background environments in the video, resulting in a lack of interpretability and semantic depth in professional scenarios. For example, in the field of financial technology, short videos for investor education or identification contain a large number of logical actions, object labels, and scene switches, and traditional methods cannot effectively disassemble and capture the risk signals therein; in medical health and elderly care scenarios, short video content often involves patient actions, drug treatment, or environmental identification, and existing technologies cannot carefully characterize the interactive relationship between key actions and background events. Therefore, there is an urgent need for a method that can decouple short video content in time and space, refine semantics, and automatically generate text descriptions to adapt to the needs of cross-modal intelligent analysis and improve the practicality and reliability of video content. Summary of the Invention

[0004] The present invention provides a text generation method, device, equipment and medium to solve the technical problem in related technologies that it is difficult to accurately distinguish dynamic actions, key entities and background environment in videos, resulting in a lack of interpretability and semantic depth in the text generated in professional scenarios.

[0005] In a first aspect, a text generation method is provided, the method comprising:

[0006] Obtaining a target video, and encoding a video frame of the target video through an encoder to obtain a discrete feature vector;

[0007] The discrete feature vectors are mapped to the target subspace respectively through a preset projection matrix to obtain three types of sub-features; wherein the target subspace includes a motion subspace, an object subspace, and a scene subspace;

[0008] Analyze the three types of sub-features through a dynamic gating network, and output fusion weight vectors corresponding to the three types of subspaces;

[0009] The analysis module is configured to analyze the three types of sub-features by using a dynamic gating network, and output a fusion weight vector corresponding to the three types of sub-spaces.

[0010] In a second aspect, a text generation apparatus is provided, and the text generation apparatus comprises:

[0011] The acquisition module is configured to acquire a target video, encode video frames of the target video by using an encoder, and obtain a discrete feature vector.

[0012] The mapping module is configured to map the discrete feature vector to a target subspace by using a preset projection matrix, and obtain three types of sub-features; the target subspace comprises a motion subspace, an object subspace, and a scene subspace.

[0013] The analysis module is configured to analyze the three types of sub-features by using a dynamic gating network, and output a fusion weight vector corresponding to the three types of sub-spaces.

[0014] The text generation module is configured to analyze the fusion weight vector by using a decoding model, and output a description text corresponding to the target video.

[0015] In a third aspect, a computer device is provided, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the text generation method when executing the computer program.

[0016] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the text generation method when executed by a processor.

[0017] The scheme realized by the text generation method, device, computer device and storage medium comprises the following steps: first, a target video is acquired, and a video frame of the target video is encoded by an encoder to obtain a discrete feature vector. Further, the discrete feature vector can be mapped to a target subspace by a preset projection matrix to obtain three types of sub-features; the target subspace comprises a motion subspace, an object subspace and a scene subspace. Thus, the three types of sub-features can be analyzed by a dynamic gating network to output a fusion weight vector corresponding to the three types of subspaces, and finally, the fusion weight vector is analyzed by a decoding model to output a description text corresponding to the target video. In the present application, the target video is encoded into a discrete feature vector, and the projection matrix is used to map the discrete feature vector to the motion, object and scene subspaces, which not only realizes the deep analysis of multi-dimensional semantics, but also improves the accuracy and richness of the description text. The dynamic gating network adaptively adjusts the fusion weight of each subspace feature according to the video content, thereby enhancing the attention ability of the model to key elements while maintaining the integrity of the semantics. The decoding model generates text by combining the weighted features, which significantly improves the semantic consistency and expression accuracy of the video-to-text conversion. Thus, the present application can adapt to the cross-modal intelligent analysis demand and improve the practicability and reliability of the video content. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without any creative labor.

[0019] Figure 1 is an application environment schematic diagram of the text generation method of an embodiment of the present application;

[0020] Figure 2 is a flowchart of the text generation method of an embodiment of the present application;

[0021] Figure 3 is Figure 1 is a specific implementation flowchart of step S10;

[0022] Figure 4 is a structural schematic diagram of the text generation device of an embodiment of the present application;

[0023] Figure 5 is a structural schematic diagram of the computer device of an embodiment of the present application;

[0024] Figure 6 is another structural schematic diagram of the computer device of an embodiment of the present application. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of the present application.

[0026] The text generation method provided by the embodiments of the present application can be applied in an application environment such as Figure 1 , wherein the client communicates with the server through the network. The server can obtain a target video through the client, encode video frames of the target video through an encoder to obtain a discrete feature vector, map the discrete feature vector to a target subspace through a preset projection matrix to obtain three types of subspace features, wherein the target subspace includes a motion subspace, an object subspace and a scene subspace, analyze the three types of subspace features through a dynamic gating network to output a fusion weight vector corresponding to the three types of subspaces, analyze the fusion weight vector through a decoding model to output a description text corresponding to the target video, and finally feedback quality analysis data to the client. In the present application, the target video is encoded into a discrete feature vector, and the projection matrix is used to map the discrete feature vector to the motion, object and scene subspaces, which not only realizes deep analysis of multi-dimensional semantics, but also improves the accuracy and richness of the description text. The dynamic gating network adaptively adjusts the fusion weight of each subspace feature according to the video content, thereby maintaining the integrity of the semantics while enhancing the attention ability of the model to key elements. The decoding model generates text in combination with the weighted features, which significantly improves the semantic consistency and expression accuracy of the video-to-text conversion. Thus, it can adapt to cross-modal intelligent analysis requirements and improve the practicality and reliability of video content. The client can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.

[0027] Please refer to Figure 2 , which is a flowchart of the text generation method provided by the embodiments of the present application. The method includes the following steps: Figure 2

[0028] S10: Obtain a target video, and encode video frames of the target video through an encoder to obtain a discrete feature vector.

[0029] ​For example, a target video can be acquired first as input data for subsequent processing. The target video can be from a financial monitoring system, a medical image platform, or a visual collection device in a smart elderly care scenario, such as video monitoring of a bank business hall, real-time operation images of an operating room, or remote nursing records of daily activities of the elderly. Second, the video frames in the target video can be input to an encoder, which can use a convolutional neural network or other deep learning structure to extract and convert features of each frame image. Through this process, the image information in the video frame is converted into a discrete feature vector with semantic expression ability, which not only retains the key spatiotemporal content but also removes redundant original pixel information, laying a foundation for subsequent video analysis, retrieval, or generation tasks.

[0030] In the financial scenario, this step can be used to automatically identify abnormal transaction behavior, such as suspicious interaction between a teller and a customer, or to assist customer behavior analysis and business process compliance review. In the medical health and elderly care field, discrete feature vectors can be used for action phase identification of surgical videos, monitoring of postoperative rehabilitation process of patients, or automatic identification of risk behaviors such as falls and abnormal activities of the elderly. Compared with directly analyzing original image data, discrete feature vectors can significantly reduce data processing amount, while enhancing the robustness and generalization ability of the model, providing a structured and quantifiable input basis for subsequent disease prediction, financial risk assessment, or elderly care intervention.

[0031] As shown in FIG. 1, the step S10, i.e., the encoding of the video frames of the target video by the encoder to obtain the discrete feature vector, includes the following steps: Figure 3

[0032] S11: frame sequence extraction is performed on the target video to obtain video frames of the target video.

[0033] S12: spatial structure features and cross-frame time sequence features in the video frames are extracted in sequence by a convolution module of the encoder to obtain a latent representation vector.

[0034] S13: vector matching is performed between the latent representation vector and a preset discrete codebook to obtain a discrete vector set.

[0035] S14: the discrete vector set is encoded to obtain the discrete feature vector.

[0036] ​For steps S11-S14, first in step S11, frame sequence extraction can be performed on the target video, and continuous video content is split into a series of static image frames, i.e., video frames, according to the time axis. This processing helps to convert dynamic information in the time domain into a frame sequence structure that is easy to analyze, ensuring that the subsequent encoder can process and capture the change information in the video frame by frame. This step is suitable for monitoring and analysis tasks in financial and medical and health care scenarios. For example, in video auditing in a bank business place, frame sequence can accurately capture customer behavior details, such as frequent looking around, abnormal body movements, etc. In the medical and smart care field, frame sequence extraction helps to build a timeline record of patient or old person activities for post-event behavior analysis or real-time health warning, especially in scenarios such as fall detection or postoperative motion anomaly warning.

[0037] Next, in step S12, the extracted video frames can be input to the convolution module of the encoder, which usually includes multiple convolution layers, activation functions, and pooling operations, and can extract spatial structure features in each frame, such as contours, textures, object boundaries, etc., while using time series modeling mechanisms (such as time convolution, recurrent structure, or attention mechanism) to extract dynamic evolution rules across frames, thereby generating a set of latent representation vectors representing the video content, reflecting the comprehensive semantic features of the video in both spatial and temporal dimensions.

[0038] Further, in step S13, the above latent representation vectors can be matched with a preset discrete codebook, which is usually generated in the training stage and contains multiple representative vectors (i.e., code words) for discretizing latent features. By minimum distance matching or other metric mechanisms, each latent representation is mapped to the closest vector in the codebook, realizing the conversion from continuous features to discrete features, and generating a set of discrete vectors. Subsequently, in step S14, a uniform encoding operation is performed on this set of discrete vectors, such as splicing, compression, or index arrangement, to form the final discrete feature vector. This discrete feature vector not only retains the key semantics and time sequence information of the original video, but also greatly reduces the data dimension and storage complexity, providing an efficient and structured input for subsequent tasks such as video retrieval, classification, or generation model. The above steps lay the foundation for intelligent video understanding, especially suitable for serving financial and health care scenarios with high security and reliability requirements.

[0039] S20: mapping the discrete feature vectors to target subspaces respectively by a preset projection matrix, to obtain three types of sub-features.

[0040] Among them, the target subspaces include a motion subspace, an object subspace, and a scene subspace.

[0041] For example, the purpose of step S20 is to perform subspace division and projection on discrete feature vectors, so as to achieve fine-grained decoupling and representation of video semantics. Specifically, through a preset projection matrix, discrete feature vectors can be mapped to three target subspaces with different semantic directivities, corresponding to the motion subspace, object subspace and scene subspace respectively. This processing is equivalent to dividing three feature components in a unified coding representation, so that in subsequent tasks, different dimensions of information such as dynamic behavior, static entities and environmental background in the target video can be focused on respectively. The motion subspace is used to capture temporal changes between video frames, such as human movements, traffic movements and other patterns that are strongly related to time evolution; the object subspace is mainly responsible for identifying and characterizing key objects and their attributes that appear in the video, such as people, vehicles or handheld tools; and the scene subspace is used to characterize the environment and background of the video, such as indoor and outdoor distinctions, scene categories or lighting structures. The projection matrix itself can be trained on large-scale semantic video data through prior learning. Its function is to deconstruct the original mixed representation into these three sub-features in a supervised manner in the high-dimensional space, so that different semantic dimensions are decoupled from each other in the feature space, and the model's ability to express multi-dimensional semantic understanding is enhanced.

[0042] For example, in financial applications, the projection of the motion subspace can effectively identify changes in customer or employee motion patterns, such as suspicious fund transfers, non-standard operating procedures, or unusual body movements. In healthcare and elderly care, the motion subspace can deeply characterize the daily activity patterns of elderly people at home, such as transitioning from sitting to standing, unstable gait during walking, and sudden falls, greatly improving the ability of elderly care systems to predict dangerous conditions.

[0043] The above-mentioned feature design based on subspace decomposition not only improves the adaptability of downstream tasks such as multi-task learning or conditional generation, but also helps to achieve more controllable and explainable video analysis and content generation.

[0044] In some embodiments, the preset projection matrix includes a motion projection matrix, an object projection matrix, and a scene projection matrix, and the three types of sub-features include motion sub-features, object sub-features, and scene sub-features. The discrete feature vectors are mapped to the target subspace through the preset projection matrix to obtain three types of sub-features, including: mapping the discrete feature vector to the motion subspace through the motion projection matrix to obtain the motion sub-feature; and, mapping the discrete feature vector to the object subspace through the object projection matrix to obtain the object sub-feature; and, mapping the discrete feature vector to the scene subspace through the scene projection matrix to obtain the scene sub-feature.

[0045] For example, to achieve accurate disassembly of various semantic elements in the target video, three different preset projection matrices are introduced, namely a motion projection matrix, an object projection matrix, and a scene projection matrix. These matrices are learned in the training stage, and each is designed to extract components in the discrete feature vector that are highly related to a specific semantic dimension. In actual operation, the original discrete feature vector can first be input into the motion projection matrix to complete the mapping from the overall feature to the motion subspace, obtaining the motion sub-feature. This part of information is mainly used to capture continuous changes and dynamic behaviors over time, such as human actions, traffic trajectories, etc. Then, the discrete feature vector is mapped to the object subspace by the object projection matrix to extract structured expressions related to static objects, thereby forming the object sub-feature. This part can represent the category, pose, or relative position of the main object in the target video. For example, in a financial scenario, this process can identify abnormal operation rhythms of bank customers when conducting business, or suspicious interaction dynamics in front of an ATM machine. In a medical and elderly care environment, the motion sub-feature helps to capture the action quality of patients during rehabilitation training or the behavior coherence and change trend of the elderly during home activities, thereby providing clues for early detection of abnormal behavior.

[0046] Further, the discrete feature vector also needs to be mapped to the scene subspace by the scene projection matrix to generate the scene sub-feature. This part is used to depict the overall environmental information of the video, such as whether it is indoors or outdoors, the degree of light, building structure, or natural background, etc. The generation process of the three sub-features is essentially a semantic disassembly of the unified encoding result, allowing different components of the video content to be extracted and modeled independently, improving the clarity and control accuracy of the above method in the understanding, generation, or retrieval process.

[0047] Through this mapping method, the recognition ability of the model for multi-dimensional video semantics can be effectively enhanced while maintaining the compactness and structure of the features, providing more detailed basic feature support for subsequent tasks.

[0048] S30: analyzing the three types of sub-features through a dynamic gating network, and outputting a fusion weight vector corresponding to the three types of subspaces.

[0049] In some embodiments, the analyzing the three types of sub-features through a dynamic gating network, and outputting a fusion weight vector corresponding to the three types of subspaces, includes: taking a decoder in the dynamic gating network as a query input to calculate attention matching degrees of the three types of sub-features; weighting the attention matching degrees of the three types of sub-features to obtain a weight distribution vector of the three types of sub-features; and concatenating and linearly transforming the weight distribution vector of the three types of sub-features to obtain the fusion weight vector corresponding to the three types of subspaces.

[0050] By way of example, step S30 can intelligently analyze the aforementioned three types of sub-features through a dynamic gating network to generate a fusion weight vector that can reflect the current video semantic requirements. Specifically, first, the decoder output in the dynamic gating network is taken as a query signal, and attention matching degree calculations are performed on the query signal with respect to the motion sub-feature, the object sub-feature, and the scene sub-feature. This attention mechanism can automatically determine which sub-feature is more valuable according to the current context or task state. For example, in a scenario involving behavior recognition, the matching degree of the motion sub-feature is more likely to be improved; and in a task of describing a scene background, the response of the scene sub-feature is more likely to be enhanced. In this way, the relative importance of each sub-space in the current semantic expression can be dynamically identified, and semantic regulation at the feature level can be achieved.

[0051] For example, in the financial field, this method can be widely applied to a multi-task video understanding system. If it is found that a customer behavior is abnormally stationary, frequently turns back, or has a sudden change in action, the matching degree of the motion sub-feature is automatically improved. In the medical health and elderly care scenarios, the dynamic gating mechanism is particularly suitable for tasks such as elderly behavior monitoring and medical operation process analysis. For example, when identifying a fall event, the model automatically improves the weight of the motion sub-feature according to the time-series behavior anomaly.

[0052] Further, the three types of sub-features can be weighted according to the matching degree results to form a weight distribution vector that reflects the current feature value distribution. Then, the weight distribution vector is concatenated to enable the semantic importance of the three types of sub-features to be combined and expressed, and a linear transformation is further performed to compress and refine the feature structure, thereby generating a fusion weight vector corresponding to the three types of sub-spaces. The fusion weight vector will serve as a control signal for subsequent multi-feature fusion or generation mechanisms, guiding the model to adapt to different types of semantic attention according to the current task requirements.

[0053] The weight generation strategy based on the attention and gating mechanisms described above not only enhances the dynamic understanding ability of the multi-dimensional semantics of video content, but also realizes a semantic-driven feature integration method, making the model have higher adaptability and intelligent processing capability.

[0054] S40: Analyzing the fusion weight vector through a decoding model to output a description text corresponding to the target video.

[0055] In some embodiments, the analyzing, by the decoding model, the fusion weight vector outputs a description text corresponding to the target video, comprising: inputting the fusion weight vector as a decoding input into the decoding model for modeling to obtain a semantic context vector; analyzing the semantic context vector through a stacked structure of the decoding model to obtain a hidden state sequence; performing linear mapping on the hidden state sequence to obtain optimal text candidate terms of the target video; and splicing a plurality of the optimal text candidate terms to obtain the description text corresponding to the target video.

[0056] By way of example, the above technical solution can utilize a decoding model to convert the fused multi-dimensional semantic information into a description text in natural language form, completing cross-modal generation from video content to language expression. Specifically, the fusion weight vector is first introduced as an input into the decoding model, which can model the fusion weight vector based on existing semantic expression capabilities to generate a semantic context vector reflecting the overall semantic trend. The context vector not only contains the weighted fusion results of the three types of sub-features of motion, object, and scene, but also integrates the semantic attention distribution output by the gating mechanism, thus having highly compressed and semantically concentrated feature expression capabilities. Then, the decoding model uses its internal stacked structure (such as a multi-layer Transformer decoder or a recurrent neural network) to deeply analyze the semantic context vector, layer by layer extracting and combining the logical, semantic, and temporal relationships between contexts to form a hidden state sequence for text generation.

[0057] Subsequently, linear mapping can be performed on the hidden state sequence to convert each hidden state into a probability distribution in the natural language vocabulary space, from which the optimal text candidate terms are selected. Each term generated at each time represents the optimal language expression of the current semantic segment, and these terms are strictly organized and constrained in terms of grammatical structure and semantic logic throughout the generation process. Finally, a plurality of optimal text candidate terms are spliced in the order of generation to obtain a target video description text that is structurally complete and semantically coherent. This text not only accurately conveys the core content and visual elements of the video, but also has readability and contextual coherence of natural language expression, enabling the entire model to achieve high-quality conversion from visual signals to language expressions, suitable for various practical application scenarios such as video content retrieval, auxiliary understanding, and automatic caption generation.

[0058] In some embodiments, the method further comprises: obtaining a training data set and a pre-trained model; wherein the training data set comprises a plurality of historical videos; labeling the training data set to obtain a labeling result, wherein the labeling result comprises historical text descriptions corresponding to the historical videos; and training the pre-trained model using the training data set and the labeling result to obtain the decoding model.

[0059] On the basis of the above-mentioned embodiments, after obtaining the decoding model, the method further comprises: iteratively training the decoding model based on the training data set and the annotation result to extract data features and calculate a loss function; iteratively training the loss function using a preset method to reduce the value of the loss function until the value of the loss function is less than an expected threshold; and obtaining an iteratively trained decoding model based on the iteratively trained loss function.

[0060] Specifically, a training data set containing a plurality of historical videos can be collected for training. The training data set can be obtained by manual collection, web crawling, or a public data set, and the present application does not limit this.

[0061] Further, each of the plurality of historical videos can be annotated to obtain an annotation result corresponding to the text of each of the plurality of historical videos. Then, the annotation result is used as a label of the training data set, and each group of training data set with the label is input into the pre-trained model for supervised learning. When the training end condition is met, such as when the number of training times reaches a threshold or the output accuracy of the model reaches a threshold, the training is ended, and a trained decoding model is obtained.

[0062] In the embodiments of the present application, the training data set and the annotation result are input into the pre-trained model for supervised learning, and then the decoding model is trained. Thus, the annotation result can be output based on the decoding model.

[0063] The above embodiments can enhance data quality and diversity during decoding model training, and improve the generalization ability and actual application effect of the model.

[0064] It can be understood that, in order to train a decoding model with higher accuracy, the decoding model can be iteratively trained in a manner that reduces the loss function until the loss function meets the expected threshold, and then a more accurate annotation result can be obtained based on the iteratively trained decoding model.

[0065] It should be noted that the present application does not limit the above-mentioned preset method and expected threshold, for example, the preset method can be gradient descent algorithm, batch gradient descent algorithm, stochastic gradient descent algorithm, etc., and the present application takes the gradient descent algorithm as an example for illustration.

[0066] The purpose of the gradient descent algorithm is to find the minimum value of the loss function through iteration, or to converge to the minimum value. Geometrically, the gradient descent algorithm is to move in the opposite direction of the vector where the function changes the fastest, and the gradient decreases the fastest, so it is easier to find the minimum value of the function. Based on this, in the embodiments of the present application, the gradient descent algorithm can be used to iteratively train the decoding model to reduce the loss function, thereby reducing the error of the calculation result.

[0067] In the embodiments of the present application, the gradient descent algorithm is used to iteratively train the decoding model to reduce the loss function to obtain the iterated decoding model, and then a more accurate annotation result can be obtained based on the iterated decoding model.

[0068] As can be seen, in the above scheme, by encoding the target video into a discrete feature vector and mapping it to the motion, object and scene subspaces using a projection matrix, not only is the multi-dimensional semantic depth analyzed, but also the accuracy and richness of the description text are improved. The dynamic gating network adaptively adjusts the fusion weights of the features in each subspace according to the video content, thereby maintaining the integrity of the semantics while enhancing the model's attention to key elements. The decoding model generates text based on the weighted features, significantly improving the semantic consistency and expression accuracy of the video-to-text conversion. As a result, it can adapt to cross-modal intelligent analysis requirements and improve the practicality and reliability of video content.

[0069] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0070] In an embodiment, a text generation apparatus is provided, which corresponds one-to-one to the text generation method described above. As shown in the figure, the text generation apparatus includes an acquisition module 101, a mapping module 102, an analysis module 103, and a text generation module 104. The functions of each module are described in detail as follows: Figure 4

[0071] The acquisition module 101 is configured to acquire a target video and encode video frames of the target video through an encoder to obtain a discrete feature vector.

[0072] The mapping module 102 is configured to map the discrete feature vector to a target subspace through a preset projection matrix to obtain three types of sub-features; wherein the target subspace includes a motion subspace, an object subspace, and a scene subspace.

[0073] ​The analysis module 103 is configured to analyze the three types of sub-features through a dynamic gating network, and output a fusion weight vector corresponding to the three types of sub-spaces.

[0074] The text generation module 104 is configured to analyze the fusion weight vector through a decoding model, and output a description text corresponding to the target video.

[0075] The acquisition module 101 is configured to extract a frame sequence of the target video to obtain video frames of the target video, sequentially extract spatial structure features and cross-frame timing features in the video frames through a convolution module of the encoder to obtain a latent representation vector, perform vector matching on the latent representation vector and a preset discrete codebook to obtain a discrete vector set, and encode the discrete vector set to obtain a discrete feature vector.

[0076] The mapping module 102 is configured to map the discrete feature vector to the motion subspace through the motion projection matrix to obtain motion sub-features, map the discrete feature vector to the object subspace through the object projection matrix to obtain object sub-features, and map the discrete feature vector to the scene subspace through the scene projection matrix to obtain scene sub-features.

[0077] The analysis module 103 is configured to take a decoder in the dynamic gating network as a query input, calculate attention matching degrees of the three types of sub-features, weight the attention matching degrees of the three types of sub-features to obtain a weight distribution vector of the three types of sub-features, and perform concatenation and linear transformation operations on the weight distribution vector of the three types of sub-features to obtain a fusion weight vector corresponding to the three types of sub-spaces.

[0078] The text generation module 104 is configured to take the fusion weight vector as a decoding input, input the decoding input into the decoding model for modeling to obtain a semantic context vector, analyze the semantic context vector through a stacked structure of the decoding model to obtain a hidden state sequence, perform linear mapping on the hidden state sequence to obtain optimal text candidate terms of the target video, and splice a plurality of the optimal text candidate terms to obtain a description text corresponding to the target video.

[0079] In an embodiment, the acquisition module 101 is further configured to acquire a training data set and a pre-trained model, wherein the training data set includes a plurality of historical videos, label the training data set to obtain a labeling result, wherein the labeling result includes historical text descriptions corresponding to the historical videos, and train the pre-trained model through the training data set and the labeling result to obtain the decoding model.

[0080] In an embodiment, the acquisition module 101 is further configured to: based on the training data set and the annotation result, iteratively train the decoding model to extract data features and calculate a loss function; iteratively train the loss function using a preset method to reduce the value of the loss function until the value of the loss function is less than a preset threshold; and based on the iteratively trained loss function, obtain an iteratively trained decoding model.

[0081] The application provides a text generation device. By encoding a target video into a discrete feature vector and mapping it to three subspaces of motion, object and scene using a projection matrix, not only is the multi-dimensional semantic depth analyzed, but also the accuracy and richness of the description text are improved. The dynamic gating network adaptively adjusts the fusion weight of each subspace feature according to the video content, thereby maintaining the semantic integrity while enhancing the attention ability of the model to key elements. The decoding model generates text by combining the weighted features, significantly improving the semantic consistency and expression accuracy of the video-to-text conversion. Thus, the cross-modal intelligent analysis demand can be adapted, and the practicability and reliability of the video content are improved.

[0082] The specific limitations of the text generation device can be referred to the limitations of the text generation method in the above, which will not be repeated here. Each module in the above text generation device can be realized by software, hardware and their combination in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.

[0083] In one embodiment, a computer device is provided, which can be a server, and its internal structure diagram can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface and a database connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external client through a network connection. The computer program is executed by the processor to implement the functions or steps of a text generation method server side.

[0084] In one embodiment, a computer device is provided, which can be a client, and its internal structure diagram can be as shown in Figure 6As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is used to communicate with the external server through the network connection. The computer program is executed by the processor to realize the function or step of the text generation method client side

[0085] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the computer program to implement the following steps:

[0086] Obtain a target video, and encode video frames of the target video through an encoder to obtain discrete feature vectors;

[0087] Map the discrete feature vectors to target subspaces through a preset projection matrix to obtain three types of sub-features; wherein the target subspaces include a motion subspace, an object subspace, and a scene subspace;

[0088] Analyze the three types of sub-features through a dynamic gating network to output a fusion weight vector corresponding to the three types of subspaces;

[0089] Analyze the fusion weight vector through a decoding model to output a description text corresponding to the target video.

[0090] In one embodiment, a computer readable storage medium is provided, having a computer program stored thereon, the computer program being executed by a processor to implement the following steps:

[0091] Obtain a target video, and encode video frames of the target video through an encoder to obtain discrete feature vectors;

[0092] Map the discrete feature vectors to target subspaces through a preset projection matrix to obtain three types of sub-features; wherein the target subspaces include a motion subspace, an object subspace, and a scene subspace;

[0093] Analyze the three types of sub-features through a dynamic gating network to output a fusion weight vector corresponding to the three types of subspaces;

[0094] Analyze the fusion weight vector through a decoding model to output a description text corresponding to the target video.

[0095] It should be noted that the functions or steps described above with respect to the computer readable storage medium or the computer device can correspond to the related descriptions of the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0096] A person of ordinary skill in the art can understand that all or part of the processes in the foregoing method embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a nonvolatile computer readable storage medium. When the computer program is executed, the processes of the foregoing embodiments of the method can be included. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include nonvolatile and / or volatile memory. The nonvolatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. The volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.

[0097] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified. In actual applications, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0098] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features. Such modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A text generation method, characterized in that: The method comprises: Obtaining a target video, and encoding a video frame of the target video through an encoder to obtain a discrete feature vector; The discrete feature vectors are mapped to the target subspace respectively through a preset projection matrix to obtain three types of sub-features; wherein the target subspace includes a motion subspace, an object subspace, and a scene subspace; Analyze the three types of sub-features through a dynamic gating network, and output fusion weight vectors corresponding to the three types of subspaces; The fusion weight vector is analyzed by a decoding model, and a description text corresponding to the target video is output.

2. The method according to claim 1, characterized in that The step of encoding the target video frame by an encoder to obtain a discrete feature vector includes: Extracting a frame sequence from the target video to obtain a video frame of the target video; Sequentially extracting spatial structural features and cross-frame temporal features in the video frames through a convolution module of the encoder to obtain a latent representation vector; Performing vector matching on the potential representation vector and a preset discrete codebook to obtain a discrete vector set; The discrete vector set is encoded to obtain the discrete feature vector.

3. The method according to claim 1, characterized in that The preset projection matrix includes a motion projection matrix, an object projection matrix, and a scene projection matrix. The three types of sub-features include motion sub-features, object sub-features, and scene sub-features. The discrete feature vectors are respectively mapped to the target subspace through the preset projection matrix to obtain three types of sub-features, including: Mapping the discrete feature vector to the motion subspace through the motion projection matrix to obtain the motion subfeature; and Mapping the discrete feature vector to the object subspace through the object projection matrix to obtain the object subfeature; and The discrete feature vector is mapped to the scene subspace through the scene projection matrix to obtain the scene subfeature.

4. The method according to claim 1, wherein The three types of sub-features are analyzed by the dynamic gating network to output the fusion weight vectors corresponding to the three types of subspaces, including: The decoder in the dynamic gating network is used as a query input to calculate the attention matching degree of the three types of sub-features; Weighting the attention matching degrees of the three types of sub-features to obtain weight distribution vectors of the three types of sub-features; The weight distribution vectors of the three types of sub-features are concatenated and linearly transformed to obtain fusion weight vectors corresponding to the three types of subspaces.

5. The method according to claim 1, wherein The step of analyzing the fusion weight vector by a decoding model and outputting a description text corresponding to the target video includes: The fusion weight vector is used as a decoding input and input into the decoding model for modeling to obtain a semantic context vector; Analyzing the semantic context vector through the stacking structure of the decoding model to obtain a hidden state sequence; Performing linear mapping on the hidden state sequence to obtain an optimal text candidate term of the target video; Several of the optimal text candidate terms are concatenated to obtain a description text corresponding to the target video.

6. The method according to claim 1, characterized in that The analyzing the structured semantic representation by a text decoding network to generate a video text description corresponding to the target network includes: Obtaining a training data set and a pre-trained model; wherein the training data set includes a plurality of historical videos; Annotating the training data set to obtain an annotation result, wherein the annotation result includes a historical text description corresponding to the historical video; The pre-training model is trained using the training data set and the annotation results to obtain the decoding model.

7. The method according to claim 6, characterized in that After obtaining the decoding model, the method further includes: Iteratively training the decoding model based on the training data set and the annotation results to extract data features, and calculate the loss function; Iteratively training the loss function using a preset method for the purpose of reducing the loss function value until the loss function value is less than an expected threshold; Based on the loss function after iterative training, an iterative decoding model is obtained.

8. A text generation device, characterized in that: The text generating device comprises: An acquisition module is used to acquire a target video and encode the video frames of the target video through an encoder to obtain a discrete feature vector; A mapping module, configured to map the discrete feature vectors to target subspaces using a preset projection matrix to obtain three types of subfeatures; wherein the target subspaces include motion subspaces, object subspaces, and scene subspaces; An analysis module is used to analyze the three types of sub-features through a dynamic gating network and output fusion weight vectors corresponding to the three types of subspaces; The text generation module is used to analyze the fusion weight vector through a decoding model and output a description text corresponding to the target video.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the text generation method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the text generation method according to any one of claims 1 to 7 are implemented.