Text generation method and device, equipment and medium
By constructing video image block representations, extracting spatiotemporal features, and combining language embedding matrices and attention mechanisms to generate structured text descriptions, the difficult problem of converting dynamic visual information into coherent text is solved, and intelligent text generation applications in the financial and medical fields are realized.
Patent Information
- Application Number
- CN202510852161.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies make it difficult to efficiently convert dynamic visual information into structured and coherent text. Especially in the fields of financial technology and healthcare and elderly care, traditional models find it difficult to achieve text generation under cross-modal and multi-time sequence conditions, resulting in fragmented and incoherent generated text.
By acquiring the target video frame sequence, constructing the video image block representation, using the hierarchical temporal network to extract spatiotemporal features, combining the language vector embedding matrix and attention mechanism to generate structured semantic representation, and finally generating the video text description through the text decoding network.
It realizes text generation under cross-modal and multi-time sequence conditions. The generated text has high semantic consistency and can accurately describe video events, behaviors and scenes, improving the efficiency of compliance inspection and fraud identification in the financial field, and the automatic recording and knowledge archiving capabilities in the medical field.
Smart Images

Figure CN120764495A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and in particular, to a text generation method, device, equipment and medium. BACKGROUND
[0002] With the wide application of short videos, image sequences and multi-modal data in various industries, how to efficiently convert dynamic visual information into structured, coherent and professional text has become an important challenge for text generation technology. Especially in the fields of financial technology and medical and health care, the text generation task faces multiple technical bottlenecks such as complex time sequence, high semantic accuracy and variable language structure.
[0003] In the financial field, market fluctuations, investment behavior and transaction processes often exist in the form of video monitoring, chart dynamics, etc., requiring accurate identification of behavior details and output of structured risk control reports or audit records. In the field of medical health and nursing care, patient activities, vital sign evolution and nursing operation processes are often implied in time sequence image data, and traditional models are difficult to clearly transcribe these changes into text content that conforms to medical semantics and logical order. Existing video text generation methods mainly rely on image feature splicing or simple pooling operations, lacking modeling capability for video time sequence evolution, resulting in fragmented and incoherent generated text. Therefore, it is necessary to propose a text generation method that can realize text generation under cross-modal and multi-time sequence conditions. SUMMARY
[0004] The present application provides a text generation method, device, equipment and medium to solve the technical problem of fragmented and incoherent text generation in related technologies.
[0005] In a first aspect, a text generation method is provided, comprising:
[0006] Obtaining a target video and extracting a frame sequence in the target video to obtain a video image block representation;
[0007] Analyzing the video image block representation through a hierarchical time sequence network to obtain a video space-time feature;
[0008] Establishing a language vector embedding matrix based on the target video and calculating an attention weight matrix of the video space-time feature and the language vector embedding matrix to generate a structured semantic representation;
[0009] Analyzing the structured semantic representation through a text decoding network to generate a video text description corresponding to the target network.
[0010] In a second aspect, a text generation device is provided, comprising:
[0011] An acquisition module is configured to acquire a target video and extract a frame sequence in the target video to obtain a video image block representation.
[0012] An analysis module is configured to analyze the video image block representation by using a hierarchical temporal network to obtain a video spatio-temporal feature.
[0013] A semantic representation module is configured to establish a language vector embedding matrix based on the target video, calculate an attention weight matrix of the video spatio-temporal feature and the language vector embedding matrix, and generate a structured semantic representation.
[0014] A text generation module is configured to analyze the structured semantic representation by using a text decoding network to generate a video text description corresponding to the target network.
[0015] In a third aspect, a computer device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the above text generation method when executing the computer program.
[0016] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the above text generation method when executed by a processor.
[0017] The scheme realized by the text generation method, device, computer device and storage medium comprises the following steps. First, a target video is acquired, and a frame sequence in the target video is extracted to obtain a video image block representation. Further, the video image block representation can be analyzed by a hierarchical temporal network to obtain a video space-time feature. Thus, a language vector embedding matrix can be established based on the target video, and an attention weight matrix of the video space-time feature and the language vector embedding matrix can be calculated to generate a structured semantic representation. Finally, the structured semantic representation can be analyzed by a text decoding network to generate a video text description corresponding to the target network. In the present application, the frame sequence in the target video is extracted and the video image block representation is constructed, so that the visual information of the target video can be fully retained. Further, the hierarchical temporal network is used to extract the space-time feature of the target video, which not only enhances the modeling capability of dynamic changes and local details, but also provides multi-scale feature support for subsequent semantic understanding. By constructing the language vector embedding matrix and combining the attention mechanism, the key alignment relationship between the video content and the language description can be accurately captured, so that the structured representation with high semantic consistency can be generated. Finally, the text decoding network is used to complete the conversion from the video content to the natural language, so that the accurate description of the video events, behaviors, scenes and other contents can be realized. In the financial field, the method can be used for intelligent analysis of transaction behaviors, counter operations, risk events and the like in the monitoring video, so as to improve the compliance checking and fraud identification efficiency. In the medical field, the method can be used to assist in generating operation step descriptions of surgical videos, remote consultation image interpretation reports and the like, so as to realize the automatic recording and knowledge archiving of medical images, thereby accelerating the structuralization and intelligent upgrading of professional information. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 is an application environment schematic diagram of a text generation method according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of a text generation method according to an embodiment of the present application;
[0021] Figure 3 is Figure 1 is a specific implementation flowchart of step S10 in the method;
[0022] Figure 4 is a structural schematic diagram of a text generation device according to an embodiment of the present application;
[0023] Figure 5 is a structural schematic diagram of a computer device in an embodiment of the present application;
[0024] Figure 6 is another structural schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of the present application.
[0026] The text generation method provided by the embodiments of the present application can be applied to, for example, Figure 1In an application environment of the application, the client communicates with the server through a network. The server can obtain a target video through the client, extract a frame sequence in the target video to obtain a video image block representation, analyze the video image block representation through a hierarchical temporal network to obtain a video space-time feature, establish a language vector embedding matrix based on the target video, calculate an attention weight matrix of the video space-time feature and the language vector embedding matrix to generate a structured semantic representation, analyze the structured semantic representation through a text decoding network to generate a video text description corresponding to the target network, and finally feed back the video text description to the client. In the application, the frame sequence in the target video is extracted and the video image block representation is constructed, which can fully retain the visual information of the target video. Further, the hierarchical temporal network is used to extract the space-time feature of the target video, which not only enhances the modeling capability of dynamic changes and local details, but also provides multi-scale feature support for subsequent semantic understanding. By constructing the language vector embedding matrix and combining the attention mechanism, the key alignment relationship between the video content and the language description can be accurately captured, so as to generate a structured representation with high semantic consistency. Finally, with the help of the text decoding network, the conversion from the video content to the natural language is completed, and the accurate description of the video events, behaviors, scenes and other contents is realized. In the financial field, the method can be used for intelligent analysis of transaction behaviors, counter operations, risk events and the like in the monitoring video, and the efficiency of compliance checking and fraud identification is improved. In the medical field, it can assist in generating operation step descriptions of surgical videos, remote consultation image interpretation reports and the like, realize automatic recording and knowledge archiving of medical images, and accelerate the structured and intelligent upgrading of professional information. The client can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be realized by an independent server or a server cluster composed of multiple servers. The application will be described in detail below through specific embodiments.
[0027] Please refer to Figure 2 , as shown in Figure 2 , as shown in
[0028] S10: Obtain a target video, and extract a frame sequence in the target video to obtain a video image block representation.
[0029] For example, the target video to be processed can be obtained first, and then the frame sequence in the target video is extracted, that is, the continuous static image frames are extracted from the video in time sequence. Further, each frame image extracted can be divided into a plurality of image blocks, so as to convert the entire video sequence into a video image block representation composed of a plurality of image blocks, thereby providing a basis for subsequent feature extraction, analysis or model processing.
[0030] The above image block representation method helps capture local image features and temporal changes, improving the accuracy and efficiency of video content understanding.
[0031] In step S10, i.e., extracting the frame sequence in the target video to obtain the video image block representation, includes the following steps: Figure 3
[0032] S11: Key frame sampling of the target video to obtain key image frames.
[0033] S12: Dividing the key image frames into image blocks of a preset size to obtain the video image blocks.
[0034] The video image blocks retain the spatial position information in the original frame.
[0035] For example, in the financial field, this step can be used for monitoring risk behavior recognition in the target video. For example, the monitoring video of a bank business hall can be used as a target video, and through key frame sampling in step S11, only the key image frames when the customer interacts with the teller or high-risk operations (such as large amount of cash withdrawal) occur are extracted, reducing the processing pressure of redundant data. Further in step S12, the key image frames can be divided into image blocks of a preset size, each image block retains its spatial position information in the original frame, so that the position of the key area (such as the customer's hand and the money cabinet) in the image can be accurately identified, providing fine input for subsequent behavior analysis and anomaly detection.
[0036] For example, in the medical field, this step is also applicable to the analysis of surgical videos or monitoring of diagnosis processes. Through key frame sampling, key moments in medical operations can be extracted, such as surgical instruments entering incisions or doctors performing diagnostic gestures. Through image block division, the accurate positions of various instruments and tissue structures in the image are retained, which helps subsequent medical behavior recognition, intraoperative risk assessment, or automatic recording of diagnosis processes. The above structured video image block representation enables the model to more accurately perceive details, effectively supporting medical quality tracking and intelligent decision-making assistance.
[0037] S20: Analyzing the video image block representation through a hierarchical temporal network to obtain video spatiotemporal features.
[0038] In some embodiments, the hierarchical temporal network comprises a conversion module and an attention module, and the analysis of the video image block representation by the hierarchical temporal network to obtain the video spatio-temporal feature comprises: analyzing the video image block by the conversion module to extract action features; performing aggregation processing on the action features in a preset time dimension to obtain an intermediate state sequence; wherein the intermediate state sequence comprises continuous action features; and analyzing the intermediate state sequence by the attention module to obtain the video spatio-temporal feature.
[0039] For example, the video image block obtained in the previous step can be analyzed in depth by the hierarchical temporal network to extract video spatio-temporal features reflecting time and space information. First, the conversion module can perform preliminary processing on the video image block to identify and extract action features therein, such as behavior changes of characters or objects in the image; then, the action features are aggregated in the time dimension to form an intermediate state sequence representing a series of continuous behaviors; finally, the attention module further analyzes the intermediate state sequence to highlight the features of key behaviors or key moments, thereby obtaining video spatio-temporal features with global understanding capability.
[0040] In the financial field, this method can be used to identify abnormal operation patterns of customers in counter services or to identify collective behavior risks in trading halls; in the medical field, it can be used to analyze whether the continuous actions of doctors in surgery conform to the specifications, or to assist in judging the action quality of patients during rehabilitation training, thereby realizing intelligent behavior tracking and evaluation.
[0041] S30: establishing a language vector embedding matrix based on the target video, calculating an attention weight matrix of the video spatio-temporal feature and the language vector embedding matrix, and generating a structured semantic representation.
[0042] For example, a corresponding language vector embedding matrix can be first constructed based on the target video. The language vector embedding matrix can be derived from text descriptions, labels or task instructions related to the target video, and represents semantic information in language in the form of a vector. Then, by calculating the attention weight matrix between the video spatio-temporal feature obtained in the previous step and the language vector embedding matrix, the spatio-temporal segments in the target video that are highly related to the language content are identified. Finally, through this attention mechanism, a structured semantic representation is generated, accurately aligning and mapping the video content to the language semantic space.
[0043] In the financial field, this processing can be used to realize intelligent retrieval or question answering of monitoring videos based on natural language, such as "whether the customer has been queuing at the counter for more than 10 minutes"; in the medical field, it can be used to align doctor operation videos with electronic medical records or medical instructions, realizing intraoperative semantic understanding and intelligent auxiliary analysis.
[0044] In some embodiments, the establishing the language vector embedding matrix based on the target video comprises: determining a language element category in the target video; wherein the language element category comprises an action type, an object category, and a scene element; performing an initialization operation on the language element category to obtain an initialized language prototype; and embedding the language prototype into an embedding vector of a preset dimension to obtain the language vector embedding matrix.
[0045] For example, in this step, to realize the deep fusion of video and language information, first, the relevant language element categories need to be identified from the target video, that is, the video content is semantically disassembled, and descriptive elements such as action types (for example, “walking” and “taking objects”), object categories (for example, “medical record book” and “cash”), and scene elements (for example, “ward” and “counter”) are extracted. Subsequently, the language element categories are initialized to generate language prototypes, that is, each semantic label is assigned a representation template with basic semantic meaning. These prototypes do not directly participate in analysis, but serve as an intermediate bridge to connect video visual content and natural language semantics.
[0046] Further, the initialized language prototype can be embedded into a preset vector space, converted into a uniform dimension vector representation through an embedding mechanism, and a language vector embedding matrix is formed. The matrix carries the structured information of each semantic element in the language, and can be used for subsequent attention matching and alignment analysis of video spatio-temporal features.
[0047] In the financial field, this processing helps to understand the language expressions corresponding to specific business behaviors such as “withdrawal” and “queuing” in the target video; in the medical field, it can be used to correspond actions such as “surgical incision” and “instrument movement” in the target video to term expressions one by one, realize the structural modeling of professional language elements, and lay a foundation for subsequent semantic reasoning and intelligent question answering.
[0048] In some embodiments, the calculating the attention weight matrix of the video spatio-temporal features and the language vector embedding matrix to generate a structured semantic representation comprises: determining a semantic similarity of each frame of the video spatio-temporal features and the language vector embedding matrix; calculating the attention weight matrix of the video spatio-temporal features and the language vector embedding matrix based on the semantic similarity; weighting and aggregating the video spatio-temporal features and the language vector embedding matrix according to the attention weight matrix to obtain a semantic path distribution; and determining the semantic path distribution as the structured semantic representation.
[0049] Exemplarily, semantic similarity between the video spatio-temporal features and the language vector embedding matrix can be calculated first to determine the matching degree between the visual information expressed by each frame or each time unit and each type of language element (such as action, object, scene). The semantic similarity measure can be realized by dot product, cosine similarity or deep network prediction to quantify the correlation between the video content and the language description. Then, an attention weight matrix is constructed based on the similarity value, which is used to highlight the video feature area highly associated with the language semantics and provide accurate guidance for further information fusion.
[0050] Further, the video spatio-temporal features and the language vector embedding can be weighted and aggregated according to the attention weight matrix to extract the most representative video-language fusion information, and a semantic path distribution is constructed accordingly. The semantic path distribution reflects the dynamic association process of language elements in the video time axis and spatial structure. Finally, the semantic path distribution is determined as a structured semantic representation as the language understanding output of the video content.
[0051] In the financial field, this method can realize accurate mapping of semantic-level action paths such as "customer hands over ID card to teller"; in the medical field, it can be used to extract the semantic relationship between "doctor completes suturing action" and "instrument storage", supporting automatic intraoperative recording or intelligent evaluation system.
[0052] S40: analyzing the structured semantic representation by a text decoding network to generate a video text description corresponding to the target network.
[0053] Exemplarily, the structured semantic representation generated in the previous step can be parsed and converted by a text decoding network to decode the information containing the action, object and scene semantics in the video into natural language text. The decoding network usually adopts recurrent neural network, Transformer or other sequence generation model, which can automatically generate video description text conforming to grammar and semantic logic according to the structured semantic content. The final output video text description not only has fluent language expression ability, but also can accurately reflect the key behaviors and events in the target video.
[0054] In the financial field, this method can automatically generate business descriptions such as "customer performs transfer operation after submitting materials at the counter"; in the medical field, it can be used to generate intraoperative process records such as "doctor removes surgical instruments after completing suturing", realizing efficient information archiving and auxiliary diagnosis.
[0055] In some embodiments, the method further comprises: obtaining a training data set and a pre-trained model; wherein the training data set comprises a plurality of historical videos; labeling the training data set to obtain a labeling result, wherein the labeling result comprises a historical video text description corresponding to the historical videos; training the pre-trained model based on the training data set and the labeling result to obtain the text decoding network.
[0056] Based on the above embodiments, after obtaining the text decoding network, the method further comprises: iteratively training the text decoding network based on the training data set and the labeling result to extract data features and calculate a loss function; iteratively training the loss function using a preset method to reduce the value of the loss function until the value of the loss function is less than an expected threshold.
[0057] The loss function after iterative training is obtained, and the text decoding network after iteration is obtained.
[0058] Specifically, a training data set comprising a plurality of historical videos can be collected for training. For example, the training data set can be obtained by manual collection, web crawling, or a public data set, and the present application does not limit this.
[0059] Further, each of the plurality of historical videos can be labeled to obtain a labeling result corresponding to each of the plurality of historical videos. Then, the labeling result is used as a label of the input data, and each set of training data set carrying the label is input into the pre-trained model for supervised learning. When the training end condition is met, such as when the number of training times reaches a number threshold or the output accuracy of the model reaches an accuracy threshold, the training is ended, and a trained text decoding network is obtained.
[0060] In the embodiments of the present application, the training data set and the labeling result can be input into the pre-trained model for supervised learning, and then the text decoding network is trained. Thus, the text decoding network can be used to output the labeling result.
[0061] The above embodiments can enhance the data quality and diversity during the training of the text decoding network, and improve the generalization ability and practical application effect of the model.
[0062] It can be understood that, in order to train a text decoding network with higher accuracy, the text decoding network can be iteratively trained in a manner that reduces the loss function until the loss function meets the expected threshold, and then a more accurate labeling result can be obtained based on the text decoding network after iteration.
[0063] It should be noted that the preset method and the expected threshold are not limited in the present application, for example, the preset method can be a gradient descent algorithm, a batch gradient descent algorithm, a stochastic gradient descent algorithm, etc., and the present application takes the gradient descent algorithm as an example for description.
[0064] The purpose of the gradient descent algorithm is to find the minimum value of the loss function by iteration, or to converge to the minimum value. Geometrically, the gradient descent algorithm is to find the minimum value of the function in the direction opposite to the vector where the function changes fastest, and the gradient decreases fastest, so it is easier to find the minimum value of the function. Based on this, in the embodiment of the present application, the gradient descent algorithm is used to iteratively train the text decoding network so that the loss function is continuously reduced, thereby reducing the error of the calculation result.
[0065] In the embodiment of the present application, the gradient descent algorithm is used to iteratively train the text decoding network so that the loss function is continuously reduced to obtain the iterated text decoding network, and then a more accurate labeled result can be obtained based on the iterated text decoding network.
[0066] As can be seen in the above scheme, by extracting the frame sequence in the target video and constructing the video image block representation, the visual information of the target video can be fully retained. Further, the hierarchical temporal network is used to extract the spatio-temporal features of the target video, which not only enhances the modeling ability of dynamic changes and local details, but also provides multi-scale feature support for subsequent semantic understanding. By constructing a language vector embedding matrix and combining an attention mechanism, the key alignment relationship between the video content and the language description can be accurately captured, thereby generating a structured representation with high semantic consistency. Finally, the text decoding network is used to complete the conversion of the video content to natural language, and accurate description of the video events, behaviors, scenes and other contents is realized. In the financial field, this method can be used for intelligent analysis of transaction behaviors, counter operations, risk events and other contents in monitoring videos, and can improve the efficiency of compliance checking and fraud identification. In the medical field, it can assist in generating operation step descriptions of surgical videos, remote consultation image interpretation reports, etc., and realize automatic recording and knowledge archiving of medical images, thereby accelerating the structuralization and intelligentization of professional information.
[0067] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0068] In an embodiment, a text generation apparatus is provided, which corresponds to the text generation method described above. As shown in Figure 4As shown, the text generation apparatus includes an acquisition module 101, an analysis module 102, a semantic representation module 103, and a text generation module 104. The functions of each module are described in detail as follows:
[0069] The acquisition module 101 is configured to acquire a target video and extract a frame sequence in the target video to obtain a video image block representation.
[0070] The analysis module 102 is configured to analyze the video image block representation by a hierarchical temporal network to obtain a video spatio-temporal feature.
[0071] The semantic representation module 103 is configured to establish a language vector embedding matrix based on the target video, calculate an attention weight matrix of the video spatio-temporal feature and the language vector embedding matrix, and generate a structured semantic representation.
[0072] The text generation module 104 is configured to analyze the structured semantic representation by a text decoding network to generate a video text description corresponding to the target network.
[0073] The acquisition module 101 is configured to sample key frames from the target video to obtain key image frames, divide the key image frames into image blocks of a preset size to obtain video image blocks, and retain spatial position information in the original frames.
[0074] The analysis module 102 is configured to analyze the video image blocks by the conversion module to extract action features, aggregate the action features in a preset time dimension to obtain an intermediate state sequence, wherein the intermediate state sequence includes continuous action features, and analyze the intermediate state sequence by the attention module to obtain the video spatio-temporal feature.
[0075] The semantic representation module 103 is configured to determine a language element category in the target video, wherein the language element category includes an action type, an object category, and a scene element, perform an initialization operation on the language element category to obtain an initialized language prototype, embed the language prototype into an embedding vector of a preset dimension to obtain the language vector embedding matrix.
[0076] The semantic representation module 103 is configured to determine a semantic similarity between each frame of the video spatio-temporal feature and the language vector embedding matrix, calculate an attention weight matrix of the video spatio-temporal feature and the language vector embedding matrix based on the semantic similarity, weight and aggregate the video spatio-temporal feature and the language vector embedding matrix according to the attention weight matrix to obtain a semantic path distribution, and determine the semantic path distribution as the structured semantic representation.
[0077] In an embodiment, the acquisition module 101 is further configured to acquire a training data set and a pre-trained model; the training data set comprises a plurality of historical videos; the training data set is labeled to obtain a labeling result, wherein the labeling result comprises a historical video text description corresponding to the historical videos; and the pre-trained model is trained based on the training data set and the labeling result to obtain the text decoding network.
[0078] In an embodiment, the acquisition module 101 is further configured to iteratively train the text decoding network based on the training data set and the labeling result to extract data features and calculate a loss function; iteratively train the loss function using a preset method to reduce the value of the loss function until the value of the loss function is less than an expected threshold; and obtain an iteratively trained text decoding network based on the iteratively trained loss function.
[0079] The present application provides a text generation device, which can fully retain the visual information of the target video by extracting the frame sequence in the target video and constructing a video image block representation. Further, the hierarchical temporal network is used to extract the spatio-temporal features of the target video, which not only enhances the modeling capability of dynamic changes and local details, but also provides multi-scale feature support for subsequent semantic understanding. By constructing a language vector embedding matrix and combining an attention mechanism, the key alignment relationship between the video content and the language description can be accurately captured, thereby generating a structured representation with high semantic consistency. Finally, the text decoding network is used to complete the conversion from the video content to the natural language, realizing accurate description of the video events, behaviors, scenes and other contents. In the financial field, the method can be used for intelligent analysis of transaction behaviors, counter operations, risk events and the like in monitoring videos, improving the efficiency of compliance checking and fraud identification; in the medical field, it can assist in generating operation step descriptions of surgical videos, remote consultation image interpretation reports and the like, realizing automatic recording and knowledge archiving of medical images, thereby accelerating the structuralization and intelligentization of professional information.
[0080] The specific limitations of the text generation device can be referred to the limitations of the text generation method in the above, which will not be repeated here. Each module in the above text generation device can be realized by software, hardware and their combination in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to call and execute the operations corresponding to the above modules by the processor.
[0081] In one embodiment, a computer device is provided, which can be a server, and its internal structure diagram can be as shown in Figure 5As shown in the figure. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes non-volatile and / or volatile storage medium, internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with the external client through the network connection. The computer program is executed by the processor to realize the function or step of the server side of the text generation method.
[0082] In one embodiment, a computer device is provided, which can be a client, and its internal structure diagram can be as shown in the figure. Figure 6 As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes non-volatile storage medium, internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with the external server through the network connection. The computer program is executed by the processor to realize the function or step of the client side of the text generation method
[0083] In one embodiment, a computer device is provided, including a memory, a processor and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to realize the following steps:
[0084] Obtain a target video, and extract a frame sequence in the target video to obtain a video image block representation;
[0085] Analyze the video image block representation through a hierarchical temporal network to obtain a video spatio-temporal feature;
[0086] Establish a language vector embedding matrix based on the target video, and calculate an attention weight matrix of the video spatio-temporal feature and the language vector embedding matrix to generate a structured semantic representation;
[0087] Analyze the structured semantic representation through a text decoding network to generate a video text description corresponding to the target network.
[0088] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to realize the following steps:
[0089] acquire a target video, and extract a frame sequence in the target video to obtain a video image block representation;
[0090] analyze the video image block representation by a hierarchical temporal network to obtain a video spatio-temporal feature;
[0091] establish a language vector embedding matrix based on the target video, and calculate an attention weight matrix of the video spatio-temporal feature and the language vector embedding matrix to generate a structured semantic representation;
[0092] analyze the structured semantic representation by a text decoding network to generate a video text description corresponding to the target network.
[0093] It should be noted that the functions or steps described above in relation to the computer-readable storage medium or the computer device can correspond to the related descriptions of the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0094] Those skilled in the art can understand that all or part of the processes in the foregoing method embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the foregoing method embodiments. In the embodiments provided in the present application, any reference to a memory, storage, database or other medium can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.
[0095] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified. In actual applications, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0096] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those ordinarily skilled in the art should understand: the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A text generation method, characterized in that: The method comprises: Acquire a target video and extract a frame sequence from the target video to obtain a video image block representation; Analyzing the video image block representation through a hierarchical temporal network to obtain video spatiotemporal features; Establishing a language vector embedding matrix based on the target video, and calculating an attention weight matrix of the video spatiotemporal features and the language vector embedding matrix to generate a structured semantic representation; The structured semantic representation is analyzed through a text decoding network to generate a video text description corresponding to the target network.
2. The method according to claim 1, characterized in that The step of extracting a frame sequence from the target video to obtain a video image block representation includes: Performing key frame sampling on the target video to obtain key image frames; The key image frame is divided into image blocks of a preset size to obtain the video image blocks; wherein the video image blocks retain the spatial position information in the original frame.
3. The method according to claim 1, characterized in that The hierarchical temporal network includes a conversion module and an attention module. The hierarchical temporal network is used to analyze the video image block representation to obtain the video spatiotemporal features, including: Analyzing the video image blocks by the conversion module to extract motion features; Aggregating the action features in a preset time dimension to obtain an intermediate state sequence; wherein the intermediate state sequence includes continuous action features; The intermediate state sequence is analyzed by the attention module to obtain the video spatiotemporal features.
4. The method according to claim 1, wherein The establishing of a language vector embedding matrix based on the target video includes: Determining language element categories in the target video; wherein the language element categories include action types, object categories, and scene elements; Initializing the language element category to obtain an initialized language prototype; The language prototype is embedded in an embedding vector of a preset dimension to obtain the language vector embedding matrix.
5. The method according to claim 1, wherein The step of calculating an attention weight matrix of the video spatiotemporal features and the language vector embedding matrix to generate a structured semantic representation includes: Determining the semantic similarity between the spatiotemporal features of the video and the language vector embedding matrix for each frame; Calculate the attention weight matrix of the video spatiotemporal features and the language vector embedding matrix based on the semantic similarity; Performing weighted aggregation on the video spatiotemporal features and the language vector embedding matrix according to the attention weight matrix to obtain a semantic path distribution; The semantic path distribution is determined as the structured semantic representation.
6. The method according to claim 1, characterized in that The method further comprises: Obtaining a training data set and a pre-trained model; wherein the training data set includes a plurality of historical videos; Annotating the training data set to obtain an annotation result, wherein the annotation result includes a historical video text description corresponding to the historical video; The pre-training model is trained using the training data set and the annotation results to obtain the text decoding network.
7. The method according to claim 6, characterized in that After obtaining the text decoding network, the method further includes: Iteratively training the text decoding network based on the training data set and the annotation results to extract data features, and calculate the loss function; Iteratively training the loss function using a preset method for the purpose of reducing the loss function value until the loss function value is less than an expected threshold; Based on the loss function after iterative training, an iterative text decoding network is obtained.
8. A text generation device, characterized in that: The text generating device comprises: An acquisition module is used to acquire a target video and extract a frame sequence in the target video to obtain a video image block representation; An analysis module, configured to analyze the video image block representation through a hierarchical temporal network to obtain video spatiotemporal features; A semantic representation module is used to establish a language vector embedding matrix based on the target video, and calculate an attention weight matrix between the spatiotemporal features of the video and the language vector embedding matrix to generate a structured semantic representation; The text generation module is used to analyze the structured semantic representation through a text decoding network to generate a video text description corresponding to the target network.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the text generation method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the text generation method according to any one of claims 1 to 7 are implemented.