Video description model training method and device and video description method and device

Through the reconstruction of the original description text and multi-task learning strategy, the flexibility and diversity problems of description text tags in video description model training are solved, and accurate summary of training video content and high-quality description generation are achieved.

CN120708122APending Publication Date: 2025-09-26INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510799216.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing video description text tags lack flexibility and diversity, which makes it difficult to accurately summarize the video content through manual annotation, affecting video screening and retrieval.

Method used

By reconstructing the original description text, using more standardized sample description text to train the video description model, and combining multi-task learning strategies and visual feature extraction, the model training effect is improved.

Benefits of technology

It achieves accurate summary of the training video content, improves the training effect of the video description model and the quality of the generated description text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708122A_ABST
    Figure CN120708122A_ABST
Patent Text Reader

Abstract

The invention discloses a video description model training method and device and a video description method and device. The method comprises the following steps: reconstructing an original description text of a sample training video to obtain a sample description text of the sample training video; and training a video description model according to the sample description text and the sample training video. According to the embodiment of the invention, the training content in the training video can be accurately summarized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a video description model training method and device, and a video description method and device. Background Art

[0002] The bank will conduct business training regularly and use recorded videos as training materials. Employees can choose training based on the descriptive text corresponding to the materials.

[0003] The existing description text tags lack flexibility and diversity, and are difficult to adapt to different video content. They need to rely on a large amount of manual annotation, which often fails to accurately summarize the video content. As a result, the description text cannot standardize the training content in the training video, affecting the screening and retrieval of the video. Summary of the Invention

[0004] The present invention provides a video description model training method and device, and a video description method and device, so as to achieve accurate summary of training content in a training video.

[0005] According to one aspect of the present invention, a method for training a video description model is provided, comprising:

[0006] Reconstructing the original description text of the sample training video to obtain a sample description text of the sample training video;

[0007] The video description model is trained based on the sample description text and the sample training video.

[0008] According to another aspect of the present invention, a video description method is provided, comprising:

[0009] Generate target description text for target training videos through the video description model;

[0010] The video description model is obtained by training using the video description model training method described in any embodiment of the present invention.

[0011] According to another aspect of the present invention, a video description model training device is provided, comprising:

[0012] A description reconstruction module is used to reconstruct the original description text of the sample training video to obtain a sample description text of the sample training video;

[0013] The model training module is used to train the video description model based on the sample description text and the sample training video.

[0014] According to another aspect of the present invention, there is provided a video description apparatus, comprising:

[0015] A video description module is used to generate target description text of the target training video through the video description model;

[0016] The video description model is obtained by training using the video description model training device described in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the video description model training method or the video description method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the video description model training method or the video description method described in any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the video description model training method or video description method described in any embodiment of the present invention when executed.

[0020] The embodiment of the present invention reconstructs the existing original description text and uses more standardized sample description text to train the video description model. By improving the text quality of the description text, the training effect of the video description model is effectively improved, and an accurate summary of the training content in the training video is achieved.

[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 is a flowchart of a method for training a video description model according to an embodiment of the present invention;

[0024] Figure 2 is a flowchart of a method for training a video description model according to another embodiment of the present invention;

[0025] Figure 3 is a flowchart of a video description method provided according to another embodiment of the present invention;

[0026] Figure 4 is a structural diagram of a video description model training device provided according to another embodiment of the present invention;

[0027] Figure 5 is a structural diagram of a video description device provided according to another embodiment of the present invention;

[0028] Figure 6 It is a schematic structural diagram of an electronic device implementing an embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," and the like in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.

[0031] Figure 1 This is a flow chart of a method for training a video description model provided by one embodiment of the present invention. This embodiment is applicable to the case where a neural network model is used instead of manual description of a training video of banking content. This method can be executed by a training device for a video description model. The device can be implemented in the form of hardware and / or software and can be configured in an electronic device with corresponding data processing capabilities. Figure 1 As shown, the method includes:

[0032] S110 : Reconstruct the original description text of the sample training video to obtain a sample description text of the sample training video.

[0033] S120: Training a video description model according to the sample description text and the sample training video.

[0034] The original description text is the description text given after manually summarizing the training content in the training video. The sample description text is a more standardized expression of the original description text with the same meaning. The video description model is a neural network model based on a specific network structure. It can recognize visual information such as objects, actions, scenes, and temporal dynamic information in the video and convert it into a coherent text description. The visual information in the video is extracted through the visual feature extraction model, and this information is generated into descriptive text through the language generation model. This technology can process complex video content and generate natural language descriptions that closely match the video and are highly accurate, improving the quality of video content analysis and automatic generation.

[0035] Specifically, training videos with corresponding description texts are collected on the training system, and these videos are used as sample training videos for model training. Since the original description texts are generally unable to express the training content in the training videos in a standardized manner, if the original description texts are used directly for model training, it will affect the model's ability to summarize the videos. For this reason, the original description texts are reconstructed before model training, and a more standardized expression, namely the sample description text, is used to replace the original description texts. For example, if the original description text is "necessary preparations before applying for a personal loan", after reconstruction, the sample description text is "credit analysis and document preparation before applying for a personal loan". The video features of the sample training video are input into the video description model to obtain the output results of the video description model. The differences between the output results and the corresponding sample description texts are compared, and the network structure and parameters of the video description model are updated accordingly until the training results of the video description model meet expectations and the training is completed.

[0036] The embodiment of the present invention reconstructs the existing original description text and uses more standardized sample description text to train the video description model. By improving the text quality of the description text, the training effect of the video description model is effectively improved, and an accurate summary of the training content in the training video is achieved.

[0037] Figure 2 This is a flow chart of a training method for a sample training video provided by another embodiment of the present invention. This embodiment is optimized and improved on the basis of the above embodiment. Figure 2 As shown, the method includes:

[0038] S210 , based on the title and original description text of the sample training video, obtaining the operation process of the training content in the sample training video from the operation process database.

[0039] S220 : Based on the large language model, reconstruct the original description text of the sample training video through the operation process to obtain a sample description text of the sample training video.

[0040] Among them, the operation process database stores the operation processes of various operations within the bank.

[0041] Specifically, the titles and original description texts of the sample training videos are summarized to infer the possible training content that may appear in the training videos, and based on this, the work flow of the training content is obtained from the work flow database. The large language model is fine-tuned in advance to enable it to have the ability to reconstruct the input original text. The work flow, original description text, and reconstruction task prompts are input into the large language model. The large language model understands the knowledge in the relevant materials and uses the professional vocabulary in the work flow to reconstruct the input original text, including concept explanation and synonym replacement, to obtain a sample description text composed of professional vocabulary. By using a standardized work flow to reconstruct the original description text, the expression accuracy of the sample description text is further improved.

[0042] Based on the above embodiment, optionally, obtaining the operation flow of the training content in the sample training video from the operation flow database based on the title and original description text of the sample training video includes:

[0043] Perform keyword extraction on the title and original description text of the sample training video to obtain summary keywords;

[0044] The summary keywords are searched in the operation process database to obtain the operation process of the training content in the sample training video.

[0045] Specifically, the title and original description text are segmented and keywords are extracted to obtain summary keywords. The summary keywords reflect the work content involved in the sample training video. The summary keywords are used as the search object to search in the work process database. The retrieved work process is the work process of the training content in the training video. By extracting and searching with keywords, the analysis efficiency of the title and original description text is improved, and the work process of the training content in the sample training video is located more quickly.

[0046] S230. Establish a multi-task learning strategy with video description generation as the main focus and video text matching as the auxiliary focus.

[0047] S240: Based on the sample description text and the sample training video, the video description model is trained using the multi-task learning strategy.

[0048] Multi-task learning (MTL) is a machine learning paradigm whose core idea is to improve the generalization performance of a model by learning multiple related tasks simultaneously. Compared with single-task learning, multi-task learning can leverage the correlation between tasks and share representation information, thereby improving the learning effect of each task.

[0049] Specifically, convolutional neural networks (RestNet, VGG) or transformer-based visual feature extraction models (VIT, Swin Transformer) are used to extract high-level video semantic features from sample training videos. Sample description text is then processed using a transformer model to obtain high-level textual semantics. A pretrained natural language generation model is loaded. During training, a primary task and an auxiliary task are designed for the pretrained natural language generation model. The primary task is to generate corresponding description text based on video features (cross-entropy loss); the auxiliary task is to determine whether the video matches the given text (binary cross-entropy loss). During training, these two tasks are trained simultaneously. Pretrained natural language generation models that meet the expected training results in both aspects are designated as video description models. The matching task requires the model to understand the video and text from the perspective of "determining relevance," while the generation task requires the model to understand them from the perspective of "constructing descriptions." This joint training enables the model to learn cross-modal relationships from different perspectives, enhancing its generalization capabilities.

[0050] The embodiment of the present invention further improves the expression accuracy of the sample description text by reconstructing the original description text using a standardized operation process.

[0051] Figure 3 This is a flowchart of a video description method provided by another embodiment of the present invention. This embodiment is applicable to situations where it is necessary to summarize internal training videos of banks. The method can be executed by a video description device, which can be implemented in the form of hardware and / or software and can be configured in an electronic device with corresponding data processing capabilities. Figure 3 As shown, the method includes:

[0052] S310: Generate target description text for the target training video through a video description model.

[0053] The video description model is obtained by training using the video description model training method described in any embodiment of the present invention.

[0054] Specifically, the target training video to be described is received and preprocessed, including video framing and video frame cropping steps. Video framing is the process of decomposing the target training video into a sequence of static images by frame. A video is actually composed of a series of rapidly played continuous image frames. The purpose of framing is to extract these individual images from the video for further analysis and processing. Video cropping is the process of spatially cropping video frames, that is, cutting out specific areas in the video frame, removing unnecessary background or noise, and focusing on key information. Feature extraction is performed on the preprocessed target training video, and the extracted video features are input into the video description model to obtain the target description text generated by the video description model.

[0055] The embodiment of the present invention does not require human intervention, automatically analyzes the video and generates detailed text descriptions, which is of great significance for large-scale video content processing and generation. For example, when recording some training videos in a bank, some training copy can be easily added, saving a lot of manpower.

[0056] Based on the above embodiment, optionally, generating a target description text of a target training video by using a video description model includes:

[0057] Based on the autoregressive strategy, a video description model is used to generate at least two description words for the target training video.

[0058] The at least two description words are aggregated through a large language model to obtain a target description text of the target training video.

[0059] Specifically, when generating video text descriptions using the autoregressive generation strategy, the model relies on the already generated word sequence and the already input video features to gradually predict the next descriptive word. Specifically, the model first extracts the temporal or global visual features of the video as an encoding vector, then initializes the production process. At each time step, the previously generated words and video features are input into the decoder transformer, gradually generating descriptive words that match the video content, ultimately obtaining multiple descriptive words at different time steps. Multiple descriptive words are input into the large model together with corresponding summary prompts. The large model organizes the input words and outputs a summary of their process specifications as the target description text for the target training video. By adopting the autoregressive generation strategy, the coherence of the target description text can be guaranteed.

[0060] Figure 4 A structural diagram of a video description model training device provided by another embodiment of the present invention. Figure 4 As shown, the device includes:

[0061] A text reconstruction module is used to reconstruct the original description text of the sample training video to obtain a sample description text of the sample training video;

[0062] The model training module is used to train the video description model based on the sample description text and the sample training video.

[0063] The video description model training device provided in the embodiment of the present invention can execute the video description model training method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0064] Optionally, the text reconstruction module includes:

[0065] A process acquisition unit is used to acquire the operation process of the training content in the sample training video from the operation process database based on the title and original description text of the sample training video;

[0066] The text reconstruction unit is used to reconstruct the original description text of the sample training video based on the large language model through the operation process to obtain the sample description text of the sample training video.

[0067] Optionally, the text reconstruction unit is specifically used to: extract keywords from the title and original description text of the sample training video to obtain summary keywords; search the summary keywords in the operation process database to obtain the operation process of the training content in the sample training video.

[0068] Optionally, the model training module is specifically used to: establish a multi-task learning strategy with video description generation as the main focus and video text matching as the auxiliary focus; and train the video description model through the multi-task learning strategy based on the sample description text and the sample training video.

[0069] The video description model training device further described can also execute the video description model training method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0070] Figure 5 FIG. 1 is a structural diagram of a video description device provided by another embodiment of the present invention. Figure 5 As shown, the device includes:

[0071] The video description module is used to generate target description text for the target training video through the video description model.

[0072] The video description device provided in the embodiment of the present invention can execute the video description method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0073] Optionally, the video description module is specifically configured to: generate at least two description words of the target training video through the video description model based on an autoregressive strategy; and aggregate the at least two description words through a large language model to obtain a target description text of the target training video.

[0074] The video description device further described can also execute the video description method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0075] Figure 6 A schematic diagram of the structure of an electronic device 60 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0076] like Figure 6 As shown, the electronic device 60 includes at least one processor 61 and a memory, such as a read-only memory (ROM) 62, a random access memory (RAM) 63, etc., which is communicatively connected to the at least one processor 61. The memory stores a computer program that can be executed by the at least one processor. The processor 61 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 62 or the computer program loaded from the storage unit 68 into the random access memory (RAM) 63. Various programs and data required for the operation of the electronic device 60 can also be stored in the RAM 63. The processor 61, ROM 62, and RAM 63 are connected to each other via a bus 64. An input / output (I / O) interface 65 is also connected to the bus 64.

[0077] Multiple components in the electronic device 60 are connected to the I / O interface 65, including an input unit 66, such as a keyboard, a mouse, etc.; an output unit 67, such as various types of displays, speakers, etc.; a storage unit 68, such as a magnetic disk, an optical disk, etc.; and a communication unit 69, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 69 allows the electronic device 60 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0078] The processor 61 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 61 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 61 executes the various methods and processes described above, such as the video description model training method or the video description method.

[0079] In some embodiments, the training method of the video description model or the video description method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 68. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 60 via the ROM 62 and / or the communication unit 69. When the computer program is loaded into the RAM 63 and executed by the processor 61, one or more steps of the training method of the video description model or the video description method described above can be performed. Alternatively, in other embodiments, the processor 61 can be configured to execute the training method of the video description model or the video description method by any other appropriate means (for example, by means of firmware).

[0080] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0081] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0082] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0083] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0084] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0085] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0086] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0087] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A video description model training method, characterized in that: The method comprises: Reconstructing the original description text of the sample training video to obtain a sample description text of the sample training video; The video description model is trained based on the sample description text and the sample training video.

2. The method according to claim 1, characterized in that The reconstructing of the original description text of the sample training video to obtain the sample description text of the sample training video includes: Based on the title and original description text of the sample training video, obtaining the operation process of the training content in the sample training video from the operation process database; Based on the large language model, the original description text of the sample training video is reconstructed through the operation process to obtain the sample description text of the sample training video.

3. The method according to claim 2, characterized in that The operation process of obtaining the training content in the sample training video from the operation process database based on the title and original description text of the sample training video includes: Perform keyword extraction on the title and original description text of the sample training video to obtain summary keywords; The summary keywords are searched in the operation process database to obtain the operation process of the training content in the sample training video.

4. The method according to claim 1, wherein The training of the video description model according to the sample description text and the sample training video includes: Establish a multi-task learning strategy with video description generation as the main focus and video text matching as the auxiliary; Based on the sample description text and the sample training video, the video description model is trained through the multi-task learning strategy.

5. A method for training a video description model, characterized in that: The method comprises: Generate target description text for target training videos through the video description model; The video description model is obtained by training the video description model according to any one of claims 1 to 4.

6. The method according to claim 5, characterized in that Generating a target description text of a target training video by using a video description model includes: Based on the autoregressive strategy, a video description model is used to generate at least two description words for the target training video. The at least two description words are aggregated through a large language model to obtain a target description text of the target training video.

7. A video description model training device, characterized in that: The device comprises: A text reconstruction module is used to reconstruct the original description text of the sample training video to obtain a sample description text of the sample training video; The model training module is used to train the video description model based on the sample description text and the sample training video.

8. A video description device, characterized in that: The device comprises: A video description module is used to generate target description text of the target training video through the video description model; The video description model is obtained by training using the video description model training device according to claim 7.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the video description model training method described in any one of claims 1 to 4 or the video description method described in any one of claims 5 to 6.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the video description model training method according to any one of claims 1 to 4 or the video description method according to any one of claims 5 to 6.