Text generation method and device, electronic equipment and storage medium
By extracting event features in videos and stitching together to generate prompt texts in the extraction order, the problem that the video description text generation method in the prior art cannot accurately perceive the content and order of video events, and achieve better text generation effect.
Patent Information
- Application Number
- CN202311695859.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-13
AI Technical Summary
Existing video description text generation methods cannot accurately perceive the content and order of different events in the video, resulting in poor text generation results.
By extracting each event in the video to be processed, and determining the target video frame corresponding to each event; extracting the frame features of each target video frame to determine each event feature; splicing each event feature in the extraction order to generate a prompt text; and then generating a description text based on the prompt text through the first language model.
By splicing event features in the order of extraction, and entering them into the first language model, the model can more accurately perceive the content and order of different events in the video, thereby improving the text generation effect.
Smart Images

Figure CN120144772A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and in particular, to a method, apparatus, electronic device, and storage medium for text generation. Background Art
[0002] The existing methods for generating video description texts usually generate description texts based on the video frame features of the entire video through a generation model. In the existing methods, the generation model cannot accurately perceive the content and order of different events in the video, and it is easy to generate incorrect descriptions, resulting in poor text generation effects. Summary of the Invention
[0003] Embodiments of the present disclosure provide a method, apparatus, electronic device, and storage medium for text generation, which can improve the text generation effect.
[0004] In a first aspect, embodiments of the present disclosure provide a method for text generation, including:
[0005] extracting each event in the video to be processed, and determining each target video frame corresponding to each event;
[0006] extracting the frame features of each target video frame, and determining each event feature according to each frame feature;
[0007] concatenating each event feature in the extraction order of the corresponding event to generate a prompt text;
[0008] generating a description text of the video to be processed based on the prompt text through a first language model.
[0009] In a second aspect, embodiments of the present disclosure further provide a text generation apparatus, including:
[0010] a video frame determination module, configured to extract each event in the video to be processed, and determine each target video frame corresponding to each event;
[0011] an event feature determination module, configured to extract the frame features of each target video frame, and determine each event feature according to each frame feature;
[0012] a prompt text generation module, configured to concatenate each event feature in the extraction order of the corresponding event to generate a prompt text;
[0013] a description text generation module, configured to generate a description text of the video to be processed based on the prompt text through a first language model.
[0014] In a third aspect, embodiments of the present disclosure further provide an electronic device, where the electronic device includes:
[0015] One or more processors;
[0016] A storage device for storing one or more programs,
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the text generation method according to any one of the embodiments of the present disclosure.
[0018] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the text generation method according to any one of the embodiments of the present disclosure when executed by a computer processor.
[0019] The technical solution of the embodiment of the present disclosure extracts each event in the video to be processed and determines each target video frame corresponding to each event; extracts the frame features of each target video frame and determines each event feature according to each frame feature; splices each event feature in the extraction order of the corresponding event to generate a prompt text; and generates a description text of the video to be processed based on the prompt text through a first language model. By splicing the event features in the extraction order of the corresponding event to obtain the prompt text and inputting the prompt text into the first language model, the first language model can perceive the content and order of different events in the video, and the text generation effect can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In combination with the drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.
[0021] Figure 1 It is a schematic flowchart of a text generation method provided by an embodiment of the present disclosure;
[0022] Figure 2 It is a schematic block diagram of a text generation method provided by an embodiment of the present disclosure;
[0023] Figure 3 It is a schematic diagram of a prompt text in a text generation method provided by an embodiment of the present disclosure;
[0024] Figure 4 It is a schematic block diagram of a text generation method provided by an embodiment of the present disclosure;
[0025] Figure 5 It is a schematic structural diagram of a text generation device provided by an embodiment of the present disclosure;
[0026] Figure 6A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0027] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0028] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0029] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0030] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions executed by these devices, modules or units or their interdependent relationships.
[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0034] Figure 1The flowchart shows a text generation method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the scenario of generating descriptive text for videos. This method can be executed by a text generation device, which can be implemented in the form of software and / or hardware, and can be configured in an electronic device, such as a computer.
[0035] As Figure 1 shown, the text generation method provided in this embodiment may include:
[0036] S110. Extract each event in the video to be processed, and determine the target video frames corresponding to each event.
[0037] In the embodiment of the present disclosure, the video to be processed can be regarded as a video for which descriptive text needs to be generated. For example, it can be a long video or a short video. After obtaining the video to be processed, the boundaries between different shots in the video to be processed can be detected based on existing shot boundary detection algorithms (such as TransNet-V2, etc.); and shot segmentation can be performed according to the boundaries between shots. Among them, the video frames of each shot obtained by segmentation can be regarded as belonging to different events, so as to complete the extraction of each event in the video to be processed.
[0038] Among them, the target video frames corresponding to the event can be determined according to the video frames belonging to the same event. For example, all the video frames belonging to the same event can be used as the target video frames of the event. Another example is that the video frames belonging to the same event can be sampled to obtain the target video frames of the event, so as to save computing resources to a certain extent and improve the text generation efficiency.
[0039] Among them, when sampling the video frames belonging to the same event, for example, the uniform sampling method can be used to sample the target video frames from the video frames belonging to the same event. Another example is that the target video frames can be sampled from the video frames belonging to the same event according to the quality parameters of the video frames (such as brightness, etc.). In addition, other ways of sampling to obtain the target video frames can also be applied here, and no specific limitation is made here.
[0040] Among them, the number of target video frames corresponding to different events can be the same or different. Among them, the number of target video frames can be set according to the actual application scenario to represent the essential information of different events based on a smaller number of target video frames. Exemplarily, through research, it can be obtained that 4 target video frames can be extracted to represent each event to balance the text generation effect and the generation efficiency.
[0041] Exemplarily, Figure 2 The block diagram shows a text generation method provided by an embodiment of the present disclosure. Refer to Figure 2 , three events can be extracted from the video to be processed, and 4 target video frames corresponding to each event can be determined.
[0042] S120. Extract the frame features of each target video frame, and determine each event feature according to each frame feature.
[0043] In the embodiments of the present disclosure, the frame features of each target video frame can be extracted based on the existing image feature extraction algorithm. Exemplarily, Figure 2 the VIT-G / 14 of the Contrastive Language-Image Pre-Training (CLIP) can be used to extract the frame features of each target video frame. At this time, the size of the input target video frame can be transformed into an image of 3×224×224, where 3 represents the number of channels of the image, and 224×224 represents the width and height of the image; and the feature extraction can be performed on the target frame image to obtain an image of the frame feature with a size of 256×1408.
[0044] Among them, the frame features of each target video frame belonging to each event can be used to determine the event feature corresponding to the event. Among them, the event feature can be considered as a feature that can characterize the essence of the event and can be decoded by the language model. Among them, processing methods such as fusion and compression can be adopted to process the frame features corresponding to the same event to obtain the event feature corresponding to the event.
[0045] In some alternative implementation manners, determining each event feature according to each frame feature may include: converting each frame feature into the input space of the first language model to obtain each converted feature; and determining each event feature according to each converted feature.
[0046] Among them, based on the existing algorithm for converting image features into text features, the frame feature can be converted into the input space of the first language model to obtain a converted feature, where the converted feature can be considered as belonging to text features. Exemplarily, Figure 2 the Querying Transformer (Q-Former) in the Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (BLIP-2) can be used to convert the frame feature into a text feature. At this time, the frame feature with a size of 256×1408 can be compressed into a converted feature with a size of 32×768.
[0047] Among them, the Q-Former can be pre-trained to learn visual representations related to text and make the visual representations interpretable by the first language model, that is, an extraction model that can effectively utilize the image features of frozen parameters and the first language model can be used to realize the conversion from image features to text features. In addition, based on other multi-modal transformer algorithms, the frame features can also be converted into text features, which will not be specifically limited here.
[0048] See again Figure 2 , the conversion features corresponding to the events can be combined in any combination manner to obtain the event features of the events. For example, in some implementation manners, determining the event features according to the conversion features may include at least one of the following: splicing the conversion features to obtain the event features; performing interactive compression on the conversion features to obtain the event features. Among them, any splicing manner can be adopted to splice the conversion features to obtain the event features. For example, the conversion features can be spliced in the head-to-tail splicing manner to obtain the event features. Among them, existing interactive compression methods can be adopted, such as performing interactive compression on the conversion features based on the transformer algorithm.
[0049] In these alternative implementation manners, by first compressing each frame feature into each text feature and then combining the text features, the event features are determined according to each frame feature. In addition, the frame features can also be combined first and then compressed to determine the event features.
[0050] S130. Splice the event features in the extraction order corresponding to the events to generate a prompt text.
[0051] In the embodiments of the present disclosure, the corresponding event features can be spliced in sequence according to the extraction sequence of the events to generate an advance text. It can be considered that the prompt text contains both event features representing different event contents and a splicing order representing the occurrence order of different events.
[0052] In some alternative implementation manners, splicing the event features in the extraction order corresponding to the events may include: splicing the event features in the extraction order corresponding to the events based on a predefined order prompt word and a generation instruction prompt word.
[0053] Among them, different order prompt words can be preset as splicing connection words representing event boundaries and orders. Any prompt word that can indicate event boundaries and orders can be set as an order prompt word. Exemplarily, Figure 2 the order prompt words in may include "The first event is <>; The second event is <>;...;".
[0054] Among them, an instruction generation prompt word can be preset to indicate the end of the text prompted by the first language model. Any prompt word that can serve as an instruction for generation can be set as a generation quality prompt word. Exemplarily, Figure 2 the generation quality prompt word in
[0055] Among them, based on the predefined sequential prompt word and the instruction generation prompt word, splicing each event feature in the extraction order corresponding to the event may include: splicing each event feature after the corresponding sequential prompt word according to the extraction order of each event to generate an event sequential prompt text; splicing the instruction generation prompt word after the event sequential prompt text to generate a prompt text.
[0056] Among them, the extraction order of each event may include the sequence of "first, second...", and the sequential prompt word may also represent the sequence of "first, second...". It can be considered that there is a corresponding relationship between the extraction order of the event and the sequential prompt word, and the event feature corresponding to the extraction order can be spliced after the sequential prompt word according to this corresponding relationship to obtain the event sequential prompt text. See Figure 2 , Figure 2 "The first event is <event feature 1>; the second event is <event feature 2>;...;" in Figure 2 is the event sequential prompt text. See again
[0057] In these optional implementation manners, based on the predefined sequential prompt word and the instruction generation prompt word, a prompt text that can represent the event content and the event order can be generated.
[0058] S140. Through the first language model, generate a description text of the video to be processed based on the prompt text.
[0059] Among them, the first language model may include an existing pre-trained language model with the ability to generate text based on text. For example, it can be the Vicuna V0-7B model, etc. By inputting the prompt text into the first language model, the first language model can output a description text that can represent the event order in the video to be processed. Exemplarily, Figure 2 the generated description text in
[0060] The technical solution of the embodiment of the present disclosure extracts each event in the video to be processed and determines each target video frame corresponding to each event; extracts the frame features of each target video frame and determines each event feature according to each frame feature; splices each event feature in the extraction order of the corresponding event to generate a prompt text; and through a first language model, generates a description text of the video to be processed based on the prompt text. By splicing the event features in the extraction order of the corresponding event to obtain a prompt text and inputting the prompt text into the first language model, the first language model can perceive the content and order of different events in the video, which can improve the text generation effect.
[0061] The embodiments of the present disclosure can be combined with each optional solution in the text generation method provided in the above embodiments. The text generation method provided in this embodiment details the generation process of the prompt text. By adding the speech text and / or audio features of the video to be processed to the prompt text, the generated content of the description text can be further enriched and the generation quality of the description text can be improved.
[0062] Figure 3 is a schematic diagram of the prompt text in a text generation method provided by an embodiment of the present disclosure. It can be considered that Figure 3 the prompt text of Figure 2 is a supplement based on Figure 2 part. For details not described in detail about its generation, reference can be made to
[0063] In the embodiment of the present disclosure, when the video to be processed includes audio data, the text generation method may further include: performing speech recognition on the audio data to obtain speech text; correspondingly, when generating the prompt text, it may further include: splicing the spliced event features with the speech text to generate the prompt text.
[0064] Among them, the audio data of the video to be processed can be subjected to speech recognition based on an existing Automatic Speech Recognize (ASR) algorithm to obtain speech text. Among them, the spliced event features can be considered to include Figure 2 the event order prompt text in Figure 3 As shown in a, the speech text can be spliced after the event order prompt text and before the generation instruction prompt word based on the speech text prompt word "The speech text is <>. Thus, in the generated description text, in addition to including the description corresponding to the vision, it also includes the description corresponding to the audio, which can further enrich the content of the description text and improve the generation effect of the description text. By using the speech text as the content of the prompt text, the first language model can directly use the speech text content to generate the description text, which can save the time-consuming of generating the description text to a certain extent.
[0065] In the embodiments of the present disclosure, when the video to be processed includes audio data, the text generation method may further include: extracting features from the audio data to obtain audio features; correspondingly, generating a prompt text further includes: splicing the spliced event features and the audio features to generate a prompt text.
[0066] Among them, based on the existing audio feature extraction algorithm, initial feature extraction can be performed on the audio data of the video to be processed, and the extracted initial features can be transformed into the input space of the first language model to obtain audio features. Among them, the spliced event features can be considered to include Figure 2 the event sequence prompt text in. As Figure 3 shown in b, the audio features can be spliced after the event sequence prompt text and before the generation instruction prompt word based on the audio feature prompt word "the audio feature is <>". Thus, in the generated description text, in addition to including the description corresponding to the vision, it can also include the description corresponding to the audio, improving the generation effect of the description text. In addition, compared with splicing the speech text, the prompt text can be refined to a certain extent.
[0067] The technical solution of the embodiments of the present disclosure describes the generation process of the prompt text in detail. By adding the speech text and / or audio features of the video to be processed to the prompt text, the generation content of the description text can be further enriched, and the generation quality of the description text can be improved. The text generation method provided by the embodiments of the present disclosure and the text generation method provided by the above embodiments belong to the same general inventive concept. The technical details not described in detail in this embodiment can be referred to the above embodiments, and the same technical features have the same beneficial effects in this embodiment and the above embodiments.
[0068] The various alternative solutions in the text generation methods provided in the embodiments of the present disclosure and the above embodiments can be combined. The text generation method provided in this embodiment describes the application scenario of the generated video description text in detail. In the video question and answer scenario, after generating the description text of the video once, for multiple questions, there is no need to repeatedly process the features of the video, and multiple answer texts can be generated only based on this description text, which can save computing costs and improve the answering efficiency. In addition, in the process of generating the answer text, since both the description text and the answer text are in the text modality, the advantages of single-modal fusion can be fully utilized to improve the accuracy of the answer, avoiding the problem of low accuracy caused by poor multi-modal feature fusion when introducing video modality features.
[0069] Figure 4 It is a block diagram schematic diagram of a text generation method provided by an embodiment of the present disclosure. As Figure 4 shown, the text generation method provided in this embodiment may include:
[0070] S410. Obtain the video to be processed and the problem text of the video to be processed.
[0071] In this embodiment, the problem text may include, for example, entity extraction and entity relationship extraction in the video.
[0072] S420. Determine whether to store the description text of the video to be processed; if yes, jump to S470, if not, jump to S430.
[0073] In this embodiment, during the first answering process, the description text has not been generated yet. At this time, it can be considered that the description text to be processed has not been stored, and the description text can be generated based on the steps of S430 - S460. If it is not the first answering, it can be considered that the description text has been generated and stored. At this time, the description text can be directly obtained to perform the answering steps. After generating the description text of the video once, for multiple questions, there is no need to repeatedly process the features of the video. Only based on this description text, multiple answer texts can be generated, which can save computational costs and improve the answering efficiency.
[0074] S430. Extract each event in the video to be processed and determine each target video frame corresponding to each event.
[0075] S440. Extract the frame features of each target video frame and determine each event feature according to each frame feature.
[0076] S450. Concatenate each event feature in the extraction order of the corresponding event to generate a prompt text.
[0077] S460. Based on the prompt text, generate the description text of the video to be processed through the first language model.
[0078] S470. Through the second language model, generate the answer text based on the description text and the problem text.
[0079] In the embodiments of the present disclosure, the second language model can be considered as a pre-trained language model with the ability to generate one text based on two texts. By inputting the description text and the problem text into the second language model, the second language model can determine the answer text corresponding to the problem text in the description text.
[0080] The technical solution of the embodiment of the present disclosure details the application scenarios of the generated video description text. In the video question and answer scenario, after generating the description text of the video once, for multiple questions, it is not necessary to repeatedly process the features of the video. Only based on this description text, multiple answer texts can be generated, which can save computing costs and improve the answering efficiency. In addition, during the generation process of the answer text, since both the description text and the answer text are in the text modality, the advantages of single-modal fusion can be fully utilized to improve the accuracy of the answer, avoiding the problem of low accuracy caused by poor multi-modal feature fusion when introducing video modality features.
[0081] In addition, the text generation method provided in the embodiment of the present disclosure and the text generation method provided in the above embodiment belong to the same general concept. The technical details not described in detail in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.
[0082] Figure 5 It is a schematic structural diagram of a text generation device provided by an embodiment of the present disclosure. The text generation device provided in this embodiment is applicable to the situation of generating the description text of a video.
[0083] As Figure 5 shown, the text generation device provided by the embodiment of the present disclosure may include:
[0084] A video frame determination module 510, configured to extract each event in the video to be processed and determine each target video frame corresponding to each event;
[0085] An event feature determination module 520, configured to extract the frame features of each target video frame and determine each event feature according to each frame feature;
[0086] A prompt text generation module 530, configured to splice each event feature in the extraction order of the corresponding event to generate a prompt text;
[0087] A description text generation module 540, configured to generate the description text of the video to be processed based on the prompt text through a first language model.
[0088] In some optional implementation manners, the event feature determination module may be configured to:
[0089] Convert each frame feature into the input space of the first language model to obtain each converted feature;
[0090] Determine each event feature according to each converted feature.
[0091] In some optional implementation manners, the event feature determination module may be configured to perform at least one of the following:
[0092] Concatenate each conversion feature to obtain each event feature;
[0093] Perform interactive compression on each conversion feature to obtain each event feature.
[0094] In some alternative implementation manners, the prompt text generation module can be used for:
[0095] Based on the predefined sequential prompt words and generation instruction prompt words, concatenate each event feature in the extraction order of the corresponding event.
[0096] In some alternative implementation manners, the prompt text generation module can be used for:
[0097] According to the extraction order of each event, concatenate each event feature after the corresponding sequential prompt word to generate an event sequential prompt text;
[0098] Concatenate the generation instruction prompt word after the event sequential prompt text to generate a prompt text.
[0099] In some alternative implementation manners, when the video to be processed contains audio data, the text generation device may further include:
[0100] A speech recognition module, configured to perform speech recognition on the audio data to obtain a speech text;
[0101] Correspondingly, the prompt text generation module can also be used for: concatenating the concatenated event features with the speech text to generate a prompt text.
[0102] In some alternative implementation manners, when the video to be processed contains audio data, the text generation device may further include:
[0103] An audio feature determination module, configured to extract features from the audio data to obtain audio features;
[0104] Correspondingly, the prompt text generation module can also be used for: concatenating the concatenated event features with the audio features to generate a prompt text.
[0105] In some alternative implementation manners, the text generation device may further include a question and answer module;
[0106] The question and answer module can be used for obtaining the question text of the video to be processed after generating the description text;
[0107] Generate an answer text based on the description text and the question text through a second language model.
[0108] The text generation device provided by the embodiments of the present disclosure can execute the text generation method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0109] It should be noted that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and are not used to limit the protection scope of the embodiments of the present disclosure.
[0110] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device (such as the terminal device or server in Figure 6 ) 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0111] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0112] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 shows the electronic device 600 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.
[0113] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the text generation method of the embodiment of the present disclosure are executed.
[0114] The electronic device provided by the embodiment of the present disclosure and the text generation method provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0115] An embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the text generation method provided by the above embodiment is implemented.
[0116] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0117] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0118] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0119] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to:
[0120] Extract each event in the video to be processed and determine each target video frame corresponding to each event; extract the frame features of each target video frame and determine each event feature according to each frame feature; splice each event feature in the extraction order of the corresponding event to generate a prompt text; and generate a description text of the video to be processed based on the prompt text through a first language model.
[0121] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, through the Internet using an Internet service provider).
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0123] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the names of the units and modules do not constitute a limitation on the units and modules themselves.
[0124] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0125] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0126] According to one or more embodiments of the present disclosure, there is provided a text generation method, the method comprising:
[0127] extracting each event in the video to be processed and determining each target video frame corresponding to each event;
[0128] extracting frame features of each of the target video frames and determining each event feature according to each of the frame features;
[0129] concatenating the event features in the extraction order of the corresponding events to generate a prompt text;
[0130] generating a description text of the video to be processed based on the prompt text through a first language model.
[0131] According to one or more embodiments of the present disclosure, there is provided a text generation method, further comprising:
[0132] In some alternative implementation manners, the determining each event feature according to each of the frame features includes:
[0133] Convert each of the frame features into the input space of the first language model to obtain each converted feature;
[0134] Determine each event feature according to each of the converted features.
[0135] According to one or more embodiments of the present disclosure, a text generation method is provided, which further includes:
[0136] In some optional implementation manners, the determining each event feature according to each of the converted features includes at least one of the following:
[0137] Concatenate each of the converted features to obtain each event feature;
[0138] Perform interactive compression on each of the converted features to obtain each event feature.
[0139] According to one or more embodiments of the present disclosure, a text generation method is provided, which further includes:
[0140] In some optional implementation manners, the concatenating each of the event features in the extraction order of the corresponding event includes:
[0141] Based on a predefined order prompt word and a generation instruction prompt word, concatenate each of the event features in the extraction order of the corresponding event.
[0142] According to one or more embodiments of the present disclosure, a text generation method is provided, which further includes:
[0143] In some optional implementation manners, the concatenating each of the event features in the extraction order of the corresponding event based on a predefined order prompt word and a generation instruction prompt word includes:
[0144] According to the extraction order of each of the events, concatenate each of the event features after the corresponding order prompt word to generate an event order prompt text;
[0145] Concatenate the generation instruction prompt word after the event order prompt text to generate a prompt text.
[0146] According to one or more embodiments of the present disclosure, a text generation method is provided, which further includes:
[0147] In some optional implementation manners, when the video to be processed includes audio data, it further includes:
[0148] Perform speech recognition on the audio data to obtain speech text;
[0149] Correspondingly, the generation of the prompt text further includes: concatenating the concatenated event features with the speech text to generate the prompt text.
[0150] According to one or more embodiments of the present disclosure, a text generation method is provided, further including:
[0151] In some alternative implementation manners, when the video to be processed includes audio data, it further includes:
[0152] Performing feature extraction on the audio data to obtain audio features;
[0153] Correspondingly, the generation of the prompt text further includes: concatenating the concatenated event features with the audio features to generate the prompt text.
[0154] According to one or more embodiments of the present disclosure, a text generation method is provided, further including:
[0155] In some alternative implementation manners, after generating the description text, it further includes:
[0156] Obtaining the question text of the video to be processed;
[0157] Generating an answer text based on the description text and the question text through a second language model.
[0158] According to one or more embodiments of the present disclosure, a text generation device is provided, and the device includes:
[0159] A video frame determination module, configured to extract each event in the video to be processed and determine each target video frame corresponding to each event;
[0160] An event feature determination module, configured to extract the frame features of each target video frame and determine each event feature according to each frame feature;
[0161] A prompt text generation module, configured to concatenate the event features in the extraction order of the corresponding events to generate a prompt text;
[0162] A description text generation module, configured to generate the description text of the video to be processed based on the prompt text through a first language model.
[0163] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0164] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0165] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
Claims
1. A text generation method, characterized in that, comprising: extracting each event in the video to be processed and determining each target video frame corresponding to each event; extracting the frame features of each target video frame and determining each event feature according to each of the frame features; concatenating each of the event features in the extraction order of the corresponding event to generate a prompt text; generating a description text of the video to be processed based on the prompt text through a first language model.
2. The method according to claim 1, characterized in that, the determining each event feature according to each of the frame features includes: converting each of the frame features into the input space of the first language model to obtain each converted feature; determining each event feature according to each of the converted features.
3. The method according to claim 2, characterized in that, the determining each event feature according to each of the converted features includes at least one of the following: concatenating each of the converted features to obtain each event feature; performing interactive compression on each of the converted features to obtain each event feature.
4. The method according to claim 1, characterized in that, the concatenating each of the event features in the extraction order of the corresponding event includes: concatenating each of the event features in the extraction order of the corresponding event based on a predefined order prompt word and a generation instruction prompt word.
5. The method according to claim 4, characterized in that, the concatenating each of the event features in the extraction order of the corresponding event based on a predefined order prompt word and a generation instruction prompt word includes: according to the extraction order of each event, concatenating each of the event features after the corresponding order prompt word to generate an event order prompt text; concatenating the generation instruction prompt word after the event order prompt text to generate a prompt text.
6. The method according to claim 1, characterized in that, when the video to be processed includes audio data, it further includes: performing speech recognition on the audio data to obtain a speech text; correspondingly, the generating the prompt text further includes: concatenating the concatenated event features with the speech text to generate a prompt text.
7. The method according to claim 1, characterized in that, when the video to be processed includes audio data, it further includes: extracting features of the audio data to obtain audio features; correspondingly, the generating the prompt text further includes: concatenating the concatenated event features with the audio features to generate a prompt text.
8. The method according to any one of claims 1-7, characterized in that, after generating the description text, it further includes: obtaining a question text of the video to be processed; generating an answer text based on the description text and the question text through a second language model.
9. A text generation device, characterized in that, comprising: a video frame determination module for extracting each event in the video to be processed and determining each target video frame corresponding to each event; An event feature determination module, configured to extract frame features of the respective target video frames, and determine respective event features according to the respective frame features; A prompt text generation module, configured to splice the respective event features in the extraction order of the corresponding events to generate a prompt text; A description text generation module, configured to generate a description text of the video to be processed based on the prompt text through a first language model.
10. An electronic device characterized in that the electronic device includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the text generation method according to any one of claims 1-8.
11. A storage medium containing computer-executable instructions, where the computer-executable instructions are used to execute the text generation method according to any one of claims 1-8 when executed by a computer processor.