Comment generation method and device, electronic equipment and storage medium

By obtaining the conversation content of the target video and using a large language model to generate comments that match the preset description style, the problem of low correlation between automated generation of comments and video content is solved, and the comment quality and user experience are improved.

CN120455790APending Publication Date: 2025-08-08SHANGHAI ZHONG YUAN NETWORK CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510777658.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the prior art, the video comments generated automatically have low correlation with video content, resulting in lower review quality and poor experience for users and video creators.

Method used

By obtaining text and/or voice information of dialogue content in the target video, first content information is generated using a large language model, and a plurality of different second content information is generated according to the preset description style modification to improve the correlation between comments and video content.

Benefits of technology

The generated comments are highly correlated with the video content, and users can quickly determine whether the video content is interested, the video creators' sense of accomplishment and creative enthusiasm have improved, and the quality of comments has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455790A_ABST
    Figure CN120455790A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a comment generation method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring target text information for representing dialogue content in a target video; the dialogue content is a text and / or voice used for reflecting an event occurring in the target video; generating first content information according to the target text information; the first content information is a text used for describing an event occurring in the target video; and according to a preset description style, modifying the first content information to generate a plurality of different second content information, and determining the second content information as a comment corresponding to the target video. By applying the method and the device, the relevance between the automatically generated comments and the video content can be improved, so that the quality of the automatically generated comments is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a comment generation method, device, electronic device and storage medium. Background Art

[0002] Currently, when comments are automatically generated for videos, the majority are meaningless, such as "Like" and "Very good." These meaningless comments are applied to all videos and have little relevance to the content of the videos for which they are automatically generated, resulting in low-quality comments. Summary of the Invention

[0003] The purpose of the embodiments of the present invention is to provide a comment generation method, apparatus, electronic device, and storage medium to improve the relevance between automatically generated comments and video content, thereby improving the quality of automatically generated comments. The specific technical solution is as follows:

[0004] In a first aspect of the present invention, a method for generating comments is provided, the method comprising:

[0005] Acquire target text information for representing the content of a conversation in a target video; the conversation content is text and / or voice for reflecting an event occurring in the target video;

[0006] Generate first content information according to the target text information; the first content information is a text for describing an event occurring in the target video;

[0007] The first content information is modified according to a preset description style to generate a plurality of different second content information, and the second content information is determined as comments corresponding to the target video.

[0008] In a possible embodiment, obtaining target text information for representing the dialogue content in the target video includes:

[0009] For each video frame in the target video, performing text recognition on the target text in the video frame to obtain the original text corresponding to the video frame; the target text is a text used to reflect the event occurring in the target video;

[0010] For any two adjacent video frames among the video frames, if the original texts corresponding to the two adjacent video frames include continuous and repeated repeated text, the repeated text included in the first video frame is removed to obtain a modified text corresponding to the first video frame, and the modified text corresponding to the first video frame and the original text corresponding to the second video frame are determined as target text information; wherein the first video frame is the video frame that is earlier in time sequence among the two adjacent video frames, and the second video frame is the video frame that is later in time sequence among the two adjacent video frames;

[0011] or,

[0012] Speech recognition is performed on a target speech in the target video to obtain target text information; the target speech is a speech used to reflect an event occurring in the target video.

[0013] In a possible embodiment, determining the modified text corresponding to the first video frame and the original text corresponding to the second video frame as target text information includes:

[0014] Performing text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame to obtain first text information;

[0015] Filtering the first text information that matches the first filtering condition in the first text information to obtain second text information as the target text information;

[0016] The performing speech recognition on the target speech in the target video to obtain target text information includes:

[0017] Performing speech recognition on the target speech in the target video to obtain a recognition result;

[0018] Performing text correction on the recognition result to obtain third text information;

[0019] The third text information that matches the first filtering condition in the third text information is filtered to obtain fourth text information as target text information.

[0020] In a possible embodiment, performing text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame to obtain the first text information includes:

[0021] Inputting the modified text corresponding to the first video frame, the original text corresponding to the second video frame, and a first prompt word into a first large language model to obtain first text information output by the first large language model under the guidance of the first prompt word; the first prompt word is used to guide the first large language model to perform text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame;

[0022] The filtering of the first text information that matches the first filtering condition in the first text information to obtain the second text information as the target text information includes:

[0023] Inputting the first text information and the second prompt word into a second language model, obtaining second text information output by the second language model under the guidance of the second prompt word as the target text information, wherein the second prompt word is used to guide the second language model to filter first text information in the first text information that matches a first filtering condition, where the first filtering condition is the condition represented by the second prompt word;

[0024] The performing text correction on the recognition result to obtain third text information includes:

[0025] inputting the recognition result and a third prompt word into a third language model to obtain third text information output by the third language model under the guidance of the third prompt word; the third prompt word is used to guide the third language model to perform text correction on the recognition result;

[0026] The filtering of the third text information that matches the first filtering condition in the third text information to obtain fourth text information as the target text information includes:

[0027] The third text information and the second prompt word are input into a second language model to obtain fourth text information output by the second language model under the guidance of the second prompt word as the target text information, wherein the second prompt word is used to guide the second language model to filter the third text information that meets a first filtering condition in the third text information, and the first filtering condition is the condition represented by the second prompt word.

[0028] In a possible embodiment, obtaining target text information for representing the dialogue content in the target video includes:

[0029] Acquire fifth text information for representing the content of the conversation in the target video;

[0030] The fifth text information that matches the first filtering condition in the fifth text information is filtered to obtain sixth text information as target text information.

[0031] In a possible embodiment, generating the first content information according to the target text information includes:

[0032] The target text information and the fourth prompt word are input into a fourth language model to obtain first content information output by the fourth language model under the guidance of the fourth prompt word; the first content information is text described in natural language and is used to represent the theme and specific content of the target video, the theme being text that provides an overall description of events occurring in the target video, and the specific content being text that provides segmented descriptions of events occurring in the target video.

[0033] In a possible embodiment, the modifying the first content information according to a preset description style to generate a plurality of different second content information includes:

[0034] The first content information, the title of the target video, and a fifth prompt word are input into a fifth language model to obtain a plurality of different second content information output by the fifth language model under the guidance of the fifth prompt word; wherein the second content information is text described in a natural language; the fifth prompt word is used to guide the fifth language model to modify the first content information according to a preset description style, and the fifth prompt word includes a comment with the preset description style.

[0035] In a possible embodiment, determining the second content information as a comment corresponding to the target video includes:

[0036] Second content information matching a second filtering condition in the second content information is filtered to obtain third content information, and the third content information is determined as a comment corresponding to the target video.

[0037] In a possible embodiment, the method further includes:

[0038] The comments corresponding to the target video are sorted according to the relevance between the comments corresponding to the target video and the content of the target video.

[0039] In a possible embodiment, filtering the second content information that matches the second filtering condition in the second content information to obtain the third content information includes:

[0040] inputting the second content information and the sixth prompt word into a sixth language model to obtain third content information output by the sixth language model under the guidance of the sixth prompt word;

[0041] The sixth prompt word is used to guide the sixth language model to filter the second content information that matches the second filtering condition in the second content information; the second filtering condition is the condition represented by the sixth prompt word;

[0042] The sorting of the comments corresponding to the target video according to the relevance between the comments corresponding to the target video and the content of the target video includes:

[0043] The comments corresponding to the target video and the seventh prompt word are input into the seventh language model, and the comments corresponding to the target video with an arrangement order are output by the seventh language model; the seventh language model is used to sort the comments corresponding to the target video according to the relevance between the comments corresponding to the target video and the content of the target video under the guidance of the seventh prompt word.

[0044] In a second aspect of the present invention, a comment generation device is provided, comprising:

[0045] A target text information acquisition module is used to acquire target text information used to represent the dialogue content in the target video; the dialogue content is text and / or voice used to reflect the events occurring in the target video;

[0046] A first content information generating module is configured to generate first content information based on the target text information; the first content information is a text describing an event occurring in the target video;

[0047] The second content information generating module is configured to modify the first content information according to a preset description style to generate a plurality of different second content information, and determine the second content information as comments corresponding to the target video.

[0048] In a third aspect of the present invention, an electronic device is provided, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0049] Memory for storing computer programs;

[0050] The processor is configured to implement any of the comment generation methods described in the first aspect above when executing a program stored in the memory.

[0051] In a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the comment generation method described in any one of the first aspects is implemented.

[0052] In a fifth aspect of the present invention, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute any of the above-mentioned comment generation methods.

[0053] Embodiments of the present invention provide a comment generation method, device, electronic device, and storage medium. These methods automatically generate comments for a target video by obtaining target text information representing the content of a conversation in a target video; the conversation content being text and / or speech that reflects events occurring in the target video; generating first content information based on the target text information; the first content information being text that describes events occurring in the target video; modifying the first content information according to a preset description style to generate multiple different second content information; and determining the second content information as comments corresponding to the target video. Because the target text information represents the content of a conversation in a target video, which is text and / or speech that reflects events occurring in the target video, generating the first content information based on the target text information can summarize the events occurring in the target video, such that the first content information can describe the events occurring in the target video. And because the second content information is generated by modifying the first content information according to a preset description style, the second content information can also describe the events that occurred in the target video. That is to say, the second content information has a high correlation with the content of the target video, and compared with the first content information, the description style of the second content information is closer to the preset description style, so that the description of the second content information is more natural and easier to be understood by users who view the comments. By determining the second content information as the comment corresponding to the target video, the comments generated for the target video have a high correlation with the content of the target video, and the description of the second content information is more natural and easier to be understood by users who view the comments. It can be seen that the embodiment of the present invention can improve the correlation between the comments automatically generated for the target video and the content of the target video, thereby improving the quality of the automatically generated comments. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art.

[0055] Figure 1 This is a schematic diagram of a first flow chart of a method for generating comments in an embodiment of the present invention;

[0056] Figure 2 This is a schematic diagram of a first process for obtaining target text information in an embodiment of the present invention;

[0057] Figure 3 This is a schematic diagram of a second process for obtaining target text information in an embodiment of the present invention;

[0058] Figure 4 This is a schematic diagram of a third process for obtaining target text information in an embodiment of the present invention;

[0059] Figure 5 This is a schematic diagram of a fourth process for obtaining target text information in an embodiment of the present invention;

[0060] Figure 6 This is a fifth flow chart of obtaining target text information in an embodiment of the present invention;

[0061] Figure 7 This is a second flow chart of the comment generation method according to an embodiment of the present invention;

[0062] Figure 8 This is a third flow chart of the comment generation method according to an embodiment of the present invention;

[0063] Figure 9 This is a fourth flow chart of the comment generation method according to an embodiment of the present invention;

[0064] Figure 10 Schematic diagram of the fifth flow chart of the comment generation method in an embodiment of the present invention;

[0065] Figure 11 A schematic diagram of the structure of a comment generating device according to an embodiment of the present invention;

[0066] Figure 12 Schematic diagram of the structure of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION

[0067] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.

[0068] In order to more clearly illustrate the comment generation method provided by the present application, the possible application scenarios of the comment generation method provided by the present application will be exemplified below. It can be understood that the following examples are only possible application scenarios of the comment generation method provided by the present application. In other possible embodiments, the comment generation method provided by the present application can be applied to other possible application scenarios, and the following examples do not impose any limitations on this.

[0069] Currently, when comments are automatically generated for videos, the majority are meaningless, such as "Like" and "Very good." These meaningless comments are common to all videos and have little relevance to the content of the videos for which they are automatically generated. This results in low-quality comments, which in turn creates a poor user experience for both users and video creators.

[0070] Specifically, when watching a video, users typically check other users' comments to determine whether the video contains content they are interested in. However, since the automatically generated comments are poorly correlated with the video content, the quality of these comments is low. Therefore, users cannot determine the content of the video from the comments, and therefore cannot determine whether the video contains content they are interested in. Only after watching the video can users determine whether the video contains content they are interested in. This makes it difficult for users to quickly watch videos that contain content they are interested in, resulting in a poor user experience.

[0071] On the other hand, in the context of video creation, the number of comments is a crucial indicator for video creators. The number of comments not only affects the income of video creators, but is also closely related to their sense of accomplishment and creative enthusiasm. As the number of comments increases, video creators will believe that their videos are widely watched and discussed, which can motivate video creators, further enhance their creative motivation, and improve their sense of accomplishment and creative enthusiasm. However, in actual scenarios, the number of real user comments is often small, which results in video creators being unable to obtain sufficient interaction and feedback from the small number of comments, which in turn reduces their sense of accomplishment and creative enthusiasm. In order to address the above issues, a method for automatically generating comments is needed to enhance video creators' sense of accomplishment and creative enthusiasm. In the current method of automatically generating comments, since it is not combined with the understanding of the video content, the automatically generated comments cannot reflect the understanding and discussion of the video content, resulting in a low correlation between the automatically generated comments and the video content. As a result, video creators cannot obtain content that is helpful for video creation from the automatically generated comments, that is, the quality of the automatically generated comments is low. As a result, the automatically generated comments not only fail to improve the video creator's sense of accomplishment and creative enthusiasm, but will have the opposite effect, resulting in a poor experience for video creators.

[0072] Based on this, in order to improve the relevance between automatically generated comments and video content, thereby improving the quality of automatically generated comments and further improving the user or video creator's experience, the present invention provides a comment generation method, such as Figure 1 As shown, the method includes:

[0073] S101, obtaining target text information for representing the dialogue content in a target video.

[0074] The dialogue content is text and / or voice that reflects the events that occurred in the target video.

[0075] S102: Generate first content information according to the target text information.

[0076] The first content information is text used to describe an event occurring in the target video.

[0077] S103: Modify the first content information according to a preset description style to generate a plurality of different second content information, and determine the second content information as comments corresponding to the target video.

[0078] By applying an embodiment of the present invention, target text information representing the content of a conversation in a target video is obtained; the conversation content is text and / or speech used to reflect the events that occurred in the target video; first content information is generated based on the target text information; the first content information is text used to describe the events that occurred in the target video; the first content information is modified according to a preset description style to generate multiple different second content information, and the second content information is determined as a comment corresponding to the target video, thereby automatically generating comments for the target video. Since the target text information is used to represent the content of a conversation in the target video, and the conversation content is text and / or speech used to reflect the events that occurred in the target video, a summary of the events that occurred in the target video can be achieved by generating the first content information based on the target text information, so that the first content information can describe the events that occurred in the target video. And because the second content information is generated by modifying the first content information according to a preset description style, the second content information can also describe the events that occurred in the target video. That is to say, the second content information has a high correlation with the content of the target video, and compared with the first content information, the description style of the second content information is closer to the preset description style, so that the description of the second content information is more natural and easier to be understood by users who view the comments. By determining the second content information as the comment corresponding to the target video, the comments generated for the target video have a high correlation with the content of the target video, and the description of the second content information is more natural and easier to be understood by users who view the comments. It can be seen that the embodiment of the present invention can improve the correlation between the comments automatically generated for the target video and the content of the target video, thereby improving the quality of the automatically generated comments.

[0079] As described above, the comments corresponding to the target video generated by the embodiment of the present invention have a high correlation with the content of the target video. Therefore, on the one hand, when the comments corresponding to the target video generated by the embodiment of the present invention are displayed to the user, the user can determine the content included in the target video by viewing the comments corresponding to the displayed target video, thereby helping the user to quickly determine whether the target video includes content that they are interested in, and then facilitating the user to quickly watch the video including the content that they are interested in, thereby improving the user experience. On the other hand, when the comments corresponding to the target video generated by the embodiment of the present invention are displayed to the video creator, the high-quality comments with a high correlation to the target video content can enhance the video creator's sense of accomplishment and creative enthusiasm. At the same time, the number of comments on the target video is increased, thereby enhancing the interactive atmosphere of the video platform displaying the comments.

[0080] The above-mentioned S101-S103 will be exemplarily described below.

[0081] In S101, the dialogue content in the target video refers to text and / or voice that reflects the events occurring in the target video. Specifically, the text that reflects the events occurring in the target video may include subtitles located below the target video, transition text located within the target video, and the like. The transition text refers to explanatory text displayed within the target video when the target video switches scenes, such as "One Year Later," "Shanghai," and the like.

[0082] The speech used to reflect the events in the target video may refer to: narration, inner monologue, character dialogue, etc. in the target video in the form of speech. The following will exemplify the method of obtaining the target text information, which will not be repeated here.

[0083] In S102, in one possible embodiment, first content information can be generated using a large language model. Specifically, the target text information and a fourth prompt word are input into a fourth large language model, and the first content information is output by the fourth large language model under the guidance of the fourth prompt word. The first content information is text described in natural language and is used to describe events occurring in the target video.

[0084] In this embodiment, the fourth large language model is any LLM (Large Language Model) that can generate the first content information according to the target text information and the fourth prompt word, for example, a model such as Meta-Llama-3-70B.

[0085] The fourth prompt can be configured based on user needs. For example, the fourth prompt can be: "Chinese answer," so that the first content information is described in Chinese. Alternatively, the fourth prompt can be: "Chinese answer, beginning with "The video is described." This fourth prompt ensures that the first content information output by the fourth language model is described in Chinese and begins with "The video is described."

[0086] It is understandable that the events occurring in the target video can generally be described in a variety of different ways. For example, the events occurring in the target video can be described by classifying the target video's subject matter, for example, the video describes the appearance characteristics and living habits of a certain animal; or the events occurring in the target video can be described by using the specific content in the target video, for example, the video describes the appearance characteristics of a certain animal as... (the content omitted here is the appearance characteristics of the animal appearing in the target video), and the living habits of the animal as... (the content omitted here is the living habits of the animal appearing in the target video).

[0087] However, using a single method will result in the generated first content information being unable to fully describe the events that occurred in the target video. For example, the subject classification method cannot reflect the specific process of the events that occurred in the target video, and the specific content method cannot reflect the brief summary of the target video. Based on this, in order to make the first content information more comprehensive in reflecting the events that occurred in the target video, the first content information is used to represent the theme and specific content of the target video. The theme is the text that describes the events that occurred in the target video as a whole, and the specific content is the text that describes the events that occurred in the target video in sections.

[0088] Since the subject is a text that provides an overall description of the events occurring in the target video, the portion of the first content information that represents the subject of the target video can be considered a comprehensive summary of the target video. Furthermore, since the specific content is a text that provides a segmented description of the events occurring in the target video, the portion of the first content information that represents the specific content of the target video can be considered a detailed, segmented summary of the target video. Therefore, in this embodiment, the fourth prompt word could also be: "Chinese answer," and the summary content includes both a detailed, segmented summary and an overall summary.

[0089] For example, assuming the target video is a video showing how to make a certain food, the first content information generated based on the target text information representing the dialogue content in the target video may be: The video describes a tutorial on how to make a certain food. To make this food, you first need to perform step 1, then step 2, and then step 3... In the first content information generated above, "The video describes a tutorial on how to make a certain food" is used to represent the theme of the target video, and "To make this food, you first need to perform step 1, then step 2, and then step 3..." is used to represent the specific content of the target video.

[0090] By selecting this embodiment, the user's requirements for the generated first content information can be reflected through the setting of the fourth prompt word, so that the fourth language model can generate the first content information that meets the user's requirements under the guidance of the fourth prompt word. Furthermore, by using the first content information to represent the theme and specific content of the target video, an overall summary and segmented detailed summary of the events occurring in the target video are performed. By using the overall summary and segmented detailed summary, the first content information can reflect both the brief theme of the target video and the specific process of the events occurring in the target video, thereby enabling the first content information to more comprehensively reflect the events occurring in the target video. This facilitates the subsequent modification of the first content information according to a preset description style to obtain second content information that is more relevant to the content in the target video, and to determine the second content information as the comment corresponding to the target video, thereby further improving the relevance between the automatically generated comment and the content of the target video and improving the quality of the automatically generated comment.

[0091] In other possible embodiments, the first content information may be generated based on the target text information by other means besides the large language model, for example, by using a deep learning model to generate the first content information based on the target text information.

[0092] In S103, the preset description style can be set according to user needs, and the preset description style can be one style or multiple styles. For example, the preset description style can refer to: humorous style, or the preset description style can refer to: humorous style, graceful style, fresh style, etc.

[0093] The modification in S103 refers to any operation that can be considered as modifying the first content information to obtain the second content information according to the preset description style. For example, the modification may include: directly modifying the first content information according to the preset description style to obtain the second content information; the modification may also include: regenerating the second content information based on the first content information according to the preset description style, etc.

[0094] In one possible embodiment, the first content information can be generated using a large language model. Specifically, the first content information and a fifth prompt word are input into a fifth large language model, resulting in the output of multiple different second content information by the fifth large language model under the guidance of the fifth prompt word. The second content information is text described in natural language. The fifth prompt word is used to guide the fifth large language model to modify the first content information according to a preset description style, and the fifth prompt word includes a comment with the preset description style.

[0095] In this embodiment, the fifth language model is any LLM that can generate the second content information based on the first content information and the fifth prompt word, for example, models such as GPT4-0.

[0096] The fifth prompt word can be set based on user needs and includes comments with a preset descriptive style. To make the description of the second content information more natural and easier for users to understand, the comments with the preset descriptive style should be manually generated. Therefore, comments with the preset descriptive style are referred to as manual comment examples below.

[0097] Specifically, the fifth prompt word may only include manual review examples; or, in addition to the manual review examples, the fifth prompt word may also include one or more of the following: requirements for the second content information, positive examples of the requirements for the second content information, and negative examples of the requirements for the second content information.

[0098] For example, examples of manual comments may include one or more of the following: 1. Is Guodanpi made from apricots? I always thought it was made from hawthorns; 2. What are the almonds for? It's urgent, when will the next episode be updated? 3. I always thought Guodanpi was made from hawthorns; 4. What's the name of the thing he ate at the beginning; 5. Are apricots sour or sweet? I've never seen or eaten one. 6. Are there apricots in Northeast China? 7. This doesn't mean you can eat unlimited apricots, it just adds ¥1. 8. It's shorter. 9. Take a whole piece and compare it. 10. It's smaller. 11. You don't appear to be smaller, but it's actually smaller. 12. He turned thrust into suction. 13. This must be the principle behind a rocket launch. 14. The principle behind pouring Coke in mid-air is that a thin string is tied between the Coke and the cup. 15. This is a dust explosion. When particles rub against each other, sparks are produced, causing them to fly into the air. 16. Chen's favorite artist is undeniable. 17. Chen is truly insightful. 18. Waiting online, what's the name of the game in the video? 19. I always need an excuse to hit you. 20. The cat said it didn't hit you, so why are you crying? 21. That girl: You're done! 22. I’m so sorry; 23. If your girlfriend doesn’t want it, you can give it to me; 24. That girl is so cute when she cries; 25. The world shouldn’t have only one way of thinking; 26. My mother has a house full of things, but I’m the only one who survived; 27. You can be a scientist!

[0099] The requirements for the second content information may include one or more of the following, and the positive examples and / or negative examples of the requirements will be exemplified in the following requirements:

[0100] 1. Style generation based on human review examples.

[0101] 2. The types of secondary content information are diverse and should not be similar. For example, the types of secondary content information can include: meaningless secondary content information; questioning secondary content information; second content information based on personal experience, etc. There is no need to generate positive secondary content information.

[0102] Since positive second content information generally praises the target video in an exaggerated manner, the credibility of the positive second content information is low, which makes it difficult for users or video creators to believe the content described in the positive second content information, resulting in poor quality of the positive second content information. Therefore, there is no need to generate positive second content information in this requirement.

[0103] 3. Avoid exaggerated emotions. For example, avoid using exaggerated words like "too much," "very," "extremely," "really," "surprisingly," "especially," "without a doubt," "first time," or "brain-opening," "unexpected," and other exaggerated emotional expressions.

[0104] 4. The quality of the comments should be moderate in language, without using fancy words, and should be colloquial and natural in expression.

[0105] 5. The second content information should be within 15 words.

[0106] 6. Don’t use idioms.

[0107] 7. Avoid unnatural, colloquial, empty and pretentious expressions, avoid using emotional adjectives to express secondary content information, and avoid redundancy of secondary content information.

[0108] Since not using emotional adjectives will make the second content information more natural, in order to make the second content information generated by the fifth language model more natural, it is necessary to avoid using emotional adjectives when generating the second content information. For example, a counterexample to this requirement can be as follows:

[0109] Counterexample: (1) I didn’t expect that fish and crayfish could play like this. How creative!

[0110] (2) This is the first time I’ve heard of cashew juice, so I’m a little curious.

[0111] (3) Is the taste of this juice really that special? I doubt it.

[0112] (4) What is a roasted hairy egg? I've never tried it, so I'm a little curious.

[0113] (5) My brother never eats roasted eggs. Haha, that’s a bit cute.

[0114] In the above example, the content before the first punctuation mark already expresses emotions such as creativity, curiosity, and suspicion, so emotional adjectives such as "really creative," "a little curious," and "a little skeptical" appear redundant and unnatural. The following is a positive example of this requirement:

[0115] Positive example: (1) I didn’t expect fish and crayfish could play like this.

[0116] (2) This is the first time I’ve heard of cashew juice.

[0117] (3) Does the taste of this juice really have such a special flavor?

[0118] (4) What is a roasted hairy egg? I've never eaten one.

[0119] (5) My brother never eats roasted eggs, haha

[0120] 8. Avoid complex sentence structures. The subject, predicate, and object of a sentence do not have to be complete.

[0121] Counterexample: This ink painting of crayfish looks really good. I want to try it next time!

[0122] Positive example: It looks good, I will try it next time.

[0123] In the above examples, compared with the negative examples, the positive examples with simple sentence structures are more natural and colloquial.

[0124] 9. Second, the content information should be fluent and logical.

[0125] Counterexample: What is six minus three? Haha

[0126] In the above example, "What is six minus three?" does not contain any funny elements, so the "haha" after the sentence has no connection with the previous text, making the above example illogical.

[0127] 10. If the target video is not a children's video, there is no need to generate relatively simple and childish second content information.

[0128] Counterexample: (1) What is six minus three? Waiting online.

[0129] (2) Xiaohua is great, come on Fat Tiger!

[0130] The question in example (1) is relatively simple, while the tone in example (2) is relatively childish.

[0131] 11. The second content information needs to be combined with the first content information and the title of the target video to generate second content information that is highly correlated with the first content information and the target video.

[0132] 12. Since the first content information and the target video title that are inconsistent with common sense are likely to be funny or ironic, no relevant second content information is generated for the first content information and the target video title that are inconsistent with common sense.

[0133] Examples of first content information and target video titles that are inconsistent with common sense include: (1) The first content information includes: "It's all about super-difficult math problems like 6 minus 3 equals what?" Since 6 minus 3 is not difficult, the first content information in the above example is inconsistent with common sense. (2) The first content information includes: "If I'm late again this morning, I'll be late for eight consecutive days this week." Since there are only 7 days in a week, the first content information in the above example is inconsistent with common sense.

[0134] 13. Do not use commas as punctuation marks in the second content message.

[0135] 14. After correctly understanding the first content information and the content in the title of the target video, generate the second content information. The generated second content information must conform to basic common sense and objective laws.

[0136] 15. Do not use the sentence structure of “a bit XX”, such as “a bit interesting”, “a bit cute”, etc. Use other more colloquial expressions.

[0137] 16. Do not use "quite XX" sentences, such as quite special, quite amazing, etc. Use other more colloquial expressions.

[0138] 17. The number of second content information may be less than 7, but the generated second content information must meet the above requirements 1-16.

[0139] 18. You can use internet buzzwords and slang in your comments to make them more natural and colloquial. The following are common buzzwords and slang expressions. In the following examples, the right side of the "-" symbol is a common buzzword or slang, and the left side is the corresponding expression of the buzzword or slang:

[0140] Arguing - a nitpicker; being obsessed with love - being in love; not working hard - lying down; careless operation error - slip of the hand; admitted - getting on board; buddy - old friend; cheering - cheering; rich - financial freedom; showing off - Versailles; competing - involution; not motivated - Buddhist; partner - buddy; naughty child - mythical beast; real or fake - Zundu, fakedu; liking someone or something - getting it; being popular and hot - going viral; self-deprecating - I'm the clown; describing something as simple - all you need are hands; expressing helplessness - convinced; being cool - street-blasting; deliberately attracting attention - a conspicuous bag.

[0141] 19. Output the second content information in descending order of quality.

[0142] In one possible embodiment, the fifth prompt word may include all the contents of the manual review example and all the contents of the second content information requirement example. In another possible embodiment, the fifth prompt word may also include 1-5 in the manual review example and 1-6 in the second content information requirement example.

[0143] When generating the second content information, the fifth language model may refer to the target video's title in addition to the fifth prompt and the first content information. In this embodiment, the second content information is generated by: inputting the first content information, the target video's title, and the fifth prompt into the fifth language model, obtaining a plurality of different second content information output by the fifth language model under the guidance of the fifth prompt; wherein the second content information is text described in natural language; and the fifth prompt is used to guide the fifth language model to modify the first content information according to a preset description style, and the fifth prompt includes comments with the preset description style. The fifth language model and the fifth prompt have been exemplarily described above and will not be repeated here.

[0144] By selecting this embodiment, the user's requirements for the generated second content information can be reflected through the setting of the fifth prompt word, so that the fifth language model can generate second content information that meets the user's requirements and has a preset description style under the guidance of the fifth prompt word. Since the fifth language model is a large language model, it has strong deep learning capabilities, enabling the fifth language model to accurately understand the content of the target video, generate natural, logically clear, and highly relevant second content information to the content of the target video, and determine the second content information as a comment corresponding to the target video, so that the comment generated for the target video has a high relevance to the content of the target video, and the language description is natural, logically clear, and coherent, thereby improving the quality of the automatically generated comments.

[0145] In other possible embodiments, the first content information may be modified according to a preset description style using other methods besides a large language model to generate multiple different second content information. For example, the first content information may be modified according to a preset description style using a deep learning model to generate multiple different second content information.

[0146] The above has been described as an example of how to generate comments corresponding to the target video. As mentioned above, the generation of the second content information depends on the first content information, and the generation of the first content information depends on the target text information. Therefore, the following will be described as an example of how to obtain the target text information in S101. The target text information can be obtained through the following methods:

[0147] See also Figure 2 , Method 1 includes:

[0148] S201 : performing text recognition on target text in each video frame in a target video to obtain original text corresponding to the video frame.

[0149] The target text is a text used to describe the events that occurred in the target video.

[0150] As previously described regarding the text used to represent events in the target video, the text used to represent events in the target video, i.e., the target text, can refer to: subtitles located below the target video, transition text located within the target video, etc. The video frames in the target video can refer to all video frames in the target video, or can refer to a portion of the video frames in the target video.

[0151] If the video frames in the target video refer to a portion of the target video, the text included in each video frame can be obtained by taking a screenshot of the target video using ffmpeg (an audio and video editing software), uploading the five pictures generated by ffmpeg per second to the cloud disk of the electronic device, and sending the cloud disk address of the picture to the QW2-VL (Qwen2-VL, a multimodal large language model) service. The QW2-VL service processes the picture using the QW2-VL model and extracts the text information in the picture. In this embodiment, the picture generated by ffmpeg is the selected portion of the video frame, and the text information in each picture is the text included in each video frame.

[0152] S202, for any two adjacent video frames in each video frame, if the original texts corresponding to the two adjacent video frames include continuous and repeated duplicate texts, then remove the duplicate text included in the first video frame to obtain the modified text corresponding to the first video frame, and determine the modified text corresponding to the first video frame and the original text corresponding to the second video frame as target text information.

[0153] The first video frame is a video frame that is earlier in time sequence among two adjacent video frames, and the second video frame is a video frame that is later in time sequence among two adjacent video frames.

[0154] It is understandable that the target texts in adjacent video frames are most likely the same or have continuous and repeated text, that is, the corresponding original texts in adjacent video frames are the same or have continuous and repeated text. Therefore, if the corresponding original texts in each video frame are directly combined to obtain the target text information, it may cause the target text information to include more repetitive sentences or texts, making it difficult for the target text information to accurately reflect the events that occurred in the target video, thereby making it difficult to accurately generate the first content information for describing the events that occurred in the target video based on the target text information, reducing the relevance of the obtained first content information with the events that occurred in the target video, that is, reducing the relevance of the first content information with the content in the target video. Based on this, in order to obtain more accurate target text information, thereby improving the relevance of the first content information with the content in the target video, it is necessary to process the original text corresponding to each video frame in the manner in S202 to obtain the target text information.

[0155] To facilitate understanding of the processing method in S202 , the processing method in S202 will be described below by taking two adjacent video frames and multiple adjacent video frames as examples.

[0156] For two adjacent video frames, the two adjacent video frames are recorded as video frame 1 and video frame 2 respectively in chronological order, the original text corresponding to video frame 1 is recorded as original text 1, and the original text corresponding to video frame 2 is recorded as original text 2. In this example, video frame 1 is the video frame with an earlier chronological order among the two adjacent video frames, that is, video frame 1 is the first video frame, and video frame 2 is the video frame with a later chronological order among the two adjacent video frames, that is, video frame 2 is the second video frame.

[0157] If the original text 1 and the original text 2 contain continuous and repeated duplicate texts, the duplicate texts in the original text 1 are removed to obtain the modified text corresponding to the video frame 1, which is recorded as modified text 1. The modified text 1 and the original text 2 are determined as the target text information. For example, assuming that the original text 1 is "I like" and the original text 2 is "I like to eat tomatoes", the original text 1 and the original text 2 contain continuous and repeated duplicate texts "I like". After removing "I like" from the original text 1, the modified text 1 is obtained. Since there is no other text in the original text 1 after removing "I like", in this embodiment, the modified text 1 can be recorded as When determining the target text information, you can Ignore it and directly determine the original text 2 "I like to eat tomatoes" as the target text information.

[0158] For another example, suppose original text 1 is "I said I like" and original text 2 is "I like to eat tomatoes." Original text 1 and original text 2 contain consecutive and repeated "I like" sentences. Remove "I like" from original text 1 to obtain modified text 1, "I said." Modified text 1 and original text 2 are then determined as target text information, meaning "I said I like to eat tomatoes" is determined as the target text information.

[0159] For multiple adjacent video frames, these are recorded in chronological order as video frames 3-10, and the original text corresponding to each video frame is recorded as original text 3-10. Assuming that original text 3 and original text 4 include continuous and repeated duplicate text 1, original text 4 and original text 5 include continuous and repeated duplicate text 2, and original text 9 and original text 10 include continuous and repeated duplicate text 3, and that the original text corresponding to the other two adjacent video frames in video frames 3-10 does not include duplicate text, in this example, duplicate text 1 in original text 3 is removed to obtain modified text 3; duplicate text 2 in original text 4 is removed to obtain modified text 4; and duplicate text 3 in original text 9 is removed to obtain modified text 9. Modified text 3, modified text 4, original texts 5-8, modified text 9, and original text 10 are then combined to determine target text information.

[0160] Method 1 obtains target text information by performing text recognition on the target text in the target video. Therefore, method 1 can be considered as a method of obtaining target text information using OCR (Optical Character Recognition) technology. Method 1 is more suitable for target videos with text in the video screen.

[0161] By selecting this embodiment, text recognition can be performed on the target text in each video frame in the target video, which is used to reflect the events occurring in the target video, to obtain the original text corresponding to the video frame; and for any two adjacent video frames in each video frame, if the original text corresponding to the two adjacent video frames includes continuous and repeated repeated text, the repeated text included in the first video frame that is earlier in time sequence among the two adjacent video frames is removed to obtain the modified text corresponding to the first video frame, and the modified text corresponding to the first video frame and the original text corresponding to the second video frame that is later in time sequence among the two adjacent video frames are determined as target text information to obtain the target text information. Through the above method, it is possible to avoid the situation where the target text information includes a large number of repetitive sentences or texts, so that the target text information can accurately reflect the events occurring in the target video, so that the first content information used to describe the events occurring in the target video can be accurately generated based on the target text information, thereby improving the correlation between the obtained first content information and the events occurring in the target video, that is, improving the correlation between the first content information and the content in the target video, and then improving the correlation between the second content information obtained based on the first content information and the content in the target video, and determining the second content information as the comment corresponding to the target video, further improving the correlation between the comments automatically generated for the target video and the content in the target video, and improving the quality of the automatically generated comments.

[0162] Method 2: Perform speech recognition on the target speech in the target video to obtain target text information; the target speech is the speech used to reflect the events occurring in the target video.

[0163] Referring to the aforementioned description of the voice used to reflect the events occurring in the target video, the voice used to reflect the events occurring in the target video, that is, the target voice, may refer to: narration, inner monologue, character dialogue, etc. presented in the form of voice in the target video.

[0164] Specifically, ffmpeg is used to extract the audio file from the target video, and the extracted audio file is uploaded to the cloud disk of the electronic device. The cloud disk address storing the audio file is sent to the Whisper (a speech recognition model) service. The Whisper service accesses the received cloud disk address, downloads the audio file, and performs speech recognition on the audio file through the Whisper model to extract the text information in the audio file to obtain the target text information.

[0165] Method 2 obtains the target text information by performing speech recognition on the audio in the target video. Therefore, Method 2 can be considered a method of obtaining the target text information through ASR (Automatic Speech Recognition) technology. Method 2 is more suitable for target videos with clear audio content.

[0166] By selecting this embodiment, the target text information can be obtained by performing voice recognition on the target voice in the target video, so that the target text information can reflect the events occurring in the target video, thereby enabling the subsequent generation of first content information for describing the events occurring in the target video based on the target text information, thereby improving the correlation between the obtained first content information and the events occurring in the target video, that is, improving the correlation between the first content information and the content in the target video, and then improving the correlation between the second content information obtained based on the first content information and the content in the target video, and determining the second content information as the comment corresponding to the target video, thereby further improving the correlation between the comments automatically generated for the target video and the content in the target video, and improving the quality of the automatically generated comments.

[0167] The above has been an exemplary description of the method of obtaining the target text information in the above S101. In order to further improve the accuracy of the obtained target text information, in a possible embodiment, see Figure 3 , the method for obtaining the target text information can be selected from method 1 and method 2 in the following manner, and the target content information can be obtained based on the obtained target text information:

[0168] S1011, when text recognition is performed on the target text in each video frame of the target video to obtain the original text corresponding to each video frame, for any two adjacent video frames in each video frame, if the original texts corresponding to the two adjacent video frames include continuous and repeated repeated texts, the repeated text included in the first video frame is removed to obtain the modified text corresponding to the first video frame, and the modified text corresponding to the first video frame and the original text corresponding to the second video frame are determined as target text information.

[0169] The target text is a text used to describe the events that occurred in the target video.

[0170] The method of obtaining the target text information in S1011 is the aforementioned method 1. If the target text in each video frame of the target video is recognized separately and the original text corresponding to each video frame is obtained, it means that the target video includes the target text and the target text information can be obtained through the aforementioned method 1.

[0171] As explained above, method 1 can be regarded as a method of obtaining target text information through OCR technology. The target text information obtained through OCR technology is generally calibrated manually. Therefore, the target text information obtained through OCR technology is more accurate. Figure 3 In the embodiment shown, the target text information is obtained preferentially through the above-mentioned method 1. The method 1 has been exemplarily described above, so it will not be repeated here.

[0172] S1012: After performing text recognition on the target text in each video frame of the target video, if the original text corresponding to each video frame is not obtained, performing speech recognition on the target speech in the target video to obtain target text information.

[0173] The target voice is a voice used to reflect the events occurring in the target video.

[0174] The method of obtaining the target text information in S1012 is the aforementioned method 2. If the target text in each video frame of the target video is recognized separately and the original text corresponding to each video frame is not obtained, it means that the target video does not include the target text and the target text information cannot be obtained by the aforementioned method 1. In this case, the target text information can be obtained by the aforementioned method 2. That is, in Figure 3 In the embodiment, the target text information is obtained by the above-mentioned method 1 first, and only when the target video does not include the target text, the target text information is obtained by the above-mentioned method 2. The above-mentioned method 2 has been exemplarily described, so it will not be repeated here.

[0175] S102: Generate first content information according to the target text information.

[0176] The first content information is text used to describe an event occurring in the target video.

[0177] S103: Modify the first content information according to a preset description style to generate a plurality of different second content information, and determine the second content information as comments corresponding to the target video.

[0178] Figure 1 S102 - S103 have been exemplarily described in the illustrated embodiment. Please refer to the aforementioned related descriptions of S102 - S103 and will not be repeated here.

[0179] By selecting this embodiment, the target text information can be obtained preferentially through the above-mentioned method 1 (i.e., OCR technology), and only when the target text is not included in the target video, the target text information can be obtained through the above-mentioned method 2 (i.e., ASR technology). Since the target text information obtained by OCR technology is generally calibrated manually, the accuracy of the target text information obtained by OCR technology is higher, so that the accuracy of the target text information obtained by this embodiment is higher, and thus the first content information obtained based on the target text information has a higher correlation with the content of the target video, and the second content information obtained based on the first content information has a higher correlation with the content of the target video, and the second content information is determined as the comment corresponding to the target video, so that the correlation between the comments automatically generated for the target video and the content of the target video can be improved, thereby improving the quality of the automatically generated comments. Moreover, in the case where the target text is not included in the target video, the target text information can also be obtained by the above-mentioned method 2, so as to ensure that the target text information can be effectively obtained in different situations, increase the flexibility and applicability of the target text information acquisition method, and improve the comprehensiveness and accuracy of obtaining the target text information.

[0180] The above has provided an exemplary explanation of how to select a method for obtaining the target text information. It is understandable that the text information used to represent the content of the dialogue in the target video may include some text information that is difficult for the machine to understand, such as garbled characters, text information in small languages, and the like, making it difficult for the machine to automatically generate the first content information based on the above text information, resulting in the obtained first content information having a low correlation with the content of the target video, thereby resulting in the second content information obtained based on the first content information having a low correlation with the content of the target video, and further resulting in the second content information being determined as the comment corresponding to the target video. The correlation between the comments automatically generated for the target video and the content of the target video is reduced, resulting in a reduction in the quality of the automatically generated comments. Based on this, in order to facilitate the generation of the first content information based on the target text information, improve the correlation between the first content information and the content of the target video, and thereby improve the quality of the automatically generated comments, in a possible embodiment, see. Figure 4 , you can get the target text information in the following ways:

[0181] S401: Acquire fifth text information for representing the content of a conversation in a target video.

[0182] The fifth text information representing the content of the conversation in the target video can be obtained through the aforementioned method 1 or method 2. Specifically, the steps for obtaining the fifth text information through method 1 are as follows: for each video frame in the target video, performing text recognition on the target text in the video frame to obtain the original text corresponding to the video frame; for any two adjacent video frames in each video frame, if the original text corresponding to the two adjacent video frames includes continuous and repeated repeated text, then removing the repeated text included in the first video frame to obtain the modified text corresponding to the first video frame, and using the modified text corresponding to the first video frame and the original text corresponding to the second video frame as the fifth text information.

[0183] The steps of obtaining the fifth text information by the second method are: performing speech recognition on the target speech in the target video to obtain the fifth text information; the target speech is the speech used to reflect the events occurring in the target video.

[0184] The fifth text information representing the content of the conversation in the target video can also be obtained by the aforementioned S1011-S1012. That is, when the target text in each video frame of the target video is respectively recognized and the original text corresponding to each video frame is obtained, for any two adjacent video frames in each video frame, if the original text corresponding to the two adjacent video frames includes continuous and repeated repeated text, the repeated text included in the first video frame is removed to obtain the modified text corresponding to the first video frame, and the modified text corresponding to the first video frame and the original text corresponding to the second video frame are determined as the fifth text information; when the target text in each video frame of the target video is respectively recognized and the original text corresponding to each video frame is not obtained, the target voice in the target video is recognized to obtain the fifth text information.

[0185] S402: Filter the fifth text information that meets the first filtering condition in the fifth text information to obtain sixth text information as target text information.

[0186] In one possible embodiment, the sixth text information can be obtained through the large language model. Specifically, the fifth text information and the second prompt word are input into the second large language model, and the sixth text information output by the second large language model under the guidance of the second prompt word is obtained as the target text information. The second prompt word is used to guide the second large language model to filter the fifth text information that meets the first filtering condition in the fifth text information, where the first filtering condition is the condition represented by the second prompt word.

[0187] The second prompt word may include: conditions required for the target text information, or conditions not required for the target text information. The filtered sixth text information needs to meet the conditions required for the target text information in the second prompt word, and may meet or not meet the conditions not required for the target text information.

[0188] Exemplarily, the conditions required for the target text information may include:

[0189] 1. It should be logical and the large model should be understandable;

[0190] 2. The language should only include Chinese and English, without any garbled characters;

[0191] 3. No lyrics or lyrics account for less than 50%;

[0192] 4. Don’t be too abstract;

[0193] 5. The text should have certain details;

[0194] 6. It cannot be a chicken soup article;

[0195] 7. It cannot be just a simple statement without details;

[0196] 8. It cannot be a letter, prose, or poem.

[0197] Conditions that are not required for target text information may include:

[0198] 1. Have clear paragraphs;

[0199] 2. Separated by punctuation marks;

[0200] 3. The content is structured.

[0201] In one possible embodiment, the second prompt word may include all of the above-mentioned conditions required for the target text information, as well as all of the above-mentioned conditions not required for the target text information. In another possible embodiment, the second prompt word may also only include all of the above-mentioned conditions required for the target text information. In yet another possible embodiment, the second prompt word may include 1 and 2 of the above-mentioned conditions required for the target text information, as well as 1 and 3 of the above-mentioned conditions not required for the target text information.

[0202] If a fifth text message meets the requirements for the target text message in the second prompt word, that is, the fifth text message meets the first filtering condition, the second language model will output a prompt message such as "does not meet the conditions" or "subtitles are illegal" for the fifth text message. If a fifth text message does not meet the requirements for the target text message in the second prompt word, that is, the fifth text message does not meet the first filtering condition, the second language model will output a prompt message such as "meets the conditions" or "subtitles are legal" for the fifth text message.

[0203] The fifth text information whose prompt information output by the second language model is "qualified" or "legal subtitles" is the filtered sixth text information, and the sixth text information is used as the target text information.

[0204] The second largest language model is any LLM that can filter the fifth text information according to the fifth text information and the second prompt word, for example, models such as GPT4-Turbo-128K.

[0205] In other possible embodiments, the fifth text information that matches the first filtering condition in the fifth text information may be filtered using other methods other than the large language model to obtain the sixth text information as the target text information. For example, the fifth text information that matches the first filtering condition in the fifth text information may be filtered using a deep learning model according to a preset description style to obtain the sixth text information as the target text information.

[0206] By selecting this embodiment, after obtaining the fifth text information used to represent the conversation content in the target video, the fifth text information that hits the first filtering condition in the fifth text information can be filtered to obtain the sixth text information as the target text information, so that the fifth text information can be filtered to remove the text information in the fifth text information that is difficult for the machine to understand, and the sixth text information can be obtained as the target text information, which facilitates the machine to automatically generate the first content information based on the target text information, improve the relevance of the first content information with the content of the target video, thereby improving the relevance of the second content information obtained based on the first content information with the content of the target video, and then when the second content information is determined as the comment corresponding to the target video, the relevance of the automatically generated comments for the target video with the content of the target video is improved, and the quality of the automatically generated comments is improved.

[0207] The above has provided an exemplary description of the method for filtering the fifth text information. It is understandable that the target text information obtained according to the above-mentioned method one or method two may contain typos, and may also include some text information that is difficult for the machine to understand. Therefore, if the modified text corresponding to the first video frame in method one and the original text corresponding to the second video frame are directly used as the target text information, or the speech recognition result in method two is directly used as the target text information, the accuracy of the target text information will be reduced, resulting in a decrease in the relevance between the first content information obtained based on the erroneous target text information and the content of the target video, thereby resulting in a decrease in the relevance between the second content information obtained based on the first content information and the content of the target video, and further resulting in a decrease in the relevance between the comments automatically generated for the target video and the content of the target video when the second content information is determined as the comment corresponding to the target video, resulting in a decrease in the quality of the automatically generated comments. Based on this, in order to further improve the accuracy of the target text information, so as to improve the relevance between the comments automatically generated for the target video and the content of the target video, thereby improving the quality of the automatically generated comments, in a possible embodiment, see. Figure 5 , the aforementioned S202 includes:

[0208] S2021, for any two adjacent video frames in each video frame, if the original texts corresponding to the two adjacent video frames include continuous and repeated repeated texts, then remove the repeated text included in the first video frame to obtain the modified text corresponding to the first video frame, and perform text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame to obtain first text information.

[0209] The above has provided an exemplary description of how to obtain the modified text corresponding to the first video frame and the original text corresponding to each video frame for any two adjacent video frames in each video frame. Please refer to the aforementioned description of S201-S202, which will not be repeated here. For ease of description below, the modified text corresponding to the first video frame and the original text corresponding to the second video frame obtained for any two adjacent video frames in each video frame are recorded as processed text. The step of performing text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame to obtain first text information is equivalent to performing text correction on the processed text to obtain the first text information.

[0210] In one possible embodiment, the first text information can be obtained through a large language model. Specifically, the processed text and a first prompt word are input into the first large language model to obtain the first text information output by the first large language model under the guidance of the first prompt word; the first prompt word is used to guide the first large language model to perform text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame.

[0211] The first prompt word may include: (1) if there are no typos and grammatical errors in the processed text, directly outputting the processed text; (2) if there are typos and grammatical errors in the processed text, correcting the typos and grammatical errors in the processed text to obtain first text information, and outputting the first text information.

[0212] That is, if there are no typos and grammatical errors in the processed text, the processed text is directly used as the first text information; if there are typos and grammatical errors in the processed text, the typos and grammatical errors in the processed text are corrected to obtain the first text information.

[0213] The first language model is any LLM that can perform text correction on the processed text based on the processed text and the first prompt word, for example, the Doubao-pro-128k model.

[0214] In other possible embodiments, the processed text may be modified by other methods other than the large language model to obtain the first text information. For example, the processed text may be modified by a deep learning model to obtain the first text information.

[0215] S2022: Filter the first text information that matches the first filtering condition in the first text information to obtain second text information as target text information.

[0216] In one possible embodiment, the second text information can be obtained through a large language model. Specifically, the first text information and the second prompt word are input into the second large language model, and the second text information output by the second large language model under the guidance of the second prompt word is obtained as the target text information. The second prompt word is used to guide the second large language model to filter the first text information in the first text information that meets a first filtering condition, where the first filtering condition is the condition represented by the second prompt word.

[0217] The method of using the second largest language model to filter the first text information to obtain the second text information is similar to the method of using the second largest language model to filter the fifth text information to obtain the sixth text information in the aforementioned S402. The only difference is that the fifth text information in the aforementioned S402 is replaced by the first text information, and the sixth text information is replaced by the second text information. Therefore, the method of using the second largest language model to filter the first text information to obtain the second text information will not be repeated here.

[0218] In other possible embodiments, the first text information that matches the first filtering condition in the first text information may be filtered using other methods other than the large language model to obtain the second text information. For example, the first text information that matches the first filtering condition in the first text information may be filtered using a deep learning model to obtain the second text information.

[0219] By selecting this embodiment, the processed text can be corrected to obtain the first text information, thereby improving the accuracy of the first text information, and filtering the first text information that hits the first filtering condition in the first text information, removing the text information that is difficult for the machine to understand in the first text information with higher accuracy, and obtaining the second text information as the target text information, so that the target text information includes neither typos or grammatical errors, nor text information that is difficult for the machine to understand, thereby making the target text information more accurate, and facilitating the machine to automatically generate the first content information based on the target text information, thereby improving the relevance of the first content information with the content of the target video, thereby improving the relevance of the second content information obtained based on the first content information with the content of the target video, and then when the second content information is determined as the comment corresponding to the target video, the relevance of the automatically generated comments for the target video with the content of the target video is improved, thereby improving the quality of the automatically generated comments.

[0220] See also Figure 6 , the aforementioned method 2 includes:

[0221] S601, performing speech recognition on a target speech in a target video to obtain a recognition result.

[0222] The method for obtaining the recognition result is the same as the method for obtaining the target text information in the aforementioned method 2. Please refer to the relevant description in the aforementioned method 2, so it will not be repeated here.

[0223] S602: Perform text correction on the recognition result to obtain third text information.

[0224] In one possible embodiment, the third text information can be obtained using a large language model. Specifically, the recognition result and the third prompt word are input into the third large language model to obtain the third text information output by the third large language model under the guidance of the third prompt word. The third prompt word is used to guide the third large language model to perform text correction on the recognition result.

[0225] The third prompt word may include: (1) if there are no typos and grammatical errors in the recognition result, directly outputting the recognition result; (2) if there are typos and grammatical errors in the recognition result, correcting the typos and grammatical errors in the recognition result to obtain third text information, and outputting the third text information.

[0226] That is to say, if there are no typos and grammatical errors in the recognition result, the recognition result is directly used as the third text information; if there are typos and grammatical errors in the recognition result, the typos and grammatical errors in the processed text are corrected to obtain the third text information.

[0227] The third language model is any LLM that can perform text correction on the recognition result based on the recognition result and the third prompt word, for example, the Doubao-pro-128k model.

[0228] The third largest language model in S602 and the first largest language model in S2022 may be the same model or different models.

[0229] In other possible embodiments, the recognition result may be modified by other methods other than the large language model to obtain the third text information. For example, the recognition result may be modified by a deep learning model to obtain the third text information.

[0230] S603: Filter the third text information that meets the first filtering condition in the third text information to obtain fourth text information as target text information.

[0231] In one possible embodiment, the fourth text information can be obtained through the large language model. Specifically, the third text information and the second prompt word are input into the second large language model, and the fourth text information output by the second large language model under the guidance of the second prompt word is obtained as the target text information. The second prompt word is used to guide the second large language model to filter the third text information that meets the first filtering condition in the third text information, where the first filtering condition is the condition represented by the second prompt word.

[0232] The method of using the second largest language model to filter the third text information to obtain the fourth text information is similar to the method of using the second largest language model to filter the fifth text information to obtain the sixth text information in the aforementioned S402. The only difference is that the fifth text information in the aforementioned S402 is replaced by the third text information, and the sixth text information is replaced by the fourth text information. Therefore, the method of using the second largest language model to filter the third text information to obtain the fourth text information will not be repeated here.

[0233] In other possible embodiments, the third text information that matches the first filtering condition in the third text information may be filtered using other methods other than the large language model to obtain the fourth text information. For example, the third text information that matches the first filtering condition in the third text information may be filtered using a deep learning model to obtain the fourth text information.

[0234] By selecting this embodiment, after performing speech recognition on the target speech in the target video to obtain a recognition result, the recognition result can be text-corrected to obtain a third text information, thereby improving the accuracy of the third text information, and filtering the third text information that hits the first filtering condition in the third text information, removing the text information that is difficult for the machine to understand from the third text information with higher accuracy, and obtaining the fourth text information as the target text information, so that the target text information includes neither typos or grammatical errors, nor text information that is difficult for the machine to understand, making the target text information more accurate, and facilitating the machine to automatically generate the first content information based on the target text information, thereby improving the relevance of the first content information with the content of the target video, thereby improving the relevance of the second content information obtained based on the first content information with the content of the target video, and then when the second content information is determined as the comment corresponding to the target video, the relevance of the comment automatically generated for the target video with the content of the target video is improved, thereby improving the quality of the automatically generated comments.

[0235] In other possible embodiments, the target text information may be obtained through method 1 and method 2. Specifically, the processed text is obtained through method 1, the recognition result is obtained through method 2, and the processed text and the recognition result are fused to obtain the target text information.

[0236] The above has provided an exemplary explanation of the method for obtaining the target text information. It is understandable that the generated second content information may include words banned by the video platform. If the second content information is used directly as a comment corresponding to the target video, there will be a problem that the words banned by the video platform in the comment cannot be displayed, or there will be a problem that the comment including words banned by the video platform cannot be displayed, that is, there will be a problem that the comment cannot be fully displayed or cannot be displayed, which reduces the practicality of the comment, resulting in lower quality of the automatically generated comments, and at the same time, it will also reduce the user or video creator's experience. In addition, the second content information may also have the problem of repeated sentences. If the second content information is used directly as a comment, the repeated sentence comments will reduce the user or video creator's experience.

[0237] Based on this, in order to avoid the situation where comments cannot be fully displayed or cannot be displayed, to improve the quality of automatically generated comments, and at the same time, to improve the user or video creator's experience, in a possible embodiment, see Figure 7 , the comment generation method provided by the present invention includes:

[0238] S101, obtaining target text information for representing the dialogue content in a target video.

[0239] The dialogue content is text and / or voice that reflects the events that occurred in the target video.

[0240] S102: Generate first content information according to the target text information.

[0241] The first content information is text used to describe an event occurring in the target video.

[0242] S101 - S102 have been described exemplarily in the foregoing text. Please refer to the aforementioned related descriptions of S101 - S102 and will not be repeated here.

[0243] S1031 : Modify the first content information according to a preset description style to generate a plurality of different second content information.

[0244] The method of generating the second content information in S1031 is the same as the method of generating the second content information in S103 mentioned above. Please refer to the above description of S103 and will not be repeated here.

[0245] S1032: Filter the second content information that matches the second filtering condition in the second content information to obtain third content information, and determine the third content information as a comment corresponding to the target video.

[0246] In a possible embodiment, the third content information can be obtained through a large language model. Specifically, the second content information and the sixth prompt are input into the sixth large language model to obtain the third content information output by the sixth large language model under the guidance of the sixth prompt; wherein, the sixth prompt is used to guide the sixth large language model to filter the second content information that hits the second filtering condition in the second content information; the second filtering condition is the condition represented by the sixth prompt.

[0247] The sixth prompt may include: (1) Do not use exaggerated words such as "too", "very", "extremely", "really", "actually", "especially", "undoubtedly", "for the first time", "truly", "really is", "mind-blowing", "unexpected". (2) Do not use sentence patterns similar to "a bit xx", such as: a bit interesting, a bit cute, etc.

[0248] The sixth prompt may also only include: (1) Do not use exaggerated words such as "too", "very", "extremely", "really", "actually", "especially", "undoubtedly", "for the first time", "truly", "really is", "mind-blowing", "unexpected". (2) Do not use words prohibited by the video platform.

[0249] If a certain second content information hits the second filtering condition represented by the above sixth prompt, the sixth large language model will filter out this second content information; if a certain second content information does not hit the second filtering condition represented by the above sixth prompt, the sixth large language model will output this second content information.

[0250] Exemplarily, assume that the second filtering condition represented by the sixth prompt is: the second content information includes the word "too", and the 100 pieces of second content information from second content information 1 to second content information 100 and the sixth prompt are input into the sixth large language model. If the sixth large language model detects that the three pieces of second content information, i.e., second content information 3, second content information 20, and second content information 85, include the word "too", then the sixth large language model will filter out the three pieces of second content information, i.e., second content information 3, second content information 20, and second content information 85. The output of the sixth large language model is: 97 pieces of second content information, namely second content information 1 - second content information 2, second content information 4 - second content information 19, second content information 21 - second content information 84, and second content information 86 - second content information 100.

[0251] The sixth large language model is any LLM that can filter the second content information according to the second content information and the sixth prompt. For example, models such as GPT4-Turbo-128K.

[0252] In other possible embodiments, the second content information that matches the second filtering condition may be filtered using other methods other than the large language model to obtain the third content information. For example, the second content information that matches the second filtering condition may be filtered using a deep learning model to obtain the third content information.

[0253] With this embodiment, after generating the second content information, the third content information can be obtained by filtering the second content information that matches the second filtering condition, so that the third content information does not include words banned by the video platform and does not contain repeated sentences. The filtered third content information is then used as the comment corresponding to the target video, avoiding the inclusion of words banned by the video platform or repeated sentences in the comments, thereby avoiding the situation where the comments cannot be fully displayed or cannot be displayed at all, improving the practicality of the comments, and thus improving the quality of the automatically generated comments, while also improving the user or video creator's experience.

[0254] It is understandable that the multiple generated second content information is different. When the second content information is determined as the comments corresponding to the target video, the multiple comments are different, and thus the relevance of each comment to the content of the target video varies, and the quality of each comment also varies. If the comments are displayed in an arbitrary order, users may preferentially see comments of lower quality than other comments. Low-quality comments are less relevant to the content of the target video, making it difficult for users to accurately determine whether the target video contains content that interests them, resulting in a poor user experience.

[0255] Based on this, in order to improve user experience, in a possible embodiment, see Figure 8 , the comment generation method provided by the present invention includes:

[0256] S101, obtaining target text information for representing the dialogue content in a target video.

[0257] The dialogue content is text and / or voice that reflects the events that occurred in the target video.

[0258] S102: Generate first content information according to the target text information.

[0259] The first content information is text used to describe an event occurring in the target video.

[0260] S103: Modify the first content information according to a preset description style to generate a plurality of different second content information, and determine the second content information as comments corresponding to the target video.

[0261] S101 - S103 have been described exemplarily in the foregoing text. Please refer to the aforementioned related descriptions of S101 - S103 and will not be repeated here.

[0262] S104: Sort the comments corresponding to the target video according to the relevance between the comments corresponding to the target video and the content of the target video.

[0263] In one possible embodiment, the comments corresponding to the target video can be sorted using a large language model. Specifically, the comments corresponding to the target video and the seventh prompt word are input into the seventh large language model, and the seventh large language model outputs the comments corresponding to the target video in a sorted order. The seventh large language model is used to sort the comments corresponding to the target video based on the relevance between the comments corresponding to the target video and the content of the target video, guided by the seventh prompt word.

[0264] The seventh cues may include: 1. Natural Tone: The target content should be fluent and authentic, not stiff or contrived, and should not use overly formal or written language. 2. Substantial Content: The target content should include specific information or insights, not just brief statements or emotional expressions. The target content should include detailed viewpoints, examples, or explanations. 3. Colloquialism: The target content should use everyday language and expressions, avoiding technical jargon or overly complex vocabulary.

[0265] In other possible embodiments, the seventh prompt word may also include: 1. Natural Tone: The target content information should be fluent and authentic, not stiff or contrived, and should not use overly formal or written language. 2. Substantial Content: The target content information should contain specific information or insights, not just brief statements or emotional expressions, and should include detailed viewpoints, examples, or explanations.

[0266] The seventh language model can score each target content information according to the requirements indicated in the seventh prompt word, obtain a score corresponding to each target content information, and output each target content information in descending order of the corresponding score. For example, assuming that comments 1-5 and the seventh prompt word are input into the seventh language model, the scores corresponding to comments 1-5 are recorded as score 1-score 5, respectively, and score 1> score 3> score 5> score 2> score 4, then the output of the seventh language model is as follows:

[0267] 1. Comment 1

[0268] 2. Comment 3

[0269] 3. Comment 5

[0270] 4. Comment 2

[0271] 5. Comment 4

[0272] The order of the comments can be determined by the serial number before each comment.

[0273] The seventh language model is any LLM that can sort comments based on the comments and the seventh prompt word, for example, models such as GPT4-Turbo-128K.

[0274] In other possible embodiments, the comments corresponding to the target video may be ranked using other methods besides the large language model, for example, by using a deep learning model.

[0275] By selecting this embodiment, the comments corresponding to the target video can be sorted based on their relevance to the target video's content, helping users assess the relevance of each comment to the target video's content, thereby helping users assess the quality of each comment. This allows the comments to be displayed according to their ranking, prioritizing users with higher-quality comments. This allows users to accurately determine whether the target video contains content of interest based on comments with higher quality, i.e., those with a higher relevance to the target video's content, thereby improving user experience.

[0276] In one possible embodiment, see Figure 9 , the comment generation method provided by this application includes:

[0277] S901: Acquire target text information for representing the dialogue content in a target video.

[0278] The dialogue content is text and / or voice that reflects the events that occurred in the target video.

[0279] S902: Generate first content information according to the target text information.

[0280] The first content information is text used to describe an event occurring in the target video.

[0281] S903: Modify the first content information according to a preset description style to generate a plurality of different second content information.

[0282] S904: Filter the second content information that matches the second filtering condition in the second content information to obtain third content information, and determine the third content information as a comment corresponding to the target video.

[0283] S905 , sorting the comments corresponding to the target video according to the relevance between the comments corresponding to the target video and the content of the target video.

[0284] Figure 9In the illustrated embodiment, S901-S902 correspond to the aforementioned S101-S102, and the relevant descriptions in S101-S102 can be referred to, and will not be repeated here. S903-S904 correspond to the aforementioned S1031-S1032, and the relevant descriptions in S1031-S1032 can be referred to, and will not be repeated here. S905 corresponds to the aforementioned S104, and the relevant descriptions in S104 can be referred to, and will not be repeated here.

[0285] The flowchart of the comment generation method provided by the present invention can also be as follows Figure 10 As shown, it includes three processes: video information extraction, video information processing, content understanding, and comment generation.

[0286] Specifically, the video information extraction process includes:

[0287] S1001, audio and video editing software (ffmpeg) extracts audio files from the video.

[0288] S1002, use audio and video editing software (ffmpeg) to capture the video, generating 5 pictures per second.

[0289] S1003, automatic speech recognition (ASR) text extraction.

[0290] S1004, optical character recognition (OCR) text extraction.

[0291] The video is the aforementioned target video, S1001 and S1003 correspond to the aforementioned method 2, S1002 and S1004 correspond to the aforementioned method 1, and the relevant descriptions of the aforementioned methods 1 and 2 can be referred to, which will not be repeated here.

[0292] Video information processing and content understanding process include:

[0293] S1005, confirm the text content, and use optical character recognition (OCR) if available, or automatic speech recognition (ASR) if not.

[0294] S1005 corresponds to the aforementioned S1011-S1012. Please refer to the relevant description of the aforementioned S1011-S1012 and will not be repeated here.

[0295] S1006, correct typos and grammar.

[0296] S1006 corresponds to the aforementioned S2021 and S602. Please refer to the relevant descriptions of the aforementioned S2021 and S602, and will not be repeated here.

[0297] S1007, filtering text information that is not suitable for generating comments.

[0298] S1007 corresponds to the aforementioned S2022 and S603. Please refer to the relevant descriptions of the aforementioned S2022 and S603, and will not be repeated here.

[0299] S1008, extract and summarize the video content.

[0300] S1008 corresponds to the aforementioned S102. Please refer to the relevant description of the aforementioned S102 and will not be repeated here.

[0301] The review generation process includes:

[0302] S1009, comment generation.

[0303] S1010, comment filtering.

[0304] S1011, comment sorting.

[0305] The comments are the aforementioned target content information. S1009-S1011 correspond to S903-S905. Please refer to the relevant descriptions of S903-S905 above and will not be repeated here. In the embodiment where steps S1006-S1011 are implemented using a large language model, the large language model used in S1006-S1011 can be deployed in series on Dify (a generative artificial intelligence application engine).

[0306] Corresponding to the aforementioned comment generation method, the embodiment of the present invention further provides a comment generation device, see Figure 11 , the device comprises:

[0307] The target text information acquisition module 111 is used to acquire target text information used to represent the dialogue content in the target video; the dialogue content is text and / or voice used to reflect the events occurring in the target video;

[0308] The first content information generating module 112 is configured to generate first content information based on the target text information; the first content information is a text describing an event occurring in the target video;

[0309] The second content information generating module 113 is configured to modify the first content information according to a preset description style, generate a plurality of different second content information, and determine the second content information as comments corresponding to the target video.

[0310] In a possible embodiment, obtaining target text information for representing the content of a conversation in a target video includes:

[0311] For each video frame in the target video, perform text recognition on the target text in the video frame to obtain the original text corresponding to the video frame; the target text is the text used to reflect the events that occurred in the target video;

[0312] For any two adjacent video frames in each video frame, if the original texts corresponding to the two adjacent video frames include continuous and repeated duplicate texts, then remove the duplicate text included in the first video frame to obtain the modified text corresponding to the first video frame, and determine the modified text corresponding to the first video frame and the original text corresponding to the second video frame as the target text information; wherein the first video frame is the video frame that is earlier in time sequence among the two adjacent video frames, and the second video frame is the video frame that is later in time sequence among the two adjacent video frames;

[0313] Alternatively, speech recognition is performed on a target speech in a target video to obtain target text information; the target speech is a speech used to reflect an event occurring in the target video.

[0314] In a possible embodiment, determining the modified text corresponding to the first video frame and the original text corresponding to the second video frame as target text information includes: performing text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame to obtain first text information; filtering first text information that matches a first filtering condition in the first text information to obtain second text information as the target text information;

[0315] Performing speech recognition on the target speech in the target video to obtain target text information includes: performing speech recognition on the target speech in the target video to obtain a recognition result; performing text correction on the recognition result to obtain third text information; filtering the third text information that hits the first filtering condition in the third text information to obtain fourth text information as the target text information.

[0316] In one possible embodiment, performing text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame to obtain first text information includes: inputting the modified text corresponding to the first video frame, the original text corresponding to the second video frame, and a first prompt word into a first large language model, thereby obtaining the first text information output by the first large language model under the guidance of the first prompt word; the first prompt word being used to guide the first large language model to perform text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame;

[0317] Filtering the first text information that matches the first filtering condition in the first text information to obtain second text information as the target text information, including: inputting the first text information and the second prompt word into a second language model, obtaining the second text information output by the second language model under the guidance of the second prompt word as the target text information, wherein the second prompt word is used to guide the second language model to filter the first text information that matches the first filtering condition in the first text information, and the first filtering condition is the condition represented by the second prompt word;

[0318] Performing text correction on the recognition result to obtain third text information, including: inputting the recognition result and the third prompt word into a third language model to obtain third text information output by the third language model under the guidance of the third prompt word; the third prompt word is used to guide the third language model to perform text correction on the recognition result;

[0319] The method includes filtering the third text information that matches the first filtering condition in the third text information to obtain fourth text information as the target text information, including: inputting the third text information and the second prompt word into the second largest language model, obtaining the fourth text information output by the second largest language model under the guidance of the second prompt word as the target text information, wherein the second prompt word is used to guide the second largest language model to filter the third text information that matches the first filtering condition in the third text information, and the first filtering condition is the condition represented by the second prompt word.

[0320] In a possible embodiment, obtaining target text information for representing the content of a conversation in a target video includes: obtaining fifth text information for representing the content of a conversation in the target video; filtering the fifth text information that meets a first filtering condition in the fifth text information to obtain sixth text information as the target text information.

[0321] In one possible embodiment, generating first content information based on target text information includes: inputting the target text information and a fourth prompt word into a fourth language model, thereby obtaining first content information output by the fourth language model under the guidance of the fourth prompt word; the first content information is text described in natural language and is used to represent the theme and specific content of the target video, where the theme is text that provides an overall description of events occurring in the target video, and the specific content is text that provides segmented descriptions of events occurring in the target video.

[0322] In one possible embodiment, first content information is modified according to a preset description style to generate multiple different second content information, including: inputting the first content information, the title of the target video, and a fifth prompt word into a fifth language model, obtaining multiple different second content information output by the fifth language model under the guidance of the fifth prompt word; wherein the second content information is text described in natural language; the fifth prompt word is used to guide the fifth language model to modify the first content information according to the preset description style, and the fifth prompt word includes a comment with the preset description style.

[0323] In a possible embodiment, determining the second content information as a comment corresponding to the target video includes: filtering the second content information that hits the second filtering condition in the second content information to obtain third content information, and determining the third content information as a comment corresponding to the target video.

[0324] In a possible embodiment, the apparatus further includes: a sorting module configured to sort the comments corresponding to the target video according to the relevance between the comments corresponding to the target video and the content of the target video.

[0325] In one possible embodiment, filtering the second content information that matches a second filtering condition in the second content information to obtain the third content information includes: inputting the second content information and a sixth prompt word into a sixth language model, and obtaining the third content information output by the sixth language model under the guidance of the sixth prompt word; wherein the sixth prompt word is used to guide the sixth language model to filter the second content information that matches the second filtering condition in the second content information; and the second filtering condition is the condition represented by the sixth prompt word.

[0326] The comments corresponding to the target video are sorted according to the relevance between the comments corresponding to the target video and the content of the target video, including: inputting the comments corresponding to the target video and the seventh prompt word into the seventh language model, obtaining the comments corresponding to the target video with an arrangement order output by the seventh language model; and the seventh language model is used to sort the comments corresponding to the target video according to the relevance between the comments corresponding to the target video and the content of the target video under the guidance of the seventh prompt word.

[0327] The embodiment of the present invention further provides an electronic device, such as Figure 12 As shown, it includes a processor 121, a communication interface 122, a memory 123 and a communication bus 124, wherein the processor 121, the communication interface 122, and the memory 123 communicate with each other through the communication bus 124.

[0328] Memory 123, for storing computer programs;

[0329] The processor 121 is configured to execute the program stored in the memory 123 by performing the following steps:

[0330] Obtaining target text information for representing the content of a conversation in a target video; the conversation content is text and / or voice for reflecting an event occurring in the target video;

[0331] Generate first content information based on the target text information; the first content information is text used to describe an event occurring in the target video;

[0332] The first content information is modified according to a preset description style to generate a plurality of different second content information, and the second content information is determined as comments corresponding to the target video.

[0333] The communication bus mentioned in the terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0334] The communication interface is used for communication between the above terminal and other devices.

[0335] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.

[0336] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0337] In another embodiment of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the comment generation method described in any one of the above embodiments is implemented.

[0338] In another embodiment of the present invention, a computer program product including instructions is provided. When the computer program product is run on a computer, the computer is enabled to execute the comment generation method described in any one of the above embodiments.

[0339] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0340] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0341] Each embodiment in this specification is described in a related manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the embodiments of the apparatus, electronic device, computer-readable storage medium, and computer program product containing instructions are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, reference can be made to the description of the method embodiments.

[0342] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of protection of the present invention.

Claims

1. A comment generation method, characterized in that: The method comprises: Acquire target text information for representing the content of a conversation in a target video; the conversation content is text and / or voice for reflecting an event occurring in the target video; Generate first content information according to the target text information; the first content information is a text for describing an event occurring in the target video; The first content information is modified according to a preset description style to generate a plurality of different second content information, and the second content information is determined as comments corresponding to the target video.

2. The method according to claim 1, characterized in that The step of obtaining target text information for representing the dialogue content in the target video includes: For each video frame in the target video, performing text recognition on the target text in the video frame to obtain the original text corresponding to the video frame; the target text is a text used to reflect the event occurring in the target video; For any two adjacent video frames among the video frames, if the original texts corresponding to the two adjacent video frames include continuous and repeated repeated text, the repeated text included in the first video frame is removed to obtain a modified text corresponding to the first video frame, and the modified text corresponding to the first video frame and the original text corresponding to the second video frame are determined as target text information; wherein the first video frame is the video frame that is earlier in time sequence among the two adjacent video frames, and the second video frame is the video frame that is later in time sequence among the two adjacent video frames; or, Speech recognition is performed on a target speech in the target video to obtain target text information; the target speech is a speech used to reflect an event occurring in the target video.

3. The method according to claim 2, characterized in that The step of determining the modified text corresponding to the first video frame and the original text corresponding to the second video frame as target text information includes: Performing text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame to obtain first text information; Filtering the first text information that matches the first filtering condition in the first text information to obtain second text information as the target text information; The performing speech recognition on the target speech in the target video to obtain target text information includes: Performing speech recognition on the target speech in the target video to obtain a recognition result; Performing text correction on the recognition result to obtain third text information; The third text information that matches the first filtering condition in the third text information is filtered to obtain fourth text information as target text information.

4. The method according to claim 3, characterized in that The performing text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame to obtain first text information includes: Inputting the modified text corresponding to the first video frame, the original text corresponding to the second video frame, and a first prompt word into a first large language model to obtain first text information output by the first large language model under the guidance of the first prompt word; the first prompt word is used to guide the first large language model to perform text correction on the modified text corresponding to the first video frame and the original text corresponding to the second video frame; The filtering of the first text information that matches the first filtering condition in the first text information to obtain the second text information as the target text information includes: Inputting the first text information and the second prompt word into a second language model, obtaining second text information output by the second language model under the guidance of the second prompt word as the target text information, wherein the second prompt word is used to guide the second language model to filter first text information in the first text information that matches a first filtering condition, where the first filtering condition is the condition represented by the second prompt word; The performing text correction on the recognition result to obtain third text information includes: inputting the recognition result and a third prompt word into a third language model to obtain third text information output by the third language model under the guidance of the third prompt word; the third prompt word is used to guide the third language model to perform text correction on the recognition result; The filtering of the third text information that matches the first filtering condition in the third text information to obtain fourth text information as the target text information includes: The third text information and the second prompt word are input into a second language model to obtain fourth text information output by the second language model under the guidance of the second prompt word as the target text information, wherein the second prompt word is used to guide the second language model to filter the third text information that meets a first filtering condition in the third text information, and the first filtering condition is the condition represented by the second prompt word.

5. The method according to claim 1, wherein The step of obtaining target text information for representing the dialogue content in the target video includes: Acquiring fifth text information for representing the content of the dialogue in the target video; The fifth text information that matches the first filtering condition in the fifth text information is filtered to obtain sixth text information as target text information.

6. The method according to claim 1, characterized in that Generating first content information according to the target text information includes: The target text information and the fourth prompt word are input into a fourth language model to obtain first content information output by the fourth language model under the guidance of the fourth prompt word; the first content information is text described in natural language and is used to represent the theme and specific content of the target video, the theme being text that provides an overall description of events occurring in the target video, and the specific content being text that provides segmented descriptions of events occurring in the target video.

7. The method according to claim 1, characterized in that The step of modifying the first content information according to a preset description style to generate a plurality of different second content information includes: The first content information, the title of the target video, and a fifth prompt word are input into a fifth language model to obtain a plurality of different second content information output by the fifth language model under the guidance of the fifth prompt word; wherein the second content information is text described in a natural language; the fifth prompt word is used to guide the fifth language model to modify the first content information according to a preset description style, and the fifth prompt word includes a comment with the preset description style.

8. The method according to claim 1, characterized in that The determining the second content information as a comment corresponding to the target video includes: Second content information matching a second filtering condition in the second content information is filtered to obtain third content information, and the third content information is determined as a comment corresponding to the target video.

9. The method according to claim 1 or 8, characterized in that The method further comprises: The comments corresponding to the target video are sorted according to the relevance between the comments corresponding to the target video and the content of the target video.

10. The method according to claim 9, characterized in that The filtering of the second content information matching the second filtering condition in the second content information to obtain the third content information includes: inputting the second content information and the sixth prompt word into a sixth language model to obtain third content information output by the sixth language model under the guidance of the sixth prompt word; The sixth prompt word is used to guide the sixth language model to filter the second content information that matches the second filtering condition in the second content information; the second filtering condition is the condition represented by the sixth prompt word; The sorting of the comments corresponding to the target video according to the relevance between the comments corresponding to the target video and the content of the target video includes: The comments corresponding to the target video and the seventh prompt word are input into the seventh language model, and the comments corresponding to the target video with an arrangement order are output by the seventh language model; the seventh language model is used to sort the comments corresponding to the target video according to the relevance between the comments corresponding to the target video and the content of the target video under the guidance of the seventh prompt word.

11. A comment generation device, characterized in that: The device comprises: A target text information acquisition module is used to acquire target text information used to represent the dialogue content in the target video; the dialogue content is text and / or voice used to reflect the events occurring in the target video; A first content information generating module is configured to generate first content information based on the target text information; the first content information is a text describing an event occurring in the target video; The second content information generating module is configured to modify the first content information according to a preset description style to generate a plurality of different second content information, and determine the second content information as comments corresponding to the target video.

12. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method steps described in any one of claims 1 to 10 when executing a program stored in a memory.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps of any one of claims 1 to 10 are implemented.

Citation Information

Cited By

  • Video comment generation method and device, electronic equipment, storage medium and computer program product

    CN122240877A