Short comment generation method and apparatus, device, and storage medium
By extracting target comments from multimedia content comment information and using a language model to generate short reviews, and eliminating non-unique short reviews, the problem of low quality of recommendation reasons in existing technologies is solved, and high-quality recommendation short reviews are generated.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2026-03-12
AI Technical Summary
In existing technologies, when recommending multimedia content by directly obtaining the original comments and generating them using traditional NLP methods, semantic incompleteness, generic categories, or negative short reviews are prone to occur, resulting in low recommendation quality.
By acquiring comment information from multimedia content, extracting target comment information with recommendation significance, segmenting it into short sentences, using a language model to generate a short comment recommendation library, and removing non-unique short comments to ensure the quality of the generated recommended short comments.
It improves the quality of recommended short reviews, avoids semantically incomplete and generic short reviews, enhances recommendation effectiveness, and increases generation efficiency.
Smart Images

Figure CN2025119439_12032026_PF_FP_ABST
Abstract
Description
Method, device and storage medium for generating short comments
[0001] The present application claims priority to the Chinese patent application No. 202411250999.5, filed on September 6, 2024, and entitled "Method, device and storage medium for generating short comments", the entire content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The present application relates to the technical field of computer, in particular to a method, device and storage medium for generating short comments. BACKGROUND
[0003] When displaying multimedia content, an application usually displays a recommended reason for the multimedia content beside the name of the multimedia content.
[0004] In related technologies, the original comments of the multimedia content are directly obtained, and traditional NLP (Natural Language Processing) methods such as key phrase extraction and summary generation are used on the original comments of the multimedia content. The frequency of words and the number of times of phrases appearing in the original comments are used as weights to extract short comments from the original comments of the multimedia content as recommended reasons for the multimedia content.
[0005] However, the short comments extracted by the above method are prone to incomplete semantic sentences, general short comments, and poor short comments, which are not suitable for being described as recommended reasons, and the quality of the short comments is low. SUMMARY
[0006] The present application provides a method, device and storage medium for generating short comments. The technical solution provided by the present application is as follows:
[0007] According to an aspect of the present application, a method for generating short comments is provided, which is executed by a computer device, and the method comprises:
[0008] obtaining a plurality of comment information of a first multimedia content;
[0009] extracting a plurality of target comment information from the plurality of comment information, the target comment information being the comment information having a recommended meaning for the first multimedia content;
[0010] segmenting the plurality of target comment information respectively to obtain a plurality of segmented sentences;
[0011] generating a short comment recommendation library of the first multimedia content by a language model according to the information of the first multimedia content and the plurality of segmented sentences, the short comment recommendation library containing a plurality of recommended short comments;
[0012] The non-unique short comment in the plurality of recommended short comments is removed, and a final recommended short comment of the first multimedia content is selected from the remaining recommended short comments, wherein the non-unique short comment refers to a recommended short comment having relevance to a plurality of other multimedia contents except the first multimedia content.
[0013] According to an aspect of an embodiment of the present application, a short comment generation device is provided, and the device comprises:
[0014] A comment acquisition module is configured to acquire a plurality of comment information of a first multimedia content.
[0015] A comment extraction module is configured to extract a plurality of target comment information from the plurality of comment information, wherein the target comment information refers to comment information having a recommended meaning for the first multimedia content.
[0016] A comment segmentation module is configured to segment the plurality of target comment information respectively to obtain a plurality of segmented short sentences.
[0017] A short comment generation module is configured to generate a short comment recommendation library of the first multimedia content according to information of the first multimedia content and the plurality of segmented short sentences by using a language model, wherein the short comment recommendation library comprises a plurality of recommended short comments.
[0018] A short comment selection module is configured to remove a non-unique short comment in the plurality of recommended short comments, and select a final recommended short comment of the first multimedia content from the remaining recommended short comments, wherein the non-unique short comment refers to a recommended short comment having relevance to a plurality of other multimedia contents except the first multimedia content.
[0019] According to an aspect of an embodiment of the present application, a computer device is provided, which comprises a processor and a memory, wherein the memory stores a computer program, the computer program is loaded and executed by the processor to implement the short comment generation method.
[0020] According to an aspect of an embodiment of the present application, a computer readable storage medium is provided, which stores a computer program, the computer program is loaded and executed by a processor to implement the short comment generation method.
[0021] According to an aspect of an embodiment of the present application, a computer program product is provided, which comprises a computer program, the computer program is loaded and executed by a processor to implement the short comment generation method.
[0022] The technical scheme provided by the embodiment of the present application can bring the following beneficial effects:
[0023] By extracting the target comment information with recommendation significance for the first multimedia content from the plurality of comment information, it is avoided that the final recommended short comment is a short comment without recommendation significance such as a poor short comment, and the recommendation quality of the final recommended short comment is ensured. And by the extraction and rewriting ability of the language model to generate the recommended short comment library, it is avoided that incomplete semantic sentences appear in the generated recommended short comment, the data noise is reduced, the quality of the recommended short comment is significantly improved, and the generation efficiency of the recommended short comment is improved. And by eliminating the non-unique short comment in the plurality of recommended short comments, the proportion of the general short comment in the recommended short comment is reduced, it is avoided that the sentence of generally praising the first multimedia content appears, and the quality of the final recommended short comment is further improved, thereby improving the recommendation effect of the final recommended short comment. BRIEF DESCRIPTION OF DRAWINGS
[0024] Fig. 1 is a schematic diagram of a scheme implementation environment provided by an embodiment of the present application;
[0025] Fig. 2 is a schematic diagram of a song display interface provided by an embodiment of the present application;
[0026] Fig. 3 is a flowchart of a short comment generation method provided by an embodiment of the present application;
[0027] Fig. 4 is a flowchart of a second vocabulary extraction process provided by an embodiment of the present application;
[0028] Fig. 5 is a flowchart of a first comment information extraction process provided by an embodiment of the present application;
[0029] Fig. 6 is a flowchart of a second comment information extraction process provided by an embodiment of the present application;
[0030] Fig. 7 is a flowchart of a non-unique short comment elimination process provided by an embodiment of the present application;
[0031] Fig. 8 is a schematic diagram of a short comment generation process provided by an embodiment of the present application;
[0032] Fig. 9 is a block diagram of a short comment generation device provided by an embodiment of the present application;
[0033] Fig. 10 is a structural block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0034] To make the purpose, technical scheme and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0035] Please refer to Fig. 1, which shows a schematic diagram of a scheme implementation environment provided by an embodiment of the present application. The scheme implementation environment can be implemented as a short comment generation system. The scheme implementation environment can include a terminal device 10 and a server 20.
[0036] The number of terminal devices 10 can be one or more. The terminal device 10 can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a game console, an e-book reader, a multimedia playback device, a wearable device, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc.
[0037] The terminal device 10 can be installed with a client of a target application, which has a function of displaying multimedia content and comment information, and a user can view multimedia content and at least one piece of comment information corresponding to each piece of multimedia content in the target application. The type of the target application is not limited in the present application, including but not limited to a music application, a video application, a reading application, a web page application, etc. Optionally, the target application can be an application that needs to be downloaded and installed, or an application that can be used immediately after being clicked, which is not limited in the present application.
[0038] The server 20 is configured to provide background services for the client of the target application installed and running in the terminal device 10. For example, the server 20 can be a background server of the application described above. The server 20 can be a physical server, or a server cluster composed of multiple servers, or a cloud computing service center. Optionally, the server 20 provides background services for multiple applications in multiple terminal devices 10. The terminal device 10 and the server 20 can communicate with each other through a network.
[0039] In the embodiment of the present application, a plurality of pieces of comment information of the first multimedia content are acquired first, a plurality of pieces of target comment information having a recommendation significance for the first multimedia content are extracted from the plurality of pieces of comment information, and then the plurality of pieces of target comment information are respectively segmented to obtain a plurality of segmented short sentences. According to the information of the first multimedia content and the plurality of segmented short sentences, a short comment recommendation library of the first multimedia content is generated by using a language model, and the short comment recommendation library contains a plurality of recommended short comments. Non-unique short comments are removed from the plurality of recommended short comments, and a final recommended short comment of the first multimedia content is selected from the remaining recommended short comments.
[0040] The first multimedia content can be displayed on a multimedia content display interface, which includes but is not limited to a multimedia content list display interface, a multimedia content recommendation interface, a multimedia content search interface, and the like. The final recommended comment is usually displayed beside the name of the first multimedia content. For example, the final recommended comment can be displayed below, above, right of, left of, right below, right above, or the like of the name of the first multimedia content. As shown in FIG. 2, which shows a schematic diagram of a song display interface, the multimedia content is a song. (1) of FIG. 2 shows a song list display interface 210, in which the names of a plurality of songs in a song list are displayed, and the final recommended comment 201 is displayed above right of each song name. (2) of FIG. 2 shows a song recommendation interface 220, in which the names of a plurality of recommended songs are displayed, and the final recommended comment 201 is displayed above right of each recommended song name. (3) of FIG. 2 shows a song search interface 230, in which the names of a plurality of songs matching the search content are displayed, and the final recommended comment 201 is displayed below each song name.
[0041] Referring to FIG. 3, which shows a flowchart of a comment generation method according to an embodiment of the present application. The execution subject of each step of the method can be a computer device. The method can include at least one of the following steps 310-350:
[0042] Step 310: Obtain a plurality of comment information of a first multimedia content.
[0043] The first multimedia content can be any multimedia content, including but not limited to a song, a video, a novel, an article, and the like. The comment information is text information published by a user for the first multimedia content, which is used to indicate the user's feelings after watching / listening to the first multimedia content.
[0044] Step 320: Extract a plurality of target comment information from the plurality of comment information, the target comment information being comment information having a recommendation significance for the first multimedia content.
[0045] The plurality of comment information includes comment information having a recommendation significance for the first multimedia content, and comment information not having a recommendation significance for the first multimedia content. The comment information having a recommendation significance for the first multimedia content usually refers to comment information having a high relevance to the first multimedia content, or having a noun expression. The comment information not having a recommendation significance for the first multimedia content usually refers to comment information having a low relevance to the first multimedia content, or comment information not meeting the content requirements, such as comment information for cheating, comment information in a multi-round dialogue style, comment information being too long or too short, and the like.
[0046] At step 330, the plurality of target review information is respectively segmented to obtain a plurality of segmented sentences.
[0047] Each target review information is segmented to obtain at least one segmented sentence corresponding to each target review information.
[0048] Optionally, the target review information can be segmented according to the punctuation contained in the target review information to obtain a plurality of segmented sentences. For example, if the target review information is "the song melody is beautiful and emotional, and XXX sings well.", the target review information can be segmented into two segmented sentences "the song melody is beautiful and emotional" and "XXX sings well" according to the punctuation contained in the target review information.
[0049] Optionally, a review segmentation model can be used to segment the target review information according to the semantic information contained in the target review information to obtain a plurality of segmented sentences. For example, if the target review information is "the song melody is beautiful and emotional, and XXX sings well.", the target review information can be segmented into three segmented sentences "the song melody is beautiful", "emotional" and "XXX sings well" by using the review segmentation model according to the semantic information contained in the target review information.
[0050] At step 340, a short review recommendation library of the first multimedia content is generated by a language model according to the information of the first multimedia content and the plurality of segmented sentences, and the short review recommendation library contains a plurality of recommended short reviews.
[0051] The language model refers to a large language model (LLM), and the large language model herein can be any publicly available large language model, such as a natural language model based on a transformer structure trained by a large amount of data, and the amount of data can reach a sample level of hundreds of millions, which is not limited in the present application.
[0052] In some embodiments, the information of the first multimedia content includes N sub-information, each sub-information is used to indicate the configuration situation of one dimension of the first multimedia content, and N is a positive integer. The information of the first multimedia content can also include content information of the first multimedia content, and the content information is a textual expression of the first multimedia content.
[0053] For example, if the first multimedia content is a song, the information of the first multimedia content is used to indicate the production information of the song, including but not limited to singer name, song name, album name, lyrics name, melody name, and lyrics information. If the first multimedia content is a video, the information of the first multimedia content is used to indicate the production information of the video, including but not limited to video title, personnel name contained in the video, video director, video editor, and video script.
[0054] According to the information of the first multimedia content and the plurality of segmented short sentences, input prompt information is constituted, the input prompt information is input into the language model, and recommended short comments of the first multimedia content are generated. Through multiple inputs to the language model, a short comment recommendation library of the first multimedia content is generated, and the short comment recommendation library contains a plurality of recommended short comments.
[0055] Optionally, the input prompt information can contain N pieces of sub-information and the plurality of segmented short sentences. For example, if the first multimedia content is a song, the input prompt information contains the singer name, the song name, the album name, the lyrics name, the music name, and the plurality of segmented short sentences. Optionally, the input prompt information can also contain at least one piece of sub-information in the N pieces of sub-information and the plurality of segmented short sentences. For example, if the first multimedia content is a song, the input prompt information can contain the singer name, the song name, and the plurality of segmented short sentences.
[0056] Exemplarily, the input prompt information can be specifically represented as "You are a song pusher, I input the singer-song-song comment, please describe the song popularity, influence, theme, emotional resonance, song listening scene, song listening feeling, song style, etc. from the comments described by the song, extract 3 eye-catching positive evaluation phrases as selling points for others, preferably with a joke, interesting, exaggerated, and the phrase must have appeared directly in the comment, do not change lines: {0}-{1}-{2}." Wherein 0, 1, 2 are respectively substituted for the singer name, the song name, and the recommended short comment. The generated recommended short comment can refer to Table 1 shown below.
[0057] Table 1
[0058] Optionally, the recommended short comment can be a short comment generated by the language model from the information of the first multimedia content and the plurality of segmented short sentences contained in the input prompt information. Optionally, the recommended short comment can also be a short comment extracted from the plurality of segmented short sentences by the language model according to the plurality of segmented short sentences contained in the input prompt information.
[0059] In some embodiments, the recommended short comment has a character number limit. The recommended short comment needs to be displayed next to the name of the first multimedia content without interfering with the display of the name of the first multimedia content, and therefore the recommended short comment has an upper limit value of the character number and a lower limit value of the character number. The size of the upper limit value of the character number and the lower limit value of the character number is related to the display information such as the display font size of the recommended short comment, the display screen size of the terminal device, and the character length of the name of the first multimedia content, which is not limited in the present application. For example, if the upper limit value of the character number of the recommended short comment is 10 and the lower limit value of the character number is 4, the language model is used to generate a recommended short comment of 4-10 characters.
[0060] In step 350, the non-unique short comment in the plurality of recommended short comments is removed, and a final recommended short comment of the first multimedia content is selected from the remaining recommended short comments. The non-unique short comment refers to a recommended short comment having relevance to a plurality of other multimedia contents except the first multimedia content.
[0061] If the non-unique short comment has relevance to a plurality of other multimedia contents except the first multimedia content, the non-unique short comment cannot produce a recommendation effect on the first multimedia content, which is likely to cause user aesthetic fatigue and reduce the conversion rate of the user to the first multimedia content. Therefore, the non-unique short comment in the plurality of recommended short comments needs to be removed, and the final recommended short comment of the first multimedia content is selected from the remaining recommended short comments.
[0062] The remaining recommended short comments include at least one recommended short comment, and a recommended short comment is randomly selected from the remaining recommended short comments as the final recommended short comment. Alternatively, a recommended short comment can be randomly selected from the remaining recommended short comments as the final recommended short comment displayed beside the name of the first multimedia content each time the name of the first multimedia content is displayed. Alternatively, a corresponding recommended short comment can be selected from the remaining recommended short comments as the final recommended short comment displayed beside the name of the first multimedia content each time the name of the first multimedia content is displayed for different display interfaces of the first multimedia content. For example, in a display list interface of the first multimedia content, the recommended short comment 1 is selected from the remaining recommended short comments as the final recommended short comment; in a recommendation interface of the first multimedia content, the recommended short comment 2 is selected from the remaining recommended short comments as the final recommended short comment; and in a search interface of the first multimedia content, the recommended short comment 3 is selected from the remaining recommended short comments as the final recommended short comment.
[0063] The technical scheme provided by the embodiments of the present application extracts a plurality of target comment information having recommendation significance for the first multimedia content from a plurality of comment information, avoids the final recommended short comment from being a short comment without recommendation significance such as a poor comment, and ensures the recommendation quality of the final recommended short comment. The recommendation short comment library is generated by the extraction and rewriting capabilities of the language model, which avoids the occurrence of incomplete sentences in the generated recommended short comment, reduces data noise, significantly improves the quality of the recommended short comment, and improves the generation efficiency of the recommended short comment. By removing the non-unique short comment in the plurality of recommended short comments, the proportion of general short comments in the recommended short comment is reduced, the sentence of generally praising the first multimedia content is avoided, and the quality of the final recommended short comment is further improved, thereby improving the recommendation effect of the final recommended short comment.
[0064] In some embodiments, step 320 includes at least one of sub-steps 321-323 (not shown in the figure).
[0065] In sub-step 321, at least one first comment information is extracted from the plurality of comment information according to the words contained in the comment information.
[0066] In some embodiments, step 321 comprises at least one of sub-steps 3211-3214 (not shown in the figure).
[0067] In sub-step 3211, a plurality of keywords are extracted from the plurality of comment information.
[0068] The keyword refers to a word with actual meaning in the comment information, and the keyword is extracted from the comment information by using a word segmentation tool. For example, the word segmentation tool can be jieba, hanlp, etc.
[0069] In some embodiments, the keywords can be extracted from the comment information according to the words in the corresponding field according to the type of the first multimedia content. For example, if the first multimedia content is a song, the keywords in the music field can be introduced to extract the keywords from the comment information, so as to improve the accuracy of keyword extraction.
[0070] In sub-step 3212, at least one first word is extracted from the plurality of keywords according to the part of speech of the keyword, and the first word includes the keyword with the adjective part of speech and the keyword with the noun part of speech.
[0071] The part of speech includes noun, adjective, verb, quantifier, pronoun, preposition, etc. At least one first word is extracted from the plurality of keywords according to the pre-set part of speech to be extracted.
[0072] Optionally, the pre-set part of speech to be extracted includes noun and adjective, and the first word includes the keyword with the adjective part of speech and the keyword with the noun part of speech. Optionally, the pre-set part of speech to be extracted includes only noun, and the first word includes the keyword with the noun part of speech.
[0073] By limiting the part of speech of the keyword, the extracted first word can be more consistent with the information description of the first multimedia content itself, so as to reduce the comment information without actual meaning and improve the quality of the first comment information.
[0074] In sub-step 3213, the occurrence frequency corresponding to the at least one first word is calculated, and the first word with the occurrence frequency greater than or equal to a first threshold value is determined as a second word. The occurrence frequency refers to the proportion of the number of occurrences of the first word in the plurality of comment information to the total number of occurrences of the plurality of keywords in the plurality of comment information.
[0075] The number of occurrences of the first vocabulary in the plurality of pieces of comment information is obtained, and the total number of occurrences of the plurality of keywords in the plurality of pieces of comment information is obtained, and the occurrence frequency of the first vocabulary = the number of occurrences of the first vocabulary in the plurality of pieces of comment information / the total number of occurrences of the plurality of keywords in the plurality of pieces of comment information.
[0076] A first vocabulary with a high occurrence frequency in the at least one first vocabulary is extracted as a second vocabulary, specifically, a first vocabulary with an occurrence frequency greater than or equal to a first threshold value is determined as a second vocabulary. The size of the first threshold value can be set by the technician according to the extraction requirement of the first comment information, and the present application is not limited.
[0077] Exemplarily, the extraction process of the second vocabulary can refer to FIG. 4, and (1) in FIG. 4 shows that there is at least one first vocabulary, and a first vocabulary with an occurrence frequency greater than or equal to a first threshold value is determined as a second vocabulary, obtaining (2) in FIG. 4, which shows that there is at least one second vocabulary. As shown in FIG. 4, “regret” and “good to listen to” are screened out because the occurrence frequency is less than the first threshold value, and are not determined as the second vocabulary.
[0078] In substep 3214, at least one piece of first comment information is obtained according to the number of likes of the comment information corresponding to the second vocabulary, and the first comment information is the comment information with a number of likes greater than or equal to a second threshold value.
[0079] The comment information corresponding to the at least one second vocabulary is obtained, and the number of likes of the comment information is obtained, which is used to indicate the approval degree of the user to the comment information. The comment information with a high number of likes is extracted, specifically, the comment information with a number of likes greater than or equal to a second threshold value is extracted, and at least one piece of first comment information is obtained. The size of the second threshold value can be set by the technician according to the extraction requirement of the first comment information, and the present application is not limited.
[0080] By extracting the comment information with a number of likes greater than or equal to the second threshold value, the obtained first comment information is the comment information with a high general approval degree, so that the final recommended short comment can be approved by the user, and the recommendation effect of the final recommended short comment can be improved.
[0081] The extraction process of the first comment information can refer to FIG. 5, first, a plurality of pieces of comment information is obtained, a plurality of keywords is extracted from the plurality of pieces of comment information, the keywords are tagged with parts of speech, at least one first vocabulary is extracted according to the parts of speech, then the occurrence frequency of the first vocabulary is calculated, the first vocabulary with an occurrence frequency greater than or equal to a first threshold value is determined as a second vocabulary, and finally, the comment information with a number of likes greater than or equal to a second threshold value is determined as a first comment information.
[0082] In substep 322, at least one piece of second comment information is extracted from the plurality of pieces of comment information according to the correlation degree of the comment information and the first multimedia content.
[0083] In some embodiments, step 322 comprises at least one of sub-steps 3221-3225 (not shown in the figure).
[0084] Sub-step 3221: scoring the relevance between each piece of comment information and the first multimedia content by using a first relevance model, to obtain a first relevance score corresponding to each piece of comment information.
[0085] The first relevance model is used to determine whether the comment information is relevant to the first multimedia content, and to obtain a first relevance score corresponding to the comment information according to the relevance between the comment information and the first multimedia content. For example, if the first multimedia content is a song, the first relevance model is used to determine whether the comment information is relevant to the song, and to obtain a first relevance score corresponding to the comment information according to the relevance between the comment information and the song.
[0086] The first relevance score ranges from 0 to 1, where a first relevance score of 0 indicates that the comment information is irrelevant to the first multimedia content, and a first relevance score of 1 indicates that the comment information is relevant to the first multimedia content. The first relevance score is used to indicate the relevance between the comment information and the first multimedia content.
[0087] Sub-step 3222: scoring the relevance between each piece of comment information and the information of the first multimedia content by using a second relevance model, to obtain a second relevance score corresponding to each piece of comment information.
[0088] In some embodiments, for each piece of comment information, a relevance score corresponding to each of the N sub-information is obtained according to the proportion of the text length of the N sub-information in the comment information; a text similarity between the comment information and the content information of the first multimedia content is calculated by using the second relevance model, to obtain a similarity score corresponding to the content information; and a second relevance score corresponding to the comment information is obtained according to the relevance scores corresponding to the N sub-information and the similarity score corresponding to the content information.
[0089] If the first sub-information is not included in the comment information, the relevance score corresponding to the first sub-information is 0.
[0090] For example, if the first multimedia content is a song, and the N sub-information includes the name of the singer, the name of the song, and the name of the album, a relevance score corresponding to the name of the singer is obtained according to the proportion of the text length of the name of the singer in the comment information, a relevance score corresponding to the name of the song is obtained according to the proportion of the text length of the name of the song in the comment information, and a relevance score corresponding to the name of the album is obtained according to the proportion of the text length of the name of the album in the comment information.
[0091] The second correlation model is used to calculate the text similarity between the comment information and the content information of the first multimedia content. For example, if the first multimedia content is a song, the content information of the song is the lyric information of the song, and the second correlation model is used to calculate the text similarity between the comment information and the lyric information. If the first multimedia content is a video, the content information of the video is the dialogue information of the video, and the second correlation model is used to calculate the text similarity between the comment information and the dialogue information.
[0092] According to the correlation scores respectively corresponding to the N pieces of sub-information and the similarity scores corresponding to the content information, and the weights respectively corresponding thereto, a second correlation score corresponding to the comment information is calculated. The correlation scores respectively corresponding to the N pieces of sub-information and the similarity scores corresponding to the content information, and the weights respectively corresponding thereto, are set by the technical personnel according to the extraction requirement of the second comment information, and the present application is not limited thereto. Generally, the weight of the similarity scores corresponding to the content information is greater than the weights of the correlation scores respectively corresponding to the N pieces of sub-information.
[0093] By comprehensively considering the information of the first multimedia content and the correlation between the first multimedia content and the comment information, the correlation between the comment content of the comment information and the first multimedia content is fully considered, so as to ensure the accuracy of the extraction of the second comment information, and help to improve the recommendation quality of the short comments.
[0094] In substep 3223, a sentiment analysis model is used to analyze the favorable degree of the plurality of pieces of comment information for the first multimedia content, respectively, to obtain a favorable degree score corresponding to each of the plurality of pieces of comment information.
[0095] The sentiment analysis model is used to analyze the favorable degree of the comment information for the first multimedia content, and the favorable degree score has a value range of 0 to 1, wherein a favorable degree score of 0 indicates that the comment information is negative for the first multimedia content, and a favorable degree score of 1 indicates that the comment information is positive for the first multimedia content.
[0096] In substep 3224, according to the first correlation score, the second correlation score, and the favorable degree score corresponding to each of the plurality of pieces of comment information, a comprehensive score corresponding to each of the plurality of pieces of comment information is obtained.
[0097] According to the weights respectively corresponding to the first correlation score, the second correlation score, and the favorable degree score corresponding to each of the plurality of pieces of comment information, a comprehensive score corresponding to each of the plurality of pieces of comment information is calculated. The comprehensive score is used to indicate the correlation between the comment information and the first multimedia content.
[0098] The weights respectively corresponding to the first correlation score, the second correlation score, and the favorable degree score are set by the technical personnel according to the extraction requirement of the second comment information, and the present application is not limited thereto.
[0099] Optionally, the weights corresponding to the first correlation degree score, the second correlation degree score and the praise degree score can be the same, all being 1 / 3. Optionally, the weights corresponding to the first correlation degree score, the second correlation degree score and the praise degree score can also be different, for example, the weight corresponding to the first correlation degree score and the weight corresponding to the praise degree score are greater than the weight corresponding to the second correlation degree score.
[0100] Sub-step 3225, the comment information with the comprehensive score greater than or equal to the third threshold value is determined as the second comment information.
[0101] The size of the third threshold value can be set by the technician according to the extraction requirement of the second comment information, and the present application is not limited.
[0102] The second comment information is determined through the above steps, so that the obtained second comment information is the comment information related to the entity characteristics of the first multimedia content itself and with positive emotional analysis, rather than the comment information praising the first multimedia content in general, thereby improving the quality of the second comment information and the quality of the final recommended short comment.
[0103] The extraction process of the second comment information can refer to FIG. 6. The first correlation degree score corresponding to the comment information is obtained based on the first correlation model, the second correlation degree score corresponding to the comment information is obtained based on the second correlation model, the second correlation degree score is obtained based on the correlation degree scores corresponding to the N pieces of sub-information and the similarity scores corresponding to the content information, and the praise degree score corresponding to the comment information is obtained based on the emotional analysis model. Based on the first correlation degree score, the second correlation degree score and the praise degree score corresponding to the comment information, the comprehensive score of the comment information is obtained, and the comment information with the comprehensive score greater than or equal to the third threshold value is determined as the second comment information.
[0104] In some embodiments, after at least one piece of first comment information is extracted from the plurality of comment information according to the words contained in the comment information, at least one piece of second comment information is extracted from the remaining comment information according to the correlation degree of the comment information and the first multimedia content.
[0105] In some embodiments, after at least one piece of second comment information is extracted from the plurality of comment information according to the correlation degree of the comment information and the first multimedia content, at least one piece of first comment information is extracted from the remaining comment information according to the words contained in the comment information.
[0106] Sub-step 323, screening the at least one piece of first comment information and the at least one piece of second comment information to obtain a plurality of target comment information.
[0107] The at least one first comment information and the at least one second comment information are integrated, and the integrated comment information is filtered to obtain a plurality of target comment information.
[0108] The at least one first comment information with high occurrence frequency and user recognition degree and the at least one second comment information with high correlation degree with the first multimedia content are extracted from the plurality of comment information, so that the obtained target comment information sufficiently includes comment information with high quality in different aspects, and a recommended short comment with high quality can be extracted therefrom, thereby improving the quality of the final recommended short comment.
[0109] In some embodiments, the step 323 includes at least one of sub-steps 3231-3233 (not shown in the figure).
[0110] In the sub-step 3231, the comment information satisfying the elimination condition is eliminated from the at least one first comment information and the at least one second comment information to obtain at least one remaining comment information.
[0111] In some embodiments, the elimination condition includes at least one of the following: the comment information contains a pre-set conflict word, the comment information is a multi-round dialogue sentence, the text length of the comment information is less than or equal to a fifth threshold value, and the text length of the comment information is greater than or equal to a sixth threshold value.
[0112] The comment information satisfying the elimination condition can refer to the following Table 2.
[0113] Table 2
[0114] The conflict word is a word irrelevant to the first multimedia content, and is usually a word used by a user to cheat likes, including but not limited to celebrity, famous school, enterprise, disease, and region-related words. The comment information containing the conflict word needs to be filtered, and therefore, the comment information of “like XX, comment XX”, “XXX, I miss you”, and “next year can be admitted to XXX” in Table 2 needs to be eliminated from the at least one first comment information and the at least one second comment information.
[0115] The multi-round dialogue sentence comment information is usually comment information irrelevant to the first multimedia content, and contains at least one round of dialogue between at least two objects. The multi-round dialogue sentence comment information needs to be filtered, and therefore, the comment information of “XXX: XXX, if you can guess how many candies I have in my hand, I will give you the two candies in my hand, how about it? XXX: OK, I guess 5.” in Table 2 needs to be eliminated from the at least one first comment information and the at least one second comment information.
[0116] Comments with a text length less than or equal to the fifth threshold are considered too short and have low content value. Therefore, these comments need to be filtered. The fifth threshold can be set by technical personnel based on the extraction requirements of the target comment information; this application does not impose any limitations. For example, the fifth threshold could be 3. Therefore, comments such as "passed the exam," "accept or not," and "to…" in Table 2 above need to be removed from at least one first comment and at least one second comment.
[0117] Comments with a text length greater than or equal to the sixth threshold are considered excessively long and have low content value. Therefore, these comments need to be filtered. The size of the sixth threshold can be set by technical personnel based on the extraction requirements of the target comment information; this application does not impose any limitations. For example, the sixth threshold could be 20.
[0118] By removing comments based on the above criteria, we can avoid including comments that are not relevant to recommendations in the remaining comments, thus improving the quality of target comments and contributing to the quality of generated short recommendations.
[0119] Sub-step 3232: The first relevance model is used to score the relevance between at least one remaining comment and the first multimedia content, so as to obtain the first relevance score corresponding to at least one remaining comment.
[0120] Alternatively, the first relevance score corresponding to at least one remaining comment obtained in sub-step 3221 can be directly obtained.
[0121] Sub-step 3233: Remove at least one remaining comment from the list of comments whose first relevance score is less than or equal to the fourth threshold, and obtain multiple target comment information.
[0122] The size of the fourth threshold can be set by the technicians themselves according to the extraction requirements of the target comment information, and this application does not impose any restrictions.
[0123] In some embodiments, the first relevance score corresponding to at least one remaining comment is multiplied by -1 to obtain the water-filling information corresponding to at least one remaining comment. Comments with negative values of the first relevance score greater than or equal to the fourth threshold are removed from at least one remaining comment to obtain multiple target comment information.
[0124] By filtering the extracted at least one first comment information and at least one second comment information, low-quality comment information without recommendation significance is avoided in the target comment information, the quality of the corpus sent into the language model in the next round is improved, and the generation quality of the recommended short comments is improved.
[0125] In some embodiments, step 350 includes at least one of sub-steps 351-355 (not shown in the figure).
[0126] In sub-step 351, the number of associated multimedia contents corresponding to each of the plurality of recommended short comments is obtained. The associated multimedia content refers to the multimedia content in the database having a first correlation degree score greater than or equal to a seventh threshold value with the recommended short comment.
[0127] The database stores a plurality of multimedia contents. The first correlation model is used to score the correlation degrees between the plurality of recommended short comments and the plurality of multimedia contents in the database, respectively, to obtain the first correlation degree scores between the plurality of recommended short comments and the plurality of multimedia contents in the database, respectively.
[0128] For each of the plurality of recommended short comments, the first correlation degree scores between the recommended short comment and the plurality of multimedia contents in the database are obtained. The multimedia content having a first correlation degree score greater than or equal to the seventh threshold value is determined as the associated multimedia content corresponding to the recommended short comment. The number of associated multimedia contents is obtained. The size of the seventh threshold value can be set by the technician according to the screening requirements of the recommended short comment, which is not limited in the present application.
[0129] For example, if the first multimedia content is a song, the number of associated songs corresponding to each of the plurality of recommended short comments is obtained.
[0130] The number of associated multimedia contents of each recommended short comment is used to indicate the concentration of multimedia contents of the recommended short comment, and represents the number of multimedia contents having similar short comments.
[0131] In sub-step 352, at least one quantile point of the number of associated multimedia contents is obtained according to the number of associated multimedia contents corresponding to each of the plurality of recommended short comments. The quantile point is used to indicate the distribution proportion of the number of associated multimedia contents arranged in descending order.
[0132] The quantile point can be any quantile mode such as quartile, median, percentile, etc. The number of associated multimedia contents corresponding to each of the plurality of recommended short comments is arranged in descending order, and at least one quantile point of the number of associated multimedia contents is obtained according to the selected quantile mode.
[0133] The number of associated multimedia contents arranged in descending order can refer to Table 3 shown below.
[0134] Table 3
[0135] If the selected quantile mode is quartiles, the first quartile in the at least one quantile point indicates that 25% of the number of the associated multimedia contents corresponding to the respective recommended short comment is below it, the second quartile indicates that 50% of the number of the associated multimedia contents corresponding to the respective recommended short comment is below it, the third quartile indicates that 75% of the number of the associated multimedia contents corresponding to the respective recommended short comment is below it, and the fourth quartile indicates that 100% of the number of the associated multimedia contents corresponding to the respective recommended short comment is below it.
[0136] If the selected quantile mode is percentiles, the first percentile in the at least one quantile point indicates that 1% of the number of the associated multimedia contents corresponding to the respective recommended short comment is below it, the second percentile indicates that 2% of the number of the associated multimedia contents corresponding to the respective recommended short comment is below it, the third percentile indicates that 3% of the number of the associated multimedia contents corresponding to the respective recommended short comment is below it, and so on.
[0137] In substep 353, at least one recommended short comment corresponding to the first quantile point in the at least one quantile point is determined as a general short comment, and the first quantile point is used to indicate the number of the first proportion of the associated multimedia contents in the ranking.
[0138] The first quantile point is a quantile point determined according to the quantile mode of the quantile point, and is used to indicate the number of the first proportion of the associated multimedia contents in the ranking. Exemplarily, if the selected quantile mode is quartiles, the first quantile point can be one of the first quartile, the second quartile, the third quartile and the fourth quartile. For example, if the first quantile point is the third quartile, the third quantile point is used to indicate the number of 25% of the associated multimedia contents in the ranking, and the recommended short comment corresponding to the number of 25% of the associated multimedia contents in the ranking is determined as a general short comment.
[0139] In combination with Table 3 above, if the number of the recommended short comment is 8 and the first quantile point is the third quartile, “a song with a story” and “a hit song” are determined as general short comments.
[0140] The size of the first proportion can be set by the technician according to the screening requirement of the recommended short comment, and the present application is not limited.
[0141] In substep 354, the non-unique short comment in the plurality of recommended short comments is removed according to the general short comment, and the remaining recommended short comment is obtained.
[0142] In some embodiments, the semantic similarity between the general short comment and each of the plurality of recommended short comments is calculated; a recommended short comment with a semantic similarity greater than or equal to an eighth threshold value is determined as a non-unique short comment; and the non-unique short comments are removed from the plurality of recommended short comments to obtain the remaining recommended short comments.
[0143] The general short comment is used as a seed for identifying non-unique short comments, and a recommended short comment with a semantic similarity greater than or equal to the eighth threshold value is determined as a non-unique short comment, so that at least one non-unique short comment corresponding to each general short comment can be obtained. The non-unique short comments corresponding to each general short comment are removed from the plurality of recommended short comments to obtain the remaining recommended short comments.
[0144] By setting the eighth threshold value, a recommended short comment with a semantic similarity greater than or equal to the eighth threshold value is determined as a non-unique short comment, which avoids removing all recommended short comments as non-unique short comments due to a too high threshold value, and avoids failing to remove a recommended short comment with a high relevance to a general short comment due to a too low threshold value, thereby ensuring the uniqueness and pertinence of the remaining recommended short comments.
[0145] For example, the semantic similarity between the general short comment "too good to listen" and the plurality of recommended short comments can be referred to Table 4 shown below.
[0146] Table 4
[0147] The size of the eighth threshold value can be set by a technician according to the screening requirements of the recommended short comments, which is not limited in the present application.
[0148] For example, if the eighth threshold value is 0.985 for Table 4, the non-unique short comments corresponding to the general short comment "too good to listen" include "too good to listen", "too good to listen", "really too good to listen" and "too good to listen".
[0149] Fig. 7 shows a process of removing non-unique short comments from the plurality of recommended short comments. First, based on the plurality of recommended short comments contained in the short comment recommendation library, the number of associated multimedia contents corresponding to each of the plurality of recommended short comments is obtained, and the first quantile point in at least one quantile point of the number of associated multimedia contents is determined according to the number of associated multimedia contents corresponding to each of the plurality of recommended short comments. The general short comment in the plurality of recommended short comments is obtained according to the first quantile point, the non-unique short comment in the plurality of recommended short comments is obtained according to the semantic similarity between the general short comment and each of the plurality of recommended short comments, and the non-unique short comment in the plurality of recommended short comments is removed to obtain the remaining recommended short comments.
[0150] Sub-step 355, selecting the final recommended short comment of the first multimedia content from the remaining recommended short comments.
[0151] By identifying the common short comments in the recommended short comments and proposing the non-unique short comments in the recommended short comments, the proportion of the common short comments in the recommended short comments can be reduced, and the quality of the final recommended short comments can be improved.
[0152] FIG. 8 shows a schematic diagram of a short comment generation process. First, based on a plurality of comment information to be input, first comment information and second comment information are extracted from the plurality of comment information. Then, comment information satisfying a rejection condition and comment information with a low first correlation score are rejected from at least one first comment information and at least one second comment information, to obtain a plurality of target comment information. The plurality of target comment information is segmented to obtain a plurality of segmented short sentences. Based on the information of the first multimedia content and the plurality of segmented short sentences, input prompt information of a language model is constructed, and a short comment recommendation library containing a plurality of recommended short comments is generated. Finally, non-unique short comments in the plurality of recommended short comments are rejected, and the final recommended short comments of the first multimedia content are selected from the remaining recommended short comments.
[0153] The following is an apparatus embodiment of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the apparatus embodiments of the present application, please refer to the method embodiments of the present application.
[0154] Please refer to FIG. 9, which shows a block diagram of a short comment generation apparatus according to an embodiment of the present application. The apparatus has the functions of the above-mentioned short comment generation method, which can be realized by hardware or corresponding software executed by hardware. The apparatus can be the computer device introduced above or can be arranged in the computer device. As shown in FIG. 9, the apparatus 900 can include a comment acquisition module 910, a comment extraction module 920, a comment segmentation module 930, a short comment generation module 940, and a short comment selection module 950.
[0155] The comment acquisition module 910 is configured to acquire a plurality of comment information of a first multimedia content.
[0156] The comment extraction module 920 is configured to extract a plurality of target comment information from the plurality of comment information, wherein the target comment information refers to comment information having a recommendation meaning for the first multimedia content.
[0157] The comment segmentation module 930 is configured to segment the plurality of target comment information respectively to obtain a plurality of segmented short sentences.
[0158] The short comment generation module 940 is configured to generate, by a language model, a short comment recommendation library of the first multimedia content according to information of the first multimedia content and the plurality of segmented short sentences, wherein the short comment recommendation library contains a plurality of recommended short comments.
[0159] The short review selection module 950 is configured to remove non-unique short reviews from the plurality of recommended short reviews, and select final recommended short reviews of the first multimedia content from the remaining recommended short reviews, wherein the non-unique short reviews refer to recommended short reviews having relevance to a plurality of other multimedia contents except the first multimedia content.
[0160] In some embodiments, the comment extraction module 920 includes:
[0161] The first extraction unit is configured to extract at least one first comment information from the plurality of comment information according to a word included in the comment information.
[0162] The second extraction unit is configured to extract at least one second comment information from the plurality of comment information according to a relevance of the comment information to the first multimedia content.
[0163] The comment screening unit is configured to screen the at least one first comment information and the at least one second comment information to obtain the plurality of target comment information.
[0164] In some embodiments, the first extraction unit is configured to:
[0165] extract a plurality of keywords from the plurality of comment information;
[0166] extract at least one first word from the plurality of keywords according to a part of speech of the keywords, wherein the first word includes a keyword with a part of speech of an adjective and a keyword with a part of speech of a noun;
[0167] calculate an occurrence frequency corresponding to the at least one first word respectively, and determine a first word with an occurrence frequency greater than or equal to a first threshold value as a second word, wherein the occurrence frequency refers to a proportion of a number of occurrences of the first word in the plurality of comment information to a total number of occurrences of the plurality of keywords in the plurality of comment information;
[0168] obtain the at least one first comment information according to a number of likes of comment information corresponding to the second word, wherein the first comment information is comment information with a number of likes greater than or equal to a second threshold value.
[0169] In some embodiments, the second extraction unit is configured to:
[0170] score a relevance between the plurality of comment information and the first multimedia content respectively by using a first relevance model to obtain a first relevance score corresponding to the plurality of comment information respectively;
[0171] The second correlation model is used to score the relevance between the plurality of pieces of comment information and the information of the first multimedia content respectively, to obtain second relevance scores respectively corresponding to the plurality of pieces of comment information;
[0172] The sentiment analysis model is used to analyze the praise degree of the plurality of pieces of comment information for the first multimedia content respectively, to obtain praise degree scores respectively corresponding to the plurality of pieces of comment information;
[0173] According to the first relevance scores, the second relevance scores and the praise degree scores respectively corresponding to the plurality of pieces of comment information, comprehensive scores respectively corresponding to the plurality of pieces of comment information are obtained.
[0174] The comment information with a comprehensive score greater than or equal to a third threshold value is determined as the second comment information.
[0175] In some embodiments, the information of the first multimedia content includes N pieces of sub-information, each of the sub-information being used to indicate a configuration situation of one dimension of the first multimedia content, N being a positive integer; and the second extraction unit is configured to:
[0176] For each piece of comment information in the plurality of pieces of comment information, a relevance score respectively corresponding to each of the N pieces of sub-information is obtained according to a text length proportion of the N pieces of sub-information in the comment information;
[0177] The second correlation model is used to calculate a text similarity between the comment information and content information of the first multimedia content, to obtain a similarity score corresponding to the content information;
[0178] According to the relevance scores respectively corresponding to the N pieces of sub-information and the similarity score corresponding to the content information, a second relevance score corresponding to the comment information is obtained.
[0179] In some embodiments, the comment screening unit is configured to:
[0180] Comment information meeting a rejection condition is rejected from the at least one piece of first comment information and the at least one piece of second comment information, to obtain at least one piece of remaining comment information;
[0181] The first correlation model is used to score the relevance between the at least one piece of remaining comment information and the first multimedia content respectively, to obtain first relevance scores respectively corresponding to the at least one piece of remaining comment information;
[0182] Comment information with a first relevance score less than or equal to a fourth threshold value in the at least one piece of remaining comment information is rejected, to obtain a plurality of pieces of target comment information.
[0183] In some embodiments, the rejection condition comprises at least one of:
[0184] The comment information contains a preset conflict vocabulary;
[0185] The comment information is a multi-round dialogue sentence;
[0186] The text length of the comment information is less than or equal to a fifth threshold value;
[0187] The text length of the comment information is greater than or equal to a sixth threshold value.
[0188] In some embodiments, the short comment selection module 950 is configured to:
[0189] Obtain the number of associated multimedia contents corresponding to the plurality of recommended short comments respectively, the associated multimedia contents being multimedia contents in the database having a first association degree score greater than or equal to a seventh threshold value with the recommended short comments;
[0190] According to the number of associated multimedia contents corresponding to the plurality of recommended short comments respectively, obtain at least one quantile point of the number of associated multimedia contents, the quantile point being used to indicate the distribution proportion of the number of associated multimedia contents arranged in descending order;
[0191] Determine at least one recommended short comment corresponding to a first quantile point in the at least one quantile point as a general short comment, the first quantile point being used to indicate the number of associated multimedia contents of a first proportion arranged in a high order;
[0192] According to the general short comment, reject the non-unique short comment in the plurality of recommended short comments to obtain the remaining recommended short comments;
[0193] Select the final recommended short comment of the first multimedia content from the remaining recommended short comments.
[0194] In some embodiments, the short comment selection module 950 is configured to:
[0195] Calculate the semantic similarity between the general short comment and the plurality of recommended short comments respectively;
[0196] Determine a recommended short comment having a semantic similarity greater than or equal to an eighth threshold value as the non-unique short comment;
[0197] Reject the non-unique short comment in the plurality of recommended short comments to obtain the remaining recommended short comments.
[0198] It should be noted that the apparatus provided by the above embodiments, in realizing its functions, is only exemplified by the above division of functional modules, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the content structure of the device is divided into different functional modules to complete all or part of the above-described functions. In addition, the apparatus and method embodiments provided by the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0199] Please refer to FIG. 10, which shows a structural block diagram of a computer device 1000 provided by an embodiment of the present application. The computer device 1000 can be any electronic device with data computing, processing and storage functions. The computer device 1000 can be used to implement the short comment generation method provided in the above embodiments.
[0200] Generally, the computer device 1000 includes a processor 1001 and a memory 1002.
[0201] The processor 1001 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1001 can be implemented in at least one of the hardware forms of a DSP (Digital Signal Processing), a FPGA (Field Programmable Gate Array), and a PLA (Programmable Logic Array). The processor 1001 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1001 can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 can also include an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.
[0202] The memory 1002 can include one or more computer-readable storage media. The computer-readable storage media can be non-transitory. The memory 1002 can also include high-speed random access memory and can include nonvolatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other nonvolatile solid-state storage devices. In some embodiments, the non-transitory computer-readable storage medium of the memory 1002 is used for storing the computer programs configured to be executed by one or more processors to implement the review generation method described above.
[0203] Those skilled in the art can understand that the structure shown in FIG. 10 does not constitute a limitation on the computer device 1000, and can include more or fewer components than shown, or combine certain components, or adopt different component arrangements.
[0204] In an illustrative embodiment, a computer-readable storage medium is also provided, and the storage medium stores a computer program, which, when executed by a processor of a computer device, implements the review generation method described above. Optionally, the computer-readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0205] In an exemplary embodiment, a computer program product is also provided, and the computer program product includes a computer program stored in a computer-readable storage medium. The processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the review generation method described above.
[0206] It should be understood that "multiple" referred to herein refers to two or more. The "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the associated objects before and after it are in an "or" relationship. In addition, the step numbers described herein only exemplarily show a possible execution order between steps. In some other embodiments, the above steps can also be executed in a different order from the numbering order, such as simultaneously executing two steps with different numbers, or executing two steps with different numbers in an order opposite to the illustration, and the embodiments of the present application do not limit this.
[0207] The above merely provides exemplary embodiments of the present application, but is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A short review generation method, the method being executed by a computer device, the method comprising: obtaining a plurality of comment information of a first multimedia content; extracting a plurality of target comment information from the plurality of comment information, the target comment information being comment information having a recommendation significance for the first multimedia content; respectively segmenting the plurality of target comment information to obtain a plurality of segmented short sentences; generating a short review recommendation library of the first multimedia content according to information of the first multimedia content and the plurality of segmented short sentences through a language model, the short review recommendation library containing a plurality of recommendation short reviews; eliminating non-unique short reviews in the plurality of recommendation short reviews, and selecting a final recommendation short review of the first multimedia content from the remaining recommendation short reviews, the non-unique short review being a recommendation short review having a correlation with a plurality of other multimedia contents except the first multimedia content.
2. The method of claim 1, wherein, The extracting a plurality of target comment information from the plurality of comment information comprises: extracting at least one first comment information from the plurality of comment information according to a vocabulary contained in the comment information; extracting at least one second comment information from the plurality of comment information according to a correlation degree between the comment information and the first multimedia content; screening the at least one first comment information and the at least one second comment information to obtain the plurality of target comment information.
3. The method of claim 2, wherein, The extracting at least one first comment information from the plurality of comment information according to a vocabulary contained in the comment information comprises: extracting a plurality of keywords from the plurality of comment information; extracting at least one first vocabulary from the plurality of keywords according to a part of speech of the keywords, the first vocabulary including a keyword with a part of speech of an adjective and a keyword with a part of speech of a noun; calculating an occurrence frequency corresponding to the at least one first vocabulary respectively, determining a first vocabulary with an occurrence frequency greater than or equal to a first threshold value as a second vocabulary, the occurrence frequency being a proportion of a number of occurrences of the first vocabulary in the plurality of comment information to a total number of occurrences of the plurality of keywords in the plurality of comment information; obtaining the at least one first comment information according to a like amount of comment information corresponding to the second vocabulary, the first comment information being comment information with a like amount greater than or equal to a second threshold value.
4. The method of claim 2 or 3, wherein, The extracting at least one second comment information from the plurality of comment information according to a correlation degree between the comment information and the first multimedia content comprises: respectively scoring a correlation degree between the plurality of comment information and the first multimedia content by using a first correlation model to obtain a first correlation degree score corresponding to the plurality of comment information respectively; respectively scoring a correlation degree between the plurality of comment information and information of the first multimedia content by using a second correlation model to obtain a second correlation degree score corresponding to the plurality of comment information respectively; respectively analyzing a good comment degree of the plurality of comment information for the first multimedia content by using a sentiment analysis model to obtain a good comment degree score corresponding to the plurality of comment information respectively; According to the first correlation degree score, the second correlation degree score and the good comment degree score corresponding to each of the plurality of comment information, a comprehensive score corresponding to each of the plurality of comment information is obtained; The comment information with the comprehensive score greater than or equal to the third threshold value is determined as the second comment information.
5. The method of claim 4, wherein, The information of the first multimedia content includes N sub-information, each of the sub-information is used to indicate the configuration of one dimension of the first multimedia content, and N is a positive integer; The second correlation model is used to score the correlation degree between the plurality of comment information and the information of the first multimedia content respectively, and a second correlation degree score corresponding to each of the plurality of comment information is obtained, including: For each of the plurality of comment information, according to the proportion of the length of the text in the N sub-information in the comment information, a correlation degree score corresponding to each of the N sub-information is obtained; The second correlation model is used to calculate the text similarity between the comment information and the content information of the first multimedia content, and a similarity score corresponding to the content information is obtained; According to the correlation degree score corresponding to each of the N sub-information and the similarity score corresponding to the content information, a second correlation degree score corresponding to the comment information is obtained.
6. The method according to any one of claims 2 to 5, wherein, The filtering of the at least one first comment information and the at least one second comment information to obtain the plurality of target comment information includes: The comment information meeting the elimination condition is eliminated from the at least one first comment information and the at least one second comment information to obtain at least one remaining comment information; A first correlation model is used to score the correlation degree between the at least one remaining comment information and the first multimedia content respectively, and a first correlation degree score corresponding to each of the at least one remaining comment information is obtained; The comment information with the first correlation degree score less than or equal to the fourth threshold value in the at least one remaining comment information is eliminated to obtain the plurality of target comment information.
7. The method of claim 6, wherein, The elimination condition includes at least one of the following: The comment information contains a pre-set conflict word; The comment information is a multi-round dialogue sentence; The length of the text of the comment information is less than or equal to a fifth threshold value; The length of the text of the comment information is greater than or equal to a sixth threshold value.
8. The method according to any one of claims 1 to 7, wherein, The elimination of the non-unique short comment in the plurality of recommended short comments, and the selection of the final recommended short comment of the first multimedia content from the remaining recommended short comments, includes: The number of associated multimedia contents corresponding to each of the plurality of recommended short comments is obtained, the associated multimedia content refers to the multimedia content in the database with a first correlation degree score greater than or equal to a seventh threshold value with the recommended short comment; At least one quantile point of the number of associated multimedia contents is obtained according to the number of associated multimedia contents corresponding to each of the plurality of recommended short comments, the quantile point is used to indicate the distribution proportion of the number of associated multimedia contents arranged in descending order; At least one recommended short comment corresponding to a first quantile point in the at least one quantile point is determined as a general short comment, the first quantile point is used to indicate the number of associated multimedia contents of a first proportion arranged in a high order. According to the general short comment, the non-unique short comment in the plurality of recommended short comments is removed to obtain the remaining recommended short comments; The final recommended short comment of the first multimedia content is selected from the remaining recommended short comments.
9. The method of claim 8, wherein, The removing the non-unique short comment in the plurality of recommended short comments according to the general short comment to obtain the remaining recommended short comments comprises: The semantic similarity between the general short comment and the plurality of recommended short comments is calculated respectively; The recommended short comment with the semantic similarity greater than or equal to an eighth threshold value is determined as the non-unique short comment; The non-unique short comment in the plurality of recommended short comments is removed to obtain the remaining recommended short comments.
10. A short comment generation device, the device comprising: a comment acquisition module configured to acquire a plurality of comment information of a first multimedia content; a comment extraction module configured to extract a plurality of target comment information from the plurality of comment information, the target comment information being comment information having a recommended meaning for the first multimedia content; a comment segmentation module configured to segment the plurality of target comment information respectively to obtain a plurality of segmented short sentences; a short comment generation module configured to generate a short comment recommendation library of the first multimedia content according to information of the first multimedia content and the plurality of segmented short sentences through a language model, the short comment recommendation library comprising a plurality of recommended short comments; a short comment selection module configured to remove a non-unique short comment in the plurality of recommended short comments, and select a final recommended short comment of the first multimedia content from the remaining recommended short comments, the non-unique short comment being a recommended short comment having a correlation with a plurality of other multimedia contents except the first multimedia content.
11. A computer device, the computer device comprising a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the short comment generation method according to any one of claims 1 to 9.
12. A computer readable storage medium, the computer readable storage medium storing a computer program, the computer program being loaded and executed by a processor to implement the short comment generation method according to any one of claims 1 to 9.
13. A computer program product, the computer program product comprising a computer program, the computer program being loaded and executed by a processor to implement the short comment generation method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Long and short film review fine-grained viewpoint mining method based on non-supervision
CN113641788A
News comment filtering method and system based on key sentence group and original news text
CN114443829A
Comment processing method, comment selection display method, electronic equipment and storage medium
CN116414977A
Short comment generation method and device, equipment and storage medium
CN119396994A
System and Method for filtering related review using key phrase extraction
KR102520248B1