Multimedia playing method and device, electronic equipment and computer readable storage medium

By generating and displaying comment multimedia within the multimedia playback interface, the problem of users' lack of patience for text comments is solved, user stickiness is improved, and the correlation between multimedia content is enhanced.

CN120653791APending Publication Date: 2025-09-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410305638.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Users have little patience for text comments and are less willing to read them, resulting in reduced user stickiness.

Method used

By displaying the comment multimedia control of the original multimedia in the multimedia playback interface and generating multimedia content based on the comment, the comment multimedia is generated by combining the multimedia materials to provide comment information in multimedia form.

Benefits of technology

It improves users' stickiness to comment information, enables users to understand comment information more vividly and intuitively, and enhances users' association with multimedia content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653791A_ABST
    Figure CN120653791A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a multimedia playing method and device, electronic equipment and a computer readable storage medium, and the method comprises the steps: displaying a playing interface of original multimedia, and the playing interface also comprises a multimedia control of comment multimedia corresponding to the original multimedia; comment multimedia corresponding to the multimedia control is played on a playing interface; the playing interface further comprises a comment area of the original multimedia, and the comment area of the original multimedia comprises comments; multimedia content of the comment multimedia is generated based on the comment. In the embodiment of the invention, the comment multimedia generated by the comments in the comment area of the original multimedia is played on the same playing interface, so that a user watching the original multimedia can know the comment information in the original multimedia in a multimedia form, the problems that the user is poor in tolerance to the text comment information and low in reading willingness are avoided, and the user experience is improved. And thus, the user stickiness is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a multimedia playback method, device, electronic device, and computer-readable storage medium. Background Art

[0002] In the prior art, when users browse information to obtain information, they often also check the comment area corresponding to the information to learn about netizens' views and opinions on the above information.

[0003] However, comments in comment sections are typically in text format. With the accelerating pace of life, many users have less patience for text-based information and are less willing to spend time reading it. Consequently, text-based comments have become a barrier to understanding opinions for some users, reducing user engagement. Summary of the Invention

[0004] The embodiments of the present application provide a multimedia playback method, device, electronic device, and computer-readable storage medium, which can improve the problem of reduced user stickiness in the prior art.

[0005] An embodiment of the present application provides a multimedia playback method, which includes: displaying a playback interface of original multimedia, the playback interface also including a multimedia control for comment multimedia corresponding to the original multimedia; playing the comment multimedia corresponding to the multimedia control on the playback interface; the playback interface also including a comment area for the original multimedia, the comment area for the original multimedia including comments; and generating multimedia content of the comment multimedia based on the comments.

[0006] An embodiment of the present application provides a multimedia playback device, the device comprising:

[0007] An interface display unit, configured to display a playback interface of the original multimedia, the playback interface also including multimedia controls for the commentary multimedia corresponding to the original multimedia;

[0008] A comment playing unit, configured to play the comment multimedia corresponding to the multimedia control on the playing interface;

[0009] The playback interface further includes a comment area for the original multimedia, and the comment area for the original multimedia includes comments; and multimedia content of the comment multimedia is generated based on the comments.

[0010] In one embodiment, the comments include a target comment and a lower-level comment of the target comment; the apparatus further includes:

[0011] The first generating unit is configured to generate the comment multimedia based on the target comment. In one embodiment, the comment includes a target comment, and the target comment does not have a next-level comment; the apparatus further includes:

[0012] The first generating unit is configured to generate the comment multimedia based on the target comment. In one embodiment, the comment includes the target comment and the next level comment of the target comment; the apparatus further includes:

[0013] The second generating unit is configured to generate the comment multimedia based on the target comment and the next level comment of the target comment.

[0014] In one embodiment, the first generating unit includes:

[0015] a target determination subunit, configured to, based on interaction data of other users on multiple comments in the comment area, select comments whose interaction data exceeds a threshold as target comments, wherein the target comments are target texts; the other users are: for each of the comments, users other than the user who posted the comment;

[0016] A multimedia material subunit, configured to obtain multimedia material corresponding to the target text;

[0017] A first generating subunit is configured to generate the comment multimedia based on the multimedia material when the multimedia material is a piece;

[0018] The second generating subunit is configured to, when there are at least two multimedia materials, splice the multimedia materials in the text order of the target text to generate the comment multimedia.

[0019] In one embodiment, the second generating unit includes:

[0020] a target comment subunit, configured to select comments whose interaction data exceeds a threshold as target comments based on interaction data of other users on multiple comments in the comment area; the other users are: for each comment, users other than the user who posted the comment;

[0021] A correlation screening subunit, configured to screen out the primary screening next-level comments that are strongly correlated with the target comment from the next-level comments of the target comment;

[0022] A sorting subunit, configured to sort the target comments and the pre-screened next-level comments to obtain a target text;

[0023] A multimedia material subunit, configured to obtain multimedia material corresponding to the target text;

[0024] The material splicing subunit is used to splice the multimedia materials according to the text order of the target text to generate the comment multimedia.

[0025] In one embodiment, the sorting subunit includes:

[0026] A sorting text sub-subunit is used to sort the target review and the pre-screened next-level reviews to obtain sorted text;

[0027] an insertion point determination sub-subunit, configured to determine at least one insertion point based on text content of the sorted text;

[0028] The text insertion sub-subunit is used to insert the corresponding preset connecting text at each insertion point to obtain the target text.

[0029] In one embodiment, the second generating unit includes:

[0030] a target comment subunit, configured to select comments whose interaction data exceeds a threshold as target comments based on interaction data of other users on multiple comments in the comment area; the other users are: for each comment, users other than the user who posted the comment;

[0031] A correlation screening subunit, configured to screen out the primary screening next-level comments that are strongly correlated with the target comment from the next-level comments of the target comment;

[0032] A sorting subunit, configured to sort the target review and the pre-screened next-level reviews, thereby obtaining sorted text;

[0033] A first target determination subunit is configured to determine that when no next-level comment exists in the initial screening of next-level comments, the sorted text is the target text;

[0034] The second target determination subunit is used for, when there is a next-level comment in the preliminary screening of the next-level comments, taking the preliminary screening of the next-level comments with the next-level comment as a new target comment, and jumping to the step: for the next-level comments of the target comment, screening out the preliminary screening of the next-level comments that are strongly correlated with the target comment, until there is no next-level comment in the preliminary screening of the next-level comments or the number of jumps reaches a set jump threshold, and recording the sorted text obtained when the jump stops as the target text;

[0035] A multimedia material subunit, configured to obtain multimedia material corresponding to the target text;

[0036] The comment multimedia generating subunit is used to splice the multimedia materials according to the text order of the target text to generate the comment multimedia.

[0037] In one embodiment, the target comment corresponds to a plurality of pre-screened next-level comments; the sorting subunit includes:

[0038] a coefficient calculation subunit, configured to calculate, for each of the pre-screened next-level comments, a logical relationship coefficient between the pre-screened next-level comment and the target comment, thereby obtaining a plurality of logical relationship coefficients;

[0039] The numerical sorting subunit is used to sort the target comment and the multiple initially screened next-level comments corresponding to the target comment based on the numerical values ​​of the multiple logical relationship coefficients.

[0040] In one embodiment, a target comment corresponds to multiple lower-level comments; the associated screening subunit includes:

[0041] The deduplication sub-sub-unit is used for, when at least two of the multiple next-level comments have text of the same preset length, to perform deduplication processing on the at least two next-level comments to obtain deduplication result comments corresponding to the at least two next-level comments; wherein the preset length text is text whose length exceeds the set length; the deduplication result comments and the next-level comments that do not need deduplication processing are recorded as secondary comments;

[0042] a weak correlation sub-subunit, configured to determine, for each of the secondary selected comments, if a preset symbol exists in the secondary selected comment or the format of the secondary selected comment is a preset format, whether the secondary selected comment is weakly correlated with the target comment;

[0043] The strong correlation sub-sub-unit is used to eliminate the secondary selected comments that are weakly correlated with the target comment from the plurality of secondary selected comments, thereby obtaining the pre-screened next-level comments that are strongly correlated with the target comment. In one embodiment, the strong correlation sub-sub-unit includes:

[0044] A weak correlation elimination unit is used to eliminate the secondary comments that are weakly correlated with the target comment from the plurality of secondary comments, thereby obtaining a result after elimination;

[0045] a text re-association unit, configured to calculate a text relevance between each of the secondary selected comments in the eliminated results and the target comment;

[0046] The threshold screening sub-unit is used to obtain secondary selected comments whose text relevance exceeds the relevance threshold from the eliminated results. The secondary selected comments whose text relevance exceeds the relevance threshold are the primary selected next-level comments that have a strong relevance to the target comment. In one embodiment, the deduplication sub-unit includes:

[0047] A same text again unit is used to obtain at least two next-level comments containing the same text of the preset length from the at least two next-level comments;

[0048] A unit for merging again is used to randomly select a next-level comment from the at least two next-level comments that have the same text of the preset length, and the randomly selected next-level comment is the comment to be merged;

[0049] The text adding unit is used to extract the other texts except the text of the preset length from at least two lower-level comments that have the same text of the preset length, and add the other texts to the comment to be merged.

[0050] In one embodiment, the multimedia material sub-unit includes:

[0051] A first material determination sub-sub-unit is configured to determine, for each sub-text of the target text, whether there is a first type of multimedia material in the multimedia material library whose similarity to the sub-text exceeds a first preset similarity value;

[0052] The first determination sub-subunit is configured to, if there is a first type of multimedia material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value, use the first type of multimedia material with similarity exceeding the first preset similarity value as the multimedia material corresponding to the subtext.

[0053] In one embodiment, the multimedia material sub-unit further includes:

[0054] a second determination sub-subunit, configured to generate voice data corresponding to the subtext if there is no multimedia material of the first category in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value;

[0055] a second material sub-sub-unit, configured to search the multimedia material library for a second type of multimedia material having a similarity with the subtext exceeding a preset similarity value, wherein the type of the second type of multimedia material is different from the type of the first type of multimedia material;

[0056] The material merging sub-sub-unit is configured to combine the found second-category multimedia material with the voice data to obtain the multimedia material corresponding to the subtext.

[0057] In one embodiment, determining whether there is a first category of multimedia material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value is implemented based on a label of the first category of multimedia material; the first category of multimedia material is a video material; and the apparatus further includes:

[0058] A sound recognition unit, configured to recognize the sound in the video material and generate a sound text corresponding to the sound;

[0059] A screenshot acquisition unit, configured to acquire multiple screenshots of the video material;

[0060] a subtitle text acquisition unit, configured to perform optical character recognition processing on each of the screenshots to acquire the subtitle text in the screenshot;

[0061] The material storage unit is configured to store the audio text, subtitle text, and the plurality of screenshots as tags corresponding to the video material, together with the video material, in the multimedia material library. In one embodiment, the comment playback unit is configured to play the comment multimedia corresponding to the multimedia control on the playback interface when the multimedia control is triggered.

[0062] In the multimedia playback method provided in the embodiment of the present application, the playback interface of the original multimedia can be displayed, and the playback interface also includes multimedia controls for comment multimedia corresponding to the original multimedia; the comment multimedia corresponding to the multimedia controls is played on the playback interface; wherein, the playback interface also includes a comment area for the original multimedia, and the comment area for the original multimedia includes comments; the multimedia content of the comment multimedia can be generated based on the comments.

[0063] In an embodiment of the present application, the content of the comment multimedia is generated based on the comments in the comment area of ​​the original multimedia. By setting the multimedia controls of the original multimedia and the comment multimedia in the same playback interface, the association between the original multimedia and the comment multimedia can be strengthened; and the comment multimedia generated by the comments in the comment area is played on this playback interface, so that users who watch the original multimedia can understand the comment information in the original multimedia in the form of multimedia, avoiding the problem that users have poor patience with text-based comment information and low willingness to read, which is conducive to improving user stickiness. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0065] Figure 1a This is a schematic diagram of an application scenario of the multimedia playback method provided by this application;

[0066] Figure 1b (1) shows a state where the comment multimedia is paused in one embodiment;

[0067] Figure 1b(2) in the figure shows a status of commenting on multimedia playback in one embodiment;

[0068] Figure 1c (1) shows an implementation in which the comment multimedia is suspended in the form of a floating window;

[0069] Figure 1c (2) shows an implementation method in which multimedia information is suspended in the form of a floating window;

[0070] Figure 1d This is a flowchart of a multimedia playback method provided by an embodiment of the present application;

[0071] Figure 2 This is a flowchart of a multimedia playback method provided in a specific embodiment of the present application;

[0072] Figure 3 This is a structural diagram of a multimedia playback device provided by an embodiment of the present application;

[0073] Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0074] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0075] Embodiments of the present application provide a multimedia playback method, device, electronic device, and computer-readable storage medium.

[0076] The multimedia playback device can be integrated into an electronic device, which can be a terminal, a server, or other device. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer (PC). The server can be a single server or a server cluster consisting of multiple servers.

[0077] In some embodiments, the multimedia playback device may also be integrated into multiple electronic devices. For example, the multimedia playback device may be integrated into multiple servers, and the multimedia playback method of the present application may be implemented by the multiple servers.

[0078] In some embodiments, the server may also be implemented in the form of a terminal.

[0079] For more details, please see Figure 1a The method provided by the embodiment of the present application may include: displaying a playback interface of the original multimedia, the playback interface also including a multimedia control for comment multimedia corresponding to the original multimedia; playing the comment multimedia corresponding to the multimedia control on the playback interface; the playback interface also including a comment area for the original multimedia, the comment area for the original multimedia including comments; the multimedia content of the comment multimedia is generated based on the comments.

[0080] In the above method, by placing multimedia controls for both the original multimedia and the commentary multimedia within the same playback interface, the relationship between the original multimedia and the commentary multimedia can be strengthened. Furthermore, by playing the commentary multimedia generated by the comments in the comment section within the same playback interface, users viewing the original multimedia can understand the commentary information within the original multimedia in a multimedia format, thus avoiding the problem of users having low patience and willingness to read text-based commentary information, thereby improving user stickiness.

[0081] The embodiment of the present application is mainly applied in the comment area of ​​multimedia information; specifically, it can be applied in the comment area of ​​short video applications. The multimedia information corresponds to the original multimedia mentioned above. Figure 1b ,exist Figure 1b The comment area shown is divided into text comments and video comments according to the type of comments. Figure 1b (1) in the video is in the paused state. Figure 1b (2) in the video is the status of video comment playback. Figure 1b The video corresponding to the video comment shown corresponds to the above-mentioned comment multimedia. Figure 1b An embodiment of the present invention is shown in which the original multimedia and the multimedia controls are located in the same playback interface. Specifically, in this embodiment, the multimedia controls are Figure 1b (1) and (2) both show virtual buttons: video comments.

[0082] Comments on multimedia Figure 1b The displayed method appears outside the comment area, and can also be suspended on the edge of the multimedia information in the form of a floating window. Figure 1c (1) in Figure 1c (1) shows an implementation method in which the comment multimedia is suspended in the form of a floating window. When the user clicks on the floating window, the video content in the floating window can be displayed in the display area of ​​the original multimedia information, and the multimedia information is suspended in the form of a floating window at the edge of the original location of the comment multimedia. For details, please see Figure 1c (2) in . Figure 1c Another embodiment of the original multimedia and multimedia controls being located in the same playback interface is shown. Specifically, in this embodiment, the multimedia controls are Figure 1cThe floating window corresponding to the comment multimedia shown in (1) is shown.

[0083] The above two video comment display methods can help users intuitively and efficiently understand the views and opinions in the comment area of ​​the original multimedia, thereby enhancing user stickiness.

[0084] Optionally, in addition to being displayed in the aforementioned manner, the multimedia commentary may also be displayed before the multimedia information and automatically redirect to the corresponding multimedia information after completion. The multimedia commentary may also cease displaying in response to a user's close operation and automatically redirect to the corresponding multimedia information after cessation. It should be understood that the display format and timing of the multimedia commentary should not be construed as limitations on this application.

[0085] Optionally, during the playback of the commentary multimedia, the original multimedia can be in a playback state or a paused state. Next, the playback state of the original multimedia is classified and explained: If the commentary multimedia is dynamic information without audio, the original multimedia can maintain the playback state. Among them, the dynamic information can be a video or a moving picture. If the commentary multimedia is dynamic information with audio, the original multimedia can be in a paused state. If the commentary multimedia covers the original multimedia, such as: the commentary multimedia automatically plays in full screen before the original multimedia plays, the original multimedia can be in a paused state. If the original multimedia becomes a small window, such as Figure 1c The floating window shown in (2) in the figure, the original multimedia can be in the playing state or in the pause state; and if the user clicks Figure 1c The floating window shown in (2) is Figure 1c The playback interface of (2) will change back to the following Figure 1c The playback interface shown in (1) in FIG. If the original multimedia and the comment multimedia appear side by side, such as Figure 1b In the cases shown in (1) and (2), the original multimedia can be in a playing state or a paused state. It should be understood that the specific presentation form of the original multimedia and the commentary multimedia, as well as the presentation state of the original multimedia should not be understood as a limitation of this application.

[0086] It is understandable that in the embodiments of this application, related data such as user information is involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0087] It should be noted that the serial numbers of the following embodiments are not intended to limit the preferred order of the embodiments.

[0088] In this embodiment, a multimedia playback method is provided. Figure 1d As shown, the multimedia playback method is applied to an electronic device. The specific process of the method may include the following steps 110 to 120:

[0089] 110. Display a playback interface of the original multimedia, wherein the playback interface further includes a multimedia control of the commentary multimedia corresponding to the original multimedia.

[0090] Original multimedia is used to present content in a multimedia format. This content can be news or daily life sharing. News can include current events, financial news, entertainment news, technology information, etc. Daily life sharing can include sharing pets, travel records, and food records.

[0091] Multimedia typically includes the integrated presentation of multiple media formats, including text, images, audio, and video. Through multimedia, users can obtain information more intuitively and vividly, improving the effectiveness of information presentation and communication efficiency.

[0092] The playback interface is an interface used in an application to play video or audio content. Figure 1b and Figure 1c This illustrates two specific forms of the playback interface. Optionally, the playback interface may also include a comment area for the original multimedia content. The comment area for the original multimedia content includes comments. The comment may be at least one comment in the comment area. A comment may be a selected comment; comments may also include a selected comment and its subordinate comments. The specific content covered by the comments will be described in detail below.

[0093] The multimedia content of the comment multimedia can be specifically generated based on the comments. The generation process of the comment multimedia will be described in detail below.

[0094] Multimedia controls are interactive elements located in the playback interface. Multimedia controls can be Figure 1b (1) and (2) in the figure show virtual buttons: video comments; it can also be Figure 1c The floating window corresponding to the comment multimedia shown in (1) is shown.

[0095] 120. Play the comment multimedia corresponding to the multimedia control on the playback interface.

[0096] In step 120, the commentary multimedia can be played even if the multimedia control is not triggered by the user. For example, the commentary multimedia can be played before the original multimedia is played, or the commentary multimedia can be played automatically after the original multimedia is played. Alternatively, in step 120, the commentary multimedia can be played even if the multimedia control is triggered by the user. It should be understood that whether the commentary multimedia is played or not can be independent of whether the multimedia control is triggered.

[0097] Optionally, step 120 may specifically be: if the multimedia control is triggered, playing the comment multimedia corresponding to the multimedia control on the playback interface.

[0098] In the above embodiments, the multimedia controls are triggered in various ways. The user can trigger the multimedia controls by clicking. For example, Figure 1b (1) and (2) show virtual buttons: video comments, users can trigger multimedia controls by clicking virtual buttons; Figure 1c In (1), the user can trigger the multimedia control by clicking on the floating window corresponding to the comment multimedia. The user can also trigger the multimedia control by means other than clicking, such as sliding, dragging, etc. It should be understood that the specific form of triggering the multimedia control should not be construed as a limitation of the present application.

[0099] Optionally, in one embodiment, the comments may include a target comment and a lower level comment of the target comment. Accordingly, before step 120, the embodiment of the present application may further include:

[0100] The comment multimedia is generated based on the target comment.

[0101] The target comment refers to a comment selected from the comment area. Optionally, the target comment can be selected based on the amount of interaction with other users, or it can be randomly selected from the comment area. The selection process of the target comment should not be understood as a limitation of this application.

[0102] The next level of comments of the target comment is: a comment that replies to the target comment. Optionally, in some embodiments, the next level of comments of the target comment can also be called: a follow-up comment of the target comment.

[0103] It is understood that a target comment may have a parent comment, i.e., a target comment is a comment on a comment, and the comment it targets is the parent comment of the target comment. Alternatively, a target comment may not have a parent comment, i.e., a comment on the original multimedia. Whether a target comment has a parent comment should not be construed as a limitation of this application.

[0104] In the above embodiment, even if the comments include a target comment and a lower-level comment of the target comment, the comment multimedia can be generated only from the target comment. In the above embodiment, the comment multimedia generation process is relatively fast, which can effectively save computing resources in the comment multimedia generation process and improve generation efficiency.

[0105] Alternatively, in another embodiment, the comment includes a target comment, and the target comment does not have a next-level comment. Accordingly, before step 120, the embodiment of the present application may further include:

[0106] The comment multimedia is generated based on the target comment.

[0107] In the above embodiment, if the target comment does not have a lower level comment, a comment multimedia can be generated based on the only target comment. Displaying the target comment in the form of multimedia can make users more vivid and intuitive to understand the comment information, thereby increasing user stickiness.

[0108] Optionally, in the above two implementations, the step of generating the comment multimedia based on the target comment may specifically include the following steps A1 to A4:

[0109] A1. Based on the interaction data of other users on multiple comments in the comment area, the comments whose interaction data exceeds a threshold are taken as the target comments, and the target comments are target texts.

[0110] Other users are: For each comment, users other than the user who posted the comment.

[0111] Interaction data reflects the level of attention a comment receives. Users can use comments with high interaction data to understand the public's views, attitudes, and opinions on a particular event or topic. Comments with high interaction data are also known as popular comments. Optionally, interaction data can include the number of likes or comments received on a comment.

[0112] The number of likes represents the degree of recognition and enjoyment of the comment by other users. Generally speaking, a high number of likes means that the comment resonates with other users or provides valuable information or insights.

[0113] Follow-up count indicates the extent to which other users are responding to and discussing a comment. When a comment generates more follow-up comments, it means that other users are interested in it and are willing to participate in related discussions. An increase in follow-up count indicates that a comment has attracted a certain level of attention.

[0114] The target text is the text for which multimedia material is to be obtained. In the above two implementations, since the comment multimedia is generated based only on the target comment, the scope of the target text is the target comment.

[0115] In addition to selecting target comments based on other users' interactive data, target comments can also be selected in other ways. For example, comments can be randomly selected from the comment area as target comments. In a specific embodiment, let's take the interactive data as the number of likes or comments on the comment as an example to describe the target comment selection process in detail. The target comment selection process can specifically include the following steps A11 to A12:

[0116] A11. For each comment, calculate the score of the comment based on the number of likes or comments corresponding to the comment.

[0117] In step A11, the score of a comment can be calculated based on the number of likes corresponding to the comment, the number of comments corresponding to the comment, or the number of likes and comments corresponding to the comment. The following describes the calculation of the score of a comment based on the number of likes and comments corresponding to the comment as an example.

[0118] Optionally, a score value is calculated based on the number of likes and comments of a comment, and can be specifically calculated in the following manner: score value = number of likes / 10,000 + number of comments / 10.

[0119] For example, let's assume that the number of likes for a comment is a and the number of comments is b, then the score of the comment is: a / 10000+b / 10.

[0120] The above calculation method can be used to calculate the score value corresponding to each comment. There are corresponding calculation formulas for calculating the score value of a comment based on the number of likes corresponding to the comment, or for calculating the score value of a comment based on the number of comments corresponding to the comment. The calculation formula can be set by the developer based on work experience. The specific form of the calculation formula should not be understood as a limitation of this application.

[0121] A12. Obtain the comment with the highest score from the multiple comments, and use the comment with the highest score as the target comment.

[0122] After calculating the score value corresponding to each comment, the highest score value may be selected and the comment with the highest score value may be determined as the target comment.

[0123] In the above embodiment, target comments are selected based on the number of likes and comments, so that the selected target comments have a certain viewable value, thereby improving the viewing efficiency of users viewing the comment area.

[0124] A2. Acquire multimedia materials corresponding to the target text.

[0125] Multimedia materials are elements such as images, audio, and video required for generating multimedia. Optionally, the step of obtaining multimedia materials corresponding to the target text may specifically include the following steps A21 to A25:

[0126] A21. For each subtext of the target text, determine whether there is a first-category multimedia material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value. If yes, proceed to step A22; if not, proceed to step A23.

[0127] Subtexts are components of the target text. Continuing with the previous example, where the target text is a target review, subtexts can include sentences within the target review. If the target review consists of only one sentence, the subtext is the target text. If the target review consists of N sentences (N is a positive integer greater than 1), the target text includes N subtexts.

[0128] The first preset similarity value is a preset value, and its specific value can be set by the developer based on work experience.

[0129] The first type of multimedia material is a specific type of multimedia material. Optionally, the first type of multimedia material can be a video type of multimedia material.

[0130] Optionally, in one embodiment, the subtext may be a line from a film or television work, and the line corresponds to a corresponding video clip material. Therefore, if a first type of multimedia material is found whose similarity to the subtext exceeds a first preset similarity value, it can be determined that the subtext is a line from a film or television work, and if a video clip material corresponding to the subtext is found, step A22 can be executed. For a subtext that has a video clip material, the dubbing character of the subtext is the character played by the actor who speaks the corresponding line in the video clip material. For example, if the similarity between a subtext x and the first type of multimedia material X is higher than the first preset similarity value, and the first type of multimedia material X is spoken by character D played by actor C, the dubbing character of the subtext x can be determined to be character D.

[0131] If no first-category multimedia material having a similarity with the subtext exceeding a first preset similarity value is found, it means that the subtext is not a line from a film or television work; or although the subtext is material from a film or television work, no corresponding first-category multimedia material is found, then step A23 may be executed.

[0132] The method of calculating the similarity between the first type of multimedia material and the subtext will be described in detail below.

[0133] A22: Use the first type of multimedia material whose similarity exceeds a first preset similarity value as the multimedia material corresponding to the subtext.

[0134] A23. Generate voice data corresponding to the subtext.

[0135] Optionally, in one embodiment, step A23 may specifically include the following steps A231 to A232:

[0136] A231. Calculate the sentiment tendency of the sub-text.

[0137] Optionally, the sentiment tendency of the sub-text may be calculated using a support vector machine (SVM), or the Bayesian formula may be used to calculate the sentiment tendency of the sub-text.

[0138] In one embodiment, the sentiment tendency of a subtext is calculated using SVM, which can be specifically achieved in the following manner:

[0139] Convert text into a numerical vector representation. This can be done using methods such as the bag-of-words model. Divide the preprocessed dataset into a training set and a test set. Use the Support Vector Machine (SVM) algorithm to train the training set to create a classifier. Input the test set into the trained model, perform predictions, and calculate the accuracy of the predictions. Once the test meets the requirements, use the trained classifier to calculate the sentiment of the subtext.

[0140] In another embodiment, the Bayesian formula is used to calculate the sentiment tendency of the sub-text, which can be specifically implemented as follows:

[0141] Convert the text into a numerical vector representation. You can use methods such as the bag-of-words model or TF-IDF to convert the text into a numerical vector. According to the task requirements, build a sentiment dictionary that includes multiple emotion types, such as surprise, curiosity, admiration, and love. This can be constructed using existing dictionaries or manual annotations. Based on the training dataset, calculate the prior probability of the text's sentiment category, that is, P(emotion category). For each feature (word) and sentiment category, calculate the conditional probability P(feature | sentiment category), that is, the probability of the feature appearing under a given sentiment category. For new text samples, calculate the posterior probability of its belonging to each sentiment category according to the Bayesian formula, and select the category with the highest posterior probability as the prediction result.

[0142] A232. Generate voice data corresponding to the sub-text based on the emotional tendency.

[0143] Optionally, a text-to-speech (TTS) engine may be used to generate corresponding voice data from the corresponding subtext according to the emotional tendency calculated in step A231.

[0144] In the above embodiment, for each subtext, voice data corresponding to the subtext can be generated based on steps A231 and A232. Optionally, the speech rate of the voice data can be adjusted based on the developer's work experience. For example, the speech rate of the voice data can be set to 3 words per second.

[0145] A24. Search the multimedia material library for a second type of multimedia material whose similarity to the subtext exceeds a preset similarity value, wherein the type of the second type of multimedia material is different from the type of the first type of multimedia material.

[0146] The second type of multimedia material is a multimedia material different from the first type of multimedia material. Optionally, the second type of multimedia material can be a picture type multimedia material.

[0147] Optionally, the similarity between the second-category multimedia material and the subtext can be determined based on the label of the second-category multimedia material. The label of the second-category multimedia material can be manually entered by a developer, or obtained by an electronic device performing optical character recognition on the second-category multimedia material. It should be understood that the specific method for obtaining the label of the second-category multimedia material should not be construed as a limitation of this application.

[0148] A25. Combine the found second-category multimedia material with the voice data to obtain multimedia material corresponding to the subtext.

[0149] By combining the second type of multimedia material found in step A24 with the voice data generated in step A23, the multimedia material corresponding to the subtext can be obtained.

[0150] In the above embodiment, it is first determined whether a first type of multimedia material with high similarity to the subtext can be found. If so, the first type of multimedia material can be directly used as the multimedia material corresponding to the subtext. If not, the emotional tendency of the subtext can be calculated and speech data with the corresponding emotional tendency can be generated. Then, a second type of multimedia material with high similarity to the subtext can be found and combined with the language data to obtain the multimedia material corresponding to the subtext. The above embodiment can generate intuitive and vivid multimedia materials from the subtext, which is easy for users to receive and understand.

[0151] Optionally, the judgment process in step A21 can be implemented based on the label of the first type of multimedia material.

[0152] In the above embodiment, the similarity between the first type of multimedia material and the subtext can be calculated based on the label of the first type of multimedia material. The process of calculating the similarity can be specifically as follows:

[0153] Process the text and subtext corresponding to the label, including removing punctuation and stop words, and convert the text into a unified format. Extract features from the two texts, using methods such as the bag-of-words model and TF-IDF. Represent the text as a vector, with each dimension representing a specific feature. Select an appropriate similarity metric, such as cosine similarity, Euclidean distance, and Jaccard similarity. Calculate the similarity between the two texts based on the selected similarity metric.

[0154] Optionally, in one embodiment, the specific text content of the subtext can be used as subtitles and added to the screen of the corresponding multimedia material, thereby further increasing the intuitiveness of the video.

[0155] Accordingly, before the step of determining whether there is a first type of multimedia material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value, the embodiment of the present application may further include the following steps X1 to X4:

[0156] X1. Identify the sound in the video material and generate a sound text corresponding to the sound.

[0157] X2. Obtain multiple screenshots of the video material.

[0158] X3. Perform optical character recognition on each screenshot to obtain the subtitle text in the screenshot.

[0159] X4. Use the audio text, subtitle text, and the plurality of screenshots as labels corresponding to the video material, and store them together with the video material in the multimedia material library.

[0160] In the above implementation, the sound in the video material can be extracted and the corresponding audio text generated. Alternatively, a screenshot of the video material can be obtained and optical character recognition processed on the screenshot to obtain the subtitle text in the screenshot, thereby obtaining multiple screenshots and audio text and subtitle text information corresponding to the video material, and using the above information as a label for the video material. The audio text and subtitle text can be used for comparison with the subtext, and the screenshot can be used by developers to quickly and intuitively understand the content of the video material.

[0161] A3. If the multimedia material is one piece, generate the comment multimedia based on the multimedia material.

[0162] If there is only one piece of multimedia material, the multimedia material can be directly displayed as comment multimedia.

[0163] A4. If there are at least two multimedia materials, the multimedia materials are spliced ​​according to the text order of the target text to generate the comment multimedia.

[0164] Because each piece of multimedia material has its own corresponding subtext, and the subtexts have their own corresponding order within the target text, at least two multimedia materials can be sorted according to the order of the subtexts within the target text, thereby generating the final review multimedia. This method can convert text reviews into video-type multimedia, increasing the interest of the reviews and allowing users to more intuitively absorb the content of the reviews.

[0165] Optionally, in another embodiment, the comments include a target comment and a lower level comment of the target comment. Accordingly, before step 120, the embodiment of the present application may further include:

[0166] The comment multimedia is generated based on the target comment and the next level comment of the target comment.

[0167] In the above implementation, the target comment and the next level comment of the target comment can form a logical dialogue, and the multimedia comments generated based on the two are more interesting, which can further increase user stickiness.

[0168] Optionally, in a specific implementation, the step of generating the comment multimedia based on the target comment and the next level comment of the target comment may specifically include the following steps B1 to B5:

[0169] B1. Based on the interaction data of other users on the multiple comments in the comment area, the comments whose interaction data exceeds a threshold are selected as the target comments.

[0170] The other users are: for each comment, users other than the user who posted the comment. Step B1 is the same as step A1 and will not be described in detail here.

[0171] B2. For the next level comments of the target comment, screen out the first-screened next level comments that have a strong correlation with the target comment.

[0172] Strong correlation means that the target comment has a high degree of correlation with the comments at the next level after the initial screening. The specific meaning of strong correlation will be described in detail below.

[0173] Optionally, the target comment may correspond to multiple lower-level comments. Accordingly, the step of screening out the initially screened lower-level comments that are strongly correlated with the target comment includes steps B21 to B23:

[0174] B21. If at least two of the multiple next-level comments have text of the same preset length, deduplication processing is performed on the at least two next-level comments to obtain deduplication result comments corresponding to the at least two next-level comments. The preset length text is text that exceeds the set length; the deduplication result comments and the next-level comments that do not require deduplication processing are recorded as secondary comments.

[0175] Deduplication refers to the process of identifying and deleting duplicate data during data processing.

[0176] In the above implementation, if the target comment corresponds to only one next-level comment, deduplication is not required. If the target comment corresponds to multiple next-level comments, deduplication can be used to reduce redundant information. For example, suppose the target comment corresponds to 10 next-level comments, 3 of which have duplicate content A; another 3 of which have duplicate content B; 2 of which have duplicate content C; and the remaining 2 have no duplicate content.

[0177] Then, the three comments with duplicate content A can be deduplicated into one deduplicated comment, Comment 1; the three comments with duplicate content B can be deduplicated into one deduplicated comment, Comment 2; and the two comments with duplicate content C can be deduplicated into one deduplicated comment, Comment 3. Deduplicated comments 1, 2, and 3, along with the remaining two comments without duplicate content, are considered secondary comments. A total of five secondary comments are obtained.

[0178] Optionally, step: performing deduplication processing on the at least two next-level comments may specifically include the following steps B211 to B213:

[0179] B211. Obtain at least two next-level comments containing text of the same preset length from the at least two next-level comments.

[0180] The preset length text is text that exceeds the set length. The set length is a pre-set text length, which can be 5 characters or 7 characters. It should be understood that the specific length value of the set length can be set by the developer based on work experience, and the specific length value of the set length should not be construed as a limitation of this application.

[0181] If at least two next-level comments have text of the same preset length, it means that the at least two next-level comments are highly repetitive. Therefore, multiple next-level comments can be compared in pairs to obtain at least two next-level comments that have text of the same preset length.

[0182] B212. Randomly select one next-level comment from the at least two next-level comments having the same text of the preset length, and the randomly selected next-level comment is the comment to be merged.

[0183] The comment to be merged is a comment to be merged with other comments with the same text of the preset length. The comment to be merged can be randomly selected from at least two lower-level comments with the same text of the preset length.

[0184] B213. For at least two lower-level comments that have the same text of the preset length, extract the other texts except the text of the preset length, and add the other texts to the comment to be merged.

[0185] After determining that there are at least two next-level comments with the same preset length of text, for the other next-level comments except the comment to be merged, the other text except the repeated text in the next-level comments can be extracted and added to the comment to be merged.

[0186] For example, let's assume that the content of three next-level comments is as follows: J1: "Wow, so cute and handsome"; J2: "So cute and handsome, I like it"; J3: "So cute and handsome, I love it." All three of the above next-level comments contain the same preset length text: "So cute and handsome."

[0187] You can randomly select one of the next-level comments as the comment to be merged. You may choose J3: "So cute and handsome, I love it" as the comment to be merged.

[0188] Then, for J1, extract the content other than "too cute and too handsome": "woo woo woo,"; for J2, similarly extract the content other than "too cute and too handsome": ", like", and add the above extracted content to J3 to get the deduplication result:

[0189] "Wow, so cute and handsome, I like it, I love it."

[0190] In the above-mentioned implementation mode, the number of next-level comments can be simplified by the above-mentioned method, thereby facilitating improvement of processing efficiency.

[0191] In another embodiment, step: performing deduplication processing on the at least two next-level comments may specifically include the following steps Y1 to Y3:

[0192] Y1. Obtain at least two next-level comments containing text of the same preset length from the at least two next-level comments.

[0193] Y2. From at least two next-level comments with text of the same preset length, select the next-level comment with the longest text length.

[0194] Y3. From at least two next-level comments that have text of the same preset length, delete the next-level comments except for the next-level comment with the longest text length.

[0195] For multiple sub-level comments with the same preset length, the text length of each sub-level comment can be obtained, and then the sub-level comment with the longest text length can be selected. The sub-level comment with the longest text length can then be retained, and the other sub-level comments can be deleted.

[0196] For example, let's set up three next-level comments as follows: J4: "My younger brother does what he says, he's a hero"; J5: "My younger brother does what he says, he's a hero!!"; J6: "My younger brother does what he says, he's a hero, it's awesome."

[0197] The text lengths of the three next-level comments can be obtained separately. The text length of J4 is 12, the text length of J5 is 14, and the text length of J6 is 18. Therefore, J6, which has the longest text length, can be retained, and J4 and J5 can be deleted to obtain the deduplication processing result:

[0198] "My brother does what he says and is a hero. It's really great."

[0199] In the above-mentioned implementation, the efficiency of deduplication can be further improved, the amount of deduplication calculation can be simplified, and thus the processing efficiency can be further improved.

[0200] B22. For each of the secondary comments, if there is a preset symbol in the secondary comment, or the format of the secondary comment is the preset format, it is determined that the secondary comment has a weak correlation with the target comment.

[0201] The preset symbol is a pre-set symbol that can be added in advance by the developer based on work experience. For example, the preset symbol can be the @ symbol. Normally, the appearance of the @ symbol indicates that the secondary comment is to remind someone to check the corresponding target comment, so it can be determined that the correlation between the secondary comment with the @ symbol and the target comment is weak, that is, there is a weak correlation. The preset symbol can also be other symbols, such as emoticons composed of symbols ":)" and ":(", etc. It should be understood that the specific symbol type of the preset symbol should not be understood as a limitation to the present application.

[0202] Preset formats are pre-defined formats that can be added by developers based on their experience. For example, a preset format could be an image format, such as jpg or gif. If the secondary comment is in one of these formats, it indicates that the secondary comment is a purely emoji comment. Typically, the correlation between purely emoji comments and the target comment is weak, indicating a weak correlation.

[0203] B23. Eliminate the secondary comments that are weakly correlated with the target comment from the plurality of secondary comments, thereby obtaining the pre-screened next-level comments that are strongly correlated with the target comment.

[0204] In the above implementation, the symbols and formats of the secondary comments can be used to identify several secondary comments with weak correlations to the target comment. These weakly correlated secondary comments are then deleted, and the remaining secondary comments with strong correlations to the target comment are used as the initial screening of the next-level comments. This approach can reduce the number of next-level comments corresponding to the target comment, thereby saving computational effort.

[0205] Optionally, step B23 may specifically include the following steps B231 to B233:

[0206] B231. Eliminate the secondary comments that are weakly correlated with the target comment from the plurality of secondary comments, thereby obtaining a result after elimination.

[0207] The result after elimination is: among multiple secondary comments, multiple secondary comments except for the secondary comments that have a weak correlation with the target comment.

[0208] B232. Calculate the text relevance between each of the secondary selected comments in the eliminated results and the target comment.

[0209] Text relevance is a metric used to measure the degree of relevance between two reviews. The text relevance between each secondary review and the target review in the eliminated results can be calculated using the following methods:

[0210] 1. Calculate the number or proportion of words that appear in both the second-selected and target comments. The more overlapping words, the higher the likelihood of correlation between the texts. 2. Convert the texts into word vector representations and then measure text correlation by calculating the similarity between the vectors. Common methods include cosine similarity. 3. Use natural language processing techniques, such as word sense disambiguation and semantic role labeling, to analyze the semantic information of the texts and determine the correlation between the texts. 4. Use machine learning models, such as text classification and text matching, to train the models to predict the correlation between texts.

[0211] B233. From the eliminated results, obtain the secondary selected comments whose text relevance exceeds the relevance threshold, wherein the secondary selected comments whose text relevance exceeds the relevance threshold are: the primary screened next-level comments that have a strong relevance to the target comment.

[0212] The relevance threshold is a pre-set critical value that reflects whether the text relevance meets the requirements. The relevance threshold can be set by developers based on their work experience. If the text relevance is less than or equal to the relevance threshold, the text relevance between the second-selected comment and the target comment is considered to be unreliable. If the text relevance is greater than the relevance threshold, the text relevance between the second-selected comment and the target comment is considered to be reliably relevant.

[0213] The secondary selected comments that meet the text correlation requirements with the target comments are determined as the primary screened next-level comments that have a strong correlation with the target comments.

[0214] In the above-mentioned implementation, the secondary comments with weak correlation can be first eliminated from the multiple secondary comments to obtain the eliminated results, and then the text correlation between each secondary comment in the eliminated results and the target comment is calculated to obtain the text correlation corresponding to each secondary comment in the eliminated results. The multiple text correlations are then compared with the correlation threshold, and the secondary comments with text correlations exceeding the correlation threshold are used as the pre-screened next-level comments that have a strong correlation with the target comment. The above-mentioned implementation can obtain the next-level comments that have a strong correlation with the target comment through two-stage screening, which is more conducive to improving the logical coherence between the target comment and the next-level comments.

[0215] B3. Sort the target comments and the initially screened next-level comments to obtain a target text.

[0216] Optionally, the target comment may correspond to multiple primary screening next-level comments. Accordingly, the step of sorting the target comment and the primary screening next-level comments may specifically include steps S1 to S2:

[0217] S1. For each of the initially screened next-level comments, calculate the logical relationship coefficient between the initially screened next-level comment and the target comment, thereby obtaining a plurality of logical relationship coefficients.

[0218] The logical relationship coefficient is a coefficient that reflects the logical relationship between the initial screening next-level comments and the target comments. The larger the logical relationship coefficient, the stronger the logical relationship between the initial screening next-level comments and the target comments. A logical relationship is a relationship that has a connection or correlation. For example, a logical relationship can be a causal relationship, similarity, correlation, etc. The specific relationship type of the logical relationship should not be understood as a limitation of this application.

[0219] Optionally, in a specific embodiment, the logical relationship coefficient is the Spearman rank correlation coefficient. Correspondingly, step S1 may specifically include the following steps S11 to S13:

[0220] S11. Perform word segmentation on each of the initially screened lower-level comments to obtain the result of follow-up comment decomposition.

[0221] Word segmentation refers to the operation of decomposing the sentence corresponding to the initially screened lower-level comment into words.

[0222] Specifically, the N-gram algorithm can be used to perform word segmentation on each initially screened lower-level comment. The process is as follows:

[0223] Determine the value of N in the N-gram algorithm, that is, N consecutive words (or characters) are processed as a unit. Common values of N include 1-gram (single word), 2-gram (adjacent two words), 3-gram (adjacent three words), etc.

[0224] Preprocess the initially screened lower-level comment to be decomposed, including operations such as removing punctuation marks and converting to lowercase. This process can eliminate interference and unify the text format.

[0225] Segment the preprocessed initially screened lower-level comment according to the value of N to obtain a sequence with N consecutive characters (or words) as a unit. For example, for the sentence "My little brother always keeps his word. He is a great hero. That's really wonderful.", when N = 2, the following 2-gram sequence can be obtained: ["My little", "little brother", "brother always", "always keeps", "keeps his", "his word", "word. He", ". He is", "is a", "a great", "great hero", "hero. That", "That's really", "really wonderful", "wonderful."].

[0226] Subsequently, count the frequency of each word in the 2-gram sequence in multiple initially screened lower-level comments, and replace the position of the original word in the 2-gram sequence with this frequency. For example, if the frequency of the word "My little" in multiple initially screened lower-level comments is 10, then the word "My little" will be replaced with "10". And so on, so as to obtain a sequence composed of frequency values, and this sequence composed of frequency values is the result of follow-up comment decomposition of the corresponding initially screened lower-level comment.

[0227] The above word segmentation process can be performed on each initially screened lower-level comment, so as to obtain the result of follow-up comment decomposition corresponding to each initially screened lower-level comment.

[0228] S12. Perform word segmentation on the target comment to obtain the result of main comment decomposition.

[0229] Continuing with the above example, the target comment can be tokenized using the N-gram algorithm. The process is as follows:

[0230] Determine the value of N in the N-gram algorithm, that is, N consecutive words (or characters) are treated as a unit. Common values of N include 1-gram (single word), 2-gram (adjacent two words), 3-gram (adjacent three words), etc.

[0231] Preprocess the target comment to be disassembled, including operations such as removing punctuation marks and converting to lowercase. This process can eliminate interference and unified the text format.

[0232] Split the preprocessed target comment according to the value of N to obtain a sequence with N consecutive words (or characters) as a unit. For example, for the sentence "The actor said the box office is 2 billion and performed a show", when N = 2, the following 2-gram sequence can be obtained: ["actor", "said by the actor", "said the box office", "box office", "box office two", "two billion", "billion", "billion performance", "performed", "performed a show", "show"].

[0233] Subsequently, count the frequency of each word in the 2-gram sequence in the range composed of the target comment and multiple preliminary screened next-level comments, and replace the position where the original word appears in the 2-gram sequence with this frequency. For example, if the frequency of the word "box office" in the range composed of the target comment and multiple preliminary screened next-level comments is 15, then the word "box office" will be replaced with "15". And so on, so as to obtain a sequence composed of frequency values, and this sequence composed of frequency values is the main comment disassembling result of the target comment.

[0234] S13. Calculate the Spearman rank correlation coefficient between each of the follow-up comment disassembling results and the main comment disassembling result, where the Spearman rank correlation coefficient is the logical relationship coefficient between the corresponding preliminary screened next-level comment and the target comment.

[0235] Taking the calculation of the Spearman rank correlation coefficient between any one of the multiple follow-up comment disassembling results and the main comment disassembling result as an example:

[0236] First, judge whether the amount of data included in the sequence of this follow-up comment disassembling result is the same as the amount of data included in the sequence of the main comment disassembling result.

[0237] If not, adjust the amount of data in the sequence of the follow-up comment disassembling result to be the same as the amount of data in the sequence of the main comment disassembling result. The specific method of adjustment is as follows: if the amount of data in the sequence of the follow-up comment disassembling result is less than the amount of data in the sequence of the main comment disassembling result, make it up by filling with 0s; if the amount of data in the sequence of the follow-up comment disassembling result exceeds the amount of data in the sequence of the main comment disassembling result, delete the excess part.

[0238] After adjusting the data volume of the two to be consistent, the difference between the same order positions of the two can be calculated. For example, the difference between the i-th position of the sequence of the main review decomposition result and the i-th position of the sequence of the follow-up review decomposition result is calculated, and recorded as di.

[0239] Calculate the sum of the squares of the above differences:

[0240] Where n is the value of the data volume.

[0241] The Spearman rank correlation coefficient between the comment decomposition result and a follow-up comment decomposition result is calculated according to the following formula:

[0242] In the above-mentioned embodiment, the above-mentioned calculation process can be performed on each of the multiple follow-up review decomposition results, so that the Spearman rank correlation coefficient between each follow-up review decomposition result and the main review decomposition result can be calculated, that is, the logical relationship coefficient between each follow-up review decomposition result and the main review decomposition result, thereby facilitating the subsequent determination of the sorted text.

[0243] S2. Sort the target comment and the multiple initially screened next-level comments corresponding to the target comment based on the values ​​of the multiple logical relationship coefficients.

[0244] After calculating the logical relationship coefficient between each pre-screened review and the target review, the target review and the pre-screened reviews can be sorted in descending order of the logical relationship coefficients. Specifically, the sorting process can be as follows: the target review is ranked first, and the pre-screened reviews are sorted starting from the second position in descending order of their corresponding logical relationship coefficients.

[0245] Optionally, step B3 may specifically include the following steps B31 to B33:

[0246] B31. Sort the target comments and the initially screened next-level comments to obtain sorted text.

[0247] The "sorting of the target comments and the initially screened next-level comments" in step B31 is the same as that in steps S1 to S2, and will not be described in detail here.

[0248] B32. Determine at least one insertion point based on the text content of the sorted text.

[0249] The text content consists of the target comment and multiple pre-screened comments, sorted. Therefore, you can select specific locations based on the text content, such as the beginning of the sorted text, the end of the sorted text, or between the target comment and pre-screened comments. These locations can all be determined as insertion points.

[0250] In addition to the above-mentioned insertion points, insertion points can also be randomly generated between multiple primary screening next-level comments. Specifically, the number of insertion points to be generated in the primary screening next-level comments can be determined based on the numerical interval into which the number of primary screening next-level comments falls. For example: if the number of primary screening next-level comments falls between 3 and 6, then one insertion point is randomly generated between 3 and 6 primary screening next-level comments; if the number of primary screening next-level comments falls between 7 and 10, then two insertion points are randomly generated between 7 and 10 primary screening next-level comments, etc. It should be understood that the specific values ​​of the upper and lower limits of the numerical interval and the number of randomly generated insertion points should not be understood as limitations on this application.

[0251] B33. Insert the corresponding preset linking text at each insertion point to obtain the target text.

[0252] Preset transition text is used to carry forward the context in a preset text template. Preset transition text can be pre-set by the developer based on work experience. Optionally, the preset transition text can include opening text, at least one transition text, and closing text.

[0253] For example, suppose the sorted text includes:

[0254] A target review;

[0255] Three primary screening next level comments: primary screening next level comment 1, primary screening next level comment 2, primary screening next level comment 3;

[0256] The default transition texts include:

[0257] Opening Text: This is how it all started;

[0258] Transition text 1: Regarding this matter, see how everyone responded;

[0259] Transition Text 2: Not only that

[0260] Ending text: This generation of netizens is really good at having fun

[0261] The insertion points can be determined as follows: the insertion point at the beginning corresponds to the insertion of the opening text; the insertion point between the target comment and the first-level comment corresponds to the insertion of transition text 1; an insertion point is randomly generated for the three first-level comments, which is used to insert transition text 2; the insertion point at the end corresponds to the insertion of the ending text, so the target text can be obtained as follows:

[0262] “Opening text: This is how it all started;

[0263] Target Comments;

[0264] Transition text 1: Regarding this matter, see how everyone responded;

[0265] Initial screening of next level comments 1;

[0266] Transition Text 2: Not only that

[0267] Initial screening of next level comments 2;

[0268] Initial screening of next level comments 3;

[0269] Ending text: This generation of netizens is really good at having fun"

[0270] Optionally, in one embodiment, the specific comment content of the target comment and the preliminary screening next-level comment may be brought in. Let the target comment be: "Actor A said the box office is 2 billion, he performed a show", the preliminary screening next-level comment 1 be: "King Zhou: You are my bravest son", the preliminary screening next-level comment 2 be: "Wow, my husband is so cute and handsome", and the preliminary screening next-level comment 3 be: "Bo Yikao: My brother does what he says and is a great hero".

[0271] The target text can be:

[0272] “Opening text: This is how it all started;

[0273] Actor A says the box office is 2 billion and he performs a show;

[0274] Transition text 1: Regarding this matter, see how everyone responded;

[0275] King Zhou: You are my bravest son;

[0276] Transition Text 2: Not only that

[0277] Woohoo, my husband is so cute and handsome;

[0278] Bo Yikao: My brother keeps his word and is a great hero.

[0279] Ending text: This generation of netizens is really good at having fun"

[0280] Optionally, in addition to adding preset linking texts between the sorted texts in the above manner, preset linking texts may also be added by importing the sorted texts into a preset text template. The preset text template includes at least one linking text.

[0281] The preset text template is a template with certain stylistic features. Different stylistic features correspond to different specific templates. Stylistic features include: scripts, news reports, etc. It should be understood that the specific content of the stylistic features and the number of stylistic features should not be construed as limiting this application.

[0282] For example, suppose the sorted text includes:

[0283] A target review;

[0284] Three primary screening next level comments: primary screening next level comment 1, primary screening next level comment 2, primary screening next level comment 3;

[0285] The default transition texts include:

[0286] Opening Text: This is how it all started;

[0287] Transition text 1: Regarding this matter, see how everyone responded;

[0288] Transition Text 2: Not only that

[0289] Ending text: This generation of netizens is really good at having fun

[0290] The specific process of importing sorted text into the preset text template is as follows:

[0291] Import the target comment between the opening text and transition text 1, import one or two of the three initially screened next-level comments between transition text 1 and transition text 2, and import the remaining initially screened next-level comments among the three initially screened next-level comments between transition text 2 and the ending text, so that the target text can be obtained.

[0292] For a target text with preset transition text, if voice data is to be generated, the preset transition text can use the same emotion and tone to generate corresponding voice data, so that the preset transition text can be used as narration.

[0293] It should be understood that the above-described preset text template is merely an example of a script template. The preset text template may also be other templates, such as a news template. A news template may include transition text such as: opening text: "News has been reported"; transition text: "Other parties express their views on this matter"; and closing text: "We will continue to pay attention to you." The specific content of the preset text template and the specific number and content of the transition text should not be construed as limiting this application.

[0294] B4. Acquire multimedia materials corresponding to the target text.

[0295] Compared with the above steps A21 to A25, step B4 is the same except that the specific content of the sub-text is changed, so it will not be repeated here. Specifically, the sub-text of step B4 can include the sentence of the target review and multiple sentences of the initially screened next-level reviews.

[0296] B5. Splicing the multimedia materials according to the text order of the target text to generate the comment multimedia.

[0297] Because each piece of multimedia material has its own corresponding subtext, and the subtexts have their own corresponding order within the target text, multiple multimedia materials can be sorted according to the order of the subtexts within the target text to generate the final review multimedia. This method can convert text reviews into video-type multimedia, making the reviews more interesting and allowing users to more intuitively absorb the content of the reviews.

[0298] Optionally, in another specific implementation, the step of generating the comment multimedia based on the target comment and the next level comment of the target comment may specifically include the following steps C1 to C7:

[0299] C1. Based on the interaction data of other users on the multiple comments in the comment area, the comments whose interaction data exceeds a threshold are taken as the target comments.

[0300] The other users are: for each comment, users other than the user who posted the comment.

[0301] C2. For the next level comments of the target comment, screen out the first-screened next level comments that are strongly correlated with the target comment.

[0302] Steps C1 to C2 correspond to the above-mentioned steps B1 to B2, and are not described in detail here.

[0303] C3. Sort the target comments and the initially screened next-level comments to obtain sorted text.

[0304] The process of "sorting the target comments and the initially screened next-level comments" in C3 is the same as that of the above-mentioned steps S1 to S2, and will not be described in detail here.

[0305] C4. If no next-level comments exist in the initial screening of the next-level comments, the sorted text is the target text.

[0306] C5. If there are next-level comments in the initially screened next-level comments, the initially screened next-level comments with next-level comments will be used as new target comments, and the process will jump to step: C2, until there are no next-level comments in the initially screened next-level comments or the number of jumps reaches the set jump threshold, and the sorted text obtained when the jump stops will be recorded as the target text.

[0307] In the above embodiment, the judgment of whether the corresponding next-level comments exist in the initial screening of the next-level comments can be used as a loop condition to obtain the situation of the target comment and the next-level comment of the target comment, as well as the situation of the target comment, the next-level comment of the target comment, and the next-level comment of the next-level comment of the target comment... thereby obtaining the target text composed of multiple levels of reply comments. The end condition of the loop can be that there are no next-level comments in the initial screening of the next-level comments, that is, traversing to the reply comments of the lowest level; or it can be that the number of jumps reaches the set jump threshold, that is, setting the specific level of reply comments to be obtained. For example, if the jump threshold is set to 1, three levels of comments will be obtained, namely: the target comment, the next-level comment of the target comment, and the next-level comment of the next-level comment of the target comment; if the jump threshold is set to 2, four levels of comments will be obtained, namely: the target comment, the next-level comment of the target comment, the next-level comment of the next-level comment of the target comment, and the next-level comment of the next-level comment of the next-level comment of the target comment. It should be understood that the end condition of the loop should not be understood as a limitation to this application.

[0308] C6. Acquire multimedia materials corresponding to the target text.

[0309] C7. Splicing the multimedia materials according to the text order of the target text to generate the comment multimedia.

[0310] Steps C6 to C7 are the same as the above steps B4 to B5, and are not described in detail here.

[0311] In an embodiment of the present application, any of the steps for generating comment multimedia as described above can be repeated to generate multiple comment multimedia corresponding to one original multimedia. The user can switch between multiple comment multimedia through an operation gesture. The operation gesture for switching comment multimedia can be tapping, sliding, dragging, etc. Let's take the operation gesture of sliding down as an example. Optionally, the user can specifically switch the comment multimedia through a sliding operation gesture.

[0312] In the multimedia playback method provided in the embodiment of the present application, a playback interface of the original multimedia can be displayed, and the playback interface also includes a multimedia control for the comment multimedia corresponding to the original multimedia; the comment multimedia corresponding to the multimedia control is played on the playback interface; wherein the playback interface also includes a comment area for the original multimedia, and the comment area for the original multimedia includes comments; the multimedia content of the comment multimedia can be generated based on the comments. In the embodiment of the present application, the content of the comment multimedia is generated based on the comments in the comment area of ​​the original multimedia. By setting the original multimedia and the multimedia control for the comment multimedia in the same playback interface, the association between the original multimedia and the comment multimedia can be strengthened; and by playing the comment multimedia generated by the comments in the comment area on this playback interface, users who watch the original multimedia can understand the comment information in the original multimedia in the form of multimedia, avoiding the problem that users have low patience and low willingness to read text-based comment information.

[0313] The embodiments of the present application are helpful in improving user stickiness.

[0314] This embodiment of the application generates a time-sequenced text script from text comments, reducing the interference of duplicate and invalid comment information and improving the efficiency of review browsing. Viewing comments with large interactive data (i.e., high popularity) and their corresponding follow-up comments in the form of video can enhance the fun of text comments, allowing users to more intuitively understand the content of high-profile comments and maximize their entertainment value.

[0315] In this embodiment, the method of the embodiment of the present application will be described in detail by taking the generation of comment multimedia based on the target comment and the next level comment of the target comment as an example.

[0316] like Figure 2 As shown, a specific process of a multimedia playback method is as follows:

[0317] 201. Based on interaction data of other users on multiple comments in the comment area, select comments whose interaction data exceeds a threshold as target comments; the other users are: for each comment, users other than the user who posted the comment.

[0318] 202. For the next level comments of the target comment, pre-screened next level comments that are strongly correlated with the target comment are screened.

[0319] If the target comment corresponds to only one next-level comment, then the next-level comment can be directly used as the pre-screened next-level comment that has a strong correlation with the target comment.

[0320] If the target comment corresponds to multiple lower-level comments, step 202 may specifically include the following steps 2021 to 2025:

[0321] 2021. If among the multiple next-level comments, there are at least two next-level comments with the same preset length text, the at least two next-level comments are deduplicated to obtain deduplicated result comments corresponding to the at least two next-level comments; wherein the preset length text is text whose length exceeds the set length; the deduplicated result comments and the next-level comments that do not require deduplication are recorded as secondary comments.

[0322] Optionally, in one embodiment, performing deduplication processing on the at least two next-level comments may specifically include the following steps:

[0323] From the at least two next-level comments, obtain at least two next-level comments with the same text of the preset length; from the at least two next-level comments with the same text of the preset length, randomly select one next-level comment, and the randomly selected next-level comment is the comment to be merged; for the at least two next-level comments with the same text of the preset length, extract the other text except the text of the preset length, and add the other text to the comment to be merged.

[0324] 2022. For each of the secondary selected comments, if there is a preset symbol in the secondary selected comment, or the format of the secondary selected comment is a preset format, it is determined that the secondary selected comment has a weak correlation with the target comment.

[0325] 2023. Eliminate the secondary comments that are weakly correlated with the target comment from the plurality of secondary comments, thereby obtaining a result after elimination.

[0326] 2024. Calculate the text relevance between each of the secondary comments in the eliminated results and the target comment.

[0327] 2025. Obtain secondary selected comments whose text relevance exceeds the relevance threshold from the eliminated results. The secondary selected comments whose text relevance exceeds the relevance threshold are: the primary screened next-level comments that have a strong relevance to the target comment.

[0328] 203. Sort the target comments and the initially screened next-level comments to obtain sorted text.

[0329] Optionally, the target comment corresponds to multiple initially screened next-level comments; accordingly, the step of sorting the target comment and the initially screened next-level comments specifically includes the following steps:

[0330] For each of the initially screened next-level comments, the logical relationship coefficient between the initially screened next-level comment and the target comment is calculated to obtain multiple logical relationship coefficients; based on the values ​​of the multiple logical relationship coefficients, the target comment and the multiple initially screened next-level comments corresponding to the target comment are sorted.

[0331] 204. Determine at least one insertion point based on the text content of the sorted text.

[0332] 205. Insert corresponding preset connecting text at each insertion point to obtain the target text.

[0333] 206. Acquire multimedia material corresponding to the target text.

[0334] Optionally, in one implementation, step 206 may specifically include the following steps:

[0335] For each subtext of the target text, it is determined whether there is any video material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value.

[0336] If there is a video material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value, the video material whose similarity exceeds the first preset similarity value is used as the multimedia material corresponding to the subtext.

[0337] If the multimedia material library does not contain any video material whose similarity to the subtext exceeds a first preset similarity value, voice data corresponding to the subtext is generated. The multimedia material library is searched for image material whose similarity to the subtext exceeds a preset similarity value, where the type of the image material is different from the type of the video material. The found image material is combined with the voice data to obtain the multimedia material corresponding to the subtext.

[0338] Optionally, the above step of determining whether there is any video material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value is implemented based on the video material's tag. Accordingly, before the step of determining whether there is any video material in the multimedia material library whose similarity to the subtext exceeds the first preset similarity value, the method may further include:

[0339] Identify the sound in the video material and generate a sound text corresponding to the sound; obtain multiple screenshots of the video material; perform optical character recognition on each of the screenshots to obtain the subtitle text in the screenshot; use the sound text, subtitle text and the multiple screenshots as labels corresponding to the video material, and store them together with the video material in the multimedia material library.

[0340] 207. Splice the multimedia materials according to the text order of the target text to generate the comment multimedia.

[0341] 208. Display a playback interface of the original multimedia, wherein the playback interface further includes a multimedia control of the comment multimedia corresponding to the original multimedia.

[0342] 209. If the multimedia control is triggered, play the comment multimedia corresponding to the multimedia control on the playback interface.

[0343] The playback interface also includes a comment area for the original multimedia, which includes comments, including target comments and next-level comments of the target comments; the multimedia content of the comment multimedia is generated based on the target comments and next-level comments of the target comments.

[0344] The specific execution process of steps 201 to 209 has been described in detail above and will not be repeated here.

[0345] In the multimedia playback method provided in the embodiment of the present application, a playback interface of the original multimedia can be displayed, and the playback interface also includes a multimedia control for the comment multimedia corresponding to the original multimedia; the comment multimedia corresponding to the multimedia control is played on the playback interface; wherein the playback interface also includes a comment area for the original multimedia, and the comment area for the original multimedia includes comments; the multimedia content of the comment multimedia can be generated based on the comments. In the embodiment of the present application, the content of the comment multimedia is generated based on the comments in the comment area of ​​the original multimedia. By setting the original multimedia and the multimedia control for the comment multimedia in the same playback interface, the association between the original multimedia and the comment multimedia can be strengthened; and by playing the comment multimedia generated by the comments in the comment area on this playback interface, users who watch the original multimedia can understand the comment information in the original multimedia in the form of multimedia, avoiding the problem that users have low patience and low willingness to read text-based comment information.

[0346] The embodiments of the present application are helpful in improving user stickiness.

[0347] In order to better implement the above method, the embodiment of the present application also provides a multimedia playback device, which can be integrated into an electronic device, which can be a terminal, a server, etc. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, or a personal computer (PC); the server can be a single server or a server cluster composed of multiple servers. For example, Figure 3 As shown, the device is applied to an electronic device, and the device includes:

[0348] An interface display unit 301 is configured to display a playback interface of the original multimedia, wherein the playback interface also includes multimedia controls for the commentary multimedia corresponding to the original multimedia;

[0349] The comment playing unit 302 is used to play the comment multimedia corresponding to the multimedia control on the playing interface;

[0350] The playback interface further includes a comment area for the original multimedia, and the comment area for the original multimedia includes comments; and multimedia content of the comment multimedia is generated based on the comments.

[0351] In one embodiment, the comments include a target comment and a lower-level comment of the target comment; the apparatus further includes:

[0352] The first generating unit is configured to generate the comment multimedia based on the target comment. In one embodiment, the comment includes a target comment, and the target comment does not have a next-level comment; the apparatus further includes:

[0353] The first generating unit is configured to generate the comment multimedia based on the target comment. In one embodiment, the comment includes the target comment and the next level comment of the target comment; the apparatus further includes:

[0354] The second generating unit is configured to generate the comment multimedia based on the target comment and the next level comment of the target comment.

[0355] In one embodiment, the first generating unit includes:

[0356] a target determination subunit, configured to, based on interaction data of other users on multiple comments in the comment area, select comments whose interaction data exceeds a threshold as target comments, wherein the target comments are target texts; the other users are: for each of the comments, users other than the user who posted the comment;

[0357] A multimedia material subunit, configured to obtain multimedia material corresponding to the target text;

[0358] A first generating subunit is configured to generate the comment multimedia based on the multimedia material when the multimedia material is a piece;

[0359] The second generating subunit is configured to, when there are at least two multimedia materials, splice the multimedia materials in the text order of the target text to generate the comment multimedia.

[0360] In one embodiment, the second generating unit includes:

[0361] a target comment subunit, configured to select comments whose interaction data exceeds a threshold as target comments based on interaction data of other users on multiple comments in the comment area; the other users are: for each comment, users other than the user who posted the comment;

[0362] A correlation screening subunit, configured to screen out the primary screening next-level comments that are strongly correlated with the target comment from the next-level comments of the target comment;

[0363] A sorting subunit, configured to sort the target comments and the pre-screened next-level comments to obtain a target text;

[0364] A multimedia material subunit, configured to obtain multimedia material corresponding to the target text;

[0365] The material splicing subunit is used to splice the multimedia materials according to the text order of the target text to generate the comment multimedia.

[0366] In one embodiment, the sorting subunit includes:

[0367] A sorting text sub-subunit is used to sort the target review and the pre-screened next-level reviews to obtain sorted text;

[0368] an insertion point determination sub-subunit, configured to determine at least one insertion point based on text content of the sorted text;

[0369] The text insertion sub-subunit is used to insert the corresponding preset connecting text at each insertion point to obtain the target text.

[0370] In one embodiment, the second generating unit includes:

[0371] a target comment subunit, configured to select comments whose interaction data exceeds a threshold as target comments based on interaction data of other users on multiple comments in the comment area; the other users are: for each comment, users other than the user who posted the comment;

[0372] A correlation screening subunit, configured to screen out the primary screening next-level comments that are strongly correlated with the target comment from the next-level comments of the target comment;

[0373] A sorting subunit, configured to sort the target review and the pre-screened next-level reviews, thereby obtaining sorted text;

[0374] A first target determination subunit is configured to determine that when no next-level comment exists in the initial screening of next-level comments, the sorted text is the target text;

[0375] The second target determination subunit is used for, when there is a next-level comment in the preliminary screening of the next-level comments, taking the preliminary screening of the next-level comments with the next-level comment as a new target comment, and jumping to the step: for the next-level comments of the target comment, screening out the preliminary screening of the next-level comments that are strongly correlated with the target comment, until there is no next-level comment in the preliminary screening of the next-level comments or the number of jumps reaches a set jump threshold, and recording the sorted text obtained when the jump stops as the target text;

[0376] A multimedia material subunit, configured to obtain multimedia material corresponding to the target text;

[0377] The comment multimedia generating subunit is used to splice the multimedia materials according to the text order of the target text to generate the comment multimedia.

[0378] In one embodiment, the target comment corresponds to a plurality of pre-screened next-level comments; the sorting subunit includes:

[0379] a coefficient calculation subunit, configured to calculate, for each of the pre-screened next-level comments, a logical relationship coefficient between the pre-screened next-level comment and the target comment, thereby obtaining a plurality of logical relationship coefficients;

[0380] The numerical sorting subunit is used to sort the target comment and the multiple initially screened next-level comments corresponding to the target comment based on the numerical values ​​of the multiple logical relationship coefficients.

[0381] In one embodiment, a target comment corresponds to multiple lower-level comments; the associated screening subunit includes:

[0382] The deduplication sub-sub-unit is used for, when at least two of the multiple next-level comments have text of the same preset length, to perform deduplication processing on the at least two next-level comments to obtain deduplication result comments corresponding to the at least two next-level comments; wherein the preset length text is text whose length exceeds the set length; the deduplication result comments and the next-level comments that do not need deduplication processing are recorded as secondary comments;

[0383] a weak correlation sub-subunit, configured to determine, for each of the secondary selected comments, if a preset symbol exists in the secondary selected comment or the format of the secondary selected comment is a preset format, whether the secondary selected comment is weakly correlated with the target comment;

[0384] The strong correlation sub-sub-unit is used to eliminate the secondary selected comments that are weakly correlated with the target comment from the multiple secondary selected comments, thereby obtaining the pre-screened next-level comments that are strongly correlated with the target comment.

[0385] In one embodiment, the strongly associated sub-subunit includes:

[0386] A weak correlation elimination unit is used to eliminate the secondary comments that are weakly correlated with the target comment from the plurality of secondary comments, thereby obtaining a result after elimination;

[0387] a text re-association unit, configured to calculate a text relevance between each of the secondary selected comments in the eliminated results and the target comment;

[0388] The threshold screening unit is used to obtain secondary selected comments whose text relevance exceeds the relevance threshold from the eliminated results. The secondary selected comments whose text relevance exceeds the relevance threshold are: the primary screened next-level comments that have a strong relevance to the target comment.

[0389] In one embodiment, the deduplication sub-unit includes:

[0390] A same text again unit is used to obtain at least two next-level comments containing the same text of the preset length from the at least two next-level comments;

[0391] A unit for merging again is used to randomly select a next-level comment from the at least two next-level comments that have the same text of the preset length, and the randomly selected next-level comment is the comment to be merged;

[0392] The text adding unit is used to extract the other texts except the text of the preset length from at least two lower-level comments that have the same text of the preset length, and add the other texts to the comment to be merged.

[0393] In one embodiment, the multimedia material sub-unit includes:

[0394] A first material determination sub-sub-unit is configured to determine, for each sub-text of the target text, whether there is a first type of multimedia material in the multimedia material library whose similarity to the sub-text exceeds a first preset similarity value;

[0395] The first determination sub-subunit is configured to, if there is a first type of multimedia material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value, use the first type of multimedia material with similarity exceeding the first preset similarity value as the multimedia material corresponding to the subtext.

[0396] In one embodiment, the multimedia material sub-unit further includes:

[0397] a second determination sub-subunit, configured to generate voice data corresponding to the subtext if there is no multimedia material of the first category in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value;

[0398] a second material sub-sub-unit, configured to search the multimedia material library for a second type of multimedia material having a similarity with the subtext exceeding a preset similarity value, wherein the type of the second type of multimedia material is different from the type of the first type of multimedia material;

[0399] The material merging sub-sub-unit is configured to combine the found second-category multimedia material with the voice data to obtain the multimedia material corresponding to the subtext.

[0400] In one embodiment, determining whether there is a first category of multimedia material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value is implemented based on a label of the first category of multimedia material; the first category of multimedia material is a video material; and the apparatus further includes:

[0401] A sound recognition unit, configured to recognize the sound in the video material and generate a sound text corresponding to the sound;

[0402] A screenshot acquisition unit, configured to acquire multiple screenshots of the video material;

[0403] a subtitle text acquisition unit, configured to perform optical character recognition processing on each of the screenshots to acquire the subtitle text in the screenshot;

[0404] The material storage unit is used to store the audio text, subtitle text and the plurality of screenshots as labels corresponding to the video material and store them together with the video material in the multimedia material library.

[0405] In one embodiment, the comment playing unit 302 is specifically configured to play the comment multimedia corresponding to the multimedia control on the playing interface when the multimedia control is triggered.

[0406] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.

[0407] In the multimedia playback method provided in the embodiment of the present application, a playback interface of the original multimedia can be displayed, and the playback interface also includes a multimedia control for the comment multimedia corresponding to the original multimedia; the comment multimedia corresponding to the multimedia control is played on the playback interface; wherein the playback interface also includes a comment area for the original multimedia, and the comment area for the original multimedia includes comments; the multimedia content of the comment multimedia can be generated based on the comments. In the embodiment of the present application, the content of the comment multimedia is generated based on the comments in the comment area of ​​the original multimedia. By setting the original multimedia and the multimedia control for the comment multimedia in the same playback interface, the association between the original multimedia and the comment multimedia can be strengthened; and by playing the comment multimedia generated by the comments in the comment area on this playback interface, users who watch the original multimedia can understand the comment information in the original multimedia in the form of multimedia, avoiding the problem that users have low patience and low willingness to read text-based comment information.

[0408] The embodiments of the present application are helpful in improving user stickiness.

[0409] The present application also provides an electronic device, which may be a terminal, a server, or the like. The terminal may be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, or the like; the server may be a single server or a server cluster consisting of multiple servers, or the like.

[0410] In some embodiments, the multimedia playback device may also be integrated into multiple electronic devices. For example, the multimedia playback device may be integrated into multiple servers, and the multimedia playback method of the present application may be implemented by the multiple servers.

[0411] In this embodiment, the electronic device of this embodiment is an electronic device as an example for detailed description, for example, Figure 4 , which shows a schematic diagram of the structure of the electronic device involved in the embodiment of the present application, specifically:

[0412] The electronic device may include one or more processing core processors 401, one or more computer-readable storage media memories 402, a power supply 403, an input module 404, and a communication module 405. Those skilled in the art will appreciate that Figure 4 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0413] Processor 401 is the control center of the electronic device, connecting all components of the electronic device using various interfaces and circuits. It executes software programs and / or modules stored in memory 402 and accesses data stored in memory 402 to perform various functions and process data. In some embodiments, processor 401 may include one or more processing cores. In some embodiments, processor 401 may integrate an application processor and a modem processor, with the application processor primarily processing the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 401.

[0414] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and multimedia playback by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0415] The electronic device also includes a power supply 403 for supplying power to various components. In some embodiments, the power supply 403 can be logically connected to the processor 401 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 403 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0416] The electronic device may further include an input module 404, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0417] The electronic device may further include a communication module 405. In some embodiments, the communication module 405 may include a wireless module. The electronic device may perform short-range wireless transmission via the wireless module of the communication module 405, thereby providing the user with wireless broadband Internet access. For example, the communication module 405 may be used to help the user send and receive emails, browse web pages, and access streaming media.

[0418] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0419] A playback interface of the original multimedia is displayed, wherein the playback interface also includes a multimedia control for comment multimedia corresponding to the original multimedia; the comment multimedia corresponding to the multimedia control is played on the playback interface; the playback interface also includes a comment area for the original multimedia, wherein the comment area for the original multimedia includes comments; and the multimedia content of the comment multimedia is generated based on the comments.

[0420] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0421] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0422] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the multimedia playback methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:

[0423] A playback interface of the original multimedia is displayed, wherein the playback interface also includes a multimedia control for comment multimedia corresponding to the original multimedia; the comment multimedia corresponding to the multimedia control is played on the playback interface; the playback interface also includes a comment area for the original multimedia, wherein the comment area for the original multimedia includes comments; and the multimedia content of the comment multimedia is generated based on the comments.

[0424] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0425] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations provided in the above embodiments.

[0426] Since the instructions stored in the storage medium can execute the steps in any multimedia playback method provided in the embodiments of the present application, the beneficial effects that can be achieved by any multimedia playback method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0427] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0428] The above is a detailed introduction to a multimedia playback method, device, electronic device and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A multimedia playback method, characterized in that: The method comprises: Displaying a playback interface of the original multimedia, the playback interface also including multimedia controls for the commentary multimedia corresponding to the original multimedia; Play the comment multimedia corresponding to the multimedia control on the playback interface; The playback interface further includes a comment area for the original multimedia, wherein the comment area for the original multimedia includes comments; and multimedia content of the comment multimedia is generated based on the comments.

2. The method according to claim 1, wherein The comments include a target comment and a lower-level comment of the target comment; Before playing the comment multimedia corresponding to the multimedia control on the playback interface, the method further includes: The comment multimedia is generated based on the target comment.

3. The method according to claim 1, wherein The comments include a target comment, and there is no next-level comment for the target comment; Before playing the comment multimedia corresponding to the multimedia control on the playback interface, the method further includes: The comment multimedia is generated based on the target comment.

4. The method according to claim 1, wherein The comments include a target comment and a lower-level comment of the target comment; Before playing the comment multimedia corresponding to the multimedia control on the playback interface, the method further includes: The comment multimedia is generated based on the target comment and the next level comment of the target comment.

5. The method according to claim 2 or 3, wherein: The step of generating the comment multimedia based on the target comment includes: Based on interaction data of other users on multiple comments in the comment area, the comments whose interaction data exceeds a threshold are taken as the target comments, where the target comments are target texts; the other users are: for each of the comments, users other than the user who posted the comment; Acquiring multimedia materials corresponding to the target text; If the multimedia material is one piece, generating the comment multimedia based on the multimedia material; If there are at least two multimedia materials, the multimedia materials are spliced ​​according to the text order of the target text to generate the comment multimedia.

6. The method according to claim 4, wherein The step of generating the comment multimedia based on the target comment and the next level comment of the target comment includes: Based on interaction data of other users on multiple comments in the comment area, taking comments whose interaction data exceeds a threshold as the target comments; the other users are: for each comment, users other than the user who posted the comment; For the next level of comments of the target comment, screening out the first-screened next level comments that have a strong correlation with the target comment; Sorting the target comments and the pre-screened next-level comments to obtain a target text; Acquiring multimedia materials corresponding to the target text; The multimedia materials are spliced ​​according to the text order of the target text to generate the comment multimedia.

7. The method according to claim 6, wherein The step of sorting the target comments and the pre-screened next-level comments to obtain a target text includes: Sorting the target comments and the initially screened next-level comments to obtain sorted text; determining at least one insertion point based on text content of the sorted text; The corresponding preset linking text is inserted at each of the insertion points to obtain the target text.

8. The method according to claim 4, wherein The step of generating the comment multimedia based on the target comment and the next level comment of the target comment includes: Based on interaction data of other users on multiple comments in the comment area, taking comments whose interaction data exceeds a threshold as the target comments; the other users are: for each comment, users other than the user who posted the comment; For the next level of comments of the target comment, screening out the first-screened next level comments that have a strong correlation with the target comment; Sorting the target comments and the initially screened next-level comments to obtain sorted text; If there is no next-level comment in the initial screening of next-level comments, the sorted text is the target text; If there are next-level comments in the initially screened next-level comments, then the initially screened next-level comments with next-level comments will be used as new target comments, and the process will jump to step: for the next-level comments of the target comments, the initially screened next-level comments with strong correlation with the target comments will be screened out, until there are no next-level comments in the initially screened next-level comments or the number of jumps reaches the set jump threshold, and the sorted text obtained when the jump stops will be recorded as the target text; Acquiring multimedia materials corresponding to the target text; The multimedia materials are spliced ​​according to the text order of the target text to generate the comment multimedia.

9. The method according to claim 6 or 8, wherein: The target comment corresponds to multiple preliminary screened next-level comments; The sorting of the target comments and the initially screened next-level comments includes: For each of the first-screened next-level comments, calculating the logical relationship coefficient between the first-screened next-level comment and the target comment, thereby obtaining a plurality of logical relationship coefficients; Based on the values ​​of the multiple logical relationship coefficients, the target comment and the multiple initially screened next-level comments corresponding to the target comment are sorted.

10. The method according to claim 6 or 8, characterized in that The target comment corresponds to multiple next-level comments; The screening of the first-level reviews that are strongly correlated with the target review includes: If there are at least two next-level comments with the same preset length text among the multiple next-level comments, deduplication processing is performed on the at least two next-level comments to obtain deduplication result comments corresponding to the at least two next-level comments; wherein the preset length text is text whose length exceeds the set length; the deduplication result comments and the next-level comments that do not require deduplication processing are recorded as secondary comments; For each of the secondary selected comments, if the secondary selected comment contains a preset symbol or the format of the secondary selected comment is a preset format, then it is determined that the secondary selected comment has a weak correlation with the target comment; From the plurality of secondary selected comments, the secondary selected comments that are weakly correlated with the target comment are eliminated, thereby obtaining the pre-screened next-level comments that are strongly correlated with the target comment.

11. The method according to claim 10, wherein Eliminating the secondary comments that are weakly correlated with the target comment from the plurality of secondary comments, thereby obtaining the pre-screened next-level comments that are strongly correlated with the target comment, includes: Eliminating the secondary comments that are weakly correlated with the target comment from the plurality of secondary comments, thereby obtaining a post-elimination result; Calculating the text relevance between each of the secondary comments in the eliminated results and the target comment; From the eliminated results, secondary selected comments with text relevance exceeding the relevance threshold are obtained, and the secondary selected comments with text relevance exceeding the relevance threshold are: the primary screened next-level comments that have a strong relevance to the target comment.

12. The method according to claim 10, wherein The deduplication processing of the at least two next-level comments includes: Obtaining, from the at least two next-level comments, at least two next-level comments having text of the same preset length; Randomly selecting one next-level comment from the at least two next-level comments that have the same text of the preset length, the randomly selected next-level comment being the comment to be merged; For at least two lower-level comments that have the same text of the preset length, other texts except the text of the preset length are extracted, and the other texts are added to the comment to be merged.

13. The method according to any one of claims 5, 6 and 8, wherein: The acquiring of multimedia material corresponding to the target text includes: For each subtext of the target text, determining whether there is a first type of multimedia material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value; If yes, the first type of multimedia material with a similarity exceeding a first preset similarity value is used as the multimedia material corresponding to the subtext.

14. The method according to claim 13, wherein After determining whether there is a first type of multimedia material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value, the method further includes: If there is no multimedia material of the first category in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value, generating voice data corresponding to the subtext; Searching, from the multimedia material library, for a second type of multimedia material having a similarity with the subtext exceeding a preset similarity value, wherein the type of the second type of multimedia material is different from the type of the first type of multimedia material; The second type of multimedia material found is combined with the voice data to obtain the multimedia material corresponding to the subtext.

15. The method according to claim 13, wherein The determining whether there is a first type of multimedia material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value is achieved based on the label of the first type of multimedia material; the first type of multimedia material is a video material; Before determining whether there is a first type of multimedia material in the multimedia material library whose similarity to the subtext exceeds a first preset similarity value, the method further includes: Identifying the sound in the video material and generating a sound text corresponding to the sound; Obtain multiple screenshots of the video material; Performing optical character recognition on each of the screenshots to obtain subtitle text in the screenshots; The audio text, subtitle text and the plurality of screenshots are used as tags corresponding to the video material and stored together with the video material in the multimedia material library.

16. The method according to claim 1, wherein Playing the comment multimedia corresponding to the multimedia control on the playback interface includes: If the multimedia control is triggered, the comment multimedia corresponding to the multimedia control is played on the playback interface.

17. A multimedia playback device, characterized in that: The device comprises: An interface display unit, configured to display a playback interface of the original multimedia, the playback interface also including multimedia controls for the commentary multimedia corresponding to the original multimedia; A comment playing unit, configured to play the comment multimedia corresponding to the multimedia control on the playing interface; The playback interface further includes a comment area for the original multimedia, and the comment area for the original multimedia includes comments; and multimedia content of the comment multimedia is generated based on the comments.

18. An electronic device, characterized in that: The system comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps in the multimedia playback method according to any one of claims 1 to 16.

19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the multimedia playback method according to any one of claims 1 to 16.

20. A computer program product, characterized in that The method comprises a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the steps in the multimedia playback method according to any one of claims 1 to 16 are implemented.