Comment interaction method and device, equipment and medium

By linking resource access portals to voice comments and playing pre-selected media resources, the problems of monotonous voice comment formats and low user engagement are solved, thereby enriching the comment formats and enhancing user interaction.

CN121967801APending Publication Date: 2026-05-01XINGIN INFORMATION TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XINGIN INFORMATION TECH (SHANGHAI) CO LTD
Filing Date
2025-12-31
Publication Date
2026-05-01

Smart Images

  • Figure CN121967801A_ABST
    Figure CN121967801A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a comment interaction method and device, equipment and a medium, and is applied to the technical field of Internet. The method comprises the following steps: displaying a first voice comment in a comment display area; the first voice comment comprises at least two of first voice data, a voice conversion text corresponding to the first voice data and a voice description tag corresponding to the first voice data; the first voice comment comprises target content and is associated with a resource access entry; in response to a trigger operation for the resource access entry, playing a predetermined media resource in the resource display page; the predetermined media resource at least comprises a predetermined voice resource, and the predetermined voice resource is obtained based on at least one piece of target voice data matched with the target content. By adopting the embodiment of the invention, the interestingness of comments can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Comment interaction methods, devices, equipment and media Technical Field

[0001] This application relates to the field of Internet technology, specifically to the field of data display technology, and in particular to a comment interaction method, apparatus, device, and medium. Background Technology

[0002] Currently, when commenting on media data (such as posts or videos), users can post either text or voice comments. For voice comments, a voice-to-text control (also known as a text conversion control) is typically provided to convert the voice comment into a text comment for display. This makes the comment format for media data relatively limited, and once a user posts a comment, they often do not have the opportunity to participate again, resulting in low user engagement. Summary of the Invention

[0003] This application provides a comment interaction method, apparatus, device, and medium, which can improve the richness and interest of comments on published content.

[0004] On one hand, embodiments of this application provide a comment interaction method, which includes:

[0005] The first voice comment is displayed in the comment display area; the first voice comment includes at least two of the following: first voice data, voice-converted text corresponding to the first voice data, and voice description tags corresponding to the first voice data; the first voice comment includes target content and is associated with a resource access entry;

[0006] In response to a triggered operation on the resource access entry, a pre-defined media resource is played on the resource display page; the pre-defined media resource includes at least a pre-defined audio resource, which is obtained based on at least one target audio data that matches the target content.

[0007] On the other hand, embodiments of this application provide a comment interaction method, which includes:

[0008] In response to the publishing operation for the first voice data, the first voice comment is displayed in the comment display area; the first voice comment includes at least two of the following: the first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data;

[0009] When the first voice comment includes the target content, a predetermined media resource is played; the predetermined media resource includes at least a predetermined voice resource, which is obtained based on at least one target voice data that matches the target content.

[0010] In another aspect, embodiments of this application provide a voice interaction method, the method comprising:

[0011] The first multimedia data is displayed on the content display page; the first multimedia data includes at least two of the following: interactive voice data, voice-converted text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is the first published content or a comment on the first published content; the first multimedia data includes target content and is associated with a resource access entry point;

[0012] In response to a triggered operation on the resource access entry, a pre-defined media resource is played on the resource display page; the pre-defined media resource includes at least a pre-defined audio resource, which is obtained based on at least one target audio data that matches the target content.

[0013] Furthermore, embodiments of this application provide a voice interaction method, the method comprising:

[0014] In response to a publishing operation for the first multimedia data, the first multimedia data is displayed on the content display page; the first multimedia data includes at least two of the following: interactive voice data, voice-converted text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or a comment on the first published content;

[0015] When the first multimedia data includes the target content, a predetermined media resource is played; the predetermined media resource includes at least a predetermined audio resource, which is obtained based on at least one target audio data that matches the target content.

[0016] On one hand, embodiments of this application provide a comment interaction device, the device comprising:

[0017] The comment interaction module is used to display the first voice comment in the comment display area; the first voice comment includes at least two of the following: first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data; the first voice comment includes target content and is associated with a resource access entry;

[0018] The resource playback module is used to play a pre-defined media resource on the resource display page in response to a trigger operation on the resource access entry. The pre-defined media resource includes at least a pre-defined audio resource, which is obtained based on at least one target audio data that matches the target content.

[0019] On the other hand, embodiments of this application provide a comment interaction device, the device comprising:

[0020] The voice publishing module is used to respond to the publishing operation for the first voice data and display the first voice comment in the comment display area; the first voice comment includes at least two of the following: the first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data;

[0021] The resource playback module is used to play a predetermined media resource when the first voice comment includes the target content; the predetermined media resource includes at least a predetermined voice resource, which is obtained based on at least one target voice data that matches the target content.

[0022] In another aspect, embodiments of this application provide a voice interaction device, the device comprising:

[0023] The data display module is used to display first multimedia data on the content display page; the first multimedia data includes at least two of the following: interactive voice data, voice-converted text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or a comment on the first published content; the first multimedia data includes target content and is associated with a resource access entry point;

[0024] The resource playback module is used to play a pre-defined media resource on the resource display page in response to a trigger operation on the resource access entry. The pre-defined media resource includes at least a pre-defined audio resource, which is obtained based on at least one target audio data that matches the target content.

[0025] Furthermore, embodiments of this application provide a voice interaction device, which includes:

[0026] The data publishing module is used to respond to the publishing operation for the first multimedia data and display the first multimedia data on the content display page; the first multimedia data includes at least two of the following: interactive voice data, voice-converted text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or comments on the first published content.

[0027] The resource playback module is used to play a predetermined media resource when the first multimedia data includes target content; the predetermined media resource includes at least a predetermined audio resource, which is obtained based on at least one target audio data that matches the target content.

[0028] On one hand, embodiments of this application provide a computer device including a processor and a memory, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is configured to invoke the program instructions to execute some or all of the steps in the above method.

[0029] On one hand, embodiments of this application provide a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, are used to perform some or all of the steps in the above-described method.

[0030] Accordingly, according to one aspect of this application, a computer program product or computer program is provided, which includes computer instructions that, when executed by a processor, can implement some or all of the steps in the above-described method.

[0031] In this embodiment, a first voice comment can be displayed in the comment display area. When the first voice comment includes target content, it is associated with a first resource access entry. This allows the social media platform to simultaneously provide a resource access entry to users when they post a voice comment or view voice comments posted by other users. This provides users with pre-selected media resources, offering a sense of surprise and enhancing the fun of comment interaction. Furthermore, this method allows users to access pre-selected media resources through the resource access entry. Whether it's their own comment data or comments posted by other users, both can be used to deepen interaction with users based on the media resources provided by the social media platform for voice comments, thereby increasing the richness of comment formats and the fun of interaction. Simultaneously, this can also increase user participation in comments to a certain extent, improving the platform performance of the social media platform. Attached Figure Description

[0032] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 is a schematic diagram of the network architecture of a comment interaction method provided in an embodiment of this application;

[0034] Figure 2 is a schematic diagram of a comment interaction scenario provided in an embodiment of this application;

[0035] Figure 3 is a flowchart illustrating a comment interaction method provided in an embodiment of this application;

[0036] Figure 4 is a schematic diagram of a resource access entry association scenario provided in an embodiment of this application;

[0037] Figure 5 is a schematic diagram of a resource access entry association scenario provided in an embodiment of this application;

[0038] Figure 6 is a schematic diagram of a resource access entry association scenario provided in an embodiment of this application;

[0039] Figure 7 is a schematic diagram of a resource playback scenario provided in an embodiment of this application;

[0040] Figure 8 is a schematic diagram of a voice processing scenario provided in an embodiment of this application;

[0041] Figure 9 is a schematic diagram of a resource triggering scenario provided in an embodiment of this application;

[0042] Figure 10 is a flowchart illustrating another comment interaction method provided in an embodiment of this application;

[0043] Figure 11 is a schematic diagram of the interaction flow of a comment interaction method provided in an embodiment of this application;

[0044] Figure 12 is a schematic diagram of a speech resource generation scenario provided in an embodiment of this application;

[0045] Figure 13 is a schematic diagram of a voice description tag format provided in an embodiment of this application;

[0046] Figure 14 is a schematic diagram of another form of voice description tag provided in an embodiment of this application;

[0047] Figure 15 is a flowchart of a voice interaction method provided in an embodiment of this application;

[0048] Figure 16 is a flowchart of another voice interaction method provided in an embodiment of this application;

[0049] Figure 17 is a schematic diagram of the structure of a comment interaction device provided in an embodiment of this application;

[0050] Figure 18 is a schematic diagram of another comment interaction device provided in an embodiment of this application;

[0051] Figure 19 is a structural schematic diagram of a voice interaction device provided in an embodiment of this application;

[0052] Figure 20 is a schematic diagram of another voice interaction device provided in an embodiment of this application;

[0053] Figure 21 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0054] The comment interaction method proposed in this application is implemented on a computer device, which can be a server or a terminal. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet computer, laptop computer, desktop computer, etc., but is not limited to these.

[0055] In the description of the embodiments of this application, the published content may refer to notes (including text and images), short videos, mid-length videos, etc., that users have pre-posted on social media platforms. Alternatively, it may also include live streams, instant (a short content that is shared and exists in real time), evaluations of products or any type of interest (such as locations, music, etc.) or media data (such as movies, music), comments in notes, bullet comments, topic discussions, points of interest (such as group chats, live streams, etc.), routes (such as cycling routes, travel guides, etc.). The specific content information included in the published content applicable to different scenarios can be adjusted accordingly, and the content information in the published content can be determined by the publisher (or the publishing object). The specific type of published content is not limited here.

[0056] The points of interest mentioned here refer to points of interest that can be displayed on a map. These could include locations, group chats, live streams, etc. There are no specific limitations.

[0057] The term "in response to" indicates a state where a corresponding event occurs or a condition is met. The timing of subsequent actions performed in response to this event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, "in response to XX operation, execute the target step" means that the target step is triggered upon detection of "XX operation".

[0058] For example, in some cases, subsequent actions can be performed immediately when the event occurs or the condition is met; while in other cases, subsequent actions can be performed some time after the event occurs or the condition is met, meaning there can be an intermediate process between responding to the corresponding event or condition and the subsequent actions performed.

[0059] The term "trigger operation" refers to an action performed on certain information (such as a control). This operation sends a corresponding instruction to the computer device, triggering the device to execute the next task. The next task triggered by different trigger operations on different information can all be pre-defined. This trigger operation can be a contact-based operation such as clicking, double-clicking, long-pressing, or swiping, or other preset event operations, or it can be a non-contact operation (such as voice instructions). In some cases, the trigger operation can also be executed by the computer device based on a pre-defined program. No further limitations are specified here.

[0060] Referring to Figure 1, which is a network architecture diagram of a comment interaction method provided in an embodiment of this application, the network architecture may include a server 101 (the number of servers is not limited here) and a cluster of terminal devices (the number of terminal devices is not limited here, such as terminal device 102a, terminal device 102b, ..., terminal device 102n). Communication connections may exist between the servers. Simultaneously, a server may have a communication connection with any terminal device, so that the server can interact with the terminal device through this communication connection. The communication connection is not limited to a specific method; it can be directly or indirectly connected via wired communication, wireless communication, or other methods. This application does not impose any limitations on this method. Furthermore, it is understood that the computer device involved in the embodiments of this application may be the terminal device shown in Figure 1, the server shown in Figure 1, or a system composed of the terminal device and the server shown in Figure 1, which can be referred to as a computer device.

[0061] It should be understood that each terminal device in the terminal device cluster shown in Figure 1 can be equipped with an application client for publishing content. This application client can be of any type, such as a social networking client, an instant messaging client (e.g., a conferencing client), an entertainment client (e.g., a live streaming client), a multimedia client (e.g., a video client), an information client (e.g., a news client), a shopping client, or any other client capable of displaying text, images, audio, and video data. The specific type of application client is not limited here; for ease of description, this application client is referred to as a social media platform.

[0062] For example, an application client refers to a client that can send and receive internet messages in real time and has information functions. Specifically, the application client on the terminal device (terminal device 102a as shown in Figure 1) can respond to a publishing operation for published content by sending a content publishing request to the server, and the server can publish the published content by the target account. Alternatively, the application client on the terminal device can also respond to a comment operation for published content by sending a comment request to the server for that published content, and the server can publish comment data for that published content.

[0063] Optionally, the aforementioned terminal devices and servers can be logically separated. Therefore, when referring to terminal devices and servers below, they may be physically the same device or different devices.

[0064] For ease of understanding, please further refer to Figure 2, which is a schematic diagram of a comment interaction scenario provided by an embodiment of this application. The content display page 200 can be a page where business objects (i.e., users) publish content and provide functions for interacting with that content. The content display page 200 can display the published content and a comment display area 210 for that published content. The published content can be composed of data in any one or more modalities, and can be content published on a social media platform that can be viewed or interacted with by other users (including but not limited to comments, favorites, and reposts). The comment display area 210 can display comment data from various business objects regarding the published content. Any comment data can be unimodal or multimodal. Unimodal data refers to data of one modality, such as text comments, voice comments, and image comments. Multimodal data refers to data composed of multiple modalities, such as comment data composed of text and voice comments, comment data composed of text and image comments, or comment data composed of image and voice comments, etc., without limitation. For example, in Figure 2, comment data 1 published by publisher 1 includes the text comment "Happy New Year" and voice comment 1. Similarly, comment data 2 published by publisher 2 includes the text comment "Time flies!", and comment data 3 published by publisher 2 includes voice comment 3, etc. Any voice comment can include at least two of the following: voice data, the corresponding voice-to-text conversion, and the corresponding voice description tag. For example, voice comment 1 includes voice data 201 and voice description tag 202, etc., and voice comment 3 includes voice data 2011 and the corresponding voice-to-text conversion 2012, etc. Of course, the voice-to-text conversion corresponding to a voice data is the text data obtained after processing the voice data into a text-to-text format.

[0065] Optionally, any comment data on the content display page 200 may be associated with at least one of the following data: the object data that posted the comment data (such as account avatar and / or account name, etc.), the comment description information of the comment data, and a comment reply control, etc. Taking the comment data posted by object 1 as an example, the object data includes at least one of the following: account avatar, account name "object 1", etc.; the comment description information 2041 is used to represent the relevant information of the corresponding comment data, and may include at least one of the following: posting time, geographical location information, etc.; the comment reply control 2042 is used to trigger a reply operation for the corresponding comment data.

[0066] Specifically, the computer device can display a first voice comment in the comment display area 210. When the first voice comment includes target content, the first voice comment is associated with a resource access entry. The resource access entry can be explicit, that is, the resource access entry is directly displayed in the comment display area 210 for the first voice comment. The resource access entry can also be implicit, that is, the resource access entry is directly associated with the first voice comment. In response to the trigger operation 203 for the resource access entry, the computer device can play a predetermined media resource in the resource display page 205. The predetermined media resource includes at least a predetermined voice resource 206. The predetermined voice resource 206 is obtained based on at least one target voice data that matches the target content. The target content is preset content used to trigger the media resource.

[0067] Therefore, the technical solution of this application can provide a resource access portal for media resources when a voice comment is published in the comment display area. This enables interaction between the social media platform and users, enriches the platform performance of the social media platform, and diversifies the comment formats for published content. Moreover, the diversity of comment formats ensures that users still have interactive functions (i.e., resource access portals) after publishing comment data. This allows users to further participate in the interaction with comment data when publishing or viewing it, thereby increasing user participation in published content and comment data, and thus improving the interactivity and performance of the social media platform.

[0068] In the specific implementation of this application, the triggering operation for any control can be any of the following operations: click operation, double-click operation, swipe operation, long press operation, etc. Of course, the above are only some possible triggering operations. For the purpose of implementation, the triggering operation can also be other operations, such as voice triggering operation, etc., which are not limited here.

[0069] In specific embodiments of this application, scenarios involving the acquisition of user information and related data, such as the acquisition of user information (e.g., published content and comment data), require user permission or consent. That is, when these embodiments are applied to specific products or technologies, the collection, use, and processing of relevant user data comply with the relevant laws, regulations, and standards of the relevant regions. For example, interactive pages can be used to provide prompts indicating which data will be collected or acquired. Specifically, lists or other methods can be used to display the types and content of this data to the user. Data collection and processing will only proceed after a confirmation or instruction to allow data collection is received on the interactive page.

[0070] The scenarios described above are merely examples and do not constitute a limitation on the application scenarios of the technical solutions provided in this application. The technical solutions of this application can also be applied to other scenarios. For example, as those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0071] Based on the above description, this application proposes a comment interaction method, which can be executed by the aforementioned computer device, specifically the terminal device shown in Figure 1, a server, or a system composed of a terminal device and a server. Please refer to Figure 3, which is a flowchart illustrating a comment interaction method provided in this application embodiment. As shown in Figure 3, the flow of the comment interaction method in this application embodiment may include the following:

[0072] S101. Display the first voice comment in the comment posting area; the first voice comment includes the target content and is associated with a resource access entry.

[0073] The comment display area is associated with the published content and is used to publish or display comment data related to that content. The published content can be multimedia data published by a business entity (i.e., a user of the social media platform). This multimedia data can be unimodal or multimodal, such as text data, video data, or data composed of text and audio data, etc., without limitation. The comment data published on this content can also be unimodal or multimodal. The first audio comment can include at least two of the following: first audio data, the corresponding speech-to-text of the first audio data, and a speech description tag corresponding to the first audio data. The speech description tag can be used to represent semantically relevant information of the first audio data.

[0074] The computer device can display a first voice comment in the comment display area, and when the first voice comment includes target content, associate a resource access entry for the first voice comment. Optionally, the resource access entry can be associated with the speech-to-text corresponding to the first voice data, or with the speech description tag corresponding to the first voice data, etc., without limitation. Optionally, a display resource access entry can also be associated with the first voice comment. The first voice data can be a complete comment data, such as voice data 2011 shown in Figure 2, or voice data in a multimodal comment data, such as voice data 201 shown in Figure 2. Optionally, the computer device can respond to a comment posting request for the first comment data and display the first comment data in the comment display area. When the first comment data includes the first voice data, it can synchronously display the speech-to-text corresponding to the first voice data and at least one of the speech description tags corresponding to the first voice data, and associate a resource access entry for the first voice comment in the first comment data when the first voice data includes target content. The resource access entry is determined based on the target content, which can be preset content, and the number of target content can be one or more. For example, the target content can be a preset keyword, or content semantically corresponding to the preset keyword, such as "Happy New Year" or "Happy New Year's Day." For instance, if the preset keyword is "Happy New Year," the target content can include, but is not limited to, "Happy New Year," "Happy New Year" in various other languages ​​(e.g., Happy New Year), and audio versions of "Happy New Year." The target content can also be content semantically corresponding to a target audio type. Target audio types can include, but are not limited to, any one or more of the following: blessing types, holiday types, welfare types, etc. Of course, the number and specific types of target audio types can be configured as needed. Optionally, when there are multiple target audio types, a single audio comment may include multiple target contents, or it may not include any single target content.

[0075] In this embodiment, the computer device can synchronously associate a resource access entry with the first voice comment when it is published, as shown in Figures 4 and 5. For example, Figure 4 is a schematic diagram of a resource access entry association scenario provided by an embodiment of this application. As shown in Figure 4, the computer device can display a content display page 400, which includes a comment display area 40a. Assuming that the first voice comment is published by a target object associated with the computer device, the comment display area 40a may include, but is not limited to, a comment input area 401 and a voice input control 402. In response to the publishing operation of the first voice data 404, the computer device can display the first voice comment in the comment display area 40a. The first voice comment may include at least two of the following: the first voice data 404, the voice-converted text corresponding to the first voice data 404, and the voice description tag 405 corresponding to the first voice data 404. The first voice comment includes target content and is associated with a resource access entry. Optionally, the publishing operation for the first voice data 404 may include a trigger operation 403 for the voice input control 402, and a publishing operation for the first voice data collected based on the trigger operation 403 for the voice input control 402.

[0076] For example, when a first voice comment is published, the computer device can simultaneously display the first voice-converted text obtained from the voice comment. Specifically, the computer device can display the first voice data in the comment display area, and also display the voice-converted text obtained by converting the first voice data into text. It can also display the voice description tags corresponding to the first voice data. That is, the first voice comment includes the first voice data, the voice-converted text corresponding to the first voice data, and the voice description tags corresponding to the first voice data, etc. When the voice-converted text indicates that the first voice comment includes target content, it serves as a resource access entry point associated with the first voice comment. For example, see Figure 5, which is a schematic diagram of a resource access entry association scenario provided by an embodiment of this application. As shown in Figure 5, the computer device can display a content display page 500, which includes a comment display area 50a. Assuming the first voice comment is published by a target object associated with the computer device, the computer device can respond to a publishing operation 501 on the first voice data 502 and display the first voice comment in the comment display area 50a. The first voice comment includes the first voice data 502, the voice-converted text 503 obtained by converting the first voice data 502 into text, and the voice description tag 504 corresponding to the first voice data 502, and associates a resource access entry for the first voice comment. Optionally, the voice description tag 504 can be displayed side by side with the first voice data 502, or it can be displayed side by side with the voice-converted text 503 (as shown in Figure 5). There is no limitation here. That is to say, the first voice data 502, the voice-converted text 503, and the voice description tag 504 are displayed in association, but there is no limitation on the relative position of the three in this application.

[0077] Alternatively, the computer device can display the corresponding voice-converted text and / or voice description tags after the text conversion of the first voice data is completed when the first voice data is published. At this time, the computer device can display the first voice data in the comment display area and display the voice-converted text obtained from the text conversion of the first voice data; this process can be actively triggered by the user. Specifically, the computer device can display the first voice data in the comment display area and display a text conversion control associated with the first voice data; in response to a trigger operation on the text conversion control, it displays the voice-converted text obtained from the text conversion of the first voice data. For example, see Figure 6, which is a schematic diagram of a resource access entry association scenario provided by an embodiment of this application. As shown in Figure 6, the computer device can display a content display page 600, which includes a comment display area 60a. In response to a publishing operation 601 for the first voice data 602, the computer device can display the first voice data 602 in the comment display area 60a and associate a text conversion control 603 with the first voice data 602. In response to a trigger operation on the text conversion control 603, the computer device can display the voice-converted text 604 obtained by converting the first voice data 602 into text, as well as the voice description tag 605 corresponding to the first voice data 602. When the voice-converted text 604 indicates that the first voice data 602 includes target content, the computer device can associate a resource access entry for the first voice comment. This resource access entry can be associated with the voice-converted text 604 or the voice description tag 605, etc. Optionally, the voice description tag 605 can be displayed side by side with the first voice data 602, as shown in the comment display area 60c; or, the voice description tag 605 can also be displayed side by side with the voice-converted text 604, as shown in the comment display area 60d, without any limitation.

[0078] In determining the specific implementation of the resource access entry point, the first voice comment includes first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data. Specifically, the computer device can display the first voice data and the voice-converted text obtained by converting the first voice data into text in the comment display area. Alternatively, the first voice data and a text conversion control for the first voice data can be displayed in the comment display area; in response to a trigger operation on the text conversion control, the first voice data is converted into text to obtain the voice-converted text. That is to say, in this embodiment, when publishing the first voice data, the first voice data can be directly converted into text and the first voice data and the voice-converted text can be displayed simultaneously, or only the first voice data can be displayed, and when a text conversion operation on the first voice data is detected, the voice-converted text can be displayed in association with the first voice data. Optionally, a voice description tag can also be displayed in association with the first voice data. Further, the first voice data includes target content and is associated with a resource access entry point. The target content can be at least one of the following: preset keywords, content semantically corresponding to the preset keywords (i.e., the predetermined semantics and / or predetermined speech corresponding to the predetermined keywords), and content semantically corresponding to the target speech type (i.e., the predetermined type). For example, a computer device can perform semantic matching between the speech-to-text and the target speech type to obtain the semantic similarity between the speech-to-text and the target speech type. If the semantic similarity is greater than or equal to the resource trigger similarity threshold, then the predetermined media resource corresponding to the target speech type is obtained, and the resource access entry bound to the predetermined media resource is associated with the first speech comment, or the resource access entry bound to the predetermined media resource is associated with the first speech comment.

[0079] Optionally, the first voice comment includes target content, comprising at least one of the following: any one of the first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data includes the target content; and / or, any one of the first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data matches the target content. Wherein, "matches" as mentioned in this application includes content matching and / or semantic matching, that is, the part consistent with the target content can be considered to match the target content, and the part consistent with the target content after translation or voice conversion can also be considered to match the target content. For example, for the target content "Happy New Year", the Chinese "Happy New Year" voice, the English "Happy New Year" voice, etc., can all be considered to match the target content "Happy New Year".

[0080] Optionally, when there are multiple target speech types, the computer device can perform semantic matching between the first speech-converted text and the multiple target speech types respectively to obtain the semantic similarity between the first speech-converted text and the multiple target speech types respectively, and determine the target speech type with a semantic similarity greater than or equal to the resource triggering similarity threshold as the associated speech type of the first speech comment, and associate the resource access entry for the first speech comment.

[0081] In this embodiment of the application, the comment display area can be displayed independently on the content display page or displayed within the content display page; no limitation is made here.

[0082] Therefore, through the various display methods described above, media resources for voice comments can be flexibly pushed, allowing business users to access media resources provided by social media platforms when publishing or viewing voice comments. This attracts business users to further interact with the comment data, thereby improving the interactivity and interest of the comment data.

[0083] S102. In response to a trigger operation on the resource access entry, play the scheduled media resource on the resource display page.

[0084] The predetermined media resources include at least predetermined audio resources, which are obtained based on at least one target audio data that matches the target content. Optionally, the predetermined media resources may also include, but are not limited to, animation resources, resource text of the predetermined audio resources, comment description information, resource auxiliary images, resource display backgrounds, etc. Among them, resource text refers to text data corresponding to a predetermined voice resource, such as resource text 207 in Figure 2; comment description information is used to describe the timing of providing media resources for the first voice comment, such as comment description information 208 in Figure 2. For example, a comment description template and the parameters to be filled included in the comment description template can be obtained, such as the comment description template "You are the yth person to say XX today", which includes the parameter to be filled "y" (indicating which user triggered the comment within a natural day, determined according to the actual triggering order of the user) and the parameter to be filled "XX" (used to represent the target content). The actual data corresponding to the parameters to be filled can be obtained, and the comment description template and the actual data corresponding to the parameters to be filled can be combined to form the comment description information; resource auxiliary image refers to an image or other image that is semantically related to the resource text, such as the resource text 207 "Happy New Year" shown in Figure 2, where the resource auxiliary image can be fireworks, etc.; resource display background refers to the page background of the resource display page. The resource display background can be a background independent of the content display page, or a mask added on top of the content display page, etc., without any restrictions. Optionally, the transparency of the resource's display background can be any value, or it can be adjusted by the business entity. In simpler terms, this pre-booked media resource can be considered a hidden surprise or bonus offered by social media resources, such as a perk or blessing.

[0085] Specifically, the triggering operation for the resource access entry can be triggered by a triggering operation at any location in the area where the first voice comment is associated with the resource access entry; or, when the resource access entry is explicit, the triggering operation for the resource access entry can be triggered by a triggering operation for the resource access entry displayed in the comment display area.

[0086] For example, if the reserved media resources include reserved audio resources, the computer device can respond to a trigger operation on the resource access entry and directly play the reserved media resources on the resource display page.

[0087] For example, the predetermined media resources also include animation resources that match the target content. For instance, if the target content is "Happy New Year," the animation resource could be an animation of fireworks; if the target content is "Happy Lantern Festival," the animation resource could be a Lantern Festival blessing animation, etc., without limitation. In this case, the computer device can respond to a trigger operation on the resource access entry point by simultaneously playing the animation resources and predetermined audio resources on the resource display page, or by sequentially playing the animation resources and predetermined audio resources on the resource display page. For example, the predetermined audio resources can be used as background sound for the animation resources on the resource display page to play the predetermined media resources. Also, referring to Figure 7, which is a schematic diagram of a resource playback scenario provided by an embodiment of this application, as shown in Figure 7, the computer device can display a first audio comment on the content display page 700. The first audio comment includes first audio data 701 and audio description tags 702, etc. Responding to a trigger operation 703 on the resource access entry point, the computer device can play the animation resource 7041 in the predetermined media resources on the resource display page 704. The animation resource 7041 can be video data or dynamic images, etc. Furthermore, when the animation resource finishes playing, the predetermined audio resource in the predetermined media resource is played, as shown in playback mode ① of Figure 7. When the animation resource 7041 finishes playing, the predetermined audio resource 705 in the predetermined media resource is played on the resource display page 704a. Alternatively, when the animation resource finishes playing, the resource text for the predetermined audio resource and the audio playback control are displayed. In response to a trigger operation on the audio playback control, the predetermined audio resource is played, as shown in playback mode ② of Figure 7. When the animation resource 7041 finishes playing, the resource text "Happy New Year" for the predetermined audio resource and the audio playback control 706 are displayed on the resource display page 704b. In response to a trigger operation 707 on the audio playback control 706, the predetermined audio resource is played.

[0088] The pre-booked media resources can include at least pre-booked audio resources. Leveraging the fact that sound conveys emotion more effectively than text, these resources can offer better emotional expression, thus enhancing the user experience. Furthermore, additional components, such as animation, can be added to the pre-booked media resources as needed to enrich their expressive forms and improve their presentation, ultimately providing a better user experience. Therefore, media resources can incentivize users to post more audio comments.

[0089] Optionally, in any of the above resource playback scenarios, the pre-selected audio resource may be provided directly by a social media platform.

[0090] Alternatively, in any of the above resource playback scenarios, the predetermined audio resource can be obtained based on at least one target audio data that matches the target content. The target audio data can be at least one of the following: a portion of second audio data that matches the target content, data synthesized based on the target content and target voice features, or data synthesized based on the target content. The second audio data belongs to a second audio comment that includes the target content. This second audio comment is a user-posted audio comment, and the user can include the poster of the first audio comment and / or other business objects (i.e., users other than the poster). For example, the second audio comment may include the first audio comment. The target voice features originate from at least one of the following: historical audio data posted by the poster of the first audio comment, platform audio data provided by the social media platform where the first audio comment is located, or third audio data posted by the business object. Platform audio data refers to data packets provided by the social media platform, such as youthful voices, mature female voices, gentle voices, etc.; the business object refers to the user using the social media platform.

[0091] Specifically, when the target voice features originate from the first voice data published by the object that published the first voice comment, the predetermined voice resources carried by the first voice comment and the first voice comment can appear as if they were published by the same business object, thereby incentivizing the business object to publish more voice comments. When the target voice features originate from the second voice data published by the business object that authorized voice collection, the business object, after publishing a voice comment, can hear voice resources as if they were published by other objects, allowing the media resources to convey more emotion. Especially when the predetermined media resource is a blessing, the user feels as if they are hearing blessings from other objects, thus enriching the emotional expression of the media resources.

[0092] Optionally, when the target voice features originate from platform voice data provided by the social media platform that published the first voice comment, the computer device can also display N candidate voice packets in response to a voice playback configuration operation at the first resource access entry; N is a positive integer. In response to a selection operation of the target voice packet among the N candidate voice packets, a predetermined voice resource possessing the target voice features of the target voice packet is synthesized. The voice playback configuration operation is used to invoke a voice playback configuration function, which can be provided at any one or more of the following locations: the overall platform configuration location of the social media platform (the data collected by the function provided here can be applied to the entire social media platform), the resource access entry, etc.

[0093] For example, please refer to Figure 8, which is a schematic diagram of a voice processing scenario provided by an embodiment of this application. As shown in Figure 8, the computer device can display a content display page 800, which includes a comment display area. A first voice comment can be displayed in the comment display area. The first voice comment includes first voice data 801 and a voice description tag 802. In response to a voice playback configuration operation for a resource access entry, the computer device can display N candidate voice packets 803, such as candidate voice packet 1 and candidate voice packet 2. In response to a selection operation for a target voice packet among the N candidate voice packets 803 (as shown in Figure 8, assuming the target voice packet is candidate voice packet 2), a predetermined voice resource 805 with the target sound features of the target voice packet is synthesized. The computer device can play the predetermined voice resource 805 in the resource display page 804.

[0094] Optionally, the number of target speech data is M, where M is a positive integer. The computer device can sequentially arrange the M target speech data to obtain an audio sequence; in the audio sequence, there are overlapping portions between adjacent target speech data; the audio sequence is then synthesized to obtain a predetermined speech resource with the overlapping portions cross-faded. This makes it seem as if the user is receiving voice greetings from multiple users simultaneously when playing the predetermined speech resource, providing an immersive experience and enhancing the expressive effect of the media resource.

[0095] Optionally, the system can also provide business users with the option to choose whether to provide resources for voice comments. Specifically, in response to a publishing operation for a third voice comment, the computer device can display a resource trigger prompt message when the third voice comment includes the target content; in response to a confirmation operation for the resource trigger prompt message, the computer can display the third voice comment in the comment display area and associate a resource access entry with the third voice comment; the resource access entry is used to obtain a pre-defined media resource that matches the target content.

[0096] For example, please refer to Figure 9, which is a schematic diagram of a resource triggering scenario provided by an embodiment of this application. As shown in Figure 9, in the content display page 900, in response to the collection and publishing operation 902 for the third voice comment, when the third voice comment includes target content, the computer device displays a resource triggering prompt message 903. When responding to the trigger confirmation operation ① for the resource triggering prompt message 903, the third voice comment is displayed in the comment display area 901. The third voice comment includes voice data 904 and voice description tags 905, and is associated with a resource access entry, as shown in the comment display area 90a. When responding to the trigger cancellation operation ② for the resource triggering prompt message 903, the third voice comment, including the voice data 904 in the third voice comment, is displayed in the comment display area 901, as shown in the comment display area 90b.

[0097] By providing the user-selectable resources, the autonomy of business users can be enhanced, giving them greater freedom to use the platform and thus improving the user experience. Of course, since the primary purpose of this embodiment is to provide media resources to users, the social media platform may not necessarily provide the functionality shown in Figure 9.

[0098] The first or second voice comment mentioned above can be any voice comment. That is, when any voice comment is published or displayed, the publishing and display process for the first or second voice comment can be referred to.

[0099] Optionally, after associating a resource access portal with the first voice comment, the pre-defined media resource bound to the resource access portal can be fixed. That is, after the first voice comment is successfully published, a resource access portal is associated with the first voice comment and a pre-defined media resource is bound to it. At this time, as long as the resource access portal associated with the first voice comment is triggered, the pre-defined media resource bound to the resource access portal can be played.

[0100] Alternatively, the predetermined media resource bound to the resource access entry associated with the first voice comment can also be generated in real time. That is, the predetermined media resource played may be different when the resource access entry associated with the first voice comment is triggered at different times. In other words, in response to the triggering operation of the resource access entry, the computer device acquires at least one target voice data that matches the target content, and synthesizes the at least one target voice data into a predetermined voice resource. This process can be referred to in other embodiments for the acquisition and generation process of target voice data and predetermined voice resources, and will not be described in detail here. For example, one possible method for real-time generation of predetermined voice resources involves a computer device acquiring a main audio segment (denoted as segment A). This main audio segment refers to second voice data including target content. The portion of the second voice data that matches the target content is extracted as the main audio segment. This main audio segment is equivalent to target voice data obtained through real-time processing. This real-time processing may include at least one of the following: voice data segmentation, normalization processing, voice fade-in / fade-out effect processing, etc. Optionally, the second voice data may be obtained from voice data published by the object of the first voice comment, or from voice data published by other users (i.e., users other than the object of the first voice comment). Sub-audio segments are obtained from the main audio segment and spliced ​​together to obtain a first audio segment (which can be denoted as segment B1). Some sub-audio segments in this first audio segment are obtained through pre-embedded processing. The system acquires target speech data with superimposed audio duration, preprocesses and superimposes the acquired target speech data with superimposed audio duration to obtain a second audio segment (which can be denoted as segment B2). This second audio segment can be considered a chorus, obtained by pre-embedding all audio segments. This second audio segment can be associated and stored through the resource access portal associated with the social media platform and the first voice comment. There can be one or more second audio segments. Therefore, if a second audio segment already exists, this process can be skipped, and the computer device can directly acquire the already stored second audio segment. When generating the predetermined speech resource, the computer device can randomly load the second audio segment to obtain a preloaded audio segment. The first audio segment and the preloaded audio segment are then audio-synthesized to obtain a pre-embedded audio segment (which can be denoted as segment B, and can be simply denoted as segment B = segment B1 + part or all of segment B2). The main audio segment and the pre-embedded audio segment are then audio-synthesized to obtain the predetermined speech resource, and the predetermined media resource including the predetermined speech resource is played. At this time, the predetermined speech resource can be considered as segment A + segment B. Optionally, since excessively long audio durations may cause user fatigue, the duration of each audio segment can be limited, such as segment A being 0.6-2.5 seconds, segment B1 being around 4 seconds, and segment B2 being 1-1.8 seconds (i.e., the combined audio duration), etc. The specific duration can be configured according to needs.

[0101] In this embodiment, a first voice comment can be displayed in the comment display area, and a first resource access entry can be associated with the first voice comment. This allows the social media platform to simultaneously provide a resource access entry (referred to here as the first resource access entry) to the user when the user posts a voice comment or views voice comments posted by other users. This provides the user with media resources, offering a sense of surprise and enhancing the fun of comment interaction. Furthermore, this method allows users to access media resources through the first resource access entry. Whether it's their own comment data or comments posted by other users, both can be used to deepen the interaction with users based on the media resources provided by the social media platform for voice comments, thereby increasing the richness of comment formats and the fun of interaction. Simultaneously, this can also increase user participation in comments to a certain extent, improving the platform performance of the social media platform.

[0102] Based on the above description, this application proposes a comment interaction method, which can be executed by the aforementioned computer device, specifically the terminal device shown in Figure 1, a server, or a system composed of a terminal device and a server. Please refer to Figure 10, which is a flowchart illustrating another comment interaction method provided by this application embodiment, used to describe the comment interaction process when displaying speech-converted text obtained from speech comments. As shown in Figure 10, the flow of the comment interaction method of this application embodiment may include the following:

[0103] S201. In response to the publishing operation for the first voice data, the first voice comment is displayed in the comment display area.

[0104] The first voice comment includes at least two of the following: first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data. However, for user convenience, the first voice comment typically includes at least the first voice data and the voice-converted text corresponding to the first voice data, and optionally, it may also include the voice description tag corresponding to the first voice data. When the first voice comment includes target content, a resource access entry can also be associated with the first voice comment. For a detailed implementation of step S201, please refer to the relevant descriptions in the above embodiments, which will not be repeated here.

[0105] S202, when the first voice comment includes the target content, play the predetermined media resource; the predetermined media resource includes at least the predetermined voice resource.

[0106] The predetermined voice resources are obtained based on at least one target voice data that matches the target content. Specific implementation details of this process can be found in the relevant descriptions of the above embodiments. The target voice data can be found in the relevant descriptions of S102 above. Optionally, the second voice comment can be restricted to be posted by a user (here referring to the user who posted the first voice data) and / or other users followed by that user, with the target audience being that user and / or other users followed by that user. This makes the predetermined media resources more consistent with the sounds users are usually exposed to, providing users with a more emotionally resonant and immersive experience, thereby improving the performance of the social media platform.

[0107] Optionally, before playing the scheduled media resource, the trigger status of the first voice data publisher (i.e., the user) for the scheduled media resource can be obtained; the playback of the scheduled media resource can be controlled based on the trigger status. The trigger status indicates whether the user has already played the scheduled media resource. If the trigger status is "triggered," the scheduled media resource may not be played; if the trigger status is "not triggered," the scheduled media resource will be played.

[0108] In this embodiment, a first voice comment can be displayed in the comment display area, and a first voice-converted text obtained from the first voice comment can be displayed synchronously. When a user publishes the first voice data, a pre-reserved media resource can be played directly for the user, so that when the user publishes the first voice comment containing the target content, the user can get the corresponding response in real time, i.e., the pre-reserved media resource, thereby improving the user experience, that is, the user will have a sense of satisfaction from being able to get a real-time reply.

[0109] Based on the above description, this application proposes a comment interaction method, which can be executed by the aforementioned computer device, which may include a front-end and a back-end. Please refer to Figure 11, which is a schematic diagram of the interaction flow of a comment interaction method provided in this application embodiment. As shown in Figure 11, the interaction flow of the comment interaction method in this application embodiment may include the following:

[0110] S1101. The front-end displays the content display page, which shows the published content; the content display page includes a comment display area.

[0111] The process can be referred to in the relevant description shown in S101 of Figure 3, and will not be repeated here.

[0112] S1102, The front end publishes the first voice data in the comment display area.

[0113] The front-end can respond to the publishing operation of the first voice data and can send a publishing request for the first voice data to the back-end.

[0114] S1103. The background process performs text conversion on the first voice data to obtain the voice-converted text.

[0115] The backend can perform text conversion on the first voice data to obtain voice-converted text. Furthermore, it can perform content recognition on the first voice data and / or the corresponding voice-converted text to obtain voice description tags corresponding to the first voice data. The first voice data, the corresponding voice-converted text, and the corresponding voice description tags together constitute the first voice comment.

[0116] S1104. The backend parses the content of the first voice comment and checks whether the first voice comment includes the target content.

[0117] The backend can parse the first voice comment to detect whether it includes the target content. This target content can be at least one of the following: preset keywords, content semantically corresponding to the preset keywords (i.e., the predetermined semantics and / or predetermined voice corresponding to the predetermined keywords), or content semantically corresponding to the target voice type (i.e., a predetermined type).

[0118] For example, if the target content is a predetermined keyword, the computer device can search for the target content from the first voice data, the speech-to-text corresponding to the first voice data, and the speech description tags corresponding to the first voice data. In this case, the first voice comment includes the target content, including at least one of the following situations: the target content is included in any one of the first voice data, the speech-to-text corresponding to the first voice data, and the speech description tags corresponding to the first voice data; and / or, a segment semantically matching the target content exists in any one of the first voice data, the speech-to-text corresponding to the first voice data, and the speech description tags corresponding to the first voice data.

[0119] For example, if the target content is a predetermined semantic and / or predetermined speech, the computer device can acquire the target language corresponding to the target content and the corresponding text data, and translate the speech-to-text data into standardized data in the target language; then, it matches the standardized data with the text data corresponding to the target content. In this case, the first speech comment includes the target content, including at least one of the following: the standardized data includes the text data corresponding to the target content; and / or, the standardized data includes segments that semantically match the target content.

[0120] For example, the target content is a predetermined type. The number of target speech types can be one or more. When there is only one target speech type, semantic matching can be performed between the speech-to-text and the target speech type to obtain the semantic similarity between them. If the semantic similarity is greater than or equal to a resource-triggered similarity threshold, the target speech type is determined as an associated speech type semantically similar to the speech-to-text, and step S1105 is executed. When there are multiple target speech types, semantic matching can be performed between the first speech-to-text and each of the multiple target speech types to obtain the semantic similarity between the speech-to-text and each of the multiple target speech types. Target speech types with semantic similarity greater than or equal to a resource-triggered similarity threshold are determined as associated speech types semantically similar to the speech-to-text, and step S1105 is executed. Further, this process can also be referred to the relevant description in the above embodiments.

[0121] In any of the above scenarios, if the semantic similarity between the speech-to-text and the target speech type is less than the resource-triggered similarity threshold, then the speech-to-text corresponding to the first speech data is sent to the front-end. The front-end can then display the first speech comment in the comment display area.

[0122] Optionally, resource voice types can be provided in different time periods. One possible resource database is shown in Table 1 below:

[0123] Table 1

[0124]

[0125] As shown in Table 1, the backend can obtain the system network time, which is the time when the request to publish the first voice comment was received; it can also obtain the target time period to which the system network time belongs, and determine the resource voice type associated with the target time period as the target voice type. The resource database can be any data format that can represent the relationships between data, such as tables, structured databases, etc.

[0126] Of course, when the target content is predetermined keywords, semantics, or audio, the predetermined keywords may differ for different time periods. For example, the predetermined keywords for the Chinese New Year period might be "Happy New Year" or "Good Luck," while for New Year's Day it might be "Happy New Year" or "Good Luck." In short, social media platforms can configure different target content for different time periods, allowing the predetermined media resources provided by the platform to adapt to the current time period requirements, thus improving the flexibility and applicability of media resource provision.

[0127] S1105, when the first voice comment includes target content, acquire target voice data that matches at least one target content.

[0128] Specifically, target speech data matching the target content can be obtained, as described in the relevant descriptions in the above embodiments. This target speech data can be at least one of the following: a portion of the second speech data matching the target content, data synthesized based on the target content and target sound features, or data synthesized based on the target content. For example, if the target content is "Happy New Year," the second speech data could be the audio segment corresponding to "Happy New Year" in the speech "It's the beginning of another year, wishing everyone a Happy New Year!" For example, see Figure 12, which is a schematic diagram of a speech resource generation scenario provided in an embodiment of this application. As shown in Figure 12, second speech data 1201, second speech data 1202, and second speech data 1203 are obtained. A portion matching the target content is obtained from each of the second speech data, resulting in at least one target speech data, including target speech data 12a, target speech data 12b, and target speech data 12c.

[0129] For example, a first initial resource corresponding to the target content can be obtained, and the target content or the first initial resource corresponding to the target content can be synthesized with the target sound features to obtain the target speech data. Optionally, the first initial resource can be a resource that can be played directly, in which case the first initial resource is the predetermined media resource. Alternatively, the first initial resource may include resource text; or the first initial resource may include resource text and default speech resources, etc., without limitation. The composition of the first initial resource can be seen in the composition of the predetermined media resource in S102 of Figure 3, except that when the first initial resource is a predetermined media resource, the predetermined speech resource in the predetermined media resource is the speech resource in the first initial resource; when the first initial resource is not a predetermined media resource, the predetermined speech resource in the predetermined media resource is generated based on the first initial resource.

[0130] For example, as shown in Table 1, the backend can obtain the first initial resource bound to the associated speech type from the resource database, and synthesize the target speech data based on the first initial resource and the target sound features.

[0131] Alternatively, a target speech type can correspond to different initial resources at different times. Another possible resource database is shown in Table 2 below:

[0132] Table 2

[0133]

[0134] As shown in Table 2, at this point, all resource voice types in the resource database can be considered as the target voice type. The backend can obtain the system network time and use the initial resource of the associated voice type within the time period of the system network time as the first initial resource. For example, when the resource voice type is "blessing type", the initial resource may be different in different holiday periods. For example, in the "Chinese New Year" period, the initial resource may be "Happy New Year", and in the "Mid-Autumn Festival" period, the initial resource may be "Happy Mid-Autumn Festival", etc. Of course, a resource voice type can also correspond to only one initial resource, such as the initial resource in different holiday periods may be "Happy Holidays", etc., without restriction.

[0135] Optionally, when generating target speech data based on the first initial resource or target content, if the target content or the first initial resource includes speech data, the target content or the first initial resource can be determined as the target speech data; or, target sound features can be obtained, and target speech data can be synthesized based on the target sound features and the resource text or target content in the first initial resource.

[0136] The target voice features can be at least one of the following: voice timbre, pronunciation habits, intonation, speech rate, accent, emotional expression, etc. Based on the target voice features, speech generation can be performed on the resource text in the first initial resource to obtain the predetermined speech resource.

[0137] Optionally, when there are multiple associated speech types, the resource texts corresponding to each of the multiple associated speech types can be merged to obtain target text data. Target speech data is then synthesized based on the target sound features and the target text data. Specifically, when merging the resource texts corresponding to multiple associated speech types to obtain target text data, a sentence generation model can be used; alternatively, text keywords can be extracted from the resource texts corresponding to multiple associated speech types, and sentences can be generated from these keywords to obtain target text data. No restrictions are imposed here.

[0138] Optionally, any of the above-mentioned resource databases can be flexibly updated to update or enrich the provision scenarios of media resources, and the update is lightweight, thereby improving the flexibility and versatility of media resource provision. For example, a computer device can respond to an update operation on the resource database and obtain the first resource information indicated by the update operation. When the update operation is an add operation, the first resource information includes complete resource information such as the resource audio type and initial resources, and the first resource information can be added to the resource database. When the update operation is a delete operation, the first resource information may include a first resource identifier, and the data associated with the first resource identifier in the resource database can be deleted. When the update operation is a replace operation, the first resource information may include a first resource identifier and data to be updated, and the second resource information associated with the first resource identifier can be found in the resource database, and the data to be updated can replace the data corresponding to the data to be updated in the second resource information.

[0139] Optionally, in order to make the predetermined media resources associated with the first voice comment more in line with the publishing habits of the target audience of the first voice comment, it is possible to restrict at least one target voice data to include at least target voice data that matches the target content obtained from the first voice data.

[0140] S1106. Send the voice-converted text corresponding to the first voice data.

[0141] The backend can send the speech-to-text corresponding to the first speech data to the frontend. Optionally, the backend can also send at least one target speech data to the frontend.

[0142] S1107. The front end displays the first voice comment in the comment display area; the first voice comment includes the target content and is associated with a resource access entry.

[0143] The process can be referred to in the relevant description of the above embodiments, and will not be repeated here.

[0144] Furthermore, the front end can respond to a trigger operation on the resource access entry and play the pre-defined media resources. In one case (1), if the front end does not receive at least one target voice data sent by the back end, it can execute S1108 to S1110.

[0145] In one case (2), the front end receives at least one target voice data sent by the back end. In response to the trigger operation for the first resource access entry, the front end can directly execute S1110 and play the reserved media resource on the resource display page.

[0146] S1108. In response to a triggered operation targeting a resource access entry, send a resource acquisition request.

[0147] The front-end can respond to a triggered operation on the resource access entry point by sending a resource acquisition request to the back-end. This resource acquisition request may include, but is not limited to, the resource identifier bound to the resource access entry point.

[0148] S1109. Obtain the reserved media resources bound to the resource access portal and send the reserved media resources.

[0149] The backend can obtain at least one target speech data bound to the resource access entry based on the resource identifier bound to the resource access entry, and generate a predetermined media resource based on the at least one target speech data. The number of target speech data can be M, where M is a positive integer, and the predetermined speech resource can be synthesized from M target speech data. As shown in Figure 12, assuming M is 3, the M target speech data can be arranged sequentially to obtain audio sequence 1204. Further, audio sequence 1204 can be directly synthesized into a single audio file to obtain the predetermined speech resource; alternatively, in audio sequence 1204, there can be overlapping portions between adjacent target speech data, such as the overlapping portion 12d between target speech data 12a and target speech data 12b, and the overlapping portion 12e between target speech data 12b and target speech data 12c, as shown in audio sequence 1205. Audio sequence 1205 can be synthesized to obtain a predetermined speech resource with the overlapping portions (including overlapping portions 12d and 12e) cross-faded. For example, the audio sequence can be subjected to 40% adjacent overlap and cross-fading, and then synthesized uniformly to obtain the predetermined speech resource. It is possible to obtain an offline crossfade-in audio stream (i.e., a pre-defined audio resource) that can be played together. This pre-defined audio resource may include the voices of one or more users and can achieve an overlay effect.

[0150] Furthermore, the backend can send the reserved media resources to the frontend.

[0151] S1110. Play the scheduled media resources on the resource display page.

[0152] In this embodiment, the front-end can play predetermined media resources on the resource display page. These predetermined media resources include resource playback logic, allowing the front-end to directly play them on the resource display page. Alternatively, the front-end can obtain the resource playback logic, sequentially read and play the predetermined media resources according to the logic, etc., without limitation. Specifically, this process can be described in the relevant description in S102 of Figure 3. Of course, if the process jumps from S1107, at least one target speech data can be synthesized to obtain predetermined speech resources, which can then be played on the resource display page.

[0153] In the above embodiments, the voice description tag may include at least one of the following data: a voice description icon (voice description tag 202 as shown in Figure 2), voice description information, etc. The voice description tag corresponding to a voice data is related to the content of the voice data. Specifically, the computer device can perform content recognition on the first voice data and / or the voice-converted text corresponding to the first voice data to obtain voice description data, such as recognizing "good luck" or "singing," etc.; the voice description information corresponding to the voice description data can be determined as the voice description tag, such as the voice description information "recognized good luck" corresponding to the voice description data "good luck," or the voice description information "recognized singing" corresponding to the voice description data "singing," etc. Optionally, an icon matching the voice description information can also be obtained as the voice description icon corresponding to the first voice data. For example, see Figure 13, which is a schematic diagram of a voice description tag format provided in an embodiment of this application. As shown in Figure 13, the computer device can display the first voice data 1302 and the corresponding voice description tag, such as voice description tag 1304 or voice description tag 1303, in the comment display area 1301.

[0154] Alternatively, the voice description tag may include voice description information and a voice description icon, for example, see Figure 14, which is a schematic diagram of another form of voice description tag provided in an embodiment of this application. As shown in Figure 14, the computer device can display the first voice data 1402 and the voice description tag 1403 corresponding to the first voice data 1402 in the comment display area 1401. The voice description tag 1403 may include voice description information 140a and a voice description icon 140b, etc. An Easter egg entry (i.e., an associated resource access entry) can be added after the tail (i.e., the voice description tag 1403) to realize personalized provision of media resources and realize emotional transmission between users.

[0155] The content displayed by the voice description tag can be determined based on the corresponding voice description data. For example, if the voice description data is "blessing", the content displayed by the voice description tag can be "Easter egg", "blessing", firework icon, etc.; or if the voice description data is "welfare", the content displayed by the voice description tag can be "welfare", "Easter egg", gift box icon, etc.

[0156] Optionally, the voice description tag can be used to represent the semantics of the corresponding voice comment, that is, to indicate the voice type of the voice comment, etc. Optionally, when the voice comment includes target content, the voice description tag corresponding to the voice comment can also be used to indicate that media resource push has been triggered. As shown in Figure 14, the voice description icon 140a "Recognized Good Fortune" in the voice description tag indicates that the first voice data 1402 triggered media resource push. For example, assuming the voice description tag "Recognized New Year's Greetings" is obtained, it indicates that the voice type of the corresponding voice comment is "Festival" type and "Blessing" type, etc.

[0157] Therefore, when a business user posts a specific voice comment (i.e., a voice comment with the target voice type), the system can recognize the specific voice comment and add a resource access entry (such as a first or second resource access entry) to it. This resource access entry can be added directly to the voice comment or after it has been transcribed into text. The business user can click on the resource access entry to hear a similar synthesized blessing voice (such as a reserved voice resource), creating a sense of surprise and ensuring that on the social media platform, not only are they blessing others, but others are also blessing them, thus enhancing the emotional impact of the social media platform. Furthermore, media resources can be updated for any situation, such as holidays, by updating the resource voice type, enriching the scenarios for providing media resources as described in this application. In other words, updating the resource database enriches the scenarios for providing media resources, enabling lightweight updates and improving the versatility of resource provision. Simultaneously, the form of this resource access entry can be diverse, allowing for content binding with media resources and enabling the functional template provision of media resources. For example, an entry point can be provided, and computer devices can synthesize predetermined media resources, binding these resources to the entry point to generate a resource access entry point. This enables the reusability of the entry point, essentially providing a reusable, template-based Easter egg capability. Furthermore, the resource access entry point and the media resources bound to it implemented in this application are combined in a relatively lightweight manner. Utilizing this Easter egg capability can encourage business users to engage in more voice comments to unlock Easter eggs, while also providing emotional value between various business users.

[0158] The solution implemented in this application can be applied to resource push for comment data, as well as resource push for published content. Specifically, please refer to Figure 15, which is a flowchart of a voice interaction method provided in an embodiment of this application. As shown in Figure 15, the flow of the voice interaction method in this embodiment may include the following:

[0159] S301, Display the first multimedia data on the content display page; the first multimedia data includes the target content and is associated with a resource access entry.

[0160] The first multimedia data includes at least two of the following: interactive voice data, speech-to-text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or a comment on the first published content; the first multimedia data includes target content and is associated with a resource access entry point. This process can be referred to the above-described processing procedure for the first voice comment, and will not be repeated here.

[0161] S302, in response to a triggered operation on the resource access entry, plays the scheduled media resource on the resource display page.

[0162] The predetermined media resources include at least predetermined audio resources, which are obtained based on at least one target audio data matching the target content. In this case, the target audio data is at least one of the following: the portion of the interactive audio data included in the second multimedia data that matches the target content; data synthesized based on the target content and target sound features; or data synthesized based on the target content. The second multimedia data is either second published content or a comment on the second published content. This process can be further referenced to the triggering process of the resource access entry associated with the first audio comment described above. The second published content may or may not include the first published content.

[0163] Referring also to Figure 16, which is a flowchart of another voice interaction method provided in an embodiment of this application, as shown in Figure 16, the flow of the voice interaction method in this embodiment may include the following:

[0164] S401, in response to the publishing operation for the first multimedia data, displays the first multimedia data in the content display page.

[0165] The first multimedia data includes at least two of the following: interactive voice data, speech-to-text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or a comment on the first published content. This process can be seen in the publishing process of the first voice data in S201 of Figure 10 above.

[0166] S402, when the first multimedia data includes the target content, the predetermined media resource is played.

[0167] The predetermined media resources include at least predetermined audio resources, which are obtained based on at least one target audio data that matches the target content.

[0168] Please refer to Figure 17, which is a schematic diagram of the structure of a comment interaction device provided in an embodiment of this application. It should be noted that the comment interaction device shown in Figure 17 is used to execute the methods of the embodiments shown in Figures 3, 10, and 11 of this application. For ease of explanation, only the parts related to the embodiments of this application are shown; specific technical details are not disclosed. Refer to the embodiments shown in Figures 3, 10, and 11 of this application. The comment interaction device 10 may include: a comment interaction module 11 and a resource playback module 12. Wherein:

[0169] The comment interaction module 11 is used to display a first voice comment in the comment display area; the first voice comment includes at least two of the following: first voice data, voice-converted text corresponding to the first voice data, and voice description tags corresponding to the first voice data; the first voice comment includes target content and is associated with a resource access entry.

[0170] The resource playback module 12 is used to play a predetermined media resource in the resource display page in response to a trigger operation on the resource access entry; the predetermined media resource includes at least a predetermined audio resource, which is obtained based on at least one target audio data that matches the target content.

[0171] The resource access entry is associated with the speech-to-text corresponding to the first speech data, or with the speech description tag corresponding to the first speech data.

[0172] The pre-selected media resources also include animation resources that match the target content; the resource playback module 12 can be used for:

[0173] In response to a triggered operation on the resource access point, the animation resources and scheduled audio resources are played simultaneously on the resource display page, or the animation resources and scheduled audio resources are played sequentially on the resource display page.

[0174] The target speech data is at least one of the following: the portion of the second speech data that matches the target content, data synthesized based on the target content and target sound features, and data synthesized based on the target content;

[0175] The second voice data is a second voice comment that includes the target content; the target voice features are derived from at least one of the following data: historical voice data published by the object that published the first voice comment, platform voice data provided by the social media platform where the first voice comment is located, and third voice data published by the business object.

[0176] The number of target speech data is M, where M is a positive integer; the device also includes:

[0177] The speech generation module 13 is used to arrange M target speech data sequentially to obtain an audio sequence; in the audio sequence, there are overlapping parts between adjacent target speech data;

[0178] The speech generation module 13 is also used to perform audio synthesis on the audio sequence to obtain a predetermined speech resource with overlapping parts cross-faded.

[0179] The first voice comment includes target content, including at least one of the following:

[0180] The first speech data, the speech-to-text corresponding to the first speech data, and the speech description tag corresponding to the first speech data all include the target content;

[0181] Any one of the following—the first voice data, the voice-to-text corresponding to the first voice data, or the voice description tag corresponding to the first voice data—matches the target content.

[0182] The device also includes:

[0183] The tag generation module 14 is used to perform content recognition on the first speech data and / or the speech-converted text corresponding to the first speech data to obtain speech description data;

[0184] The tag generation module 14 is also used to determine the voice description information corresponding to the voice description data as voice description tags.

[0185] The specific implementation methods of the comment interaction module, resource playback module, voice generation module, and tag generation module can be found in the description of the above embodiments, and will not be repeated here. It should be understood that the beneficial effects obtained by using the same method will also not be repeated here.

[0186] Please refer to Figure 18, which is a schematic diagram of another comment interaction device provided in an embodiment of this application. It should be noted that the comment interaction device shown in Figure 18 is used to execute the method of the embodiment shown in Figure 10 of this application. For ease of explanation, only the parts related to the embodiment of this application are shown; specific technical details are not disclosed. Refer to the embodiment shown in Figure 10 of this application. The comment interaction device 20 may include: a voice publishing module 21 and a resource playback module 22. Wherein:

[0187] The voice publishing module 21 is used to display the first voice comment in the comment display area in response to the publishing operation of the first voice data; the first voice comment includes at least two of the following: the first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data;

[0188] The resource playback module 22 is used to play a predetermined media resource when the first voice comment includes the target content; the predetermined media resource includes at least a predetermined voice resource, which is obtained based on at least one target voice data that matches the target content.

[0189] The resource playback module 22 can be used for:

[0190] Before playing the scheduled media resource, obtain the trigger status of the publishing object for the scheduled media resource.

[0191] Whether to play pre-defined media resources is controlled based on the trigger state.

[0192] The specific implementation methods of the voice publishing module, resource playback module, etc., can be found in the description of the above embodiments, and will not be repeated here. It should be understood that the beneficial effects obtained by using the same method will also not be repeated here.

[0193] Please refer to Figure 19, which is a schematic diagram of the structure of a voice interaction device provided in an embodiment of this application. It should be noted that the voice interaction device shown in Figure 19 is used to execute the method of the embodiment shown in Figure 15 of this application. For ease of explanation, only the parts related to the embodiment of this application are shown; specific technical details are not disclosed. Refer to the embodiment shown in Figure 15 of this application. The voice interaction device 30 may include: a data display module 31 and a resource playback module 32. Wherein:

[0194] The data display module 31 is used to display first multimedia data on the content display page; the first multimedia data includes at least two of the following: interactive voice data, voice-converted text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or a comment on the first published content; the first multimedia data includes target content and is associated with a resource access entry.

[0195] The resource playback module 32 is used to play a predetermined media resource in the resource display page in response to a trigger operation on the resource access entry; the predetermined media resource includes at least a predetermined audio resource, which is obtained based on at least one target audio data that matches the target content.

[0196] The target speech data is at least one of the following: the portion of the interactive speech data included in the second multimedia data that matches the target content, data synthesized based on the target content and target sound features, and data synthesized based on the target content;

[0197] The second multimedia data refers to the second published content or comments on the second published content.

[0198] The specific implementation methods of the data display module, resource playback module, etc., can be found in the description of the above embodiments, and will not be repeated here. It should be understood that the beneficial effects obtained by using the same method will also not be repeated here.

[0199] Please refer to Figure 20, which is a schematic diagram of another voice interaction device provided in an embodiment of this application. It should be noted that the voice interaction device shown in Figure 20 is used to execute the method of the embodiment shown in Figure 16 of this application. For ease of explanation, only the parts related to the embodiment of this application are shown; specific technical details are not disclosed. Refer to the embodiment shown in Figure 16 of this application. The voice interaction device 40 may include: a data publishing module 41 and a resource playback module 42. Wherein:

[0200] Data publishing module 41 is used to display the first multimedia data on the content display page in response to the publishing operation of the first multimedia data; the first multimedia data includes at least two of the following: interactive voice data, voice-converted text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or comments on the first published content.

[0201] The resource playback module 42 is used to play a predetermined media resource when the first multimedia data includes target content; the predetermined media resource includes at least a predetermined audio resource, which is obtained based on at least one target audio data that matches the target content.

[0202] The specific implementation methods of the data publishing module, resource playback module, etc., can be found in the description of the above embodiments, and will not be repeated here. It should be understood that the beneficial effects obtained by using the same method will also not be repeated here.

[0203] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0204] Please refer to Figure 21, which is a schematic diagram of the structure of a computer device provided in an embodiment of this application. As shown in Figure 21, the computer device 2100 includes at least one processor 2101 and a memory 2102. Optionally, the computer device may also include a network interface. The processor 2101, memory 2102, and network interface can exchange data. The network interface, controlled by the processor 2101, is used to send and receive messages. The memory 2102 stores computer programs, including program instructions. The processor 2101 executes the program instructions stored in the memory 2102. The processor 2101 is configured to invoke the program instructions to execute the aforementioned method. The memory 2102 may include volatile memory, such as random-access memory (RAM); the memory 2102 may also include non-volatile memory, such as flash memory, solid-state drive (SSD), etc.; the memory 2102 may also include combinations of the above types of memory.

[0205] Processor 2101 may be a central processing unit (CPU). In one embodiment, processor 2101 may also be a graphics processing unit (GPU). Processor 2101 may also be a combination of CPU and GPU. Processor 2101 can be used to call the device control application stored in memory 2102 to execute the comment interaction method described in the embodiments corresponding to Figures 3, 10, and 11 above, and can also execute the comment interaction device described in the embodiments corresponding to Figures 17 and 18 above, and can also execute the voice interaction method described in the embodiments corresponding to Figures 15 and 16 above, and can also execute the voice interaction device described in the embodiments corresponding to Figures 19 and 20 above, which will not be repeated here. In addition, the beneficial effects of using the same method will not be repeated here.

[0206] In specific implementations, the devices, processors, memory, etc., described in the embodiments of this application can execute the implementation methods described in the above method embodiments, or they can execute the implementation methods described in the embodiments of this application, which will not be repeated here.

[0207] This application also provides a computer-readable storage medium storing a computer program. The computer program includes program instructions, which, when executed by a processor, enable the processor to perform some or all of the steps described in the above method embodiments. Optionally, the computer storage medium can be volatile or non-volatile. The computer-readable storage medium may primarily include a program storage area and a data storage area. The program storage area may store an operating system, at least one application program required for a given function, etc.; the data storage area may store data created based on the use of blockchain nodes, etc.

[0208] This application provides a computer program product, which may include a computer program. When the computer program is executed by a processor, it can implement some or all of the steps in the above method, which will not be elaborated here.

[0209] In this article, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0210] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer storage medium, which can be a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0211] The above-disclosed embodiments are merely some of the embodiments of this application, and should not be construed as limiting the scope of this application. Those skilled in the art can understand that all or part of the processes for implementing the above embodiments, and equivalent changes made in accordance with the claims of this application, still fall within the scope of this application.

Claims

1. A comment interaction method, characterized in that, The method includes: displaying a first voice comment in a comment display area; the first voice comment includes at least two of the following: first voice data, voice-converted text corresponding to the first voice data, and voice description tags corresponding to the first voice data; the first voice comment includes target content and is associated with a resource access entry; in response to a trigger operation on the resource access entry, playing a predetermined media resource on a resource display page; the predetermined media resource includes at least a predetermined voice resource, which is obtained based on at least one target voice data that matches the target content.

2. The method according to claim 1, characterized in that, The resource access entry is associated with the speech-to-text corresponding to the first speech data, or with the speech description tag corresponding to the first speech data.

3. The method according to claim 1, characterized in that, The predetermined media resources also include animation resources that match the target content; the step of playing the predetermined media resources on the resource display page in response to a trigger operation on the resource access entry includes: synchronously playing the animation resources and the predetermined audio resources on the resource display page in response to a trigger operation on the resource access entry, or sequentially playing the animation resources and the predetermined audio resources on the resource display page.

4. The method according to claim 1, characterized in that, The target voice data is at least one of the following: a portion of the second voice data that matches the target content, data synthesized based on the target content and target voice features, and data synthesized based on the target content; wherein, the second voice data belongs to a second voice comment that includes the target content; the target voice features originate from at least one of the following: historical voice data published by the publisher of the first voice comment, and platform voice data provided by the social media platform where the first voice comment is located.

5. The method according to claim 1, characterized in that, The number of target speech data is M, where M is a positive integer; the method further includes: arranging the M target speech data sequentially to obtain an audio sequence; in the audio sequence, there is an overlapping part between adjacent target speech data; performing audio synthesis on the audio sequence to obtain the predetermined speech resource with the overlapping part cross-faded.

6. The method according to claim 1, characterized in that, The first voice comment includes target content, which includes at least one of the following: any one of the first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data includes target content; any one of the first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data matches the target content.

7. The method according to claim 1, characterized in that, The method further includes: performing content recognition on the first speech data and / or the speech-converted text corresponding to the first speech data to obtain speech description data; and determining the speech description information corresponding to the speech description data as the speech description tag.

8. A comment interaction method, characterized in that, The method includes: in response to a publishing operation for first voice data, displaying a first voice comment in a comment display area; the first voice comment includes at least two of the following: the first voice data, voice-converted text corresponding to the first voice data, and voice description tags corresponding to the first voice data; when the first voice comment includes target content, playing a predetermined media resource; the predetermined media resource includes at least a predetermined voice resource, the predetermined voice resource being obtained based on at least one target voice data that matches the target content.

9. The method according to claim 8, characterized in that, The step of playing a predetermined media resource when the first voice comment includes target content includes: before playing the predetermined media resource, obtaining the trigger state of the publishing object that published the first voice data for the predetermined media resource; and controlling whether to play the predetermined media resource based on the trigger state.

10. A voice interaction method, characterized in that, The method includes: displaying first multimedia data on a content display page; the first multimedia data includes at least two of the following: interactive voice data, speech-to-text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or a comment on the first published content; the first multimedia data includes target content and is associated with a resource access entry; in response to a trigger operation on the resource access entry, playing a predetermined media resource on the resource display page; the predetermined media resource includes at least a predetermined voice resource, which is obtained based on at least one target voice data that matches the target content.

11. The method according to claim 10, characterized in that, The target voice data is at least one of the following: the portion of the interactive voice data included in the second multimedia data that matches the target content, data synthesized based on the target content and target voice features, and data synthesized based on the target content; wherein, the second multimedia data is second published content or a comment on the second published content.

12. A voice interaction method, characterized in that, The method includes: in response to a publishing operation for first multimedia data, displaying the first multimedia data on a content display page; the first multimedia data includes at least two of interactive voice data, speech-to-text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or a comment on the first published content; when the first multimedia data includes target content, playing a predetermined media resource; the predetermined media resource includes at least a predetermined voice resource, the predetermined voice resource being obtained based on at least one target voice data matching the target content.

13. A comment interaction device, characterized in that, The device includes: a comment interaction module for displaying a first voice comment in a comment display area; the first voice comment includes at least two of the following: first voice data, voice-converted text corresponding to the first voice data, and voice description tags corresponding to the first voice data; the first voice comment includes target content and is associated with a resource access entry; and a resource playback module for playing a predetermined media resource in a resource display page in response to a trigger operation on the resource access entry; the predetermined media resource includes at least a predetermined voice resource, which is obtained based on at least one target voice data that matches the target content.

14. A comment interaction device, characterized in that, The device includes: a comment publishing module, configured to display a first voice comment in a comment display area in response to a publishing operation on the first voice data; the first voice comment includes at least two of the following: the first voice data, the voice-converted text corresponding to the first voice data, and the voice description tag corresponding to the first voice data; and a resource playback module, configured to play a predetermined media resource when the first voice comment includes target content; the predetermined media resource includes at least a predetermined voice resource, the predetermined voice resource being obtained based on at least one target voice data that matches the target content.

15. A voice interaction device, characterized in that, The device includes: a data display module for displaying first multimedia data on a content display page; the first multimedia data includes at least two of the following: interactive voice data, voice-converted text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or a comment on the first published content; when the first multimedia data includes target content, the first multimedia data is associated with a resource access entry; and a resource playback module for playing a predetermined media resource on the resource display page in response to a trigger operation on the resource access entry; the predetermined media resource includes at least a predetermined voice resource, which is obtained based on at least one target voice data that matches the target content.

16. A voice interaction device, characterized in that, The device includes: a data publishing module, configured to display the first multimedia data on a content display page in response to a publishing operation for the first multimedia data; the first multimedia data includes at least two of the following: interactive voice data, speech-to-text corresponding to the interactive voice data, and voice description tags corresponding to the interactive voice data; the first multimedia data is first published content or a comment on the first published content; and a resource playback module, configured to play a predetermined media resource when the first multimedia data includes target content; the predetermined media resource includes at least a predetermined voice resource, the predetermined voice resource being obtained based on at least one target voice data matching the target content.

17. A computer device, characterized in that, The method includes a processor and a memory, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is configured to invoke the program instructions to perform the method as described in any one of claims 1-12.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-12.

19. A computer program product, characterized in that, The method includes a computer program, which includes program instructions that, when executed by a processor, implement the method according to any one of claims 1-12.