Voice evaluation display method and device, computer equipment and readable storage medium
By generating non-semantic evaluation labels through speech recognition, the problem of insufficient communication of evaluation information in existing technologies is solved, and the diversity and accuracy of evaluation content are realized. Users can better perceive the real emotions and environmental background of the evaluator, thereby improving the communication power and decision-making assistance value of evaluation information.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XINGIN INFORMATION TECH (SHANGHAI) CO LTD
- Filing Date
- 2026-03-11
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, the ability to convey evaluation information is insufficient. The text and image descriptions in user feedback cannot effectively convey the user's true emotions and environmental context, resulting in a lack of diversity and accuracy in the evaluation content.
By using speech recognition technology, non-semantic information, such as emotion, acoustic environment, and the behavior of the speaker, is generated. This information is then used to generate target evaluation tags and displayed on the evaluation page. Combined with text recognition results, this enables the diversification and precise filtering of evaluation content.
It enhances the communicative power of evaluation information, allowing users to better perceive the evaluator's true emotions and environmental context, thereby increasing the diversity and accuracy of evaluations, reducing redundant content interference, and enhancing decision-making support value.
Smart Images

Figure CN122024735A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a voice evaluation display method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Technology
[0002] With the development of mobile internet technology, evaluating various objects to achieve information exchange, experience feedback, and content recommendation has become an important function of service platforms. User-posted reviews can provide references for other users and also help businesses improve service quality.
[0003] In related technologies, review pages primarily rely on user-inputted text and uploaded images to convey information. This method can only statically present consumption results and brief descriptions, lacking sufficient information delivery capability. Summary of the Invention
[0004] Therefore, it is necessary to provide a voice evaluation display method, device, computer equipment, computer-readable storage medium, and computer program product that can improve the communication power of evaluation information in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a method for displaying voice evaluation, including:
[0006] The evaluation page of the target object is displayed; the evaluation page includes target evaluation tags and evaluation content for the target object, the target evaluation tags are used to filter the evaluation content, and the evaluation content includes at least one voice evaluation.
[0007] In response to the triggering operation of the target evaluation tag, target voice evaluations related to the target evaluation tag are filtered on the evaluation page; the target voice evaluation includes target recognition information for generating the target evaluation tag, the target recognition information is generated based on the speech recognition of the target voice evaluation, and the target recognition information is different from the text recognition result of the target voice evaluation.
[0008] In one embodiment, the target recognition information includes at least one of emotion recognition results, acoustic environment recognition results, voice subject behavior recognition results, and voice evaluation recording status recognition results; the target evaluation label is generated based on at least one recognition result.
[0009] In one embodiment, the emotion recognition result includes at least one of the recognition results of emotion polarity and emotion expression mode;
[0010] The acoustic environment recognition results include at least one of the recognition results of environmental noise level, spatial characteristics, and environmental sound characteristics;
[0011] The identification results of the speaking subject behavior include at least one of the identification results of the number of speakers and the speaking method;
[0012] The recognition result of the voice evaluation recording status includes at least one of the recognition results of recording timing and voice naturalness.
[0013] In one embodiment, the target evaluation label is generated based on at least one recognition result and the number of corresponding target speech evaluations.
[0014] In one embodiment, there are multiple target evaluation tags. After the evaluation page filters out target speech evaluations related to the target evaluation tags, the method further includes:
[0015] At the associated location of the target speech evaluation, the target recognition information corresponding to the generated target evaluation label, as well as the target recognition information corresponding to other evaluation labels, are simultaneously displayed; the other evaluation labels refer to the target evaluation labels that have not been triggered.
[0016] In one embodiment, the method further includes:
[0017] In response to each playback operation of the target speech evaluation, the number of times the target speech evaluation has been played is counted;
[0018] The play count is displayed.
[0019] In one embodiment, the evaluation page displays the text recognition results of the target voice evaluation, and the text recognition results include highlighted target keywords.
[0020] Secondly, this application also provides a voice evaluation display device, comprising:
[0021] A page display module is used to display the evaluation page of the target object; the evaluation page includes target evaluation tags and evaluation content for the target object, the target evaluation tags are used to filter the evaluation content, and the evaluation content includes at least one voice evaluation.
[0022] The evaluation filtering module is used to filter out target voice evaluations related to the target evaluation tag on the evaluation page in response to the trigger operation of the target evaluation tag; the target voice evaluation includes target recognition information for generating the target evaluation tag, the target recognition information is generated based on the speech recognition of the target voice evaluation, and the target recognition information is different from the text recognition result of the target voice evaluation.
[0023] In one embodiment, the target recognition information includes at least one of emotion recognition results, acoustic environment recognition results, voice subject behavior recognition results, and voice evaluation recording status recognition results; the target evaluation label is generated based on at least one recognition result.
[0024] In one embodiment, the emotion recognition result includes at least one of the recognition results of emotion polarity and emotion expression mode; the acoustic environment recognition result includes at least one of the recognition results of environmental noise level, spatial characteristics and environmental sound characteristics; the voice subject behavior recognition result includes at least one of the recognition results of the number of voices and the voice mode; and the voice evaluation recording status recognition result includes at least one of the recognition results of recording timing and voice naturalness.
[0025] In one embodiment, the target evaluation label is generated based on at least one recognition result and the number of corresponding target speech evaluations.
[0026] In one embodiment, there are multiple target evaluation tags, and the page display module is also used to simultaneously display the target recognition information corresponding to the target evaluation tag and the target recognition information corresponding to other evaluation tags at the associated position of the target voice evaluation; the other evaluation tags refer to target evaluation tags that have not been triggered.
[0027] In one embodiment, the device further includes:
[0028] The playback statistics module is used to count the number of times the target voice evaluation is played in response to each playback operation of the target voice evaluation;
[0029] A playback count display module is used to display the playback count.
[0030] In one embodiment, the evaluation page displays the text recognition results of the target voice evaluation, and the text recognition results include highlighted target keywords.
[0031] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0032] The evaluation page of the target object is displayed; the evaluation page includes target evaluation tags and evaluation content for the target object, the target evaluation tags are used to filter the evaluation content, and the evaluation content includes at least one voice evaluation.
[0033] In response to the triggering operation of the target evaluation tag, target voice evaluations related to the target evaluation tag are filtered on the evaluation page; the target voice evaluation includes target recognition information for generating the target evaluation tag, the target recognition information is generated based on the speech recognition of the target voice evaluation, and the target recognition information is different from the text recognition result of the target voice evaluation.
[0034] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0035] The evaluation page of the target object is displayed; the evaluation page includes target evaluation tags and evaluation content for the target object, the target evaluation tags are used to filter the evaluation content, and the evaluation content includes at least one voice evaluation.
[0036] In response to the triggering operation of the target evaluation tag, target voice evaluations related to the target evaluation tag are filtered on the evaluation page; the target voice evaluation includes target recognition information for generating the target evaluation tag, the target recognition information is generated based on the speech recognition of the target voice evaluation, and the target recognition information is different from the text recognition result of the target voice evaluation.
[0037] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0038] The evaluation page of the target object is displayed; the evaluation page includes target evaluation tags and evaluation content for the target object, the target evaluation tags are used to filter the evaluation content, and the evaluation content includes at least one voice evaluation.
[0039] In response to the triggering operation of the target evaluation tag, target voice evaluations related to the target evaluation tag are filtered on the evaluation page; the target voice evaluation includes target recognition information for generating the target evaluation tag, the target recognition information is generated based on the speech recognition of the target voice evaluation, and the target recognition information is different from the text recognition result of the target voice evaluation.
[0040] The aforementioned voice evaluation display method, apparatus, computer device, computer-readable storage medium, and computer program product, in response to an operation to view evaluation information for a target object, display an evaluation page for the target object. The evaluation page displays target evaluation tags and voice evaluations for the target object. The target evaluation tags are obtained based on target recognition information generated after speech recognition of the target voice evaluation. This target recognition information differs from the text recognition result of the target voice evaluation. The text recognition result is obtained by converting the target voice evaluation into text, reflecting the textual information contained in the target voice evaluation. The target recognition information, however, differs from the text recognition result, reflecting the non-semantic information of the target voice evaluation. Therefore, the target evaluation tags can reflect evaluation information in the evaluation content that the text recognition result of the target voice evaluation fails to represent. Displaying the target evaluation tags on the evaluation page expands the evaluation dimensions, enhances the diversity and comprehensiveness of the evaluation dimensions for the target object, and makes the evaluation information more communicative. Furthermore, when a trigger operation for a target evaluation tag is received, the target voice evaluations related to the target evaluation tag are filtered out and displayed on the evaluation page. This filters out irrelevant evaluations, avoids redundant evaluation content from interfering with the user, and improves the accuracy and focus of the evaluation information display. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a schematic diagram illustrating the application environment of the voice evaluation display method in one embodiment;
[0043] Figure 2 This is a flowchart illustrating a voice evaluation display method in one embodiment;
[0044] Figure 3 This is a schematic diagram of the evaluation page in one embodiment;
[0045] Figure 4 This is a schematic diagram of the evaluation page in another embodiment;
[0046] Figure 5 This is a schematic diagram of a page when playing voice evaluation in one embodiment;
[0047] Figure 6 This is a flowchart illustrating the interaction of a voice evaluation display method in one embodiment;
[0048] Figure 7 This is a structural block diagram of a voice evaluation display device in one embodiment;
[0049] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0051] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The terms "comprising" and "having," and any variations thereof, as used in this application, are intended to cover non-exclusive inclusion. The term "at least one" as used in this application refers to one or more. The term "and / or" as used in this application refers to one of the solutions, or any combination of multiple solutions.
[0052] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with relevant regulations. For example, when a user posts a review, a prompt message may be provided to the user, informing them that their review posting operation will require the acquisition and use of the user's personal information, and the user can choose whether to agree, such as choosing to comment anonymously or to disclose their information.
[0053] The following describes some technical terms used in the embodiments of this application:
[0054] "In response to" indicates the state in which a corresponding event occurs or a condition is met. It's understandable that the timing of subsequent actions performed in response to this event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the event occurs or the condition is met.
[0055] A trigger operation refers to an action performed on information (such as controls) provided by a computer device. This operation sends a corresponding instruction to the computer device, triggering it to execute the next task. The next task triggered by different trigger operations on different information can all be pre-set in a program. This trigger operation can be manually executed by the user; for example, clicks, double-clicks, long presses, and swipes performed by the user on content displayed on the computer screen are all trigger operations. In some cases, the trigger operation can also be executed by the computer device based on a pre-defined program.
[0056] Text recognition results: These are obtained by performing text recognition on speech, meaning converting speech into corresponding text. In other words, the text recognition result represents the text containing the characters in the speech. Text recognition is used to convert speech into a corresponding text sequence, aiming to reconstruct each character in the speech. Text recognition can be implemented using acoustic models, language models, etc. The model input is audio, and the output is a text sequence (or text string), such as "This restaurant is very good." Model training can use paired data, meaning the training data includes speech plus precise text.
[0057] Target recognition information: This is obtained through speech recognition, specifically the recognition of non-semantic information in speech. It involves extracting information from the speech signal beyond the text content, such as prosodic features like tone and rate of speech, environmental features like background noise level and type, the characteristics of the speaker (e.g., the number of speakers), and the naturalness of the speech. Speech recognition can be achieved using techniques such as affective computing, voiceprint recognition, and prosodic analysis. The model input is audio, and the output is structured labels, such as emotion = angry, environment = noisy, etc. Model training requires labeled data; that is, the training data includes audio plus emotion / environment labels, etc.
[0058] Text recognition can only identify the text content within speech. For example, if the speech review says, "The bread at this restaurant is so delicious," text recognition will accurately identify this text. However, if the user says it sarcastically (meaning it's actually not very good), text recognition will completely fail to recognize it. Speech recognition, on the other hand, ignores the specific text and instead analyzes the acoustic features of the audio. For instance, when performing speech recognition on the same review, "The bread at this restaurant is so delicious," it identifies the speech rate as slower than normal and includes pauses, indicating hesitation. It also identifies the pitch as abnormal fundamental frequency fluctuations and a trembling sound, suggesting an unnatural tone, thus outputting a sarcastic tone. Therefore, even if the text is positive, non-semantic speech recognition can identify it as a potentially sarcastic negative review.
[0059] Traditional voice rating systems often rely solely on speech-to-text recognition technology, converting speech into text for analysis. This approach has significant drawbacks: it cannot capture the user's sarcastic tone, perceive the user's true emotional state during the rating, or reconstruct the environmental context in which the rating occurs. For example, a rating that reads "pretty good" may have the opposite meaning if the user says it with non-semantic characteristics such as frustration or depression.
[0060] This application expands the dimensions of voice evaluation by performing non-semantic speech recognition on voice evaluations and generating target evaluation labels based on the recognition information. By extracting target evaluation labels such as emotion labels (e.g., surprise, disappointment, anger) and environment labels (e.g., quiet, noisy, outdoor) through non-semantic recognition, other users can not only see the evaluation text, but also "perceive" the evaluator's true state, thereby greatly improving the decision-making assistance value of evaluation information.
[0061] Based on this, in an exemplary embodiment, this application provides a voice evaluation display method, which can be applied to a computer device, such as... Figure 1 The terminal or server in the application environment shown. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted displays, etc. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0062] It should be noted that, Figure 1 The illustrated application environment diagram of the voice evaluation display method is merely an example. The voice evaluation display method and application environment described in this application embodiment are for the purpose of more clearly illustrating the technical solutions of this application embodiment and do not constitute a limitation on the technical solutions provided in this application embodiment. As those skilled in the art will know, with the emergence of new business scenarios, the technical solutions provided in this application embodiment are also applicable to similar technical problems.
[0063] For ease of description, the voice evaluation display method in this application embodiment is illustrated using a terminal as an example. The terminal may include a touchscreen display and a processor (of course, the terminal may also use peripherals such as a mouse or keyboard as input devices; here, only a touchscreen display is used as an example). The touchscreen display is used to present a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. When the user operates the GUI through the touchscreen display, the GUI can control local content on the terminal in response to the received operation commands, or it can control content on a peer server in response to the received operation commands. For example, the operation commands generated by the user interacting with the GUI may include commands to launch an application. The processor is configured to launch the application after receiving the user's command to launch the application. Furthermore, the processor is configured to render and draw the GUI associated with the application on the touchscreen display. The touchscreen display is a multi-touch sensitive screen capable of sensing touch or swipe operations performed simultaneously on multiple points on the screen. When the user performs a touch operation on the GUI using their finger, the GUI, upon detecting the touch operation, controls the execution of the action corresponding to the touch operation within the GUI.
[0064] refer to Figure 2 This is a flowchart illustrating the voice evaluation display method provided in this application. In this embodiment, the method includes the following steps:
[0065] Step S210: Display the evaluation page of the target object; the evaluation page includes target evaluation tags and evaluation content for the target object. The target evaluation tags are used to filter the evaluation content, and the evaluation content includes at least one voice evaluation.
[0066] The target object refers to the object being evaluated. The target object can be a real entity or behavior, such as a restaurant, a tourist attraction, or a point of interest (POI), or a product, or a virtual or conceptual object, such as an account, a virtual service, an article, a topic, or an opinion that can be evaluated.
[0067] For example, when the target object is a product or a point of interest (POI), the review page can be an embedded review area within the product details page (POI details page) (which can display some review content), or it can be a separate review page. Users can access the corresponding review page by triggering the review entry while browsing product information (POI information).
[0068] When the target is a user account or a main account, the rating page can be accessed by triggering the "Rating" tab in the navigation bar of the account's homepage. The rating page can also be embedded in the account's homepage.
[0069] When the target is an article, topic, opinion, or content information, the evaluation page can be embedded in the content reading page, topic aggregation page, or displayed as an independent evaluation page to present other users' attitudes and evaluation information about the article, topic, or opinion.
[0070] When the target is a virtual service or an online service, the evaluation page can be triggered during the service usage process, or an evaluation entry can be provided on the service introduction page, which can then be used to display the evaluation page and show relevant evaluation information.
[0071] The evaluation page displays evaluation content for a specific target object. For example, the object details page might display an evaluation area containing the target evaluation tag and part of the evaluation content. Users can access the evaluation page by triggering a viewing action within this area to see the complete evaluation content. Access to the evaluation page can be achieved by triggering a separately configured evaluation entry point, or by triggering the target evaluation tag or evaluation area. In other words, triggering any location within the evaluation trigger area can lead to the evaluation page displaying the complete evaluation content.
[0072] The evaluation page can include an evaluation label display area and an evaluation content display area. When there are many evaluations, users can swipe down to view each evaluation. While swiping to view the evaluation content, the evaluation label display area can remain fixed or collapse to increase the page display space for the evaluation content.
[0073] refer to Figure 3 This is a schematic diagram of an evaluation page provided in an embodiment of this application. In response to an operation to view the evaluation information of a target object, an evaluation page 30 for the target object is displayed. The evaluation page 30 includes an evaluation tag display area 31a and an evaluation content display area 31b. The evaluation tag display area 31a displays the target evaluation tag 32a.
[0074] Specifically, the target evaluation tag 32a can be triggered to filter the evaluation content in the evaluation content display area 31b, thereby achieving accurate presentation of the evaluation content.
[0075] like Figure 3As shown, each comment on the evaluation page can include not only voice evaluation but also text evaluation and uploaded images. For example, when publishing an evaluation, voice recording can be performed through the evaluation publishing page on the terminal, and a voice evaluation is obtained after publication. For instance, a voice input control can be set up on the evaluation publishing page. In response to a trigger operation on the voice input control, the user's voice input can be received. When a voice recording end operation is received, voice is generated; when an evaluation publishing operation is received, the corresponding evaluation content is generated. It can be understood that the evaluation publishing page can simultaneously include a text input entry and an image upload entry for text evaluation and image upload. In one implementation, text evaluation and voice evaluation can be independent of each other, meaning the user can input text content and record voice separately to achieve text evaluation and voice evaluation. In another implementation, the user's voice input can also be converted into text recognition results and displayed in the text input entry as a text evaluation, i.e., text evaluation and voice evaluation correspond to each other.
[0076] Step S220: In response to the triggering operation of the target evaluation label, the target speech evaluation related to the target evaluation label is filtered on the evaluation page; the target speech evaluation includes target recognition information used to generate the target evaluation label. The target recognition information is generated based on the speech recognition of the target speech evaluation. The target recognition information and the text recognition result of the target speech evaluation are different.
[0077] The target evaluation label is generated based on the target speech evaluations within the evaluation content of the target object. In practice, this can be achieved by first performing speech recognition on the speech evaluations included in the evaluation content to obtain recognition information, and then generating the target evaluation label based on this recognition information. For example, all evaluation content of the target object can be pre-acquired, and speech recognition can be performed on the speech evaluations within all evaluation content to obtain the recognition information for each speech evaluation. Alternatively, a predetermined number of evaluation contents of the target object can be pre-acquired, and speech recognition can be performed on the speech evaluations within that predetermined number of evaluation contents to obtain the recognition information for each speech evaluation. The recognition information for each speech evaluation can be divided into multiple types, and each type can correspond to a generated evaluation label. The target evaluation label can be determined based on the number of speech evaluations within each type.
[0078] The classification of voice evaluation information into multiple types refers to the classification within the same dimension. For example, in the environmental dimension, it can be classified into types such as noisy environment and quiet environment; similarly, in the emotional dimension, it can be classified into types such as happy mood and slightly angry mood. In one implementation, a target evaluation label can be generated for each dimension. In this case, for each dimension, the number of voice evaluations of each type under that dimension can be counted, and the evaluation label corresponding to the type of voice evaluation with the largest number or highest proportion among all types can be used as the target evaluation label for that dimension. Taking the environmental dimension as an example, assuming it includes two types: noisy environment and quiet environment, after performing voice recognition on each voice evaluation, the first number of voice evaluations identified as noisy environment and the second number of voice evaluations identified as quiet environment are counted. If the first number is greater than the second number, the target evaluation label for the environmental dimension is noisy environment; conversely, if the first number is less than the second number, the target evaluation label for the environmental dimension is quiet environment. It can be understood that in another implementation, the evaluation label corresponding to the type with a smaller number and lower proportion can also be used as the target evaluation label for that dimension. In another embodiment, target evaluation labels can be generated for each type of evaluation label. Those skilled in the art can determine this according to actual needs, and this application does not make any specific limitations on this.
[0079] In the specific implementation of this solution, the terminal responds to the user's command to view the evaluation information of the target object and displays the evaluation page of the target object. The evaluation page displays at least one voice evaluation of the target object, as well as target evaluation tags for filtering evaluation content. When a trigger operation is received for a target evaluation tag, the terminal responds to this trigger operation by filtering out target voice evaluations related to the target evaluation tag from the evaluation content and displaying them on the evaluation page. This filters out irrelevant evaluations, avoids redundant evaluation content from interfering with the user, and improves the accuracy and focus of the evaluation information display.
[0080] For example, such as Figure 3 As shown, when a trigger operation is received for the target evaluation tag "Environment Index: 80% Noisy", voice evaluations indicating a noisy environment are selected from all evaluation information. Evaluations of other environment types, such as a quiet environment, are filtered out from the environment dimension to obtain the target voice evaluation, which is then displayed in the evaluation content display area 31b of the evaluation page 30. Similarly, if a trigger operation is received for the target evaluation tag "Mood Index: 77% Pleasant", voice evaluations indicating a pleasant mood are selected from all evaluation information. Evaluations of other emotional types, such as a neutral mood or a slightly angry mood, are filtered out from the emotion dimension to obtain the target voice evaluation, which is then displayed in the evaluation content display area 31b of the evaluation page 30.
[0081] In some embodiments, in addition to target evaluation tags generated based on target recognition information of the target speech evaluation, the evaluation page may also include text recognition result tags generated based on the text recognition results of the target speech evaluation, and text evaluation tags generated based on the text evaluation. It can be understood that target evaluation tags are generated based on non-semantic information (acoustic environment recognition results, emotion recognition results, speaker behavior recognition results, and recognition results of the speech evaluation recording status, etc.), while text recognition result tags are generated based on text information (text recognition results of the speech evaluation). For example, keywords can be extracted from the text recognition results, and text recognition result tags can be generated based on keyword clustering of each speech evaluation, such as generating praise tags ("the steak is delicious") and complaint tags ("the food is not fresh"), etc. Similarly, text evaluation tags can be generated by extracting keywords from the text evaluation and clustering them.
[0082] refer to Figure 4 Taking a restaurant as an example, the example shows a restaurant review page. The review tag display area 31a on the review page includes not only the target review tag 32a, but also text recognition result tags or text review tags such as "worth going back repeatedly," "Good manager," and "Best eel." Similarly, when a trigger operation is received on a text recognition result tag or text review tag, review information related to the triggered tag is filtered from the review content and displayed in the review content display area 31b.
[0083] In some embodiments, such as Figure 4 As shown, the evaluation label display area 31a can be divided into a first label display area 31a-1 and a second label display area 31a-2. The first label display area 31a-1 displays comprehensive evaluation labels such as "Latest Reviews," "Post-Purchase Reviews," "Not Good Enough," and "Recommended." The second label display area 31a-2 displays finer-grained labels such as "Environmental Index: 80% Noisy," "Mood Index: 77% Pleasant," "Delicious Steak," and "Worth Going Back Again." Among them, "Environmental Index: 80% Noisy" and "Mood Index: 77% Pleasant" are recognition information based on voice evaluation, that is, target evaluation labels generated from non-semantic information. The others are text recognition result labels or text evaluation labels generated from semantic information based on voice evaluation.
[0084] like Figure 3 and Figure 4As shown, the evaluation page also includes a comprehensive score of 4.5 for the target object determined based on all evaluation content. Each evaluation includes at least one evaluation message, which may include voice feedback, text feedback, images, etc. Each evaluation message also displays its final recommendation result (recommended or not recommended) below the user ID. In some embodiments, each evaluation message on the evaluation page also displays interactive controls such as reply and like. In response to a trigger operation on the reply control, a reply input page is displayed for inputting reply information. After receiving a posting operation for the reply information, the reply information is displayed below the evaluation message. In response to a trigger operation on the like control, the like count for that evaluation message is incremented by one to update its like count. Furthermore, the evaluation page also displays the posting time and location of each evaluation message.
[0085] In an exemplary embodiment, the target recognition information includes at least one of emotion recognition results, acoustic environment recognition results, speaker behavior recognition results, and voice evaluation recording status recognition results; the target evaluation label is generated based on at least one recognition result.
[0086] Among them, the emotion recognition result refers to the set of information about the speaker's emotional state obtained after performing emotion analysis on the speech evaluation.
[0087] In some embodiments, the emotion recognition result includes at least one of the recognition results of emotion polarity and emotion expression mode.
[0088] Among them, emotional polarity is used to characterize specific emotional categories, such as pleasant, satisfied, unpleasant, dissatisfied, and angry. Emotional expression style mainly describes the speaker's manner of expression, such as calm narration or excited expression, which can reflect the degree of emotional polarity.
[0089] The target evaluation label generated based on the emotion recognition results is an emotion label that includes emotional polarity and / or emotional expression mode.
[0090] In one implementation, target evaluation labels are generated based on the emotional polarity in the emotion recognition results. The generated target evaluation labels can include emotional labels such as happy, neutral, and angry.
[0091] In another implementation, target evaluation labels are generated based on the emotional expression patterns in the emotion recognition results. The generated target evaluation labels can include emotional labels such as calm statements and excited expressions.
[0092] In another implementation, an emotional label is generated based on the emotional polarity and emotional expression style in the emotion recognition results, serving as the target evaluation label. For example, an emotional polarity of pleasant and an emotional expression style of calm statement can be combined to form a composite emotional label of "calm and satisfied"; another example is an emotional polarity of pleasant and an emotional expression style of excited expression, which can be combined to form a composite emotional label of "strongly recommended"; yet another example is an emotional polarity of unpleasant and an emotional expression style of excited expression, which can be combined to form a composite emotional label of "strongly dissatisfied".
[0093] When generating target evaluation labels based on sentiment recognition results, keywords from the sentiment recognition results can be used directly, or candidate keywords can be preset and matched with the sentiment recognition results to generate target evaluation labels.
[0094] Emotional tags generated by emotional polarity and / or emotional expression methods allow other users to understand the emotional attitude of user evaluations, providing users with emotional dimension decision-making references.
[0095] Among them, the acoustic environment recognition result refers to the set of environmental feature information obtained after analyzing the acoustic environment in which the speech is located in the speech evaluation, which is used to characterize the environmental background of the target speech evaluation.
[0096] In one exemplary embodiment, the acoustic environment identification result includes at least one of the identification results of environmental noise level, spatial characteristics, and environmental sound characteristics.
[0097] Among them, the level of environmental noise characterizes the quietness of the environment in which the speech is located, and can be divided into three levels, such as quiet (no obvious background noise), slightly noisy (with faint background noise), and noisy (strong background noise).
[0098] Spatial features are used to characterize the type of physical space in which the target speech is located, or the type of recording location, such as a private room, a lobby, or an open space.
[0099] Ambient sound characteristics characterize the types of sounds in the environment, such as human voices (the sound of others talking, noise, etc.), background music (such as light music, pop music, instrumental music, etc.), and ambient sounds (such as traffic sounds, wind sounds, etc.).
[0100] The target evaluation label generated based on the acoustic environment recognition results is an environment label that includes at least one of the recognition results of environmental noise level, spatial characteristics, and environmental sound characteristics.
[0101] In one implementation, a target evaluation label is generated based on the level of environmental noise in the acoustic environment recognition results. The generated target evaluation label can be an environmental label such as noisy environment, slightly noisy environment, or quiet environment.
[0102] In another implementation, a target price label is generated based on the spatial features in the acoustic environment recognition results. The generated target evaluation label can include environmental labels such as private room, lobby, and outdoor.
[0103] In another implementation, a target value label is generated based on the ambient sound features in the acoustic environment recognition results. The generated target evaluation label can be an environmental label such as having music, vehicle sounds, or construction sounds.
[0104] In another implementation, environmental labels are generated based on the recognition results of environmental noise level, spatial features, and ambient sound features from the acoustic environment recognition results, serving as target evaluation labels. For example, if the environmental noise level is "quiet environment," the spatial feature is "private room," and the ambient sound feature is "music present," then a composite environmental label of "quiet private room environment, pleasant music" can be generated. Similarly, if the environmental noise level is "slightly noisy environment," the spatial feature is "hall," and the ambient sound feature is "music present," then a composite environmental label of "slightly noisy lobby environment, pleasant music" can be generated.
[0105] It is understandable that a composite environmental label can be formed based on any two of the following: environmental noise level, spatial characteristics, and ambient sound characteristics. For example, if the environmental noise level is "noisy" and the spatial characteristic is "hall," then the composite environmental label "hall noisy" can be generated.
[0106] The level of environmental noise, spatial characteristics, and ambient sound characteristics reflect the volume of sound, spatial attributes, and background sound type, respectively. Identifying these characteristics can provide users with references regarding environmental comfort and privacy.
[0107] Among them, the speech subject behavior recognition result represents the number of people evaluating the target speech evaluation. In some embodiments, the speech subject behavior recognition result can also represent the speech mode between the speech subjects (such as each person evaluating a segment, or multiple people evaluating in a conversational manner).
[0108] In one exemplary embodiment, the voice subject behavior recognition result includes at least one of the recognition results of the number of voices and the voice mode.
[0109] Among them, the number of speakers indicates whether it is a single-person evaluation, a two-person evaluation, or a multi-person evaluation, corresponding to a single-person experience, a two-person group evaluation, or a multi-person group evaluation. The mode of speech is used to distinguish between a single person's continuous monologue and multi-person interaction, mainly describing the expression patterns between the speakers.
[0110] The target evaluation label generated based on the behavioral recognition results of the speaker is a voice label that includes the number of speakers and / or the mode of speaking.
[0111] In one implementation, a target evaluation label is generated based on the number of speakers in the speaker behavior recognition result. The target evaluation label can be a speaker behavior label such as multiple people, two people, or one person.
[0112] In another implementation, a target evaluation label is generated based on the vocalization mode in the vocalization behavior recognition result. The target evaluation label can be a vocalization behavior label such as a single person's continuous monologue or interaction between multiple people.
[0113] In another implementation, a speaker behavior label is generated based on the number of speakers and the speaking method in the speaker behavior recognition results, serving as a target evaluation label. For example, if the speaker is a single person and the speaking method is a single person's continuous monologue, a composite speaker behavior label of "single person: continuous monologue" can be formed; or if the speaker is multiple people and the speaking method is interaction between multiple people, a composite speaker behavior label of "multiple people: interactive evaluation" can be formed.
[0114] Identifying the number of people speaking and the manner of speaking can provide a reference for users' choice of social interaction. For example, when experiencing a social interaction alone, users can choose the evaluation corresponding to a single person speaking, while users dining in groups can mainly choose the evaluation corresponding to multiple people speaking, making the evaluation content more in line with their own social interaction needs.
[0115] Among them, the recording status of speech evaluation represents the recording timing and / or the naturalness of the speech evaluation.
[0116] In one exemplary embodiment, the recognition result of the voice evaluation recording status includes at least one of the recognition results of recording timing and voice naturalness.
[0117] The recording timing can be identified by the recording time and scene, including on-site real-time feedback and post-departure feedback, for example... Figure 3 The "from on-site evaluation" shown in the figure; the recording time reflects the real-time nature of the voice evaluation; the naturalness of the voice reflects whether the voice evaluation is a genuine expression of experience or a deliberate reading or template expression, which can characterize the authenticity and credibility of the evaluation.
[0118] The target evaluation label generated based on the recognition results of the voice evaluation recording status is a recording status label that includes the recording time and / or the naturalness of the voice.
[0119] In one implementation, a target evaluation label is generated based on the recording time in the recognition result of the voice evaluation recording status. The generated target evaluation label may include recording status labels such as on-site real-time evaluation and post-departure evaluation.
[0120] In another implementation, a target evaluation label is generated based on the recording timing in the recognition result of the voice evaluation recording status. The generated target evaluation label can be a recording status label such as natural and fluent, or deliberately templated.
[0121] In another implementation, the recording time and naturalness of the voice in the recognition result of the voice evaluation recording status are used together to generate a recording status label as the target evaluation label. For example, if the recording time is on-site real-time recording and the naturalness of the voice is natural and fluent, then a composite recording status label of "on-site evaluation: fluent and natural" can be formed; as another example, if the recording time is after leaving the store and the naturalness of the voice is deliberately templated, then a composite recording status label of "after leaving the store evaluation: deliberately templated" can be formed.
[0122] Real-time, on-site recorded reviews more closely reflect the actual consumption scenario and are more suitable for users who want to understand the real-time experience and the on-site environment. Reviews recorded after leaving the store are more focused on an overall summary of the consumption experience and are suitable for users who are concerned about the overall feeling. Furthermore, reviews with natural and authentic voices reflect the user's true experience and are more credible. Users can prioritize these reviews and avoid deliberately read-out or templated low-credibility content, making the reviews more reliable and reducing interference from templated or manipulated reviews.
[0123] In one implementation, a target evaluation label can be generated based on each type of target recognition information. The target evaluation label generated based on the emotion recognition result, acoustic environment recognition result, voice subject behavior recognition result, and voice evaluation recording status recognition result may include emotion label, acoustic environment label, voice subject behavior label, and recording status label.
[0124] In another implementation, a target evaluation label can be generated based on multiple target recognition information. For example, a combined target evaluation label can be generated based on the emotion recognition result and the voice subject behavior recognition result. This target evaluation label includes two dimensions: emotion and voice subject behavior. For example, if the emotion recognition result is "angry" and the voice subject behavior recognition result is "multiple people", then the evaluation label "multiple people are angry" can be formed as the target evaluation label.
[0125] By recognizing non-semantic information such as emotion, acoustic environment, speaker behavior, and recording status of voice evaluations, non-semantic evaluation tags are generated, which can enhance the diversity and comprehensiveness of evaluations of target objects and provide richer references for user decision-making. Furthermore, through multi-dimensional mining of the aforementioned contextual information, a more complete and granular consumer context tagging system can be formed. Compared to traditional methods that generate evaluation tags solely based on text keyword extraction, this method can improve the reference value and practicality of evaluation information.
[0126] Additionally, it should be noted that the recognition information obtained from speech recognition of voice evaluations differs from the text recognition results of voice evaluations. Text recognition results are obtained by directly converting speech into text. For example, if a user's voice is "The dishes at this restaurant are very fresh, I recommend checking it out," the text recognition result will also be "The dishes at this restaurant are very fresh, I recommend checking it out." However, the target recognition information in this application refers to non-semantic information, such as acoustic environment recognition results, emotion recognition results, speaker behavior recognition results, and voice evaluation recording status recognition results. For example, acoustic environment recognition results (such as noisy or quiet environment) are obtained by recognizing background noise in the voice evaluation, and emotion recognition results (such as happy or angry mood) are obtained by recognizing the prosodic features of the human voice in the voice evaluation (such as pitch, speech rate, and volume).
[0127] Since the meaning of the text recognition results and the recognition information in speech evaluation are different, the methods for obtaining the text recognition results and recognition information in speech evaluation are also different. The following explains the acquisition of text recognition results and the acquisition of each type of recognition information.
[0128] For example, the text recognition results of speech evaluations can be converted from speech to text using Automatic Speech Recognition (ASR) technology. Non-semantic recognition information in the speech evaluations can be recognized using machine learning models.
[0129] Regarding the acquisition of recognition information, when the recognition information is an acoustic environment recognition result, background noise can be separated, acoustic scene features can be extracted, and environment classification can be performed based on the acoustic scene features to obtain the acoustic environment recognition result. For example, a machine learning model for environment recognition, such as a Convolutional Neural Network (CNN), can be pre-trained as an environment recognition model. This model can then be used to recognize speech evaluations to obtain the acoustic environment recognition result. The training process of the environment recognition model may include: firstly, acquiring several audio samples representing various environment types (such as noisy environments, quiet environments, and pleasant music) to form a training set, and then training the environment recognition model using this training set. During training, the training audio samples are input into the environment recognition model, which extracts acoustic scene features (such as spectrum and signal-to-noise ratio) for recognition, outputs the environment recognition result, and further trains the environment recognition model based on the loss between the environment recognition result and the actual environment type.
[0130] When the identified information is an emotion recognition result, the prosodic features of the human voice in the speech evaluation can be extracted for emotion classification to obtain the emotion recognition result. Similar to the method of obtaining acoustic environment recognition results, a machine learning model for emotion recognition can be pre-trained as an emotion recognition model, and the emotion recognition model can be used to classify the emotion of the speech. For example, firstly, several audio recordings including various emotion types (such as happy, neutral, angry, etc.) are acquired to form a training set, and the emotion recognition model is trained using this training set. During training, the training audio is input into the emotion recognition model, which extracts acoustic scene features (such as spectrum, signal-to-noise ratio, etc.) for recognition, outputs the emotion recognition result, and further trains the emotion recognition model based on the loss between the emotion recognition result and the true emotion type.
[0131] In some embodiments, only one multi-task speech recognition model can be trained to simultaneously recognize text and non-semantic recognition results. Taking non-semantic recognition results including acoustic environment recognition and emotion recognition results as an example, the recognition process of this multi-task speech recognition model may include: First, separating the speech into human voice and background noise. For human voice, semantic feature extraction and prosodic feature extraction are performed separately. Based on the semantic features, text recognition results are obtained; based on the prosodic features, emotion classification is performed to obtain emotion recognition results. For background noise, acoustic scene features are extracted, and based on the acoustic scene features, environment classification is performed to obtain acoustic environment recognition results. Finally, text recognition results, acoustic environment recognition results, and emotion recognition results are output simultaneously.
[0132] When the identification information is the result of the speaker's behavior, the speech can be divided into multiple segments (each segment has only one speaker). Each speech segment is encoded as a "sound fingerprint" through a machine learning model. The sound fingerprints of all segments are clustered, with similar segments in one class. Each class corresponds to one speaker, thus obtaining the number of speakers and the manner of speaking.
[0133] When the recognition information is the result of a recorded voice review, it can be combined with information from the text recognition result of the voice review and acoustic environment information in the voice to determine whether it is an immediate on-site review or a review after leaving the store. For example, it can be determined whether the text recognition result of the voice review contains words such as "currently in the store," "just ordered," "home," or "after leaving" to determine whether it is an immediate on-site review or a review after leaving the store. Alternatively, it can be determined whether it is an immediate on-site review or a review after leaving the store based on the background sound characteristics in the voice. For instance, if the background sound includes ambient sounds from the store (such as restaurant noise, waiter conversations, clinking cutlery, mall background music, other customers' conversations, etc.), it can be determined as an immediate on-site review; if the background sound is more like a quiet home environment (such as television sound, quiet indoor sounds, etc.), without any store background sound, it can be determined as a review after leaving the store. In some implementations, it is also possible to determine whether it is an immediate on-site review or a review after leaving the store by obtaining the publication time and departure time of the voice review.
[0134] In one exemplary embodiment, the target evaluation label is generated based on at least one recognition result and the number of corresponding target speech evaluations.
[0135] For example, a target evaluation label can be generated separately based on each recognition result, or a single target evaluation label can be generated based on a combination of multiple recognition results, such as "strongly recommended by multiple people." This application does not impose specific limitations on this. For ease of explanation, the following description uses the example of generating a target evaluation label for each recognition result.
[0136] It's understandable that different users will have different feelings about a target object even on the same dimension. Therefore, for the same target evaluation label, there will be multiple candidate label types. For example, for emotion recognition results, the generated target evaluation label is an emotion label, and its candidate label types could be strong dissatisfaction, strong recommendation, neutral recommendation, etc. When generating each target evaluation label, it can be obtained by clustering the target recognition information corresponding to all voice evaluations of the target object. Specifically, the number of voice evaluations under each candidate label type can be counted, and the target evaluation label can be generated based on the number of voice evaluations corresponding to each candidate label type.
[0137] In some embodiments, the target evaluation label may include: a target label type and the number of evaluations corresponding to that target label type. The number of evaluations may be the number of voice evaluations under the target label type, such as "Environmental Index: Noisy (20)" (meaning 20 voice evaluations reflect a noisy environment, others may reflect a quiet environment, a generally noisy environment, etc.); or it may be the percentage of voice evaluations under the target label type, such as "Environmental Index: 80% Noisy" (meaning 80% of voice evaluations reflect a noisy environment, others may reflect a quiet environment, a generally noisy environment, etc.). The target label type may be the label type whose voice evaluation count meets the requirements (e.g., the highest number, the highest percentage, etc.); or it may be all label types, i.e., one target evaluation label is generated for each candidate label type.
[0138] For example, if there are three candidate label types under the environment dimension: noisy environment, quiet environment, and beautiful music, then the target evaluation label for the environment dimension can be the type with the most or the highest proportion of speech evaluations among the three types, or a target evaluation label can be generated for each of the three types, that is, there are three target evaluation labels for the environment dimension, such as "Environment Index: Noisy (20)", "Environment Index: Quiet (15)", and "Environment Index: Beautiful Music (14)". When the label type corresponding to the target evaluation label is the type with the highest proportion among multiple candidate label types, the type of target recognition information included in the target speech evaluation selected based on the target evaluation label is the type with the highest proportion, and the number of target speech evaluations is the number of speech evaluations of the type with the highest proportion.
[0139] By generating target evaluation labels through at least one recognition result and the corresponding number of target speech evaluations, the number of evaluations can be intuitively represented, thereby providing users with a reference for the number of evaluations corresponding to the target evaluation labels, improving information transmission efficiency, and assisting in improving users' decision-making efficiency.
[0140] In one exemplary embodiment, there are multiple target evaluation tags; after filtering out the target speech evaluations related to the target evaluation tags on the evaluation page, the method further includes: simultaneously displaying the target recognition information corresponding to the generated target evaluation tag, as well as the target recognition information corresponding to other evaluation tags, at the associated position of the target speech evaluation.
[0141] Among them, other evaluation labels refer to target evaluation labels that have not been triggered, and the target identification information corresponding to other evaluation labels refers to the target identification information corresponding to the target evaluation labels that have not been triggered.
[0142] In one implementation, the target recognition information obtained by performing speech recognition on the target speech evaluation is also displayed in the target speech evaluation display area.
[0143] It is understandable that when there is only one target evaluation label, if a trigger operation is received for that target evaluation label, when displaying the target voice evaluation related to that target evaluation label on the evaluation page, only the target recognition information used to generate that target evaluation label can be displayed in the display area of the target voice evaluation.
[0144] However, when there are multiple target evaluation labels, for example, two or more of the following: emotion label, environment label, vocal subject behavior label, and recording status label. Only one target evaluation label can be triggered at a time. Therefore, with both triggered and untriggered target evaluation labels present, the target speech evaluation display area can show not only the target recognition information corresponding to the triggered label but also the target recognition information corresponding to the untriggered label.
[0145] For example, if the target evaluation label includes environmental label and emotion label, and the triggered target evaluation label is the environmental label, then in the display area of the target speech evaluation, not only will the target recognition information corresponding to the environmental label be displayed (i.e., the acoustic environment recognition result obtained by performing environmental dimension speech recognition on the target speech evaluation), but also the target recognition information corresponding to the emotion label will be displayed at the same time (i.e., the emotion recognition result obtained by performing emotion dimension speech recognition on the target speech evaluation).
[0146] For example, the target evaluation label includes three labels: environment label, emotion label, and recording status label. If the triggered target evaluation label is the environment label, the target recognition information corresponding to the environment label, emotion label, and recording status label will be displayed simultaneously in the target speech evaluation display area.
[0147] In one implementation, when a trigger operation for a target evaluation label is received and the target speech evaluation corresponding to the target evaluation label is displayed, the target recognition information corresponding to the target evaluation label and the target recognition information corresponding to other evaluation labels can be displayed simultaneously.
[0148] In another implementation, the target recognition information corresponding to each target evaluation tag can be continuously displayed while displaying the evaluation content. That is, after the user completes the evaluation and the voice evaluation is performed to obtain the recognition information, the obtained recognition information is displayed in the evaluation content.
[0149] Specifically, in the display area of the target speech evaluation, the target recognition information obtained by performing speech recognition on the target speech evaluation is displayed, including: in the associated position of the target speech evaluation, the target recognition information obtained by performing speech recognition on the target speech evaluation is displayed.
[0150] The associated location for the target speech evaluation can be a region adjacent to the user ID, such as below the user ID, or a region adjacent to the speech, such as behind, below, or above the speech. It is understood that other locations can also be set as the associated location for displaying the target recognition information of the target speech evaluation; this application does not specifically limit this.
[0151] For example, after filtering target speech evaluations based on target evaluation tags, the filtered target speech evaluations are displayed on the evaluation page. For each target speech evaluation, at the associated location of the target speech evaluation, such as the neighboring area of the speech or the neighboring area of the user ID, the target recognition information corresponding to the generated target evaluation tag, as well as the target recognition information corresponding to other evaluation tags, are displayed simultaneously.
[0152] refer to Figure 3 If the triggered target evaluation label is "Environmental Index: 80% Noisy," then the target voice evaluations related to the environment recognition result of "noisy environment" will be filtered from the evaluation content and displayed in the evaluation content display area 31b. Specifically, in the display area corresponding to the target voice evaluation, in addition to displaying the target recognition information corresponding to the environment label "Noisy environment detected" after the voice, the target recognition information corresponding to the emotion label "Pleasant mood" and the target recognition information corresponding to the recording status label "From on-site evaluation" will also be displayed below the user ID. In other words, regardless of which target evaluation label the user triggers, the target recognition information corresponding to all target evaluation labels can be displayed simultaneously in the evaluation content.
[0153] In one implementation, the target recognition information corresponding to each target evaluation label can be displayed in the same location, for example, all displayed after the voice or below the user ID. In another implementation, the target recognition information corresponding to each target evaluation label can also be displayed in different locations, for example, the target recognition information corresponding to the emotion label is displayed below the user ID, and the target recognition information corresponding to the environment label is displayed after the voice. In yet another implementation, the display of target recognition information corresponding to each target evaluation label can also involve displaying the target recognition information corresponding to one target evaluation label in some locations and displaying the target recognition information corresponding to multiple target evaluation labels in other locations, such as... Figure 3 and Figure 4 As shown, the target recognition information corresponding to the emotion tag and the target recognition information corresponding to the recording status tag are both displayed below the user ID, while the target recognition information corresponding to the environment tag is displayed behind the speech.
[0154] In this embodiment, after receiving a trigger operation for a target evaluation tag and filtering out the target voice evaluation on the evaluation page, in addition to displaying the target recognition information corresponding to the triggered target evaluation tag, the target recognition information corresponding to other evaluation tags is also displayed. This allows users to understand the recognition results under the target evaluation tag, as well as the recognition results in other dimensions, thus improving the efficiency of information transmission.
[0155] Based on the above, the method further includes: in response to each playback operation of the target speech evaluation, counting the number of times the target speech evaluation has been played; and displaying the number of times the playback has been played.
[0156] The playback operation can be triggered by clicking or double-clicking the voice in the target voice evaluation, or by setting a playback control corresponding to the voice in the target voice evaluation, and triggering the playback operation of the target voice evaluation by clicking the playback control.
[0157] For example, displaying a target voice review on the review page may also include displaying the number of plays for the target voice review. For instance, such as... Figure 3 and Figure 4 As shown in position 33a, the evaluation page also displays a playback icon and a number for the target voice evaluation, with the displayed number indicating the number of times the target voice evaluation has been played.
[0158] When a playback operation is received for a target voice evaluation, the voice in the target voice evaluation is played in audio form. At the same time, the playback count of the target voice evaluation is incremented by one to obtain the new playback count, and the current playback count is updated and displayed.
[0159] This embodiment helps users quickly filter out voice evaluations with higher quality or more meaningful reference by statistically analyzing and displaying playback data, thus reducing the time spent playing low-quality or unreliable voice evaluations.
[0160] Based on the above, the evaluation page can also display the text recognition results of the target speech evaluation, including the highlighted target keywords.
[0161] The text recognition result refers to the text content corresponding to the speech evaluated by the target speech.
[0162] Highlighting refers to visually enhancing the target keywords in the text recognition results to distinguish them from ordinary text, making it easier for users to quickly capture them. The highlighting effect should be clear, eye-catching, and should not affect the normal viewing of other text.
[0163] The target keywords can be determined based on at least one of the target object, target evaluation tags, or target recognition information. For example, words in the text recognition results that directly describe or evaluate the target object can be used as target keywords, such as slow service, long wait time, fresh food, or beautiful scenery. Alternatively, words in the text recognition results related to target evaluation tags or target recognition information can be excluded before determining the target keywords. This means identifying target keywords from the text other than words related to target evaluation tags or target recognition information to avoid duplication of target keywords with target recognition information displayed on the evaluation page, reducing redundant information on the evaluation page. Furthermore, since there is no need to highlight words that overlap with target recognition information, unnecessary resource consumption can also be reduced.
[0164] For example, the target voice evaluation can be subjected to text recognition to obtain the corresponding text recognition results. The text recognition results are then displayed in the associated position of each target voice evaluation on the evaluation page, allowing users to understand the content of the target voice evaluation without playing the audio.
[0165] Furthermore, when displaying the text recognition results, the target keywords can be highlighted, for example, by enlarging, highlighting, adding background color, or bolding the target keywords, so that users can quickly grasp the key points of the evaluation when viewing the text recognition results of the target voice evaluation.
[0166] In some embodiments, the text recognition results of the target speech evaluation can be displayed synchronously while the target speech evaluation is played; and the keywords can be highlighted when the playback reaches the keywords.
[0167] Among them, synchronous scrolling display means that the voice playback and the text recognition result display are kept in real time. The voice playback progress and the text display progress correspond one-to-one. When the voice plays a certain sentence or word, the text recognition result scrolls to the corresponding position synchronously, ensuring that the content heard by the user is completely consistent with the text seen, with no time difference, or the time difference is controlled within an acceptable range, so as to meet the synchronous perception needs of human eyes and ears.
[0168] In some embodiments, when playing the target voice evaluation, a playback progress synchronization indicator can be displayed simultaneously. The playback progress synchronization indicator refers to an indicator (such as a progress bar, cursor, highlighted underline, etc.) used to associate the voice playback progress with the text recognition result. It moves in real time with the voice playback and synchronously points to the text corresponding to the current playback position, helping the user to quickly locate the correspondence between the text and the voice.
[0169] For example, after converting the target speech evaluation into text to obtain the text recognition result, the text recognition result is timestamped with the speech, so that each text segment and each word corresponds to a specific playback time node in the speech. At the same time, each keyword is matched with the words in the text recognition result to determine the keywords that need to be highlighted in the text recognition result, thereby determining the speech playback time corresponding to the keywords in the text recognition result.
[0170] refer to Figure 5 When a playback command for a target speech evaluation is received, the evaluation is played, and the playback progress bar advances synchronously, displaying the current playback time in real time. Based on the timestamp of the current playback progress, the system automatically locates the corresponding content in the text recognition result and begins synchronous scrolling. As the speech continues to play, the text recognition result continues to scroll synchronously, and the playback progress indicator (progress bar) always follows the playback progress, pointing to the currently playing text, ensuring that the content the user hears and the text they see are completely synchronized. When the system detects that the keyword "good" is about to be played, it is highlighted, for example, by enlarging, highlighting, adding background color, or bolding. After playback, the highlighting style is automatically canceled, and the text returns to normal. When the next keyword is played, the highlighting continues. If a repeated keyword appears in the speech (such as "good" appearing again), the highlighting is triggered every time that keyword is played, ensuring that the user can quickly capture the playback of each keyword. After the audio playback ends, the text recognition results stop scrolling, and the progress bar positions itself at the end of the text recognition results; all keyword highlighting styles are automatically canceled, and the text returns to normal display; if the user replays the audio, the entire synchronous scrolling and keyword highlighting process is retried, maintaining consistency with the first playback.
[0171] In this embodiment, when playing the target voice evaluation, the text recognition results of the target voice evaluation are displayed in a scrolling manner, and when the keyword is played, the keyword is highlighted, so that users can visually capture the key points while listening to the voice evaluation.
[0172] The voice evaluation display method proposed in this application is applied to scenarios where evaluation information of a target object is viewed. As can be seen from the above, in the solution provided in this application, in response to the operation of viewing evaluation information of a target object, the evaluation page of the target object is displayed. The evaluation page displays target evaluation tags and voice evaluations of the target object. The target evaluation tags are obtained based on target recognition information generated after speech recognition of the target voice evaluation. Moreover, the target recognition information is different from the text recognition result of the target voice evaluation. The text recognition result is obtained by converting the target voice evaluation into text and reflects the text information contained in the target voice evaluation. However, the target recognition information is different from the text recognition result and reflects the non-semantic information of the target voice evaluation. Therefore, the target evaluation tags can reflect the evaluation information in the evaluation content that the text recognition result of the target voice evaluation cannot represent. Thus, displaying the target evaluation tags on the evaluation page expands the evaluation dimensions, improves the diversity and comprehensiveness of the evaluation dimensions of the target object, and makes the evaluation information more communicative. Furthermore, when a trigger operation for a target evaluation tag is received, the target voice evaluations related to the target evaluation tag are filtered out and displayed on the evaluation page. This filters out irrelevant evaluations, avoids redundant evaluation content from interfering with the user, and improves the accuracy and focus of the evaluation information display.
[0173] This application also provides a voice evaluation display method, which can be implemented through the interaction between a terminal and a server, see reference. Figure 6 This is a flowchart illustrating the interaction of a voice evaluation display method provided in an embodiment of this application. Figure 6 As shown, the voice evaluation display method of this embodiment may include the following steps:
[0174] In step S610, the server obtains the evaluation content of the target object, performs speech recognition on the speech evaluation in the evaluation content, and obtains recognition information of at least one dimension.
[0175] Among them, at least one dimension includes the emotional dimension, the environmental dimension, the behavioral dimension of the speaker, and the voice evaluation recording status dimension.
[0176] The identification information for the emotion dimension is the emotion recognition result, the identification information for the environment dimension is the acoustic environment recognition result, the identification information for the speaker's behavior dimension is the speaker's behavior recognition result, and the identification information for the voice evaluation recording status dimension is the voice evaluation recording status recognition result.
[0177] The server can perform speech recognition on a certain proportion of the evaluation content, or it can perform speech recognition on all the evaluation content and generate target evaluation labels based on the obtained recognition information.
[0178] The recognition information for each dimension can be divided into multiple types. For example, under the environment dimension, it can be divided into noisy environment, quiet environment, etc., and under the emotion dimension, it can be divided into happy mood, slightly angry mood, etc. Specifically, for each dimension, the number of voice evaluations of each type under that dimension can be counted, and the evaluation label corresponding to the type of voice evaluation with the largest number or highest proportion among each type of voice evaluation can be used as the target evaluation label for that dimension.
[0179] Step S620: Generate target evaluation labels based on identification information of at least one dimension.
[0180] In some embodiments, target evaluation labels can be generated based on the recognition information of each dimension, that is, emotion labels are generated based on the recognition information of the emotion dimension, environment labels are generated based on the recognition information of the environment dimension, behavior labels of the speaker are generated based on the recognition information of the speaker behavior dimension, and recording status labels are generated based on the recognition information of the voice evaluation recording status dimension.
[0181] In other embodiments, combined labels can be generated based on identification information from multiple dimensions as target evaluation labels. For example, target evaluation labels can be generated based on identification information from the emotion dimension and identification information from the behavior dimension of the speaker. If the identification information from the emotion dimension is "happy mood" and the identification information from the behavior dimension of the speaker is "three people", then the generated target evaluation labels could be "multiple people: happy mood", "multiple people recommend", etc.
[0182] In step S630, the terminal responds to the evaluation information viewing operation by sending a page retrieval request for the evaluation page to the server.
[0183] In step S640, the server retrieves the target evaluation tag and the evaluation content for the target object based on the page retrieval request, and returns them to the terminal.
[0184] In step S650, the terminal renders and displays the evaluation page based on the target evaluation tag and the evaluation content for the target object.
[0185] In step S660, the terminal responds to the trigger operation of the target evaluation label by sending an evaluation acquisition request to the server.
[0186] In step S670, the server responds to the evaluation retrieval request by filtering the evaluation content of the target object to obtain the target voice evaluation related to the target evaluation tag.
[0187] In step S680, the server returns the selected target voice evaluations related to the target evaluation tags to the terminal.
[0188] In step S690, the terminal displays the selected target voice evaluation in the evaluation content display area.
[0189] When displaying the target voice evaluation, the terminal also displays the target recognition information corresponding to the target evaluation tag, as well as the target recognition information corresponding to other evaluation tags, in the associated location of the target voice evaluation; other evaluation tags refer to target evaluation tags that have not been triggered. Additionally, the playback count of the target voice evaluation is also displayed.
[0190] In this embodiment, in response to an operation to view evaluation information for a target object, the terminal displays the evaluation page for the target object. The evaluation page displays target evaluation tags and voice evaluations of the target object. The target evaluation tags are generated based on target recognition information produced after speech recognition of the target voice evaluation. This target recognition information differs from the text recognition result of the target voice evaluation. The text recognition result is obtained by converting the target voice evaluation into text, reflecting the textual information contained in the target voice evaluation. The target recognition information, however, differs from the text recognition result and reflects the non-semantic information of the target voice evaluation. Therefore, the target evaluation tags can reflect evaluation information in the evaluation content that the text recognition result of the target voice evaluation cannot represent. Displaying the target evaluation tags on the evaluation page expands the evaluation dimensions, enhancing the diversity and comprehensiveness of the evaluation dimensions for the target object, making the evaluation information more communicative. Furthermore, when a trigger operation for a target evaluation tag is received, target voice evaluations related to the target evaluation tag are filtered and displayed on the evaluation page. This filters out irrelevant evaluations, avoids redundant evaluation content from interfering with the user, and improves the accuracy and focus of the evaluation information display.
[0191] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0192] Based on the same inventive concept, this application also provides a voice evaluation display device for implementing the voice evaluation display method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more voice evaluation display device embodiments provided below can be found in the limitations of the voice evaluation display method described above, and will not be repeated here.
[0193] In one exemplary embodiment, such as Figure 7 As shown, a voice evaluation display device is provided, including: a page display module 710 and an evaluation filtering module 720, wherein:
[0194] The page display module 710 is used to display the evaluation page of the target object; the evaluation page includes target evaluation tags and evaluation content for the target object, the target evaluation tags are used to filter the evaluation content, and the evaluation content includes at least one voice evaluation.
[0195] The evaluation filtering module 720 is used to filter out target speech evaluations related to the target evaluation tag on the evaluation page in response to the trigger operation of the target evaluation tag. The target speech evaluation includes target recognition information used to generate the target evaluation tag. The target recognition information is generated based on the speech recognition of the target speech evaluation. The target recognition information is different from the text recognition result of the target speech evaluation.
[0196] In one embodiment, the target recognition information includes at least one of emotion recognition results, acoustic environment recognition results, speaker behavior recognition results, and voice evaluation recording status recognition results; the target evaluation label is generated based on at least one recognition result.
[0197] In one embodiment, the emotion recognition result includes at least one of the recognition results of emotion polarity and emotion expression mode; the acoustic environment recognition result includes at least one of the recognition results of environmental noise level, spatial characteristics and environmental sound characteristics; the voice subject behavior recognition result includes at least one of the recognition results of the number of voices and the voice mode; and the voice evaluation recording status recognition result includes at least one of the recognition results of recording timing and voice naturalness.
[0198] In one embodiment, the target evaluation label is generated based on at least one recognition result and the number of corresponding target speech evaluations.
[0199] In one embodiment, there are multiple target evaluation tags. The page display module 710 is also used to simultaneously display the target recognition information corresponding to the generated target evaluation tag, as well as the target recognition information corresponding to other evaluation tags, at the associated position of the target voice evaluation; other evaluation tags refer to target evaluation tags that have not been triggered.
[0200] In one embodiment, the device further includes a playback statistics module for counting the number of times the target voice evaluation is played in response to each playback operation of the target voice evaluation; and a playback count display module for displaying the number of times the playback is played.
[0201] In one embodiment, the evaluation page displays the text recognition results of the target voice evaluation, and the text recognition results include highlighted target keywords.
[0202] Each module in the aforementioned voice evaluation display device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0203] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown. For example, computer device 800 can be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0204] Reference Figure 8 The computer device 800 may include one or more of the following components: processing component 802, memory 804, power supply component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.
[0205] Processing component 802 typically controls the overall operation of computer device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0206] Memory 804 is configured to store various types of data to support the operation of computer device 800. Examples of such data include instructions for any application or method operating on computer device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, optical disk, or graphene storage.
[0207] Power supply component 806 provides power to various components of computer device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to computer device 800.
[0208] Multimedia component 808 includes a screen that provides an output interface between the computer device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the computer device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0209] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when computer device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0210] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0211] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of computer device 800. For example, sensor assembly 814 may detect the on / off state of computer device 800, the relative positioning of components such as the display and keypad of computer device 800, changes in position of computer device 800 or its components, the presence or absence of user contact with computer device 800, orientation or acceleration / deceleration of device 800, and temperature changes of computer device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0212] Communication component 816 is configured to facilitate wired or wireless communication between computer device 800 and other devices. Computer device 800 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0213] In an exemplary embodiment, the computer device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0214] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of a computer device 800 to perform the above-described method. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0215] In an exemplary embodiment, a computer program product is also provided, the computer program product including instructions that can be executed by a processor 820 of a computer device 800 to perform the above-described method.
[0216] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processing components involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0217] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0218] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for displaying voice evaluation, characterized in that, The method includes: The evaluation page of the target object is displayed; the evaluation page includes target evaluation tags and evaluation content for the target object, the target evaluation tags are used to filter the evaluation content, and the evaluation content includes at least one voice evaluation. In response to the triggering operation of the target evaluation tag, target voice evaluations related to the target evaluation tag are filtered on the evaluation page; the target voice evaluation includes target recognition information for generating the target evaluation tag, the target recognition information is generated based on the speech recognition of the target voice evaluation, and the target recognition information is different from the text recognition result of the target voice evaluation.
2. The method according to claim 1, characterized in that, The target recognition information includes at least one of the following: emotion recognition results, acoustic environment recognition results, voice subject behavior recognition results, and voice evaluation recording status recognition results; the target evaluation label is generated based on at least one recognition result.
3. The method according to claim 2, characterized in that, The emotion recognition result includes at least one of the recognition results of emotion polarity and emotion expression mode; The acoustic environment recognition results include at least one of the recognition results of environmental noise level, spatial characteristics, and environmental sound characteristics; The identification results of the speaking subject behavior include at least one of the identification results of the number of speakers and the speaking method; The recognition result of the voice evaluation recording status includes at least one of the recognition results of recording timing and voice naturalness.
4. The method according to claim 2, characterized in that, The target evaluation label is generated based on at least one recognition result and the number of corresponding target speech evaluations.
5. The method according to claim 2, characterized in that, There are multiple target evaluation tags; after filtering out target speech evaluations related to the target evaluation tags on the evaluation page, the following is also included: At the associated location of the target speech evaluation, the target recognition information corresponding to the generated target evaluation label, as well as the target recognition information corresponding to other evaluation labels, are simultaneously displayed; the other evaluation labels refer to the target evaluation labels that have not been triggered.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: In response to each playback operation of the target speech evaluation, the number of times the target speech evaluation has been played is counted; The play count is displayed.
7. The method according to claim 1, characterized in that, The evaluation page displays the text recognition results of the target voice evaluation, and the text recognition results include highlighted target keywords.
8. A voice evaluation display device, characterized in that, The device includes: A page display module is used to display the evaluation page of the target object; the evaluation page includes target evaluation tags and evaluation content for the target object, the target evaluation tags are used to filter the evaluation content, and the evaluation content includes at least one voice evaluation. The evaluation filtering module is used to filter out target voice evaluations related to the target evaluation tag on the evaluation page in response to the trigger operation of the target evaluation tag; the target voice evaluation includes target recognition information for generating the target evaluation tag, the target recognition information is generated based on the speech recognition of the target voice evaluation, and the target recognition information is different from the text recognition result of the target voice evaluation.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the voice evaluation display method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the voice evaluation display method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the voice evaluation display method according to any one of claims 1 to 7.