A method, device, equipment and medium for processing voice data

By generating a voice identification queue on the application client and automatically sending a request, the server performs voice conversion processing, which solves the problems of low conversion efficiency caused by manual operation and network instability in the existing technology, and realizes efficient automatic conversion of voice messages.

CN113393842BActive Publication Date: 2025-09-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011295049.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-18
Publication Date
2025-09-16
Estimated Expiration
2040-11-18

AI Technical Summary

Technical Problem

Existing instant messaging application clients require users to manually initiate requests during the voice-to-text conversion process, and voice data upload fails when the network is unstable, resulting in low conversion efficiency.

Method used

When the application client obtains a voice message, a voice identifier is generated and added to the initial identifier queue. A voice conversion request is generated based on the queue position. The server performs the conversion and automatically outputs the converted text information on the conversation interface to avoid local upload.

Benefits of technology

It realizes active and efficient conversion of voice messages, improves conversion efficiency, and solves the problem of upload failure under unstable network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113393842B_ABST
    Figure CN113393842B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a method, apparatus, device and medium for processing voice data, which relates to the field of artificial intelligence. The method includes: when an application client obtains a voice message on a conversation interface, obtaining a voice identifier corresponding to the voice message, adding the voice identifier to an initial identifier queue, and using the initial identifier queue with the added voice identifier as a target identifier queue; based on the queue position of the voice identifier in the target identifier queue, generating a voice conversion request carrying the voice identifier, sending the voice conversion request to a server so that the server obtains the converted text information corresponding to the voice identifier; receiving the converted text information returned by the server, and outputting the converted text information to the location area where the voice message is located in the conversation interface; the voice message in the location area has an associated relationship with the converted text information. By adopting the present application, active access to the converted text information can be achieved, and the conversion efficiency of the voice message can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a method, apparatus, device, and medium for processing voice data. Background Art

[0002] When converting voice messages to text, existing instant messaging applications (e.g., social networking clients) require the target user (e.g., User A) to manually initiate a voice conversion request for a voice message (e.g., Voice Message 1) in order to obtain the text message (e.g., Text Message 1) corresponding to Voice Message 1. This means that existing voice-to-text solutions are passive and make it difficult to achieve active text delivery.

[0003] Furthermore, existing social client voice-to-text solutions rely on the social client's local data when converting the voice message 1 received by user A. When initiating a voice conversion request, the social client must upload the local voice data of voice message 1 to the server. Obviously, in unstable network environments, this can result in slow voice data upload speeds, making it impossible to quickly provide text information to user A. The voice data upload may even fail, further reducing the efficiency of voice message conversion. Summary of the Invention

[0004] The embodiments of the present application provide a voice data processing method, apparatus, device, and medium, which can realize active access to converted text information and improve the conversion efficiency of voice messages.

[0005] An embodiment of the present application provides a method for processing voice data, including:

[0006] When the application client obtains the voice message of the conversation interface, it obtains the voice identifier corresponding to the voice message, adds the voice identifier to the initial identifier queue, and uses the initial identifier queue with the added voice identifier as the target identifier queue;

[0007] generating a voice conversion request carrying the voice identifier based on the queue position of the voice identifier in the target identifier queue, and sending the voice conversion request to the server so that the server obtains converted text information corresponding to the voice identifier based on the voice conversion request;

[0008] The converted text information returned by the server is received, and the converted text information is output to the location area where the voice message is located in the conversation interface; there is an association relationship between the voice message in the location area and the converted text information.

[0009] An embodiment of the present application provides a voice data processing device, including:

[0010] The voice acquisition module is used to obtain the voice identifier corresponding to the voice message when the application client obtains the voice message of the conversation interface, add the voice identifier to the initial identifier queue, and use the initial identifier queue with the added voice identifier as the target identifier queue;

[0011] a request sending module, configured to generate a voice conversion request carrying the voice identifier based on the queue position of the voice identifier in the target identifier queue, and send the voice conversion request to the server, so that the server obtains converted text information corresponding to the voice identifier based on the voice conversion request;

[0012] The text receiving module is used to receive the converted text information returned by the server and output the converted text information to the location area where the voice message is located in the conversation interface; there is an association relationship between the voice message in the location area and the converted text information.

[0013] The conversation interface includes a second user associated with the first user; the initial identification queue includes a first subqueue and a second subqueue; the first subqueue is used to store a first voice identification; the first voice identification is used to represent the identification of a first voice message for which a voice conversion request is to be sent in the application client; the second subqueue is used to store a second voice identification; the second voice identification is used to represent the identification of a second voice message for which a voice conversion request has been sent in the application client;

[0014] The voice acquisition module includes:

[0015] A voice receiving unit, configured for the application client corresponding to the first user to receive the voice message forwarded by the second user through the server, and to receive the voice identifier configured by the server for the voice message;

[0016] a timestamp determination unit, configured to obtain a voice conversion condition associated with the conversation interface, determine the received voice identifier as a target voice identifier based on the voice conversion condition, determine the voice message received by the application client as a target voice message, and record a reception timestamp corresponding to the target voice message as a target reception timestamp;

[0017] an identifier adding unit, configured to determine, based on a target reception timestamp, a queue position of a target voice identifier of a target voice message in a first subqueue containing an identifier of the first voice message, and add the target voice identifier to the first subqueue based on the queue position to obtain an initial first subqueue;

[0018] The queue determining unit is configured to determine a target identification queue based on the initial first subqueue and the second subqueue containing the identification of the second voice message.

[0019] The request priority of the second subqueue is greater than the request priority of the first subqueue;

[0020] The voice acquisition module also includes:

[0021] A first triggering unit is configured to respond to a triggering operation on the conversation interface where the second user is located, output a target voice message to the conversation interface, and obtain an initial level adjustment instruction in the voice conversion condition;

[0022] a first adjustment unit, configured to determine, based on the initial level adjustment instruction, a queue position of the target voice identifier in the initial first subqueue as a first position, and adjust the queue position of the target voice identifier in the initial first subqueue from the first position to a second position, thereby obtaining an adjusted initial first subqueue; wherein the request priority of the identifier corresponding to the second position is greater than the request priority of the identifier corresponding to the first position;

[0023] The first updating unit is configured to update the target identification queue based on the adjusted initial first sub-queue and the second sub-queue.

[0024] Among them, the voice acquisition module also includes:

[0025] A second triggering unit is configured to respond to a triggering operation on a target voice message in a conversation interface and obtain a target level adjustment instruction in a voice conversion condition;

[0026] a second adjustment unit, configured to, based on the target level adjustment instruction, determine the adjusted initial first subqueue as the target first subqueue, and adjust the queue position of the target voice identifier in the target first subqueue from the second position to a third position, thereby obtaining an adjusted target first subqueue; wherein the request priority of the identifier corresponding to the third position is greater than the request priority of the identifier corresponding to the second position;

[0027] The second updating unit is configured to update the updated target identification queue based on the adjusted target first sub-queue and second sub-queue.

[0028] The target identifier queue includes a pending identifier queue and a requested identifier queue; the voice identifier is located in the pending identifier queue; the requested identifier queue includes M queue positions; one queue position in the requested identifier queue is used to store an identifier of a voice message to be converted; M is the total number of identifiers of voice messages to be converted for which voice conversion requests have been sent;

[0029] The request sending module includes:

[0030] An information receiving unit is configured to receive conversion success information returned by the server for M voice messages to be converted for which voice conversion requests have been sent, and record the number of conversions received in the conversion success information as N; N is a positive integer less than or equal to M;

[0031] a position determination unit, configured to obtain a queue position of the voice identifier in the queue of pending identifiers of the target identifier queue, and determine a target queue position of the voice identifier in the queue of requested identifiers when the queue position of the voice identifier meets a voice conversion condition;

[0032] The request generating unit is used to add the voice identifier to the requested identifier queue based on the target queue position, generate a voice conversion request carrying the voice identifier based on the requested identifier queue with the added voice identifier, and send the voice conversion request to the server.

[0033] The device further comprises:

[0034] The identifier deletion module is used to obtain target conversion success information for the voice message when receiving the converted text information returned by the server, and delete the voice identifier from the target identifier queue based on the target conversion success information.

[0035] An embodiment of the present application provides a method for processing voice data, including:

[0036] When a voice message from an application client is obtained, a voice identifier corresponding to the voice message is generated, and the voice message and the voice identifier are sent to a user terminal, so that the user terminal adds the voice identifier to an initial identifier queue and uses the initial identifier queue with the added voice identifier as a target identifier queue;

[0037] receiving a voice conversion request sent by a user terminal and obtaining a voice identifier from the voice conversion request; the voice conversion request is generated based on a queue position of the voice identifier in a target identifier queue;

[0038] When a voice message corresponding to the voice identifier is found, the voice message is converted to obtain converted text information corresponding to the voice message;

[0039] The converted text information is returned to the user terminal, so that the user terminal outputs the converted text information to the location area where the voice message is located in the conversation interface of the application client.

[0040] An embodiment of the present application provides a voice data processing device, including:

[0041] The voice sending module is used to generate a voice identifier corresponding to the voice message when obtaining a voice message from the application client, and send the voice message and the voice identifier to the user terminal so that the user terminal adds the voice identifier to the initial identifier queue and uses the initial identifier queue with the added voice identifier as the target identifier queue;

[0042] a request receiving module, configured to receive a voice conversion request sent by a user terminal and obtain a voice identifier from the voice conversion request; the voice conversion request is generated based on a queue position of the voice identifier in a target identifier queue;

[0043] The text acquisition module is used to convert the voice message to obtain the converted text information corresponding to the voice message when the voice message corresponding to the voice identifier is queried;

[0044] The text sending module is used to return the converted text information to the user terminal, so that the user terminal outputs the converted text information to the location area where the voice message is located in the conversation interface of the application client.

[0045] One aspect of an embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the method in one aspect of the embodiment of the present application.

[0046] On one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the steps of the method in one aspect of the embodiment of the present application are executed.

[0047] In one aspect, an embodiment of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in various optional embodiments of the above-mentioned aspect.

[0048] In an embodiment of the present application, when the application client obtains a voice message on the conversation interface, the user terminal can obtain the voice identifier corresponding to the voice message, add the voice identifier to the initial identifier queue, and use the initial identifier queue with the added voice identifier as the target identifier queue. The voice message on the conversation interface can be a voice message sent by the application client, or a voice message received by the application client. Furthermore, the user terminal can generate a voice conversion request carrying the voice identifier based on the queue position of the voice identifier in the target identifier queue, and send the voice conversion request to the server, so that the server queries the voice message corresponding to the voice identifier based on the voice conversion request, and obtains the converted text information after converting the voice message. Furthermore, the user terminal can receive the converted text information returned by the server, and output the converted text information to the location area where the voice message is located in the conversation interface. There is an association relationship between the voice message and the converted text information in the location area. For example, the voice message and the converted text information can have an adjacent position relationship in the conversation interface. It should be understood that by introducing the target identification queue, when a voice message and a voice identifier corresponding to the voice message are obtained, the first user corresponding to the user terminal does not need to perform a trigger operation. The user terminal can output the converted text information corresponding to the voice message in the conversation interface of the application client, and then automatically convert the voice message into its corresponding converted text information, so as to achieve active access to the converted text information. Among them, when converting and processing the voice message based on the voice identifier in the target identification queue, the embodiment of the present application does not require the user terminal to upload the voice message in the local memory to the server. Instead, the server intelligently queries the voice message corresponding to the voice identifier based on the uploaded voice identifier and converts and processes the queried voice message. In this way, problems such as failure to upload voice messages can be solved in the case of an unstable network environment, thereby effectively improving the conversion efficiency of voice messages. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0050] Figure 1 This is a schematic diagram of a network architecture provided by an embodiment of the present application;

[0051] Figure 2 This is a schematic diagram of a data interaction scenario provided by an embodiment of the present application;

[0052] Figure 3This is a flow chart of a method for processing voice data provided by an embodiment of the present application;

[0053] Figure 4 This is a schematic diagram of a scenario for adding a voice identifier provided in an embodiment of the present application;

[0054] Figure 5 This is a schematic diagram of a scenario in which a user opens a session, provided by an embodiment of the present application;

[0055] Figure 6 This is a schematic diagram of a scenario in which a user makes a selection provided in an embodiment of the present application;

[0056] Figure 7 This is a schematic diagram of a scenario for receiving converted text information provided by an embodiment of the present application;

[0057] Figure 8 This is a flow chart of a method for processing voice data provided by an embodiment of the present application;

[0058] Figure 9 This is a schematic diagram of a scenario for forwarding a voice message provided by an embodiment of the present application;

[0059] Figure 10 This is a flow chart of a speech-to-text solution provided in an embodiment of the present application;

[0060] Figure 11 This is a schematic diagram of a scenario for performing speech-to-text conversion provided by an embodiment of the present application;

[0061] Figure 12 This is a structural diagram of a voice data processing device provided in an embodiment of the present application;

[0062] Figure 13 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application;

[0063] Figure 14 This is a structural diagram of a voice data processing device provided in an embodiment of the present application;

[0064] Figure 15 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application;

[0065] Figure 16 This is a voice data processing system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0066] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0067] For details, see Figure 1 , Figure 1 This is a schematic diagram of a network architecture provided by an embodiment of the present application. Figure 1 As shown, the network architecture may include a server 3000 and a user terminal cluster. The user terminal cluster may specifically include one or more user terminals, and the number of user terminals in the user terminal cluster is not limited. Figure 1 As shown, the multiple user terminals may specifically include user terminal 3000a, user terminal 3000b, user terminal 3000c, ..., user terminal 3000n. User terminal 3000a, user terminal 3000b, user terminal 3000c, ..., user terminal 3000n may be respectively connected to server 3000 directly or indirectly via a wired or wireless communication, so that each user terminal may exchange data with server 3000 via the network connection.

[0068] Among them, such as Figure 1 The server 3000 shown can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0069] Among them, Figure 1 Each user terminal in the user terminal cluster shown may include: a smart phone, a tablet computer, a laptop computer, or other intelligent terminal with voice data processing capabilities. Figure 1 Each user terminal in the user terminal cluster shown can be integrated with an application client. When the application client runs in each user terminal, it can be respectively connected to the above-mentioned client terminal based on the client / server (C / S) architecture. Figure 1The application client can be understood as an instant messaging client that can load and display voice data. For example, the application client here can specifically include: a social client (for example, a WeChat client), an office client (for example, an enterprise WeChat client), an entertainment client (for example, a game client), and a car client, etc.

[0070] The speech data processing method provided in the embodiment of the present application may involve the speech technology direction in the field of artificial intelligence. It is understood that the so-called artificial intelligence (AI) refers to the use of digital computers or computer equipment controlled by digital computers (for example, Figure 1 Artificial intelligence is a new discipline that uses the theories, methods, technologies, and application systems of computer science to simulate, extend, and expand human intelligence. In other words, artificial intelligence is a comprehensive field of computer science that seeks to understand the essence of intelligence and create new intelligent machines that can respond in a manner similar to human intelligence. Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0071] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0072] Key speech technologies include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction.

[0073] It is understandable that the voice data processing method provided in the embodiment of the present application may also involve the field of cloud technology. The so-called cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network in a wide area network or a local area network to realize the calculation, storage, processing and sharing of data. Among them, cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. applied based on the cloud computing business model, which can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the rapid development and application of the Internet industry, each item may have its own identification mark in the future, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately, and all kinds of industry data require strong system backing support to be realized through cloud computing.

[0074] For ease of understanding, the embodiment of the present application can select any user terminal from the above user terminal cluster as the first terminal as an example to illustrate the specific process of voice data processing in the first terminal. For example, the embodiment of the present application can select any user terminal from the above user terminal cluster as the first terminal. Figure 1 User terminal 3000c in the user terminal cluster shown is the first terminal. It should be understood that in this embodiment of the application, the user who logs in to the application client using the first account information (e.g., account information 1) in the first terminal may be referred to as the first user, that is, the first user may be the user using the first terminal. It should be understood that in this embodiment of the application, the first user may be the user who receives the voice message through the application client in the first terminal, that is, the message recipient.

[0075] It is understandable that in the embodiment of the present application, the user who logs in to the application client through the second account information (for example, account information 2) can be referred to as the second user, and the user terminal corresponding to the second user can be referred to as the second terminal, that is, the second user can be the user who uses the second terminal. In the embodiment of the present application, any user terminal in the above user terminal cluster can be selected as the second terminal. For example, in the embodiment of the present application, Figure 1 The user terminal 3000a in the user terminal cluster shown is used as the second terminal. It is understandable that the second user in the embodiment of the present application can be a user who sends a voice message through the application client in the second terminal, that is, the message sender.

[0076] It should be understood that the first user in the embodiments of the present application can serve as both the message receiver and the message sender. For example, the first user can become the message receiver through the application client in the first terminal, and the first user can also become the message sender through the application client in the first terminal. Similarly, the second user in the embodiments of the present application can serve as both the message sender and the message receiver. For example, the second user can become the message sender through the application client in the second terminal, and the second user can also become the message receiver through the application client in the second terminal.

[0077] It is understandable that the message sender and the message receiver can be connected through a server (for example, the aforementioned server 3000), and the server synchronizes the voice message from the user terminal corresponding to the message sender (for example, the second terminal corresponding to the second user) to the user terminal corresponding to the message receiver (for example, the first terminal corresponding to the first user), and in subsequent steps, the voice message can be converted so that the first user and the second user can directly obtain the converted text information corresponding to the voice message. In particular, the first terminal and the second terminal both run the application client corresponding to the server, and the application client can realize the sending and receiving of voice messages between the first terminal and the second terminal.

[0078] For easier understanding, see Figure 2 , Figure 2 This is a schematic diagram of a data interaction scenario provided by an embodiment of the present application. Figure 2 The server shown can be the above Figure 1 The server 3000 in the corresponding embodiment is as follows: Figure 2 The user terminal X shown can be the above Figure 1 For ease of understanding, any user terminal in the user terminal cluster of the corresponding embodiment is described in the embodiment of the present application. Figure 1 The user terminal 3000c shown is taken as the user terminal X as an example to illustrate Figure 2 The specific process of data interaction between the user terminal X and the server is shown.

[0079] It is understandable that if Figure 2 The application database shown may include multiple databases, and the multiple databases may include Figure 2Databases 10a, 10b, ..., and 10n are shown. This means that the application database can be used to store voice content 1 corresponding to different voice messages in an application client (e.g., an office client). For example, database 10a can be used to store the voice content corresponding to voice message x1, database 10b can be used to store voice content 2 corresponding to voice message x2, ..., and database 10n can be used to store voice content n corresponding to voice message xn (not shown).

[0080] The application database can be simply referred to as a database. A database can be thought of as a digital filing cabinet—a place where electronic files are stored, allowing users to add, query, update, and delete data. A "database" is a collection of data stored in a specific manner, shared by multiple users, with minimal redundancy, and independent of the application.

[0081] like Figure 2 As shown, when user terminal X receives a voice message on conversation interface 2a, it can add the voice identifier corresponding to the voice message to the initial identifier queue. It should be understood that in conversation interface 2a, user terminal X can receive one or more voice messages, and the specific number of received voice messages is not limited here.

[0082] For ease of understanding, here we take the number of received voice messages as an example. Specifically, the voice messages here may include but are not limited to Figure 2 As shown in the voice message x1, voice message x2, voice message x3 and voice message x4, at this time, the user terminal X can also receive the voice identifiers of these voice messages at the same time, that is, one voice message corresponds to one voice identifier. Then, the user terminal X can add the voice identifier corresponding to the received voice message x1 as X1, the voice identifier corresponding to the voice message x2 as X2, the voice identifier corresponding to the voice message x3 as X3, and the voice identifier corresponding to the voice message x4 as X4 in the order of the received reception timestamps to the initial identifier queue, for example, it can be added to the first subqueue of the initial identifier queue to obtain the target identifier queue. Among them, the initial identifier queue here may include a first subqueue and a second subqueue, the first subqueue can be used to store the voice identifier of the voice conversion request to be sent, and the second subqueue can be used to store the voice identifier of the voice conversion request that has been sent. It should be understood that in the embodiment of the present application, the first subqueue corresponding to the currently added new voice identifier can be collectively referred to as the pending identifier queue, and the second subqueue corresponding to the voice identifier of the voice message (i.e., the voice message to be converted) that has currently sent the voice conversion request and is still in the pending conversion state can be collectively referred to as the requested identifier queue.

[0083] It should be understood that if the user terminal adds the four voice identifiers to the first subqueue at time T2 to obtain the queue of identifiers to be requested, then at the moment before time T2 (for example, time T1, before the four voice identifiers are added to the first subqueue), the first subqueue may specifically include L queue positions, each of which may be used to store a voice identifier for a voice conversion request to be sent. Here, L is a positive integer, and the value of L is not limited. For ease of understanding, for example, at time T1, if the first subqueue currently stores 6 voice identifiers for voice conversion requests to be sent, it means that 6 of the L queue positions of the first subqueue are occupied, and (L-6) positions are unoccupied. Therefore, when the above-mentioned 4 voice identifiers are added to the first subqueue at time T2, 10 (i.e., 6+4) of the L queue positions of the first subqueue will be occupied. In this embodiment of the present application, the first subqueue at time T2 can be used as the above-mentioned queue of pending request identifiers.

[0084] For another example, the second subqueue may specifically include M queue positions, and each queue position in the second subqueue may correspond to a voice identifier for which a voice conversion request has been sent. M is a positive integer. For ease of understanding, the second subqueue is taken as an example where it includes 5 (for example, M is equal to 5) queue positions at the aforementioned time T1. This means that 5 voice identifiers for which voice conversion requests have been sent are currently stored in the M positions of the second subqueue, for example, voice identifier Y1, voice identifier Y2, voice identifier Y3, voice identifier Y4, and voice identifier Y5. If, at time T2, the voice identifiers of the 5 voice messages to be converted for which voice conversion requests have been sent are still in a pending conversion state, then at time T2, the second subqueue corresponding to the 5 voice identifiers for which voice conversion requests have been sent and are in a pending conversion state may be used as the aforementioned requested identifier queue.

[0085] It should be understood that, optionally, when the server successfully obtains the converted text information of three (e.g., N equals 3) voice identifiers that have sent voice conversion requests in the requested identifier queue at the next moment after time T2 (e.g., time T3) (e.g., converted text information 1 of voice identifier Y1, converted text information 2 of voice identifier Y2, and converted text information 3 of voice identifier Y3), it can send conversion success information for these three voice identifiers to user terminal X. At this time, user terminal X can output the converted text information of these three voice identifiers that have sent voice conversion requests to the conversation interface of user terminal X for display. It should be understood that when user terminal X obtains the conversion success information for these three voice identifiers, it can further release the queue positions occupied by these three voice identifiers (i.e., voice identifier Y1, voice identifier Y2, and voice identifier Y3) in the requested identifier queue (i.e., the target identifier queue). This means that at this time, there are currently three queue positions in the requested identifier queue that are unoccupied. In this way, the user terminal can adjust the queue position of each currently stored voice identifier in the target identifier queue according to the preset voice conversion conditions. For example, in the pending identifier queue of the target identifier queue, according to the order of the receiving timestamps of these 10 voice identifiers, three voice identifiers with the highest priority (for example, voice identifier X1, voice identifier X2, voice identifier X3) can be selected from the pending identifier queue and added to the aforementioned requested identifier queue, so that voice conversion requests for these three voice identifiers (for example, voice identifier X1, voice identifier X2, voice identifier X3) can be generated, and then these three voice conversion requests can be sent to Figure 2 The server shown.

[0086] like Figure 2 As shown, the server can receive voice conversion requests for these three voice identifiers (for example, voice identifier X1, voice identifier X2, and voice identifier X3), and can distribute the voice content corresponding to the voice identifiers in these three voice conversion requests to the voice processing server cluster to improve the voice conversion efficiency through distributed processing. For example, the language processing server cluster can include one or more voice processing servers. For ease of understanding, the multiple voice processing servers here can specifically include voice processing server 100a, voice processing server 100b, and voice processing server 100c. For example, Figure 2 The server shown can forward the voice content 1 corresponding to the voice identifier X1 found in the database 10a to the voice processing server 100a, so that the voice processing server 100a can convert the voice content 1 and then convert the converted text information (for example, Figure 2 The text information shown in 1) is returned to Figure 2The server shown in FIG. 1 can transmit the converted text information (eg, Figure 2 The text information shown in 1) is output to Figure 2 The conversation interface of the user terminal X shown (for example, Figure 2 As shown in the conversation interface 2b). For example, for example, Figure 2 The server shown can forward the voice content 2 corresponding to the voice identifier X2 found in the database 10b to the voice processing server 100b, so that the voice processing server 100b can convert the voice content 2 and then convert the converted text information (for example, Figure 2 The text information shown in 2) is returned to Figure 2 The server shown in FIG. 1 can transmit the converted text information (eg, Figure 2 The text information shown in 2) is output to Figure 2 The conversation interface of the user terminal X shown (for example, Figure 2 The conversation interface 2b shown in FIG. 2b). Similarly, for example, Figure 2 The server shown can forward the voice content 3 corresponding to the voice identifier X3 found in the database 10c to the voice processing server 100c, so that the voice processing server 100c can convert the voice content 3. It should be understood that in this embodiment of the application, the converted text information can be collectively referred to as converted text information.

[0087] The speech processing server 100a, the speech processing server 100b and the speech processing server 100c here can be the same speech processing server for providing conversion processing services, or can be independent speech processing servers for providing conversion processing services, which will not be limited here. Figure 2 As shown, one or more speech processing servers with conversion processing services can run on Figure 2 The server shown can also be independent of Figure 2 The server exists only for the server shown and will not be limited here.

[0088] The voice conversion conditions here may include one or more of the following conversion conditions: a first conversion condition, a second conversion condition, and a third conversion condition. It is understood that, for the voice identifier of the same voice message currently in the first sub-queue, if the queue position of the voice identifier (e.g., queue position 1) is adjusted using the three conversion conditions respectively, the request priority of the new queue position (e.g., queue position A) of the voice identifier obtained by adjusting using the first conversion condition will be higher than the request priority of the new queue position (e.g., queue position B) of the voice identifier obtained by adjusting using the second conversion condition; at the same time, the request priority of the new queue position (e.g., queue position B) of the voice identifier obtained by adjusting using the second conversion condition will also be higher than the request priority of the new queue position (e.g., queue position C) of the voice identifier obtained by adjusting using the third conversion condition. It is understood that in the first sub-queue, the request priority of queue position C will be higher than the request priority of the original queue position 1 of the voice identifier.

[0089] For example, the first conversion condition can be understood as when the user corresponding to user terminal X performs a trigger operation on a certain voice message (e.g., voice message 1) in the conversation interface, the queue position 1 of the voice identifier of voice message 1 can be adjusted in the target identifier queue according to the first conversion condition of the preset voice conversion condition, so that it can be added to the requested identifier queue as quickly as possible. For another example, the second conversion condition can be understood as when the user opens the current conversation interface through user terminal X, user terminal X can adjust the queue position 2 of the voice identifier of the voice message read when opening the current conversation interface (e.g., changing the reading status of voice message 2 from unread to read) in the target identifier queue according to the second conversion condition of the preset voice conversion condition, so that it can be added to the requested identifier queue relatively quickly. For another example, the third conversion condition can be understood as user terminal X adjusting the queue position 3 of the voice identifier of the voice message (for example, voice message 3) that meets the third conversion condition in the target identifier queue according to the order of the reception timestamps of the received voice messages, in accordance with the third conversion condition of the preset voice conversion condition, so that it can be added to the requested identifier queue relatively quickly.

[0090] In an embodiment of the present application, the user terminal can send a voice conversion request to the server based on the voice identifier in the target identifier queue. The user terminal does not need to upload the voice message in the local memory to the server. Instead, the server intelligently queries the voice message corresponding to the voice identifier based on the uploaded voice identifier and converts the queried voice message, thereby improving the conversion efficiency of the voice message.

[0091] Among them, in the embodiment of the present application, the user terminal integrated with the application client obtains the voice message, and converts the voice message through the user terminal and the server to obtain the specific process of converting the text information. Figures 3 to 11 The corresponding embodiment.

[0092] For further information, see Figure 3 , Figure 3 This is a flow chart of a method for processing voice data provided by an embodiment of the present application. Figure 3 As shown, the method can be executed by a computer device, which can be a user terminal installed with the above-mentioned office client, and the user terminal can be the above-mentioned Figure 2 The user terminal X in the corresponding embodiment; optionally, the computer device may also be a server corresponding to the office client, which may be the above-mentioned Figure 2 The server in the corresponding embodiment. In other words, the method involved in the embodiment of the present application can be executed by the user terminal, or by the server, or by the user terminal and the server together. For ease of understanding, this embodiment is described by taking the method executed by the user terminal as an example to illustrate the specific process of obtaining the converted text information corresponding to the voice message in the user terminal. Among them, the method can at least include the following steps S101-step S103:

[0093] Step S101: When the application client obtains a voice message on the conversation interface, it obtains the voice identifier corresponding to the voice message, adds the voice identifier to the initial identifier queue, and uses the initial identifier queue with the added voice identifier as the target identifier queue;

[0094] Specifically, an application client in a user terminal corresponding to a first user can receive a voice message forwarded by a second user via a server and receive a voice identifier configured by the server for the voice message. The conversation interface includes the second user associated with the first user. An initial identifier queue includes a first subqueue and a second subqueue. The first subqueue is configured to store a first voice identifier, which represents the identifier of a first voice message for which a voice conversion request is pending from the application client; the second subqueue is configured to store a second voice identifier, which represents the identifier of a second voice message for which a voice conversion request has been sent from the application client. Furthermore, the user terminal can obtain a voice conversion condition associated with the conversation interface, determine the received voice identifier as a target voice identifier based on the voice conversion condition, determine the voice message received by the application client as a target voice message, and record the reception timestamp corresponding to the target voice message as a target reception timestamp. Furthermore, the user terminal can determine, based on the target reception timestamp, the queue position of the target voice identifier of the target voice message within the first subqueue containing the identifier of the first voice message, and add the target voice identifier to the first subqueue based on the queue position, thereby obtaining an initial first subqueue. Furthermore, the user terminal may determine a target identification queue based on the initial first subqueue and the second subqueue containing the identification of the second voice message.

[0095] For easier understanding, see Figure 4 , Figure 4 This is a schematic diagram of a scenario for adding a voice identifier provided by an embodiment of the present application. Figure 4 As shown, the initial identifier queue includes a first subqueue and a second subqueue. The first subqueue can be used to store the voice identifiers of pending voice conversion requests, and the second subqueue can be used to store the voice identifiers of already sent voice conversion requests. The voice identifiers of pending voice conversion requests can be collectively referred to as the identifiers of first voice messages, and the voice identifiers of already sent voice conversion requests can be collectively referred to as the identifiers of second voice messages.

[0096] The first sub-queue may specifically include L queue positions (for example, Figure 4 6 queue positions shown, that is, L is equal to 6), the second sub-queue may specifically include M queue positions (for example, Figure 4 Assume that L and M are positive integers. Figure 4 The time corresponding to the initial identification queue shown is time T1. At this time, the first sub-queue can store one voice identification for which a voice conversion request is to be sent, for example, voice identification Z1; the second sub-queue can store three voice identifications for which voice conversion requests have been sent, for example, voice identification Y1, voice identification Y2, voice identification Y3 and voice identification Y4.

[0097] It should be understood that the user terminal can obtain the voice message of the conversation interface, and at the next moment after moment T1 (for example, moment T2), the voice message is determined as the target voice message, the voice identifier corresponding to the target voice message is determined as the target voice identifier, and the target voice identifiers (for example, voice identifier X1, voice identifier X2, voice identifier X3 and voice identifier X4) are added to the first subqueue in the order of the target reception timestamp (i.e., reception timestamp) to obtain the initial first subqueue {Z1, X1, X2, X3, X4}. It should be understood that the embodiment of the present application can collectively refer to the first subqueue (for example, the initial first subqueue) to which the new voice identifier has been added as the pending identifier queue, and the above-mentioned second subqueue and the initial first subqueue as the target identifier queue, i.e. Figure 4 The number of voice messages obtained by the user terminal can be one or more, and the specific number of voice messages obtained is not limited here. Furthermore, in subsequent steps, the user terminal can send a voice conversion request carrying the target voice identifier to the server based on the voice conversion condition.

[0098] Among them, it can be understood that the earlier the receiving timestamp of the voice identifier obtained by the user terminal, the higher the request priority of the queue position where the target voice identifier is located when the target voice identifier is added to the first sub-queue. For example, the target receiving timestamp of the target voice identifier X1 is obtained as the T1 moment, and the receiving timestamp of the target voice identifier X2 is obtained as the target T2 moment. If the T1 moment is a moment before the T2 moment, then after the target voice identifier X1 and the target voice identifier X2 are added to the first sub-queue, the request priority of the queue position where the target voice identifier X1 is located is greater than the request priority of the queue position where the target voice identifier X2 is located. Figure 4 As shown, by the same token, the request priority of the queue position where the target voice identifier X2 is located is greater than the request priority of the queue position where the target voice identifier X3 is located, and the request priority of the queue position where the target voice identifier X3 is located is greater than the request priority of the queue position where the target voice identifier X4 is located.

[0099] It is understandable that the user terminal can respond to a trigger operation (e.g., a first trigger operation and a second trigger operation) performed on a first user (the first user here can be the user using the user terminal) and change the request priority of the target voice identifier in the initial first subqueue of the target identifier queue, that is, adjust the queue position of the target voice identifier in the first subqueue so that the following step S102 can be continued. The first trigger operation and the second trigger operation can include contact operations such as clicking, long pressing, and sliding, and can also include non-contact operations such as voice and gestures, which are not limited in this application.

[0100] It should be understood that the user terminal can respond to the trigger operation for the conversation interface where the second user is located (the trigger operation here can be the first trigger operation), output the target voice message to the conversation interface, and obtain the initial level adjustment instruction in the voice conversion condition. Among them, the request priority of the second subqueue is greater than the request priority of the first subqueue, that is, the request priority of the second subqueue is greater than the request priority of the initial first subqueue. Further, the user terminal can determine the queue position of the target voice identifier as the first position in the initial first subqueue based on the initial level adjustment instruction, and adjust the queue position of the target voice identifier from the first position to the second position in the initial first subqueue to obtain the adjusted initial first subqueue. Among them, the request priority of the identifier corresponding to the second position is greater than the request priority of the identifier corresponding to the first position. Further, the user terminal can update the target identifier queue based on the adjusted initial first subqueue and second subqueue.

[0101] For easier understanding, see Figure 5 , Figure 5 This is a schematic diagram of a scenario in which a user opens a session provided by an embodiment of the present application. Figure 5 As shown, the identification queue 5a can be the target identification queue corresponding to the conversation interface 50a, and the identification queue 5b can be the target identification queue corresponding to the conversation interface 50b. The identification queue 5a can be the target identification queue corresponding to the conversation interface 50b. Figure 4 It is understood that when the user terminal obtains the voice message sent by user "BBB" (and opens the conversation interface corresponding to user "BBB", where the voice message can be voice message z1), the conversation interface 50a can be displayed. At this time, if the target voice message sent by user "AAA" is received (where the target voice message can be voice message x1, voice message x2, voice message x3 and voice message x4), the user terminal can obtain Figure 5 In the identification queue 5a shown, the second sub-queue of the identification queue 5a may include voice identifiers {Y1, Y2, Y3}, and the initial first sub-queue may include voice identifiers {Z1, X1, X2, X3, X4}. Voice identifier X1 is the voice identifier corresponding to voice message x1, voice identifier X2 is the voice identifier corresponding to voice message x2, voice identifier X3 is the voice identifier corresponding to voice message x3, and voice identifier X4 is the voice identifier corresponding to voice message x4.

[0102] It is understandable that if the first user Figure 5The session of the user "AAA" (user "AAA" can be referred to as the second user) shown in the figure performs a first trigger operation (for example, the first trigger operation can be a click operation), then the user terminal can respond to the click operation, open the session interface corresponding to the user "AAA" (that is, the session interface 50b), and adjust the queue positions of the voice identifier X1, voice identifier X2, voice identifier X3 and voice identifier X4 in the identifier queue 5a to change the request priority of the voice identifier X1, voice identifier X2, voice identifier X3 and voice identifier X4, and obtain Figure 5 The identification queue 5b shown, wherein the second sub-queue of the identification queue 5b may include voice identifications {Y1, Y2, Y3}, and the adjusted initial first sub-queue may include voice identifications {X1, X2, X3, X4, Z1}.

[0103] Optionally, it is understandable that if the first user receives a target voice message (the target voice message here may be voice message x5, not shown in the figure) sent by user "CCC" (user "CCC" may be referred to as the third user), then the second sub-queue may include voice identifiers {Y1, Y2, Y3}, and the initial first sub-queue may include voice identifiers {X1, X2, X3, X4, Z1, X5}. Among them, voice identifier X5 is the voice identifier corresponding to voice message x5. If the second user targets Figure 5 After the first trigger operation is performed on the session of user "BBB" shown in the figure, another first trigger operation is performed on the session of user "CCC" (for example, the another first trigger operation can be a click operation). Then, the user terminal can respond to the click operation, open the session interface corresponding to user "CCC", and adjust the queue position of the voice identifier X5 corresponding to the voice message x5 in the identifier queue 5b to change the request priority of the voice identifier X5, wherein the second subqueue of the identifier queue 5b can include the voice identifiers {Y1, Y2, Y3}, and the adjusted initial first subqueue can include the voice identifiers {X5, X1, X2, X3, X4, Z1}. Optionally, the second subqueue of the identifier queue 5b can include the voice identifiers {Y1, Y2, Y3}, and the adjusted initial first subqueue can include the voice identifiers {X1, X2, X3, X4, X5, Z1}.

[0104] It should be understood that the user terminal can respond to the trigger operation for the target voice message in the conversation interface (the trigger operation here can be the second trigger operation) to obtain the target level adjustment instruction in the voice conversion condition. Furthermore, the user terminal can determine the adjusted initial first subqueue as the target first subqueue based on the target level adjustment instruction, and adjust the queue position of the target voice identifier from the second position to the third position in the target first subqueue to obtain the adjusted target first subqueue. Among them, the request priority of the identifier corresponding to the third position is greater than the request priority of the identifier corresponding to the second position. Further, the user terminal can update the updated target identifier queue based on the adjusted target first subqueue and the second subqueue.

[0105] For easier understanding, see Figure 6 , Figure 6 This is a schematic diagram of a scenario in which a user makes a selection provided by an embodiment of the present application. Figure 6 As shown, the identification queue 6a and the identification queue 6b can be the target identification queues corresponding to the conversation interface 60, and the identification queue 6a can be the above Figure 5 It is understandable that the user terminal obtains the target voice message sent by user "AAA" (user "AAA" can be called the second user) (and opens the conversation interface corresponding to user "AAA", where the target voice message can be the above Figure 5 When the voice message x1, voice message x2, voice message x3 and voice message x4 in the corresponding embodiment are displayed, a conversation interface 60 can be displayed, wherein the second sub-queue of the identification queue 6a can include voice identifications {Y1, Y2, Y3}, and the target first sub-queue can include voice identifications {X1, X2, X3, X4, Z1}.

[0106] It is understandable that if the first user performs a second trigger operation on the voice message x3 (for example, the second trigger operation may be a click operation performed after performing a right-click operation), the user terminal may respond to the click operation and adjust the queue position of the voice identifier X3 corresponding to the voice message x3 in the identifier queue 6a to change the request priority of the voice identifier X3, thereby obtaining Figure 6 As shown in the identification queue 6b, the second sub-queue of the identification queue 6b may include voice identifiers {Y1, Y2, Y3}, and the adjusted target first sub-queue may include voice identifiers {X3, X1, X2, X4, Z1}.

[0107] Optionally, it can be understood that if the first user performs a second trigger operation on voice message x3 and then performs another second trigger operation on voice message x4 (for example, the another second trigger operation can be a click operation performed after performing a right-click operation), the user terminal can respond to the click operation and adjust the queue position of the voice identifier X4 corresponding to the voice message x4 in the identification queue 6b to change the request priority of the voice identifier X4, wherein the second subqueue of the identification queue 6b can include voice identifiers {Y1, Y2, Y3}, and the adjusted target first subqueue can include voice identifiers {X4, X3, X1, X2, Z1}. Optionally, the second subqueue of the identification queue 6b can include voice identifiers {Y1, Y2, Y3}, and the adjusted target first subqueue can include voice identifiers {X3, X4, X1, X2, Z1}.

[0108] Optionally, it is understandable that the first user can perform a second trigger operation (for example, the second trigger operation can be a click operation performed after performing a multiple selection operation) on multiple target voice messages (for example, voice message x3 and voice message x4) in the conversation interface 60 at the same time, so that the user terminal can respond to the click operation and adjust the queue position of voice identifier X3 and voice identifier X3 in the identifier queue 6a at the same time to change the request priority of voice identifier X3 and voice identifier X4, and thereby obtain the adjusted target first subqueue {X3, X4, X1, X2, Z1}.

[0109] It is understandable that the first user can perform the second trigger operation for the target voice message directly on the current session interface without having to perform the first trigger operation on the session interface. In this case, the current session interface of the application client may be the session interface corresponding to the second user. In this case, the voice message of the session interface obtained by the user terminal is the voice message of the current session interface corresponding to the second user. In this case, the user does not need to perform the first trigger operation on the session interface of the second user. Since the session interface of the second user has already been opened, the same effect as the first trigger operation can be achieved, that is, the "unread" message is converted to a "read" message, so that the second trigger operation can be performed on the target voice message in the session interface of the second user in the subsequent steps.

[0110] Among them, it can be understood that the embodiment of the present application can collectively refer to the above-mentioned initial first subqueue, the adjusted initial first subqueue (i.e., the target first subqueue) and the adjusted target first subqueue as the pending identification queue, and can collectively refer to the second subqueue (i.e., the above-mentioned second subqueue) corresponding to the voice identification of the voice message that has sent a voice conversion request and is still in the pending conversion state as the requested identification queue. Therefore, the pending identification queue and the requested identification queue can be collectively referred to as the target identification queue. Based on this, the second subqueue and the initial first subqueue can be collectively referred to as the target identification queue, the second subqueue and the adjusted initial first subqueue (i.e., the target first subqueue) can be collectively referred to as the target identification queue, and the second subqueue and the adjusted target first subqueue can be collectively referred to as the target identification queue.

[0111] It is understood that the voice conversion condition may include one or more of the following conversion conditions: a first conversion condition, a second conversion condition, and a third conversion condition, wherein the second conversion condition corresponds to the aforementioned initial level adjustment instruction, and the first conversion condition corresponds to the aforementioned target level adjustment instruction. It is understood that for the voice identifier of the same voice message currently in the first sub-queue (or pending request identifier queue), if the queue position of the voice identifier (e.g., queue position 1) is adjusted using the three conversion conditions respectively, the request priority of the new queue position (e.g., queue position A) of the voice identifier obtained by adjusting using the first conversion condition will be higher than the request priority of the new queue position (e.g., queue position B) of the voice identifier obtained by adjusting using the second conversion condition; at the same time, the request priority of the new queue position (e.g., queue position B) of the voice identifier obtained by adjusting using the second conversion condition will also be higher than the request priority of the new queue position (e.g., queue position C) of the voice identifier obtained by adjusting using the third conversion condition. It is understood that in the first sub-queue, the request priority of queue position C will be higher than the request priority of the original queue position 1 of the voice identifier.

[0112] For example, the first conversion condition can be understood as when the first user corresponding to the user terminal performs a trigger operation on a certain voice message (for example, voice message 1) in the conversation interface, the queue position 1 of the voice identifier of the voice message 1 can be adjusted in the target identifier queue according to the first conversion condition of the preset voice conversion condition, so that it can be added to the requested identifier queue as quickly as possible later. For another example, the second conversion condition can be understood as when the first user opens the current conversation interface through the user terminal, the user terminal can adjust the queue position 2 of the voice identifier of the voice message read when opening the current conversation interface (for example, changing the reading status of the voice message from an unread state to a read state of voice message 2) in the target identifier queue according to the second conversion condition of the preset voice conversion condition, so that it can be added to the requested identifier queue relatively quickly later. For example, the third conversion condition can be understood as the user terminal adjusting the queue position 3 of the voice identifier of the voice message (for example, voice message 3) that meets the third conversion condition in the target identifier queue according to the order of the receiving timestamps of the received voice messages, in accordance with the third conversion condition of the preset voice conversion condition, so that it can be added to the requested identifier queue relatively quickly.

[0113] Step S102: generating a voice conversion request carrying the voice identifier based on the queue position of the voice identifier in the target identifier queue, and sending the voice conversion request to the server, so that the server obtains converted text information corresponding to the voice identifier based on the voice conversion request;

[0114] Specifically, a user terminal may receive successful conversion information returned by the server for M voice messages to be converted for which a voice conversion request has been sent, and record the number of conversions in the received successful conversion information as N. The target identifier queue includes a pending identifier queue and a requested identifier queue. The voice identifier is located in the pending identifier queue, and the requested identifier queue includes M queue positions. One queue position in the requested identifier queue is used to store the identifier of a voice message to be converted, where M may be the total number of identifiers of voice messages to be converted for which a voice conversion request has been sent. N may be a positive integer less than or equal to M. Furthermore, the user terminal may obtain the queue position of the voice identifier in the pending identifier queue of the target identifier queue. When the queue position of the voice identifier meets the voice conversion condition, the user terminal may determine the target queue position of the voice identifier in the requested identifier queue. Furthermore, the user terminal may add the voice identifier to the requested identifier queue based on the target queue position, generate a voice conversion request carrying the voice identifier based on the requested identifier queue with the added voice identifier, and send the voice conversion request to the server.

[0115] Among them, when receiving the conversion success information of N voice messages to be converted, the user terminal can send a voice conversion request to the server based on the voice identifiers of the voice messages to be requested in the queue of identifiers to be requested. It can be understood that the user terminal can obtain the voice identifiers of N voice messages to be requested in the queue of identifiers to be requested, add the N voice identifiers to the target queue position of the queue of identifiers to be requested, and send a voice conversion request to the server based on the N voice identifiers. Optionally, if the queue of identifiers to be requested does not contain the voice identifiers of N voice messages to be requested, for example, the queue of identifiers to be requested can contain the voice identifiers of K voice messages to be requested, where K can be a positive integer less than N, then the user terminal can obtain the voice identifiers of K voice messages to be requested in the queue of identifiers to be requested, add the K voice identifiers to the target queue position of the queue of identifiers to be requested, and send a voice conversion request to the server based on the K voice identifiers.

[0116] Optionally, when a user terminal obtains a voice message, it can directly send a voice conversion request to the server based on the voice identifier of the voice message to be requested in the queue of the identifier to be requested. At this time, the M queue positions in the queue of the requested identifier are unoccupied. The user can obtain the voice identifiers of the M voice messages to be requested in the queue of the identifier to be requested, add the M voice identifiers to the target queue position of the queue of the requested identifier, and send a voice conversion request to the server based on the M voice identifiers. Optionally, if the queue of the requested identifier contains voice messages to be converted, for example, the requested identifier queue can contain (ML) voice identifiers of voice messages to be converted, where L can be a positive integer less than M, and at this time, the L queue positions in the queue of the requested identifier are unoccupied, the user terminal can obtain the voice identifiers of the L voice messages to be requested in the queue of the identifier to be requested, add the L voice identifiers to the target queue position of the queue of the requested identifier, and send a voice conversion request to the server based on the L voice identifiers. Optionally, if the queue of pending identifiers does not contain the voice identifiers of the above-mentioned M or L pending voice messages, for example, the queue of pending identifiers may contain K voice identifiers of pending voice messages, where K may be a positive integer less than M or L, then the user terminal may obtain the voice identifiers of the K pending voice messages in the queue of pending identifiers, add the K voice identifiers to the target queue position of the requested identifier queue, and send a voice conversion request to the server based on the K voice identifiers.

[0117] For ease of understanding, see Figure 7 , Figure 7 This is a schematic diagram of a scenario for receiving and converting text information provided by an embodiment of the present application. Figure 7 As shown, the identification queue 7a can be the above Figure 4 The identification queue 4 in the corresponding embodiment is Figure 4 At the next moment (for example, T3 moment) after the moment T2 corresponding to the identification queue 4 shown, that is, at Figure 7 At the next moment (for example, T3) after the moment T2 corresponding to the identification queue 7a shown, the server can return a conversion success message to the user terminal based on the three voice identifications in the requested identification queue that have sent voice conversion requests (for example, the conversion success message returned for voice identification Y1 and voice identification Y2). At this time, the user terminal can output the conversion text information corresponding to voice identification Y1 and voice identification Y2 to the conversation interface of the user terminal. It should be understood that when the user terminal obtains the conversion success message returned for voice identification Y1 and voice identification Y2, it can release the queue positions occupied by these two voice identifications in the requested identification queue of the target identification queue, and obtain Figure 7 As shown in the identification queue 7b, this means that at this time, there are currently two queue positions in the requested identification queue that are unoccupied. Based on this, the user terminal can select two voice identifiers with the highest request priority (for example, voice identifier Z1 and voice identifier X1) in the identification queue to be requested of the identification queue 7b, and add these two voice identifiers to the target queue position of the above-mentioned requested identification queue, wherein, depending on the target queue position, the requested identification queue may include voice identifiers {Y3, Z1, X1}, and the voice identifiers in the requested identification queue are sorted according to the time when the voice conversion request is sent. Optionally, the requested identification queue may include voice identifiers {Z1, X1, Y3}, and the voice identifiers in the requested identification queue do not need to be sorted according to the time of the voice conversion request, and there is no restriction on the queue position of the identifiers in the requested identification queue.

[0118] Optionally, the user terminal may adjust the queue position of the voice identifier in the queue of pending identifiers of the identifier queue 7b in response to a trigger operation (e.g., a first trigger operation and a second trigger operation) performed on the first user to change the request priority of the voice identifier so that step S102 and step S103 can be performed while step S101 is performed. The first trigger operation and the second trigger operation may include contact operations such as clicking, long pressing, and sliding, and may also include non-contact operations such as voice and gestures, which are not limited in this application.

[0119] Step S103: receiving the converted text information returned by the server, and outputting the converted text information to the location area where the voice message is located in the conversation interface.

[0120] The voice message and the converted text information in the location area are associated with each other. For example, the voice message and the converted text information may be located adjacent to each other in the conversation interface (for example, the converted text information may be located below the voice message).

[0121] It is understood that the converted text information returned by the server and received by the user terminal may be the complete converted text information corresponding to a single voice message, or may be the complete converted text information corresponding to multiple voice messages. For example, a voice conversion request may include voice identifier X1, voice identifier X2, and voice identifier X3. In this case, when the server returns the converted text information, it may return text information 1 corresponding to voice identifier X1 and text information 2 corresponding to voice identifier X2 to the user terminal at a first timestamp, and may also return text information 3 corresponding to voice identifier X3 to the user terminal at a second timestamp.

[0122] Similarly, it is understood that the converted text information returned by the server and received by the first terminal may also be the converted text information corresponding to a portion of a voice message. For example, when the server converts the voice message x2 corresponding to the voice identifier X2, it may convert the content of text message 2 in batches and return the obtained partial text information to the user terminal in batches. For example, text message 2 may be "I sorted out last week's documents this morning!" At time T11, the server may return "I sorted out last week's documents this morning" to the user terminal. At this time, the user terminal's conversation interface may include the converted text information of text message y2, namely, "I sorted out last week's documents this morning..."; at time T22, the server may return "I sorted out last week's documents" to the user terminal. At this time, the user terminal's conversation interface may include the converted text information of text message y2, namely, "I sorted out last week's documents this morning..."; at time T33, the server may return "I sorted out!" to the user terminal. At this time, the user terminal's conversation interface may include the complete converted text information of text message 2, namely, "I sorted out last week's documents this morning!" Among them, time T11 is earlier than time T22, and time T22 is earlier than time T33.

[0123] Furthermore, it is understood that if the converted text information received by the user terminal is the partial converted text information corresponding to the voice message, the partial converted text information is output to the session interface of the application client, and the complete converted text information corresponding to the voice message is received. If the converted text information received by the user terminal is the complete converted text information corresponding to the voice message, or the user terminal has received the complete converted text information corresponding to one or more voice messages, the conversion success information for one or more voice messages sent by the server is obtained (for example, the target conversion success information for the target voice message), and based on the conversion success information, the voice identifier is deleted from the requested identifier queue of the target identifier queue (for example, the target voice identifier is deleted from the requested identifier queue of the target identifier queue for the target conversion success information), so that the user terminal can continue to send voice conversion requests to the server in subsequent steps.

[0124] Among them, it can be understood that when the user terminal receives the converted text information returned by the server, it can store the converted text information in the local memory. When the first user views the converted text information corresponding to the historically received voice message in the application client (assuming that the converted text information corresponding to the historically received voice message will not be directly displayed in the conversation interface, or the converted text information has been hidden by the first user), the user terminal does not need to re-convert the historically received voice message, but directly obtains the converted text information corresponding to the voice message from the local memory, and outputs the converted text information to the location of the voice message in the conversation interface of the application client.

[0125] Optionally, it is understood that when the first user is viewing the converted text information corresponding to historically received voice messages, the first user can perform a trigger operation on one or more of the historically received voice messages, so that the user terminal can generate a voice conversion request based on the voice identifiers corresponding to the one or more voice messages and send the voice conversion request to the server. In this way, when the server-side voice conversion algorithm is updated, the first user can obtain the latest converted text information to improve the accuracy of the converted text information obtained by the first user. At this time, the user terminal can use this latest converted text information to update the converted text information in local memory.

[0126] It should be understood that the voice identifiers sent by the user terminal at adjacent moments and the converted text information received do not necessarily correspond. For example, at time T1, the user terminal may send a voice conversion request to the server based on voice identifier X1 and voice identifier X2, and at time T2, send a voice conversion request to the server based on voice identifier X3. At time T3, the user terminal may receive the converted text information returned by the server, which may be the converted text information corresponding to voice identifier X1 (or voice identifier X2) or the converted text information corresponding to voice identifier X3. Among them, time T1, time T2, and time T3 may be adjacent moments sorted in time, with time T1 being earlier than time T2, and time T2 being earlier than time T3.

[0127] In an embodiment of the present application, when the application client obtains a voice message on the conversation interface, the user terminal can obtain the voice identifier corresponding to the voice message, add the voice identifier to the initial identifier queue, and use the initial identifier queue with the added voice identifier as the target identifier queue. The voice message on the conversation interface can be a voice message sent by the application client, or a voice message received by the application client. Furthermore, the user terminal can generate a voice conversion request carrying the voice identifier based on the queue position of the voice identifier in the target identifier queue, and send the voice conversion request to the server, so that the server queries the voice message corresponding to the voice identifier based on the voice conversion request, and obtains the converted text information after converting the voice message. Furthermore, the user terminal can receive the converted text information returned by the server, and output the converted text information to the location area where the voice message is located in the conversation interface. There is an association relationship between the voice message and the converted text information in the location area. For example, the voice message and the converted text information can have an adjacent position relationship in the conversation interface. It should be understood that by introducing the target identification queue, when a voice message and a voice identifier corresponding to the voice message are obtained, the first user corresponding to the user terminal does not need to perform a trigger operation. The user terminal can output the converted text information corresponding to the voice message in the conversation interface of the application client, and then automatically convert the voice message into its corresponding converted text information, so as to achieve active access to the converted text information. Among them, when converting and processing the voice message based on the voice identifier in the target identification queue, the embodiment of the present application does not require the user terminal to upload the voice message in the local memory to the server. Instead, the server intelligently queries the voice message corresponding to the voice identifier based on the uploaded voice identifier and converts and processes the queried voice message. In this way, problems such as failure to upload voice messages can be solved in the case of an unstable network environment, thereby effectively improving the conversion efficiency of voice messages.

[0128] For further information, see Figure 8 , Figure 8This is a flow chart of a method for processing voice data provided by an embodiment of the present application. Figure 8 As shown, the method can be executed by a computer device, which can be a user terminal installed with the above-mentioned office client, and the user terminal can be the above-mentioned Figure 2 The user terminal X in the corresponding embodiment; optionally, the computer device may also be a server corresponding to the office client, which may be the above-mentioned Figure 2 The server in the corresponding embodiment. In other words, the method involved in the embodiment of the present application can be executed by the user terminal, or by the server, or by the user terminal and the server together. For ease of understanding, this embodiment is described by taking the method executed by the user terminal and the server together as an example. Among them, the method may include the following steps:

[0129] Step S201: When the server obtains a voice message from the application client, it generates a voice identifier corresponding to the voice message and sends the voice message and the voice identifier to the user terminal, so that the user terminal adds the voice identifier to the initial identifier queue and uses the initial identifier queue with the added voice identifier as the target identifier queue;

[0130] It can be understood that the user terminal in the embodiment of the present application can be a first terminal, and the user corresponding to the first terminal is called a first user, that is, the first user can be a user using the first terminal, and the first terminal can be the above-mentioned Figure 1 Similarly, the embodiment of the present application may refer to the user corresponding to the second terminal as the second user, that is, the second user may be the user using the second terminal, and the second terminal may be the above-mentioned Figure 1 The user terminal 3000a in the user terminal cluster of the corresponding embodiment. The first user may be a user who logs in to the office client using the first account information (for example, account information 1) in the first terminal, and the second user may be a user who logs in to the office client using the second account information (for example, account information 2) in the second terminal.

[0131] It is understandable that the first user in the embodiment of the present application can be a user who receives a voice message through the application client in the first terminal, that is, a message receiver; the second user in the embodiment of the present application can be a user who sends a voice message through the application client in the second terminal, that is, a message sender. It should be understood that the first user in the embodiment of the present application can serve as both the above-mentioned message receiver and the above-mentioned message sender. For example, the first user can become a message receiver through the first terminal, and the first user can also become a message sender through the first terminal. Similarly, the second user in the embodiment of the present application can serve as both the above-mentioned message sender and the above-mentioned message receiver. For example, the second user can become a message sender through the second terminal, and the second user can also become a message receiver through the second terminal.

[0132] It is understandable that the message sender and the message receiver can be connected through a server, and the server synchronizes the voice message from the user terminal corresponding to the message sender (for example, the second terminal corresponding to the second user) to the user terminal corresponding to the message receiver (for example, the first terminal corresponding to the first user), and in subsequent steps, the voice message can be converted so that the first user and the second user can directly obtain the converted text information corresponding to the voice message. In particular, the first terminal and the second terminal both run the application client corresponding to the server, and the application client can realize the sending and receiving of voice messages between the first terminal and the second terminal.

[0133] Among them, it can be understood that if the second user is the message sender, the server can receive the voice message sent by the application client of the second terminal, generate a voice identifier corresponding to the voice message (that is, configure the voice identifier corresponding to the voice message), store the voice content and voice identifier corresponding to the voice message in the application database, and then forward the voice message and voice identifier to the first terminal so that the first terminal can output the voice message in the conversation interface of the application client. At the same time, the server can return the voice message sent by the second terminal to the second terminal so that the second terminal can output the voice message in the conversation interface of the application client. At the same time, the server can return the voice identifier corresponding to the voice message to the second terminal so that the second terminal can send a voice conversion request to the server based on the voice identifier corresponding to the voice message sent by the second user.

[0134] For easier understanding, see Figure 9 , Figure 9 This is a schematic diagram of a scenario for forwarding a voice message provided by an embodiment of the present application. Figure 9As shown, the second user (i.e., user "AAA") uses the second terminal to send a voice message to the first user (i.e., user "FFF"), and the voice message can be forwarded by the server. The server can receive the voice message sent by the second terminal, generate a voice identifier corresponding to the voice message (i.e., configure the voice identifier corresponding to the voice message), and store the voice identifier and the voice content corresponding to the voice message in the application database. Furthermore, the server can send a voice message and a voice identifier to the first terminal, so that the first terminal outputs the voice message to the conversation interface of the first terminal, so that in a subsequent step the first terminal can send a voice conversion request to the server based on the voice identifier. At the same time, Figure 9 When the server shown sends a voice message and a voice identifier to the first terminal, it can send a voice message and a voice identifier to the second terminal, so that the second terminal outputs the voice message to the conversation interface of the second terminal, so that in a subsequent step the second terminal can send a voice conversion request to the server based on the voice identifier.

[0135] Among them, it can be understood that, for the voice message sent by the second user, the second terminal can add the voice identifier to the initial identifier queue when obtaining the voice identifier corresponding to the voice message, so that the second terminal can automatically send a voice conversion request to the server based on the voice identifier. Optionally, when the second terminal obtains the voice identifier corresponding to the voice message, it does not need to add the voice identifier to the initial identifier queue, and when the second user performs a trigger operation on the voice message (the trigger operation here can be the above-mentioned Figure 3 When the second trigger operation in the corresponding embodiment is executed, the acquired voice identifier is added to the initial identifier queue, so that the second terminal can send a voice conversion request to the server based on the voice identifier.

[0136] Step S202: When the application client obtains the voice message of the conversation interface, the user terminal obtains the voice identifier corresponding to the voice message, adds the voice identifier to the initial identifier queue, and uses the initial identifier queue with the added voice identifier as the target identifier queue;

[0137] The user terminal may be the first terminal in step S201 above, and the voice message acquired by the user terminal may be a voice message received by the first user as a message recipient, or may be a voice message sent by the first user as a message sender.

[0138] It can be understood that the initial identification queue and the target identification queue (referred to as the identification queue or queue) can be used to store voice identification. The initial identification queue in the embodiment of the present application may include a first sub-queue and a second sub-queue, and the target identification queue may include a requested identification queue and a requested identification queue. Generally speaking, a queue is a group of people or things waiting to be served or processed in an arranged order.

[0139] It should be understood that the queue position of the voice message in the first subqueue (or queue with pending request identifier) ​​is determined by the reception timestamp of the voice message. When the user terminal responds to the trigger operation performed on the first user, the queue position of the voice message in the first subqueue (or queue with pending request identifier) ​​can be adjusted so that the request priority of the queue position after the adjustment is greater than the request priority of the queue position before the adjustment. Specifically, when adding a voice identifier to the initial identifier queue, an enqueue operation can be performed on the first subqueue, and when sending a voice conversion request to the server, a dequeue operation can be performed on the queue with pending request identifier.

[0140] It should be understood that the queue position of the voice message in the second subqueue (or the requested identifier queue) is determined by the timestamp of the voice conversion request. The M queue positions included in the second subqueue (or the requested identifier queue) indicate that the total number of identifiers that can be accommodated in the second subqueue is M. The user terminal can send voice conversion requests for M voice identifiers to the server. If the value of M is too large, this will cause too much pressure on the server. If the value of M is too small, this will cause the voice message conversion processing speed to be too slow. Therefore, in this embodiment of the application, M can be set to 5. When sending a voice conversion request to the server, an enqueue operation can be performed on the second subqueue, and when the user terminal receives the converted text information, a dequeue operation can be performed on the requested identifier queue.

[0141] For easier understanding, see Figure 10 , Figure 10 This is a flow chart of a speech-to-text solution provided in an embodiment of the present application. Figure 10 The application client shown may be an office client, which may be a client installed on a user terminal, such as Figure 10 The target user shown may be the first user using the user terminal, for example, the user "FFF" mentioned above. When the application client receives a voice message (i.e., obtains the voice message on the display interface of the application client) and a voice identifier, the voice identifier corresponding to the voice message may be added to the initial identifier queue to obtain a target identifier queue, and the target identifier queue is sorted by priority (i.e., request priority), i.e., the voice identifier is added to the initial identifier queue according to the reception timestamp of the voice message.

[0142] It should be understood that Figure 10 As shown, when the target user opens a session (that is, the target user executes the above Figure 5 The first trigger operation in the corresponding embodiment) or click on voice to text (ie the target user performs the above Figure 6When the second trigger operation in the corresponding embodiment is executed, the application client may adjust the queue position of the voice message in the target identification queue so that the request priority of the adjusted queue position is greater than the request priority of the queue position before the adjustment, that is, update the request priority of the voice message. Among them, the voice message corresponding to the second trigger operation may have a first priority, the voice message corresponding to the first trigger operation may have a second priority, and other voice messages (that is, voice messages other than the first and second priorities) may have a third priority. Among them, the request priority of the first priority is greater than the request priority of the second priority, and the request priority of the second priority is greater than the request priority of the third priority.

[0143] Among them, it can be understood that the first trigger operation can be a trigger operation performed by the target user on the conversation interface of the group message. At this time, the user terminal can obtain the voice identifier corresponding to the voice message forwarded by multiple users (for example, the second user and the third user) through the server, and add the voice identifier to the initial identifier queue to obtain the target identifier queue. When the user terminal responds to the first trigger operation and outputs the voice message to the conversation interface, the user terminal can adjust the queue position of the voice identifier corresponding to the voice message in the target identifier queue. Among them, the specific implementation method of the user terminal adjusting the queue position of the voice identifier corresponding to the voice message of the group message in the target identifier queue can refer to the description of the user terminal adjusting the queue position of the voice identifier corresponding to the voice message of the second user in the target identifier queue, which will not be repeated here.

[0144] The specific implementation method of dynamically adjusting the queue position of the voice identifier in the target identifier queue according to the request priority of the user terminal can be found in the above Figure 3 The description of step S101 in the corresponding embodiment will not be repeated here.

[0145] Step S203: The user terminal generates a voice conversion request carrying the voice identifier based on the queue position of the voice identifier in the target identifier queue, and sends the voice conversion request to the server, so that the server obtains converted text information corresponding to the voice identifier based on the voice conversion request;

[0146] Among them, Figure 10 As shown, when the voice conversion conditions are met, the application client can initiate a text conversion request, that is, send a voice conversion request to the server.

[0147] The specific implementation method of the user terminal sending the voice conversion request to the server can be found in the above Figure 3 The description of step S102 in the corresponding embodiment will not be repeated here.

[0148] Step S204: The server receives the voice conversion request sent by the user terminal and obtains the voice identifier from the voice conversion request;

[0149] The voice conversion request is generated based on the queue position of the voice identifier in the target identifier queue.

[0150] It can be understood that when the voice identifier in the voice conversion request is the voice identifier corresponding to the voice message in the group message, the server can receive multiple voice conversion requests sent by multiple user terminals (for example, the second terminal and the third terminal), and obtain the same voice identifier from the multiple voice conversion requests, so that after querying the same voice content corresponding to the same voice identifier in subsequent steps, the same voice content can be converted multiple times.

[0151] Optionally, to increase the speed of the conversion process, the server may, upon receiving a voice message, convert the voice content corresponding to the voice message and store the resulting converted text information in the application database. Upon receiving a voice conversion request from a user terminal, the server may query the application database for the converted text information corresponding to the voice identifier based on the voice identifier carried in the voice conversion request. Similarly, optionally, to increase the speed of the conversion process, upon first receiving a voice identifier, the server may query the application database for the voice content corresponding to the voice identifier based on the voice identifier and convert the voice content to store the resulting converted text information in the application database. Upon receiving the voice identifier for the second time, the server may query the application database for the converted text information corresponding to the voice identifier based on the voice identifier.

[0152] Step S205: When the voice message corresponding to the voice identifier is found, the server converts the voice message to obtain converted text information corresponding to the voice message;

[0153] It should be understood that the speech conversion algorithm used in the embodiment of the present application to convert the speech content corresponding to the speech message into text information can be a pattern matching method, that is, in the training phase, each word in the vocabulary is spoken once, and its feature vector is stored as a template in the template library; in the recognition phase, the feature vector of the input speech is compared with each template in the template library in turn for similarity, and the one with the highest similarity is output as the recognition result. Optionally, the speech conversion algorithm in the embodiment of the present application can be a method based on a hidden Markov model (HMM) of a parameter model, or an algorithm based on dynamic time warping (DTW), or a method based on a vector quantization (VQ) of a non-parametric model. The embodiment of the present application does not limit the specific type of speech conversion algorithm. Among them, the process of converting the speech message into text information can be called speech recognition, and the above-mentioned speech conversion algorithm can be called a speech recognition method.

[0154] For easier understanding, see Figure 11 , Figure 11 This is a schematic diagram of a scenario for converting speech to text provided by an embodiment of the present application. Figure 11 As shown, when receiving a voice conversion request sent by a user terminal based on a voice identifier (for example, voice identifier X1 and voice identifier X2), the server can query the application database corresponding to the server based on the voice identifier carried in the voice conversion request for voice content 1 and voice content 2 corresponding to the voice identifier X1 and voice identifier X2, and forward the voice content 1 and voice content 2 to the voice processing server, so that the voice processing server can convert the voice content 1 and voice content 2, and then send the converted text information (for example, Figure 11 The text information 1 and text information 2 shown in FIG. 1 (i.e., the converted text information) are returned to the server so that the server can output the converted text information 1 and text information 2 to the server. Figure 11 The user terminal shown.

[0155] The speech processing server here can be the same speech processing server used to provide conversion processing services, or it can be a speech processing server cluster that is independent of each other and used to provide conversion processing services. For example, the speech processing server cluster can include speech processing server 100a, speech processing server 100b, ..., speech processing server 100n (that is, the above-mentioned speech content 1 can be converted and processed by speech processing server 100a, the above-mentioned speech content 2 can be converted and processed by speech processing server 100b, and similarly, speech content 3 can be converted and processed by speech processing server 100c). It will not be limited here. Optionally, one or more speech processing servers that provide conversion processing services can run on Figure 11 The server shown can also be independent of Figure 11 The server exists only for the server shown and will not be limited here.

[0156] Step S206: The server returns the converted text information to the user terminal, so that the user terminal outputs the converted text information to the location area where the voice message is located in the conversation interface of the application client;

[0157] Among them, Figure 10 As shown, when the server converts the voice message, if the conversion is successful, the text information corresponding to the voice message (i.e., the converted text information) can be returned to the application client, so that the application client can display the converted text information, and then the target user can intuitively view the converted text information received by the application client in the conversation interface of the application client. Similarly, if the conversion of the voice message fails, the server can return a rejection prompt message to the application client, so that the target user can resend the voice conversion request to the server. There are many reasons for the failure of the voice message conversion process, for example, the voice message is spoken too fast, the voice message is in a dialect, the noise of the voice message is too loud, and the voice message is in an unsupported language type.

[0158] It is understood that due to network instability, the server may fail to return the converted text information to the user terminal. In this case, the user terminal may also receive a rejection prompt message from the server. Similarly, due to network instability, the user terminal may fail to send a voice conversion request to the server. In this case, the user terminal may also receive a rejection prompt message from the server.

[0159] The server may return a rejection prompt to the user terminal by popping up a sub-interface independent of the original session interface on the application client's session interface. The sub-interface may display the message "Voice conversion failed, please try again." It is understood that the prompt message on the sub-interface may vary depending on the reason for returning the rejection prompt.

[0160] Step S207: the user terminal receives the converted text information returned by the server, and outputs the converted text information to the location area where the voice message is located in the conversation interface.

[0161] There is an association relationship between the voice message in the location area and the converted text information.

[0162] It is understood that, in the embodiment of the present application, the first user can also select one or more users of interest in the group list corresponding to the above-mentioned group message conversation interface. In this way, when the user terminal receives voice messages sent by these users, the converted text information corresponding to the voice messages of these users can be output on this conversation interface. In addition, the user terminal can also optionally display the converted text information corresponding to the voice messages of these selected users on other display interfaces, so that only the voice messages of these selected users can be heard and only the converted text information corresponding to the voice messages of these selected users can be viewed on other display interfaces.

[0163] It is understood that when the server obtains a voice message (for example, a voice message A sent by the second user to the first user), it can configure a unique voice identifier for the voice message A, so that the voice identifier and the voice message A can be distributed to the user terminal corresponding to the first user (i.e., the first terminal) and the user terminal corresponding to the second user (i.e., the second terminal). In this way, the second user can view the voice message A they sent on the conversation interface of the second terminal, and similarly, the first user can view the voice message A sent by the other party on the conversation interface of the first terminal. At this time, the first terminal and the server can obtain the converted text information corresponding to the voice message A through the data exchange method described in steps S201 to S207 above.

[0164] Optionally, for ease of understanding, the embodiment of the present application takes the user terminal that obtains the voice identification of the voice message A as the user terminal corresponding to the first user (i.e., the above-mentioned first terminal) as an example to illustrate another implementation method in which the first terminal automatically receives the converted text information corresponding to the voice message A sent by the server.

[0165] For example, considering that the voice content corresponding to voice message A can be stored in the server, in order to improve the speed of conversion processing, the server in the embodiment of the present application can further convert the voice content corresponding to voice message A locally on the server while sending the voice message A and the voice identifier corresponding to the voice message A to the user terminal corresponding to the above-mentioned first user (i.e., the first terminal), so that there is no need to receive the voice conversion request sent by the user terminal based on the queue position of the above-mentioned voice identifier in the target identifier queue. Subsequently, when the server completes the conversion processing of the voice content of the voice message A, the server can directly and intelligently send the converted text information of the voice message A obtained by the conversion processing to the user terminal corresponding to the first user, so that the converted text information corresponding to the voice message A can be displayed in the conversation interface of the user terminal corresponding to the first user.

[0166] Optionally, during the process of converting the voice content of the voice message A, the server may also perform semantic analysis on the voice content corresponding to the voice message A. When detecting the presence of semantic information of a preset keyword in the voice message A, the server may further identify the preset keyword in the converted text information corresponding to the voice message A (for example, highlighting the preset keyword in the converted text information), and then return the identified converted text information carrying the keyword to the user terminal. It is understood that when the user terminal receives the converted text information carrying the keyword after the above identification processing, it may display the identified converted text information carrying the keyword and the voice message in the current conversation interface.

[0167] Optionally, if the type of the keyword belongs to a specific type of keyword in the group conversation, the user terminal can also output and display the converted text information carrying the keyword after the received identification processing on another display interface independent of the conversation interface. For example, the converted text information carrying the keyword after the identification processing can be displayed in a pop-up window independent of the current conversation interface. The embodiment of the present application will not limit the specific display interface for displaying the converted text information carrying the keyword after the identification processing.

[0168] It should be understood that by introducing the target identification queue, when a voice message and a voice identifier corresponding to the voice message are obtained, the first user corresponding to the user terminal does not need to perform a trigger operation. The user terminal can output the converted text information corresponding to the voice message in the conversation interface of the application client, and then automatically convert the voice message into its corresponding converted text information, so as to achieve active access to the converted text information. Among them, when converting and processing the voice message based on the voice identifier in the target identification queue, the embodiment of the present application does not require the user terminal to upload the voice message in the local memory to the server. Instead, the server intelligently queries the voice message corresponding to the voice identifier based on the uploaded voice identifier and converts and processes the queried voice message. In this way, problems such as failure to upload voice messages can be solved in the case of an unstable network environment, thereby effectively improving the conversion efficiency of voice messages.

[0169] For further information, see Figure 12 , Figure 12 The structure diagram of a voice data processing device provided by the embodiment of the present application is shown in FIG. The voice data processing device 1 can be applied to the above-mentioned user terminal, which can be the above-mentioned Figure 1 The user terminal 3000c in the corresponding embodiment. The voice data processing device 1 may include: a voice acquisition module 10, a request sending module 20, and a text receiving module 30; further, the voice data processing device 1 may also include: an identification deletion module 40;

[0170] The voice acquisition module 10 is used to obtain the voice identifier corresponding to the voice message when the application client obtains the voice message of the conversation interface, add the voice identifier to the initial identifier queue, and use the initial identifier queue with the added voice identifier as the target identifier queue;

[0171] The conversation interface includes a second user associated with the first user; the initial identification queue includes a first subqueue and a second subqueue; the first subqueue is used to store a first voice identification; the first voice identification is used to represent the identification of a first voice message for which a voice conversion request is to be sent in the application client; the second subqueue is used to store a second voice identification; the second voice identification is used to represent the identification of a second voice message for which a voice conversion request has been sent in the application client;

[0172] The voice acquisition module 10 includes: a voice receiving unit 101, a timestamp determining unit 102, an identifier adding unit 103, and a queue determining unit 104; optionally, the voice acquisition module 10 may further include: a first triggering unit 105, a first adjusting unit 106, a first updating unit 107, a second triggering unit 108, a second adjusting unit 109, and a second updating unit 110;

[0173] The voice receiving unit 101 is configured to receive, via the application client corresponding to the first user, a voice message forwarded by the second user through the server, and receive a voice identifier configured by the server for the voice message;

[0174] The timestamp determining unit 102 is configured to obtain a voice conversion condition associated with the conversation interface, determine the received voice identifier as a target voice identifier based on the voice conversion condition, determine the voice message received by the application client as a target voice message, and record a reception timestamp corresponding to the target voice message as a target reception timestamp;

[0175] an identifier adding unit 103 configured to determine, based on a target reception timestamp, a queue position of a target voice identifier of a target voice message in a first subqueue containing an identifier of the first voice message, and add the target voice identifier to the first subqueue based on the queue position to obtain an initial first subqueue;

[0176] The queue determining unit 104 is configured to determine a target identification queue based on the initial first subqueue and the second subqueue containing the identification of the second voice message.

[0177] Optionally, the request priority of the second subqueue is greater than the request priority of the first subqueue;

[0178] The first triggering unit 105 is configured to respond to a triggering operation on the conversation interface where the second user is located, output a target voice message to the conversation interface, and obtain an initial level adjustment instruction in the voice conversion condition;

[0179] A first adjustment unit 106 is configured to determine, based on the initial level adjustment instruction, a queue position of the target voice identifier in the initial first subqueue as a first position, and adjust the queue position of the target voice identifier in the initial first subqueue from the first position to a second position, thereby obtaining an adjusted initial first subqueue; wherein the request priority of the identifier corresponding to the second position is greater than the request priority of the identifier corresponding to the first position;

[0180] The first updating unit 107 is configured to update the target identification queue based on the adjusted initial first sub-queue and second sub-queue.

[0181] Optionally, a second triggering unit 108 is configured to respond to a triggering operation on the target voice message in the conversation interface and obtain a target level adjustment instruction in the voice conversion condition;

[0182] The second adjustment unit 109 is configured to determine, based on the target level adjustment instruction, the adjusted initial first subqueue as the target first subqueue, and adjust the queue position of the target voice identifier in the target first subqueue from the second position to the third position, thereby obtaining an adjusted target first subqueue; wherein the request priority of the identifier corresponding to the third position is greater than the request priority of the identifier corresponding to the second position;

[0183] The second updating unit 110 is configured to update the updated target identification queue based on the adjusted target first sub-queue and second sub-queue.

[0184] The specific implementation of the voice receiving unit 101, the timestamp determining unit 102, the identification adding unit 103 and the queue determining unit 104 can be found in the above Figure 3 The description of step S101 in the corresponding embodiment will not be repeated here. Optionally, the specific implementation of the first trigger unit 105, the first adjustment unit 106, the first update unit 107, the second trigger unit 108, the second adjustment unit 109 and the second update unit 110 can be found in the above Figure 3 The description of step S101 in the corresponding embodiment will not be repeated here.

[0185] A request sending module 20 is configured to generate a voice conversion request carrying the voice identifier based on the queue position of the voice identifier in the target identifier queue, and send the voice conversion request to the server so that the server obtains converted text information corresponding to the voice identifier based on the voice conversion request;

[0186] The target identifier queue includes a pending identifier queue and a requested identifier queue; the voice identifier is located in the pending identifier queue; the requested identifier queue includes M queue positions; one queue position in the requested identifier queue is used to store an identifier of a voice message to be converted; M is the total number of identifiers of voice messages to be converted for which voice conversion requests have been sent;

[0187] The request sending module 20 includes: an information receiving unit 201, a position determining unit 202, and a request generating unit 203;

[0188] The information receiving unit 201 is configured to receive conversion success information returned by the server for M voice messages to be converted for which voice conversion requests have been sent, and record the number of conversions received in the conversion success information as N; N is a positive integer less than or equal to M;

[0189] The position determination unit 202 is configured to obtain a queue position of the voice identifier in the queue of the target identifier, and determine a target queue position of the voice identifier in the queue of the requested identifier when the queue position of the voice identifier meets the voice conversion condition;

[0190] The request generating unit 203 is configured to add the voice identifier to the requested identifier queue based on the target queue position, generate a voice conversion request carrying the voice identifier based on the requested identifier queue with the added voice identifier, and send the voice conversion request to the server.

[0191] The specific implementation of the information receiving unit 201, the location determining unit 202 and the request generating unit 203 can be found in the above Figure 3 The description of step S102 in the corresponding embodiment will not be repeated here.

[0192] The text receiving module 30 is used to receive the converted text information returned by the server and output the converted text information to the location area where the voice message is located in the conversation interface; there is an association relationship between the voice message in the location area and the converted text information.

[0193] Optionally, the identifier deletion module 40 is configured to obtain target conversion success information for the voice message when receiving the converted text information returned by the server, and delete the voice identifier from the target identifier queue based on the target conversion success information.

[0194] The specific implementation of the voice acquisition module 10, the request sending module 20 and the text receiving module 30 can be found in the above Figure 3 The description of steps S101 to S103 in the corresponding embodiment will not be repeated here. Figure 3 The description of step S103 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0195] See Figure 13 , Figure 13 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 13As shown, the computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the above-mentioned computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. Optionally, the network interface 1004 may include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage. Optionally, the memory 1005 may also be at least one storage device located away from the aforementioned processor 1001. As Figure 13 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application.

[0196] In such Figure 13 In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an interface for user input; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0197] When the application client obtains the voice message of the conversation interface, it obtains the voice identifier corresponding to the voice message, adds the voice identifier to the initial identifier queue, and uses the initial identifier queue with the added voice identifier as the target identifier queue;

[0198] generating a voice conversion request carrying the voice identifier based on the queue position of the voice identifier in the target identifier queue, and sending the voice conversion request to the server so that the server obtains converted text information corresponding to the voice identifier based on the voice conversion request;

[0199] The converted text information returned by the server is received, and the converted text information is output to the location area where the voice message is located in the conversation interface; there is an association relationship between the voice message in the location area and the converted text information.

[0200] It should be understood that the computer device 1000 described in the embodiment of the present application can execute the above Figure 3 or Figure 8 The description of the data processing method in the corresponding embodiment can also be performed as described above. Figure 12 The description of the data processing device 1 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0201] In addition, it should be noted that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the data processing device 1 mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the above-mentioned Figure 3 or Figure 8 The description of the voice data processing method in the corresponding embodiment will therefore not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0202] For further information, see Figure 14 , Figure 14 The structure diagram of a voice data processing device provided by the embodiment of the present application is shown in FIG. The voice data processing device 2 can be applied to the above-mentioned server, which can be the above-mentioned Figure 1 The server 3000 in the corresponding embodiment. The speech data processing device 2 may include: a speech sending module 100, a request receiving module 200, a text obtaining module 300, and a text sending module 400;

[0203] The voice sending module 100 is used to generate a voice identifier corresponding to the voice message when obtaining a voice message from the application client, and send the voice message and the voice identifier to the user terminal, so that the user terminal adds the voice identifier to the initial identifier queue and uses the initial identifier queue with the added voice identifier as the target identifier queue;

[0204] The request receiving module 200 is configured to receive a voice conversion request sent by a user terminal and obtain a voice identifier from the voice conversion request; the voice conversion request is generated based on the queue position of the voice identifier in the target identifier queue;

[0205] The text acquisition module 300 is used to convert the voice message to obtain the converted text information corresponding to the voice message when the voice message corresponding to the voice identifier is found;

[0206] The text sending module 400 is configured to return the converted text information to the user terminal, so that the user terminal outputs the converted text information to the location area where the voice message is located in the conversation interface of the application client.

[0207] The specific implementation of the voice sending module 100, the request receiving module 200, the text obtaining module 300 and the text sending module 400 can be found in the above Figure 8The description of steps S201 to S207 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0208] See Figure 15 , Figure 15 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 15 As shown, the computer device 2000 may include: a processor 2001, a network interface 2004 and a memory 2005. In addition, the above-mentioned computer device 2000 may also include: a user interface 2003, and at least one communication bus 2002. The communication bus 2002 is used to realize the connection and communication between these components. The user interface 2003 may include a display screen (Display), a keyboard (Keyboard), and the optional user interface 2003 may also include a standard wired interface and a wireless interface. Optionally, the network interface 2004 may include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 2005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage. Optionally, the memory 2005 may also be at least one storage device located away from the aforementioned processor 2001. As Figure 15 As shown, the memory 2005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application.

[0209] In such Figure 15 In the computer device 2000 shown, the network interface 2004 can provide network communication functions; the user interface 2003 is mainly used to provide an interface for user input; and the processor 2001 can be used to call the device control application stored in the memory 2005 to achieve:

[0210] When a voice message from an application client is obtained, a voice identifier corresponding to the voice message is generated, and the voice message and the voice identifier are sent to a user terminal, so that the user terminal adds the voice identifier to an initial identifier queue and uses the initial identifier queue with the added voice identifier as a target identifier queue;

[0211] receiving a voice conversion request sent by a user terminal and obtaining a voice identifier from the voice conversion request; the voice conversion request is generated based on a queue position of the voice identifier in a target identifier queue;

[0212] When a voice message corresponding to the voice identifier is found, the voice message is converted to obtain converted text information corresponding to the voice message;

[0213] The converted text information is returned to the user terminal, so that the user terminal outputs the converted text information to the location area where the voice message is located in the conversation interface of the application client.

[0214] It should be understood that the computer device 2000 described in the embodiment of the present application can execute the above Figure 8 The description of the voice data processing method in the corresponding embodiment can also be performed as described above. Figure 14 The description of the data processing device 2 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0215] In addition, it should be noted that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the data processing device 2 mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the above-mentioned Figure 8 Therefore, the description of the data processing method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0216] For further information, see Figure 16 , Figure 16 The embodiment of the present application also provides a voice data processing system. The voice data processing system 3 may include a user terminal 1 and a server 2, wherein the user terminal 1 may be the aforementioned Figure 12 The speech data processing device 1 in the corresponding embodiment; the server 2 can be the aforementioned Figure 14 The speech data processing device 2 in the corresponding embodiment. It is understandable that the description of the beneficial effects of adopting the same method will not be repeated.

[0217] In addition, it should be noted that: the embodiment of the present application also provides a computer program product or computer program, which may include computer instructions, which may be stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor may execute the computer instructions, so that the computer device performs the above Figure 3 or Figure 8 The description of the voice data processing method in the corresponding embodiment will therefore not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the computer program product or computer program embodiments involved in this application, please refer to the description of the method embodiments of this application.

[0218] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0219] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A method for processing speech data, characterized in that: include: When the application client obtains a voice message of the session interface, the voice identifier corresponding to the voice message is obtained, the voice identifier is added to the initial identifier queue, and the initial identifier queue to which the voice identifier is added is used as the target identifier queue; wherein the session interface includes a second user associated with the first user, the first user is the recipient of the voice message, and the second user is the sender of the voice message; the initial identifier queue includes a first subqueue and a second subqueue, and the request priority of the second subqueue is greater than the request priority of the first subqueue; the first subqueue is used to store a first voice identifier, which is used to represent the identifier of the first voice message to be sent in the application client for voice conversion request, and the queue position of the first voice identifier in the first subqueue is determined according to the reception timestamp corresponding to the first voice message; the second subqueue is used to store a second voice identifier, which is used to represent the identifier of the second voice message to be sent in the application client for voice conversion request, and the queue position of the second voice identifier in the second subqueue is determined according to the sending timestamp of the voice conversion request corresponding to the second voice message; In response to a first trigger operation on the conversation interface where the second user is located, determining at least one voice message sent by the second user in the conversation interface as a target voice message, outputting the target voice message to the conversation interface, and obtaining an initial level adjustment instruction in a voice conversion condition associated with the conversation interface; Based on the initial level adjustment instruction, the queue position of the target voice identifier corresponding to the target voice message is determined to be the first position in the initial first subqueue, and the queue position of the target voice identifier is adjusted from the first position to the second position in the initial first subqueue to obtain an adjusted initial first subqueue; the request priority of the identifier corresponding to the second position is greater than the request priority of the identifier corresponding to the first position; updating the target identification queue based on the adjusted initial first sub-queue and the second sub-queue; In response to a second trigger operation on at least one voice message in the conversation interface, determining the at least one voice message corresponding to the second trigger operation as a target voice message, and obtaining a target level adjustment instruction in the voice conversion condition; Based on the target level adjustment instruction, the adjusted initial first subqueue is determined as the target first subqueue, and the queue position of the target voice identifier corresponding to the target voice message is adjusted from the second position to the third position in the target first subqueue, to obtain an adjusted target first subqueue; the request priority of the identifier corresponding to the third position is greater than the request priority of the identifier corresponding to the second position; Based on the adjusted target first sub-queue and the second sub-queue, updating the updated target identification queue; generating a voice conversion request carrying the voice identifier based on a queue position of the voice identifier in the target identifier queue, and sending the voice conversion request to a server, so that the server obtains converted text information corresponding to the voice identifier based on the voice conversion request; The converted text information returned by the server is received, and the converted text information is output to the location area where the voice message is located in the conversation interface; the voice message in the location area is associated with the converted text information.

2. The method according to claim 1, characterized in that The conversation interface includes a second user associated with the first user; When the application client obtains a voice message on the conversation interface, obtaining a voice identifier corresponding to the voice message, adding the voice identifier to an initial identifier queue, and using the initial identifier queue to which the voice identifier is added as a target identifier queue includes: The application client corresponding to the first user receives the voice message forwarded by the second user through the server, and receives the voice identifier configured by the server for the voice message; Obtaining a voice conversion condition associated with the conversation interface, determining the received voice identifier as a target voice identifier based on the voice conversion condition, determining the voice message received by the application client as a target voice message, and recording a reception timestamp corresponding to the target voice message as a target reception timestamp; determining, based on the target reception timestamp, a queue position of the target voice identifier of the target voice message in the first subqueue containing the identifier of the first voice message, and adding the target voice identifier to the first subqueue based on the queue position to obtain an initial first subqueue; A target identification queue is determined based on the initial first subqueue and the second subqueue containing the identification of the second voice message.

3. The method according to claim 1, characterized in that The target identifier queue includes a pending identifier queue and a requested identifier queue; the voice identifier is located in the pending identifier queue; the requested identifier queue includes M queue positions; one queue position in the requested identifier queue is used to store an identifier of a voice message to be converted; M is the total number of identifiers of the voice messages to be converted for which voice conversion requests have been sent; The generating a voice conversion request carrying the voice identifier based on the queue position of the voice identifier in the target identifier queue, and sending the voice conversion request to the server, comprises: receiving conversion success information returned by the server for the M voice messages to be converted for which the voice conversion request has been sent, and recording the number of conversions in the received conversion success information as N; wherein N is a positive integer less than or equal to M; Obtaining a queue position of the voice identifier in a queue of pending identifiers of the target identifier queue, and determining a target queue position of the voice identifier in the queue of requested identifiers when the queue position of the voice identifier meets a voice conversion condition; The voice identifier is added to the requested identifier queue based on the target queue position, a voice conversion request carrying the voice identifier is generated based on the requested identifier queue to which the voice identifier is added, and the voice conversion request is sent to a server.

4. The method according to claim 1, wherein The method further comprises: When the converted text information returned by the server is received, target conversion success information for the voice message is acquired, and the voice identifier is deleted from the target identifier queue based on the target conversion success information.

5. A method for processing speech data, characterized in that: include: When a voice message from an application client is acquired, a voice identifier corresponding to the voice message is generated, and the voice message and the voice identifier are sent to a user terminal, so that the user terminal adds the voice identifier to an initial identifier queue and uses the initial identifier queue with the voice identifier added as a target identifier queue. The initial identifier queue includes a first subqueue and a second subqueue, and the request priority of the second subqueue is greater than the request priority of the first subqueue. The first subqueue is used to store a first voice identifier, which is used to identify a first voice message to be sent in the application client for a voice conversion request, and the queue position of the first voice identifier in the first subqueue is determined according to a reception timestamp corresponding to the first voice message. The second subqueue is used to store a second voice identifier, which is used to identify a second voice message to be sent in the application client for a voice conversion request, and the queue position of the second voice identifier in the second subqueue is determined according to a transmission timestamp of the voice conversion request corresponding to the second voice message. The target identifier queue is updated by the user terminal according to a voice conversion condition associated with a session interface of the application client, where the session interface includes a second user associated with the first user, the first user being the recipient of the voice message, and the second user being the sender of the voice message. receiving a voice conversion request sent by the user terminal, and obtaining the voice identifier from the voice conversion request; the voice conversion request is generated based on a queue position of the voice identifier in the target identifier queue; When the voice message corresponding to the voice identifier is found, converting the voice message to obtain converted text information corresponding to the voice message; Returning the converted text information to the user terminal, so that the user terminal outputs the converted text information to the location area where the voice message is located in the conversation interface of the application client; The step of the user terminal updating the target identification queue includes: In response to a first trigger operation on a conversation interface where the second user is located, the user terminal determines at least one voice message sent by the second user in the conversation interface as a target voice message, outputs the target voice message to the conversation interface, and obtains an initial level adjustment instruction in a voice conversion condition associated with the conversation interface; Based on the initial level adjustment instruction, the queue position of the target voice identifier corresponding to the target voice message is determined to be the first position in the initial first subqueue, and the queue position of the target voice identifier is adjusted from the first position to the second position in the initial first subqueue to obtain an adjusted initial first subqueue; the request priority of the identifier corresponding to the second position is greater than the request priority of the identifier corresponding to the first position; updating the target identification queue based on the adjusted initial first sub-queue and the second sub-queue; In response to a second trigger operation on at least one voice message in the conversation interface, determining the at least one voice message corresponding to the second trigger operation as a target voice message, and obtaining a target level adjustment instruction in the voice conversion condition; Based on the target level adjustment instruction, the adjusted initial first subqueue is determined as the target first subqueue, and the queue position of the target voice identifier corresponding to the target voice message is adjusted from the second position to the third position in the target first subqueue, to obtain an adjusted target first subqueue; the request priority of the identifier corresponding to the third position is greater than the request priority of the identifier corresponding to the second position; Based on the adjusted target first sub-queue and the second sub-queue, the updated target identification queue is updated.

6. A voice data processing device, characterized in that: include: A voice acquisition module is configured to, when an application client obtains a voice message of a conversation interface, obtain a voice identifier corresponding to the voice message, add the voice identifier to an initial identifier queue, and use the initial identifier queue to which the voice identifier is added as a target identifier queue; wherein the conversation interface includes a second user associated with a first user, the first user is a recipient of the voice message, and the second user is a sender of the voice message; the initial identifier queue includes a first subqueue and a second subqueue, the request priority of the second subqueue being greater than the request priority of the first subqueue; the first subqueue is configured to store a first voice identifier, the first voice identifier being used to represent the identifier of a first voice message to be sent in the application client for a voice conversion request, and the queue position of the first voice identifier in the first subqueue is determined according to the reception timestamp corresponding to the first voice message; the second subqueue is configured to store a second voice identifier, the second voice identifier being used to represent the identifier of a second voice message to which a voice conversion request has been sent in the application client, and the queue position of the second voice identifier in the second subqueue is determined according to the sending timestamp of the voice conversion request corresponding to the second voice message; a request sending module, configured to generate a voice conversion request carrying the voice identifier based on a queue position of the voice identifier in the target identifier queue, and send the voice conversion request to a server, so that the server obtains converted text information corresponding to the voice identifier based on the voice conversion request; a text receiving module, configured to receive the converted text information returned by the server, and output the converted text information to a location area where the voice message is located in the conversation interface; the voice message in the location area is associated with the converted text information; Wherein, the voice acquisition module includes: a first triggering unit configured to, in response to a first triggering operation on a conversation interface where the second user is located, determine at least one voice message sent by the second user in the conversation interface as a target voice message, output the target voice message to the conversation interface, and obtain an initial level adjustment instruction in a voice conversion condition associated with the conversation interface; a first adjustment unit, configured to determine, based on the initial level adjustment instruction, a queue position of the target voice identifier corresponding to the target voice message in the initial first subqueue as a first position, and adjust the queue position of the target voice identifier in the initial first subqueue from the first position to a second position, to obtain an adjusted initial first subqueue; wherein the request priority of the identifier corresponding to the second position is greater than the request priority of the identifier corresponding to the first position; a first updating unit, configured to update the target identification queue based on the adjusted initial first sub-queue and the second sub-queue; a second triggering unit configured to respond to a second triggering operation on at least one voice message in the conversation interface, determine the at least one voice message corresponding to the second triggering operation as a target voice message, and obtain a target level adjustment instruction in the voice conversion condition; a second adjustment unit, configured to, based on the target level adjustment instruction, determine the adjusted initial first subqueue as the target first subqueue, and adjust the queue position of the target voice identifier corresponding to the target voice message from the second position to a third position in the target first subqueue, to obtain an adjusted target first subqueue; wherein the request priority of the identifier corresponding to the third position is greater than the request priority of the identifier corresponding to the second position; The second updating unit is configured to update the updated target identification queue based on the adjusted target first sub-queue and the second sub-queue.

7. A voice data processing device, characterized in that: include: A voice sending module is used to generate a voice identifier corresponding to the voice message when a voice message of the application client is obtained, and send the voice message and the voice identifier to the user terminal so that the user terminal adds the voice identifier to the initial identifier queue, and uses the initial identifier queue with the voice identifier added as the target identifier queue; wherein, the initial identifier queue includes a first subqueue and a second subqueue, and the request priority of the second subqueue is greater than the request priority of the first subqueue; the first subqueue is used to store the first voice identifier, and the first voice identifier is used to represent the identifier of the first voice message to be sent in the voice conversion request in the application client, and the first voice in the first subqueue The queue position of the identifier is determined according to the receiving timestamp corresponding to the first voice message; the second subqueue is used to store the second voice identifier, the second voice identifier is used to represent the identifier of the second voice message for which the voice conversion request has been sent in the application client, and the queue position of the second voice identifier in the second subqueue is determined according to the sending timestamp of the voice conversion request corresponding to the second voice message; the target identifier queue is updated by the user terminal according to the voice conversion condition associated with the session interface of the application client, the session interface includes a second user associated with the first user, the first user is the recipient of the voice message, and the second user is the sender of the voice message; a request receiving module, configured to receive a voice conversion request sent by the user terminal and obtain the voice identifier from the voice conversion request; the voice conversion request is generated based on a queue position of the voice identifier in the target identifier queue; A text acquisition module is used to convert the voice message when the voice message corresponding to the voice identifier is found, and obtain converted text information corresponding to the voice message; a text sending module, configured to return the converted text information to the user terminal, so that the user terminal outputs the converted text information to the location area where the voice message is located in the conversation interface of the application client; The step of the user terminal updating the target identification queue includes: In response to a first trigger operation on a conversation interface where the second user is located, the user terminal determines at least one voice message sent by the second user in the conversation interface as a target voice message, outputs the target voice message to the conversation interface, and obtains an initial level adjustment instruction in a voice conversion condition associated with the conversation interface; Based on the initial level adjustment instruction, the queue position of the target voice identifier corresponding to the target voice message is determined to be the first position in the initial first subqueue, and the queue position of the target voice identifier is adjusted from the first position to the second position in the initial first subqueue to obtain an adjusted initial first subqueue; the request priority of the identifier corresponding to the second position is greater than the request priority of the identifier corresponding to the first position; updating the target identification queue based on the adjusted initial first sub-queue and the second sub-queue; In response to a second trigger operation on at least one voice message in the conversation interface, determining the at least one voice message corresponding to the second trigger operation as a target voice message, and obtaining a target level adjustment instruction in the voice conversion condition; Based on the target level adjustment instruction, the adjusted initial first subqueue is determined as the target first subqueue, and the queue position of the target voice identifier corresponding to the target voice message is adjusted from the second position to the third position in the target first subqueue, to obtain an adjusted target first subqueue; the request priority of the identifier corresponding to the third position is greater than the request priority of the identifier corresponding to the second position; Based on the adjusted target first sub-queue and the second sub-queue, the updated target identification queue is updated.

8. A computer device, characterized in that: include: Processor, memory, network interface; The processor is connected to a memory and a network interface, wherein the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program to execute the method according to any one of claims 1 to 4.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the program instructions are used to be called by a processor to execute the method according to any one of claims 1 to 4.

10. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads and executes the computer instructions from the computer-readable storage medium, so that the computer device performs the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Voice recognition method and voice recognition system

    CN104700836A

  • Instant messaging message reading and replying method and device and electronic equipment

    CN111490928A

  • Systems and methods for prioritizing messages for conversion from text to speech based on predictive user behavior

    US20190028421A1