Multi-person dialog speech transcription method and apparatus, and device, medium and program product

By identifying and transmitting the state attributes of the dialogue text on the server side, and updating the view on the client side to achieve error correction and intelligent line breaks, the problem of poor transcription effect of multi-person dialogues is solved, improving user experience and application stability.

WO2025236614A1PCT designated stage Publication Date: 2025-11-20MIGU CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/136020
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-11
Filing Date
2024-12-02
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Existing multi-person dialogue transcription technologies suffer from poor transcription results in scenarios such as multi-person online meetings, including issues such as unreasonable sentence breaks and line breaks, disordered dialogue order, and repeated content display.

Method used

By identifying the state attributes of the dialogue text on the server side and sending them to the client along with the speaker's identity, the client updates the dialogue page view based on these attributes, enabling error correction display and intelligent line breaks of the dialogue text, and distinguishing between the intermediate and final states of different speakers' text.

Benefits of technology

Ensuring that multi-person dialogue text is displayed in a reasonable and orderly manner according to the speaker and status improves transcription efficiency, enhances the user's visual and reading experience, and avoids lag and wasted memory resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136020_20112025_PF_FP_ABST
    Figure CN2024136020_20112025_PF_FP_ABST
Patent Text Reader

Abstract

A multi-person dialog speech transcription method and apparatus, and a device, a medium and a program product, which relate to the field of information, and are used for solving the problem of the transcription effect of an existing multi-person dialog transcription technique being relatively poor. The method comprises: receiving an instant message in a multi-person dialog scenario which is sent by a server, wherein the instant message includes senderId, first content and msgType of the first content, the first content is content corresponding to dialog speech information received by the server, and the msgType is used for indicating the statement integrity of the content (201); and when there is content corresponding to the senderId in multi-person dialog data displayed in a dialog page view, on the basis of msgType of pre-stored second content, using the first content to update the content corresponding to the senderId in the multi-person dialog data, so as to update the dialog page view, wherein the second content is the last content corresponding to the senderId in the multi-person dialog data before the first content is received (202).
Need to check novelty before this filing date? Find Prior Art

Description

Multi-person conversation voice transcription method, device, equipment, medium and program product

[0001] Cross-reference to Related Applications

[0002] The present application claims priority from Chinese Patent Application No. 202410581494.0 filed on May 11, 2024, the contents of which are incorporated herein by reference in their entirety. TECHNICAL FIELD

[0003] The present application relates to the field of information, and in particular to a multi-person conversation voice transcription method, device, equipment, medium and program product. BACKGROUND

[0004] In some application scenarios of real-time text transcription of multi-person voice conversation, for example, in scenarios such as multi-person online conference, it is necessary to transcribe multi-person conversation voice into text form of conference minutes. The existing technical solution is to simply convert multi-person conversation voice into text according to receiving time, and put the text content into a list view component piece by piece. In this way, a large number of multi-person real-time conversation texts displayed on the client exist many problems such as unreasonable sentence breaking and line breaking, chaotic multi-person conversation order, and repeated display of part of the content. It can be seen that the transcription effect of the existing multi-person conversation transcription technology is poor. SUMMARY

[0005] Embodiments of the present application provide a multi-person conversation voice transcription method, device, equipment, medium and program product to solve the problem of poor transcription effect of the existing multi-person conversation transcription technology.

[0006] In a first aspect, the embodiments of the present application provide a multi-person conversation voice transcription method, comprising:

[0007] receiving an instant message in a multi-person conversation scenario sent by a server, wherein the instant message contains a speaker identity identifier, a first conversation text, and a state attribute of the first conversation text, the first conversation text is a conversation text corresponding to conversation voice information received by the server, and the state attribute is used to indicate the sentence integrity of the conversation text;

[0008] in a case where the conversation text corresponding to the speaker identity identifier exists in multi-person conversation data displayed in a conversation page view, updating the conversation text corresponding to the speaker identity identifier in the multi-person conversation data by using the first conversation text according to a pre-stored state attribute of a second conversation text to update the conversation page view, wherein the second conversation text is the last conversation text corresponding to the speaker identity identifier existing in the multi-person conversation data before the first conversation text is received.

[0009] Optionally, after receiving the instant message under the multi-person conversation scenario sent by the server, the method further comprises:

[0010] In a case where the speaker identity identifier does not correspond to the conversation text in the multi-person conversation data, loading the speaker identity identifier and the first conversation text as a newly added conversation data to the multi-person conversation data.

[0011] Optionally, the updating of the conversation text corresponding to the speaker identity identifier in the multi-person conversation data by using the first conversation text according to the state attribute of the second conversation text stored in advance comprises:

[0012] In a case where the state attribute of the second conversation text is an intermediate state, replacing the second conversation text by using the first conversation text, and recording the state attribute of the first conversation text, wherein the first conversation text contains the second conversation text, and the intermediate state indicates that a sentence of the conversation text is incomplete.

[0013] In a case where the state attribute of the second conversation text is a final state, loading the speaker identity identifier and the first conversation text as a newly added conversation data to the multi-person conversation data, and recording the state attribute of the first conversation text, wherein the final state indicates that a sentence of the conversation text is complete.

[0014] Optionally, before receiving the instant message under the multi-person conversation scenario sent by the server, the method further comprises:

[0015] Creating a speaker set, a newly added conversation list and a final conversation list, wherein the newly added conversation list is used to store newly added conversation data, and the final conversation list is used to store the multi-person conversation data.

[0016] After receiving the instant message under the multi-person conversation scenario sent by the server, before the updating of the conversation text corresponding to the speaker identity identifier in the multi-person conversation data by using the first conversation text according to the state attribute of the second conversation text stored in advance, the method further comprises:

[0017] In a case where the instant message of the server is not received for the first time, emptying the newly added conversation list.

[0018] In a case where the conversation text corresponding to the speaker identity identifier exists in the multi-person conversation data displayed in the conversation page view, the updating of the conversation text corresponding to the speaker identity identifier in the multi-person conversation data by using the first conversation text according to the state attribute of the second conversation text stored in advance comprises:

[0019] in a case that the speaker set is not empty and the speaker set contains the speaker identity, updating the second dialogue text associated with the state attribute of the speaker identity stored in the speaker set with the first dialogue text to update the dialogue text corresponding to the speaker identity in the final dialogue list.

[0020] Optionally, before receiving the instant message under the multi-person dialogue scenario sent by the server, the method further comprises:

[0021] creating a speaker set, an added dialogue list and a final dialogue list, wherein the added dialogue list is used to store added dialogue data, and the final dialogue list is used to store the multi-person dialogue data;

[0022] after receiving the instant message under the multi-person dialogue scenario sent by the server, before loading the speaker identity and the first dialogue text as a piece of added dialogue data into the multi-person dialogue data, the method further comprises:

[0023] in a case that the instant message of the server is not received for the first time, emptying the added dialogue list;

[0024] in a case that the dialogue text corresponding to the speaker identity does not exist in the multi-person dialogue data, loading the speaker identity and the first dialogue text as a piece of added dialogue data into the multi-person dialogue data, comprising:

[0025] in a case that the speaker set is empty, or in a case that the speaker set is not empty and the speaker set does not contain the speaker identity, storing the speaker identity and the first dialogue text as a piece of added dialogue data in the added dialogue list, and storing the state attribute of the speaker identity and the first dialogue text in the speaker set, and adding the added dialogue list to the final dialogue list.

[0026] Optionally, in a case that the state attribute of the second dialogue text is an intermediate state, replacing the second dialogue text with the first dialogue text, comprising:

[0027] in a case that the state attribute of the second dialogue text is an intermediate state, replacing the second dialogue text with the first dialogue text according to the timing update interval corresponding to the intermediate state;

[0028] in a case that the state attribute of the second dialogue text is a final state, loading the speaker identity and the first dialogue text as a piece of added dialogue data into the multi-person dialogue data, comprising:

[0029] In a case where the state attribute of the second dialogue text is a final state, the speaker identity and the first dialogue text are loaded as a new dialogue data to the multi-person dialogue data according to a timing update interval time corresponding to the final state.

[0030] Optionally, the method further comprises:

[0031] In a case where a size of the multi-person dialogue data currently exceeds a preset threshold, a first N dialogue data in the multi-person dialogue data are deleted in a chronological order of dialogue time of the dialogue data, N being a positive integer.

[0032] Optionally, the method further comprises:

[0033] In a case where an operation of loading historical dialogue data for the dialogue page view is received, a preset number of historical dialogue data before a time point of dialogue data at a top of the dialogue page view are obtained from the server based on the time point, and the preset number of historical dialogue data are displayed and loaded in the dialogue page view.

[0034] In a second aspect, the embodiments of the present application further provide another multi-person dialogue speech transcription method, executed by a server, comprising:

[0035] Receiving dialogue speech information collected in a multi-person dialogue scenario and sent by a first client;

[0036] Converting the dialogue speech information into an instant message, wherein the instant message comprises a speaker identity, a first dialogue text corresponding to the dialogue speech information, and a state attribute of the first dialogue text, the state attribute being used to indicate sentence integrity of dialogue text;

[0037] Sending the instant message to a second client.

[0038] Optionally, the converting the dialogue speech information into an instant message comprises:

[0039] Converting the dialogue speech information into a first dialogue text, and determining a speaker identity corresponding to the dialogue speech information;

[0040] Analyzing the dialogue speech information to determine a state attribute of the first dialogue text, wherein the state attribute comprises an intermediate state or a final state, the intermediate state indicating that a sentence of dialogue text is incomplete, and the final state indicating that a sentence of dialogue text is complete;

[0041] Generating an instant message comprising the speaker identity, the first dialogue text, and the state attribute of the first dialogue text.

[0042] In a third aspect, the embodiments of the present application further provide a multi-person conversation voice transcription device, arranged at a client, comprising:

[0043] The first receiving module is configured to receive an instant message sent by the server in a multi-person conversation scenario, wherein the instant message comprises a speaker identity, a first conversation text, and a state attribute of the first conversation text, the first conversation text is a conversation text corresponding to the conversation voice information received by the server, and the state attribute is used to indicate the sentence integrity of the conversation text.

[0044] The updating module is configured to, in a case where the multi-person conversation data displayed in the conversation page view comprises the conversation text corresponding to the speaker identity, update the conversation text corresponding to the speaker identity in the multi-person conversation data by using the first conversation text according to the state attribute of the second conversation text stored in advance, to update the conversation page view, wherein the second conversation text is the last conversation text corresponding to the speaker identity in the multi-person conversation data before the first conversation text is received.

[0045] In a fourth aspect, the embodiments of the present application further provide another multi-person conversation voice transcription device, arranged at a server, comprising:

[0046] The second receiving module is configured to receive the conversation voice information collected in a multi-person conversation scenario and sent by the first client.

[0047] The conversion module is configured to convert the conversation voice information into an instant message, wherein the instant message comprises a speaker identity, a first conversation text corresponding to the conversation voice information, and a state attribute of the first conversation text, and the state attribute is used to indicate the sentence integrity of the conversation text.

[0048] The sending module is configured to send the instant message to the second client.

[0049] In a fifth aspect, the embodiments of the present application further provide a terminal device, comprising a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the multi-person conversation voice transcription method according to the first aspect when executing the computer program, or implements the steps in the multi-person conversation voice transcription method according to the first aspect.

[0050] In a sixth aspect, the embodiments of the present application further provide a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executable on a processor to implement the steps in the multi-person conversation voice transcription method according to the first aspect, or to implement the steps in the multi-person conversation voice transcription method according to the first aspect.

[0051] In the embodiment of the present application, the server receives the conversation voice information collected in the multi-person conversation scene sent by the first client; converts the conversation voice information into an instant message, wherein the instant message contains a speaker identity identifier, a first conversation text corresponding to the conversation voice information, and a state attribute of the first conversation text, the state attribute being used to indicate the sentence integrity of the conversation text; and sends the instant message to the second client.

[0052] The client receives the instant message in the multi-person conversation scene sent by the server, wherein the instant message contains a speaker identity identifier, a first conversation text, and a state attribute of the first conversation text, the first conversation text being a conversation text corresponding to the conversation voice information received by the server, and the state attribute being used to indicate the sentence integrity of the conversation text; in the case that the conversation data displayed in the conversation page view contains the conversation text corresponding to the speaker identity identifier, the first conversation text is used to update the conversation text corresponding to the speaker identity identifier in the multi-person conversation data according to the state attribute of the second conversation text stored in advance, so as to update the conversation page view, wherein the second conversation text is the last conversation text corresponding to the speaker identity identifier in the multi-person conversation data before the first conversation text is received.

[0053] In this way, the server identifies the state attribute of the multi-person conversation text, and sends the conversation text, the state attribute, and the speaker identity identifier to the client, and the client updates the corresponding conversation data in the conversation page view based on the received state attribute of the conversation text and the speaker identity identifier, so as to ensure that the multi-person conversation text is displayed in a reasonable and orderly manner without repetition according to the speaker and the speaking state, and the effect of the multi-person conversation transcription text is improved. BRIEF DESCRIPTION OF DRAWINGS

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0055] FIG. 1 is an example diagram of the conversation text transcription effect of the prior ordinary scheme;

[0056] FIG. 2 is a flowchart of a multi-person conversation voice transcription method according to an embodiment of the present application;

[0057] FIG. 3a is a schematic diagram of the display effect of the intermediate state text in real-time transcription according to an embodiment of the present application;

[0058] Fig. 3b is a schematic diagram of a real-time transcription final state text display effect provided by an embodiment of the present application;

[0059] Fig. 3c is a schematic diagram of a real-time transcription complete conversation text display effect provided by an embodiment of the present application;

[0060] Fig. 4 is a process flow diagram of displaying conversation information by one-time dynamic error correction provided by an embodiment of the present application;

[0061] Fig. 5 is a second flow chart of a multi-person conversation voice transcription method provided by an embodiment of the present application;

[0062] Fig. 6 is a structure diagram of a multi-person conversation voice transcription device provided by an embodiment of the present application;

[0063] Fig. 7 is a second structure diagram of a multi-person conversation voice transcription device provided by an embodiment of the present application;

[0064] Fig. 8 is a structure diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0065] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0066] To make the embodiments of the present application clearer, the related technical knowledge involved in the embodiments of the present application will be introduced as follows:

[0067] At present, when a large number of multi-person real-time texts are displayed on a client, the existing technical solution is to simply put the converted text content into a list component piece by piece, and simply accumulate the amount of data.

[0068] The related technical solution has the following related problems when displaying multi-person real-time conversation texts:

[0069] From the perspective of visual experience, the existing solution is simply displayed piece by piece in a list view component, and will not be sorted and distinguished.

[0070] From the perspective of performance, the simple accumulation of a large amount of data in the related technical solution will cause the list view component to lag, increase the memory resource consumption, and reduce the application stability. From the perspective of interaction, when the number of data accumulates a lot, it is not convenient for the user to view the speech content of a specific speaker in the past period of time.

[0071] FIG. 1 is a schematic diagram of a simulated transcription effect of a conversation transcript between friends about holiday arrangements according to a scheme in the related art, from which it can be seen that the transcription effect of the general scheme has many problems such as illogical sentence breaks and line breaks, chaotic order of multi-person conversation, only cumulative text line by line, and repeated display of part of the content.

[0072] To solve the above problems, the embodiments of the present application realize text correction display and intelligent line break by distinguishing different speakers and judging the state attribute of the speaking content, which is different from the simple display of the general scheme. The embodiments of the present application display the text in the intermediate state on a single line and display the final result on a new line, which significantly improves the visual and reading experience of the user.

[0073] In addition, when displaying conversation data, if new data is only added to the list continuously, it will cause lag after a long time, and due to the large amount of data storage, performance problems will occur. Some embodiments of the present application also provide a processing scheme for a large amount of real-time conversation data. By setting a threshold capacity and using a scheme to pull up historical data, the size of the data processed by the list view at a time is limited to avoid the problem of lag, and the function of querying historical conversation data is also met.

[0074] To help better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are defined and explained as follows:

[0075] The embodiments of the present application are realized by interaction between the client and the server.

[0076] In specific implementation, a service component can be designed on the server, which has the following functions: when a multi-person conference is in progress, the server establishes a long connection with the client through the network to transmit conversation information in real time. When a user participates in the conference, the client integrates a speech software development kit (SDK) to collect the speech information of the participants and transmit it to the server; on the one hand, the server generates a unique identity code for the user when it establishes a connection with the client, so that different speakers can be accurately distinguished even in a complex context of multi-person speaking, i.e. the senderId in the above-mentioned term; on the other hand, the server processes the audio data, analyzes the tone, pause, sentence break, duration, etc. of the participants through a trained machine model, judges whether the conversation text is in an intermediate state or a final state, and puts it as an attribute value, i.e. msgType in the above-mentioned term, into the text model of the conversation, and transmits it to the client in real time.

[0077] In the client, a general function module is designed, which internally implements a timer function and simultaneously calls the list view component of the system. In operation, it receives and processes IM messages from the server as msgData at regular intervals, and at the same time, it refreshes the conversation view component at regular intervals (Unit Interval, UI). The internal principle of the module is as follows: if it is intermediate state data, the existing content is replaced according to the senderId of the speaker, and the application programming interface (API) method of the system list component is used to refresh the view, which visually forms the effect of error correction display; if it is the final state result, a new data is directly added, and the API method is called to refresh the view and make it scroll, thereby forming the effect of intelligent line breaking. In addition, the timer interval can be distinguished according to the message type, such as setting the error correction display interval to 200 milliseconds and the intelligent line breaking interval to 1000 milliseconds, so that the rhythm of text changes can more realistically simulate the actual conversation process and the display effect is better.

[0078] The above client function module is widely applicable to multiple platforms such as Android, IOS, HarmonyOS, and Web client. For example, in an Android application, the system Recyclerview list component is called according to the above principle, the view is refreshed using the notifyDataSetChanged() method in the Android API, and the smoothScrollToPosition() method is used for automatic scrolling. In this way, in the transcription process of the conversation, intermediate state data is corrected and final state data is displayed in line. Similarly, on the IOS system, we can also optimize the display of real-time conversation text. After integrating the above function module, the system list component UITableView is called, and data with state from the server is received, and the "error correction display and intelligent line breaking" display effect can be achieved.

[0079] In actual scenarios, there are many dialogue persons participating. If the dialogue is relatively intense or lasts for a long time, the amount of dialogue data generated will be large. The technical points of the present solution are in addition to the display data of the above-mentioned state types, and are as follows: first, how to distinguish different dialogue persons and whether their speaking content is in the intermediate state or the final state, which can be realized by repairMap; second, how to improve the application stability under a large amount of dialogue data, avoid performance problems such as lag and memory overflow, which is realized by limiting the total amount of data (i.e. the size of allList); third, on the basis of limiting the total amount, how to consult historical dialogue data, which depends on the pull-to-refresh mechanism: when the pull operation is performed at the top of the view, the timestamp at this place is used to query historical information from the server, and the queryList is stored to refresh the view.

[0080] The multi-person dialogue voice transcription method provided by the embodiments of the present application will be described in detail below in combination with the accompanying drawings, specific embodiments and application scenarios.

[0081] Referring to FIG. 2, which is a flowchart of the multi-person dialogue voice transcription method provided by the embodiments of the present application, executed by a client, as shown in FIG. 2, the embodiments of the present application include the following steps:

[0082] Step 201, receiving an instant message in a multi-person dialogue scenario sent by a server, wherein the instant message contains a speaker identity identifier, a first dialogue text and a state attribute of the first dialogue text, the first dialogue text is a dialogue text corresponding to dialogue voice information received by the server, and the state attribute is used to indicate the sentence integrity of the dialogue text.

[0083] In the embodiments of the present application, the client can initiate or join a multi-person dialogue scenario, such as a multi-person online conference, in which the client and the server establish a long link, such as a websocket long link, through a network to transmit dialogue information in real time.

[0084] In the process of multi-person dialogue, the client can collect dialogue voice information of participating users and upload it to the server for the server to parse into dialogue text and determine the dialogue text state attribute. For example, when a user participates, the client integrates a voice SDK to collect the voice information of the participants and transmits it to the server.

[0085] It should be noted that, in order to distinguish different speakers, the server can generate a unique identity code / identifier of the client user when establishing a connection with the client, so that even in the complex context of multiple people speaking, different speakers can still be accurately distinguished based on the unique identity code / identifier carried in the uploaded dialogue voice information, i.e. the senderId in the above-mentioned terminology.

[0086] The server side processes each of the received client uploaded dialogue voice information in real time, converts each of the dialogue voice information into an instant message IM, specifically, can convert the dialogue voice information into text, binds the unique identity code / identifier of the dialogue voice information as the speaker identity identifier of the voice information, and can also determine the state attribute of the corresponding text of the dialogue voice information through comprehensive analysis of the dialogue voice information, such as whether the voice is finished, whether it is a complete sentence, etc.

[0087] More specifically, the server side can perform text conversion and state analysis processing on each of the dialogue voice information, and can determine whether the dialogue voice information is finished or a complete sentence through comprehensive analysis of the speaker's tone, pause, punctuation, duration, etc. by using a trained machine model, and further determine whether the corresponding dialogue text is in an intermediate state or a final state, where the intermediate state indicates that the current dialogue text sentence is not complete and the speaker has not finished speaking, and the final state indicates that the current dialogue text sentence is complete and the speaker has finished speaking; and can generate an instant message with the determined state as an attribute value, i.e. msgType in the above-mentioned terminology, and the dialogue text together with the speaker identity identifier, and transmit it in real time to the participating client that needs real-time transcription.

[0088] Correspondingly, the client receives the instant message IM sent by the server side, which contains the speaker identity identifier, the first dialogue text and the state attribute of the first dialogue text, i.e. the dialogue text converted by the server side from the currently received dialogue voice information.

[0089] It should be noted that the client can convert the received IM message into the data type msgData required by the list view component of the client page, including the speaker identity identifier (senderId), the state attribute (msgType), the dialogue text (content), etc.

[0090] Step 202, in the case that the dialogue text corresponding to the speaker identity identifier exists in the multi-person dialogue data displayed in the dialogue page view, the first dialogue text is used to update the dialogue text corresponding to the speaker identity identifier in the multi-person dialogue data according to the pre-stored state attribute of the second dialogue text, to update the dialogue page view, wherein the second dialogue text is the last dialogue text corresponding to the speaker identity identifier existing in the multi-person dialogue data before the first dialogue text is received.

[0091] The above-mentioned dialogue page view can be understood as a page list view component displaying dialogue transcription text, which displays the multi-person dialogue data after transcription into text, i.e. the dialogue text data finally presented on the client.

[0092] Specifically, in this step, according to the speaker identity in the current IM message, it can be judged whether the speaker is a newly participating speaker or a historical speaker with existing conversation record, that is, it is equivalent to judging whether the conversation text corresponding to the speaker identity exists in the multi-person conversation data displayed in the current conversation page view.

[0093] If it exists, it indicates that the current page has displayed part of the unfinished speech or the complete last speech of the speaker indicated by the speaker identity, so the state attribute corresponding to the second conversation text stored by the client when receiving the IM message corresponding to the second conversation text of the speaker identity can be obtained, and the second conversation text is the latest historical conversation text of the speaker identity in the currently displayed multi-person conversation data. By determining how to update the conversation text corresponding to the speaker identity in the multi-person conversation data in the current conversation page view based on the state attribute corresponding to the second conversation text, the multi-person conversation data in the conversation page view is updated.

[0094] Exemplarily, according to whether the state attribute of the second conversation text is an intermediate state or a final state, it can be determined whether to update the second conversation text in the multi-person conversation data or to add the first conversation text currently corresponding to the speaker identity in the multi-person conversation data, so as to ensure the coherence and readability of the conversation text.

[0095] Optionally, the updating of the conversation text corresponding to the speaker identity in the multi-person conversation data by using the first conversation text according to the state attribute of the second conversation text comprises:

[0096] In the case where the state attribute of the second conversation text is an intermediate state, the second conversation text is replaced by the first conversation text, and the state attribute of the first conversation text is recorded, wherein the first conversation text contains the second conversation text, and the intermediate state indicates that the sentence of the conversation text is incomplete.

[0097] In the case where the state attribute of the second conversation text is a final state, the speaker identity and the first conversation text are loaded as a new conversation data to the multi-person conversation data, and the state attribute of the first conversation text is recorded, wherein the final state indicates that the sentence of the conversation text is complete.

[0098] That is, in some embodiments, in the case where it is judged that the last piece of dialogue text of the current speaker is intermediate state data, the intermediate state text needs to be replaced, specifically, the first received dialogue text can be used as the corrected data to replace the intermediate state text of the current speaker, i.e., the second text data, and then the view is refreshed using the API method of the system list view component, so that the visual effect of the correction display is formed.

[0099] It should be noted that in the case where the last piece of dialogue text of the same speaker is in the intermediate state, the next piece of dialogue text will contain the last piece of intermediate state dialogue text, i.e., the server will process the received dialogue voice, and when it is judged that the current dialogue text is a continuation of the last piece of dialogue text, the last piece of incomplete dialogue text and the current dialogue text will be combined to form a complete piece of dialogue text and sent to the client.

[0100] For example, the current received dialogue text of the speaker "Xiaozhang" is "I have made an appointment with Xiaowang", and the last piece of dialogue text of the speaker "Xiaozhang" displayed in the dialogue page view is "I have made an appointment", and the state attribute of the text is stored as intermediate state, so the current dialogue text "I have made an appointment with Xiaowang" of "Xiaozhang" can be used to replace the original intermediate state text "I have made an appointment" displayed in the dialogue page view.

[0101] In another case where it is judged that the last piece of dialogue text of the current speaker is final state data, a new piece of dialogue data can be directly added, i.e., the current speaker identity and the first dialogue text are loaded as a new piece of dialogue data into the current dialogue page view, specifically, the current dialogue page view can be intelligently line-wrapped, and the new piece of dialogue data is displayed in the new line, and then the API method is called to refresh the view and make it scroll, thereby forming the effect of intelligent line-wrapping.

[0102] For example, the current received dialogue text of the speaker "Xiaozhang" is "What do you plan to do?", and the last piece of dialogue text of the speaker "Xiaozhang" displayed in the dialogue page view is "I have made an appointment with Xiaowang", and the state attribute of the text is stored as final state, so the current dialogue text "I have made an appointment with Xiaowang" of "Xiaozhang" can be added as a new piece of dialogue text data to the dialogue page view.

[0103] It should be noted that the state attribute of the received first dialogue text can also be recorded, so that when the next piece of dialogue text of the speaker is received, the state attribute of the last piece of dialogue text of the speaker is used to judge how to update the dialogue text of the speaker in the page view.

[0104] The embodiment realizes the transcription display effect of dynamic error correction and intelligent line breaking by displaying the single-line correction of the intermediate-state text and the line breaking of the final result, and enhances the user visual experience.

[0105] Optionally, in the case that the state attribute of the second dialogue text is in the intermediate state, the first dialogue text is used to replace the second dialogue text, including:

[0106] In the case that the state attribute of the second dialogue text is in the intermediate state, the first dialogue text is used to replace the second dialogue text according to the timing update interval time corresponding to the intermediate state.

[0107] In the case that the state attribute of the second dialogue text is in the final state, the speaker identity and the first dialogue text are loaded as a new dialogue data into the multi-person dialogue data, including:

[0108] In the case that the state attribute of the second dialogue text is in the final state, the speaker identity and the first dialogue text are loaded as a new dialogue data into the multi-person dialogue data according to the timing update interval time corresponding to the final state.

[0109] In some embodiments, a timer function can also be designed on the client side. When working, on the one hand, it will receive instant messages (IM) from the server in a timely manner and process them as msgData, and on the other hand, it will refresh the dialogue view component in the UI thread in a timely manner.

[0110] More specifically, the timer interval time can be determined according to the received IM message type msgType, i.e. the state attribute. For example, the interval time for displaying the intermediate-state dialogue data for error correction is set to 200 milliseconds, and the interval time for intelligent line breaking of the final-state dialogue data is set to 1000 milliseconds. When refreshing the dialogue data in the dialogue page view, the dialogue text can be updated and processed for error correction display and intelligent line breaking according to the timing interval time. In this way, the rhythm of text changes can more realistically simulate the actual dialogue process, and the dialogue transcription display effect is better.

[0111] Optionally, after the step 201, the method further includes:

[0112] In the case that the speaker identity corresponding dialogue text does not exist in the multi-person dialogue data, the speaker identity and the first dialogue text are loaded as a new dialogue data into the multi-person dialogue data.

[0113] In some embodiments, in the case that it is determined that the current speaker is a new participant, i.e., the conversation text of the speaker is not displayed in the current conversation page view, the first conversation text of the current speaker and the speaker identity identifier thereof can be directly loaded into the current conversation page view as a piece of new conversation data, i.e., the new conversation data is displayed in a new line in the current conversation page view.

[0114] In this way, the conversation of the new participant can be displayed separately in a new line, so that the display order of the conversations of different speakers is reasonable and not disordered and overlapped.

[0115] Optionally, before the step 201, the method further includes:

[0116] creating a speaker set, a new conversation list and a final conversation list, wherein the new conversation list is used to store new conversation data, and the final conversation list is used to store the multi-person conversation data;

[0117] After the step 201 and before the step 202, the method further includes:

[0118] in the case that the instant message of the server is not received for the first time, emptying the new conversation list;

[0119] The step 202 includes:

[0120] in the case that the speaker set is not empty and the speaker set contains the speaker identity identifier, updating the conversation text corresponding to the speaker identity identifier in the final conversation list by using the first conversation text according to the state attribute of the second conversation text associated with the speaker identity identifier stored in the speaker set.

[0121] In some embodiments, the client can first perform function module initialization, including data preparation: creating necessary data sets, such as a speaker set (e.g., repairMap), a new conversation list (e.g., newList), and a final conversation list (e.g., allList). The repairMap is used to store a speaker identity identifier (senderId) and a state attribute (msgType), the new conversation list is used to store new conversation data, and the final conversation list is used to store multi-person conversation data to be finally displayed in a conversation page view.

[0122] The client can also empty the new conversation list newList in the case that the instant message of the server is not received for the first time, so as to prepare for processing data.

[0123] Then, it can be determined whether the repairMap is empty; if the repairMap is not empty, it is further determined whether the repairMap contains the senderId of the currently received msgData. If the repairMap contains the senderId, it indicates that the current speaker originally exists and has transcribed dialogue text, and thus it can be further determined that the msgType of the msgData is final or intermediate.

[0124] If the msgType is final, i.e., indicating that the last dialogue text of the current senderId is the final result, the current msgData can be directly stored in the newList, and the processed newList is loaded into the allList. If the msgType is intermediate, i.e., indicating that the last dialogue text of the current senderId is not the final result, the allList set needs to be traversed, and the data with the same senderId in the allList is found according to the senderId of the msgData, and the content of the latest msgData is used to replace the old content.

[0125] It should be noted that when updating the allList according to the current msgData, the state attribute msgType of the current senderId stored in the repairMap can also be updated synchronously, so as to determine how to update the allList when the next msgData of the senderId is received.

[0126] In this way, in the embodiment, through the data sets such as the repairMap, the newList, and the allList, the correction display and the ordered update display of the dialogue data of the existing speaker in the dialogue page view can be well implemented.

[0127] Optionally, before the step 201, the method further includes:

[0128] creating a speaker set, a new dialogue list, and a final dialogue list, wherein the new dialogue list is used to store new dialogue data, and the final dialogue list is used to store the multi-person dialogue data;

[0129] After the step 201, before the speaker identity and the first dialogue text are loaded as a new dialogue data into the multi-person dialogue data, the method further includes:

[0130] In the case of not receiving the instant message of the server for the first time, the new dialogue list is emptied;

[0131] In the case that the speaker identity identifier does not correspond to conversation text in the multi-person conversation data, the speaker identity identifier and the first conversation text are loaded as a new conversation data to the multi-person conversation data, including:

[0132] In the case that the speaker set is empty, or in the case that the speaker set is not empty and the speaker identity identifier is not included in the speaker set, the speaker identity identifier and the first conversation text are stored as a new conversation data in the new conversation list, and the speaker identity identifier and the first conversation text are associated with a state attribute and stored in the speaker set, and the new conversation list is added to the final conversation list.

[0133] In some embodiments, as introduced in the above embodiment, the client can first initialize the functional modules, including data preparation: creating necessary data sets, such as a speaker set (e.g., repairMap), a new conversation list (e.g., newList), and a final conversation list (e.g., allList). The client can empty the new conversation list newList in the case of receiving an instant message from the server for the first time, and prepare for data processing.

[0134] Then, it can be determined whether the repairMap is empty; if the repairMap is empty, it indicates that the current speaker is the first speaker, and the currently received msgData is directly stored in the newList, and the senderId of the new data msgData is stored as a new key and the msgType is stored as a new value in the repairMap; if the repairMap is not empty, it is further determined whether the repairMap includes the senderId of the currently received msgData; if the repairMap does not include the senderId, it indicates that a new speaker participates in the conversation, and the new data msgData is still stored in the newList, and the senderId of the msgData is stored as a new key and the msgType is stored as a new value in the repairMap. Then, the processed newList can be loaded into the allList.

[0135] In this way, in this embodiment, through the created repairMap, newList, allList, and other data sets, the intelligent line break display and ordered update display of the new speaker conversation data in the conversation page view can be well achieved.

[0136] Further, the creation of the speaker set, the new conversation list, and the final conversation list includes:

[0137] creating a speaker set, a new conversation list, a final conversation list and a history conversation list, and creating a conversation page view and an adapter;

[0138] After adding the new conversation list to the final conversation list, or updating the conversation text corresponding to the speaker identity in the final conversation list according to the state attribute of the second conversation text associated with the speaker identity stored in the speaker set, the method further comprises:

[0139] The multi-person conversation data in the final conversation list is transmitted into the adapter to adaptively adjust the format of the multi-person conversation data, and the multi-person conversation data after the format adjustment is transmitted to the conversation page view to refresh the display of the conversation page view.

[0140] That is, in some embodiments, in addition to creating data sets repairMap, newList, allList, a history conversation list (such as queryList) can also be created, and view preparation is performed, that is, a list view component is created on a page, that is, a conversation page view is created, and an adapter (adpater) required for filling data is created, and the adapter is bound to the view.

[0141] Specifically, after loading or updating the conversation data in allList, the allList array can be transmitted into the adapter, and the conversation data is transmitted to the list view component. Then the API of the list view component is called to achieve the dynamic error correction effect, and the automatic scrolling effect is achieved as the conversation information is wrapped.

[0142] For example, the Recyclerview list component of the system can be called according to the foregoing principle, the notifyDataSetChanged() method in the Android API is used to refresh the view, and the smoothScrollToPosition() method is used to automatically scroll; in this way, in the transcription process of the conversation, the intermediate state data is corrected, and the final state data is displayed in a wrapped manner. Similarly, on the IOS system, we can also optimize the display of real-time conversation text. After integrating the above-mentioned functional modules, the list component UITableView of the system is called, and the data with states from the server is received, and the display effect of error correction and intelligent wrapping can also be achieved.

[0143] In this way, it can be ensured that the continuously loaded transcription conversation data can be automatically adapted to the page view format and displayed normally and reasonably.

[0144] Optionally, the method further comprises:

[0145] If the size of the current multi-person dialogue data exceeds a preset threshold, delete the first N dialogue data in the multi-person dialogue data according to the dialogue time sequence of each dialogue data in the multi-person dialogue data, where N is a positive integer.

[0146] In some embodiments, it is also possible to consider how to improve application stability under large amounts of dialogue data, avoid performance issues such as lag and memory overflow, specifically by limiting the total amount of data in the dialogue page view, i.e., the size of the allList mentioned above.

[0147] Specifically, after loading the processed `newList` into `allList`, the processing of one IM message's data is complete. At this point, the size of `allList` can be checked. If the size of `allList` exceeds a set threshold, the earliest items added to `allList` are removed to ensure that the size of `allList` does not exceed the set threshold. This improves application stability under large amounts of dialogue data and avoids performance issues such as lag and memory overflow.

[0148] Optionally, the method further includes:

[0149] Upon receiving an operation to load historical dialogue data for the dialogue page view, based on the time point of the dialogue data at the top of the dialogue page view, a preset number of historical dialogue data prior to the time point is obtained from the server, and the preset number of historical dialogue data is loaded and displayed in the dialogue page view.

[0150] In some embodiments, in addition to limiting the total amount of data, consideration can also be given to how to retrieve historical dialogue data. This can be achieved by relying on a pull-to-refresh mechanism. For example, when a pull-down operation is performed at the top of the view, historical dialogue data will be retrieved from the server based on the timestamp at that location, stored in a queryList, and then the dialogue page view will be refreshed.

[0151] Specifically, if the user swipes up the list view component to the top, historical data below a set threshold is retrieved from the server based on the time point of the conversation data at the top, stored in queryList, and passed to the adapter. The conversation page view is then refreshed to display the historical conversation data.

[0152] As can be seen, this embodiment also provides a solution for processing large amounts of real-time dialogue data. When displaying dialogue data, simply adding new data to the list continuously will cause lag over time, and the large amount of data stored will lead to performance issues. By setting a data threshold capacity and using a pull-to-load historical data approach, the size of data processed in a single list view is limited, avoiding lag issues while still meeting the functional requirement of querying historical dialogue data.

[0153] Please refer to FIG. 3a, FIG. 3b and FIG. 3c, also taking the conversation between friends about holiday arrangement as an example, the "error correction display and intelligent line breaking effect" in the technical solution of the present application is shown. Among them, FIG. 3a shows that the text display effect in the middle state is different from the final state in the conversation process, and they will be replaced when the server's final state text is received subsequently; FIG. 3b shows the text effect after replacement; FIG. 3c shows the complete conversation effect. It can be seen that compared with the transcription effect of the technical solution shown in FIG. 1, the transcription effect in the conversation process and the final transcription record provided by the technical solution of the present application are obviously superior to the ordinary solution.

[0154] In order to more clearly illustrate the embodiment of the present application, the following will introduce the processing flow of dynamic error correction display conversation information combined with FIG. 4. Specifically, the technical solution is composed of the following four parts with fourteen steps.

[0155] The first part includes the following first to fifth steps, which is mainly the initialization of the function module. Its role is to prepare the required data and view, and establish a good connection with the server, and convert the conversation data into view data type.

[0156] The second part includes the following sixth to tenth steps, which is the key part of the present application. In the data layer, it is prepared for error correction display and intelligent line breaking. For the existing speaker and the state is the intermediate state, the content is replaced, and the error correction data is prepared; for the existing speaker but the state is the final state, or the new speaker, the intelligent line breaking data is prepared, that is, the newList adds a data. Finally, these data will be stored in allList, and from the performance point of view, the size of allList is limited, so the required data source for display is prepared.

[0157] The third part includes the following eleventh and twelfth steps, which is to call the corresponding API in the view layer to realize the error correction display and intelligent line breaking effect.

[0158] The fourth part includes the following thirteenth and fourteenth steps, which shows how to consult the historical conversation content mechanism and how to process the new conversation data mechanism in the query process.

[0159] The detailed steps are as follows:

[0160] The first step is the initialization of the function module, including data preparation: creating necessary data sets repairMap, newList, allList and queryList; view preparation: creating list view in the page and creating adapter adapter required to fill data, and binding the adapter with the view.

[0161] Second, establish a websocket long link with the server, used to get the server's IM message.

[0162] Third, new timing task, when entering the dialogue scene, in the background ready to receive server data, and will be in the UI thread refresh view regularly.

[0163] Fourth, IM message conversion to list view required data type msgData.

[0164] Fifth, empty newlist, ready to handle data.

[0165] Sixth, judge repairMap is empty. If empty, the data into the newlist, and the new data senderId as a new key, msgType as a new value into repairMap; if not empty, step seven.

[0166] Seventh, on the basis of the sixth step repairMap is not empty, judge repairMap whether contains the senderId of msgData in the fourth step, if not containing the senderId, indicating that there are new speakers involved in the dialogue, at this time still need to store the new data into the newlist, and the new data senderId as a new key, msgType as a new value into repairMap; if repairMap contains msdData senderId, step eight.

[0167] Eighth, on the basis of the seventh step repairMap contains msgData senderId, judge msgData msgType type. If msgType is the final result, then directly into the newlist. If msgType is not the final result, then traverse the entire allList set, according to msgData senderId find allList with the same senderId that data, with the latest msgData content to replace the old content.

[0168] Ninth, the newlist into allList. Thus, a IM related data processing is completed.

[0169] Tenth, check allList size, if allList size exceeds the set threshold, the earliest into allList data removed.

[0170] The tenth step is to pass the allList array into the adapter to pass the conversation data to the list view component.

[0171] The twelfth step is to call the API of the list view component to achieve dynamic error correction and automatic scrolling as the conversation information is wrapped.

[0172] The thirteenth step is to pull the historical information from the server if the user swipes the list view component to the top, and store it in queryList and pass it into the adapter to refresh the view and display the historical information.

[0173] The fourteenth step is to continuously receive IM messages in the background when the user previews the historical information in the thirteenth step, and repeat the fifth to twelfth steps when the user swipes to the bottom.

[0174] The embodiments of the present application enhance the user's visual experience through the software implementation scheme of dynamic error correction and intelligent wrapping. The embodiments of the present application optimize the processing scheme of a large amount of conversation data, reduce the use of memory resources, reduce the drawing pressure of the system view list component, and enhance the stability of the application.

[0175] The embodiments of the present application can be used in conference systems, forums and other scenarios to display a large number of real-time text conversations, which improves the user experience. At the same time, based on the optimized data storage scheme, the resource consumption is reduced, and the stability of the application is improved.

[0176] The client in the embodiments of the present application refers to a program that provides local services for clients corresponding to the server. It is usually installed on a common client and needs to cooperate with the server to run. Common clients include web browsers, email clients, instant messaging clients, etc. With the rapid development of mobile Internet, smart phones have been widely popularized in society. Smart phones can access wireless networks through mobile communication networks, and users can install various APP clients and use services provided by third-party service providers. Smart phones are equipped with independent operating systems, such as Android system, IOS system, HarmonyOS system, etc.

[0177] The multi-person conversation voice transcription method of the embodiment of the application, a client receives instant messages sent by a server in a multi-person conversation scenario, wherein the instant messages contain a speaker identity identifier, a first conversation text, and a state attribute of the first conversation text, the first conversation text is a conversation text corresponding to conversation voice information received by the server, and the state attribute is used to indicate the sentence integrity of the conversation text; in the case that the conversation text corresponding to the speaker identity identifier exists in multi-person conversation data displayed in a conversation page view, the first conversation text is used to update the conversation text corresponding to the speaker identity identifier in the multi-person conversation data according to a pre-stored state attribute of a second conversation text, so as to update the conversation page view, wherein the second conversation text is the last conversation text corresponding to the speaker identity identifier existing in the multi-person conversation data before the first conversation text is received. In this way, the state attribute of the multi-person conversation text is identified by the server, and the conversation text, the state attribute thereof, and the speaker identity identifier are sent to the client, the client updates the corresponding conversation data in the conversation page view based on the received state attribute of the conversation text and the speaker identity identifier, so that the multi-person conversation text can be displayed in a reasonable and orderly manner without repetition according to the speaker and the speaking state, and the effect of the multi-person conversation transcription text is improved.

[0178] Referring to FIG. 5, FIG. 5 is a flowchart of a multi-person conversation voice transcription method provided by the embodiment of the application, which is executed by a server, as shown in FIG. 5, and includes the following steps:

[0179] Step 501, receiving conversation voice information collected in a multi-person conversation scenario and sent by a first client.

[0180] Step 502, converting the conversation voice information into instant messages, wherein the instant messages contain a speaker identity identifier, a first conversation text corresponding to the conversation voice information, and a state attribute of the first conversation text, and the state attribute is used to indicate the sentence integrity of the conversation text.

[0181] Step 503, sending the instant messages to a second client.

[0182] The first client and the second client can be the same client or different clients participating in the same multi-person conversation, the first client is a client that speaks in the conversation, and the second client is a client that requests real-time transcription from the server.

[0183] Optionally, the step 502 includes:

[0184] converting the conversation voice information into a first conversation text and determining a speaker identity identifier corresponding to the conversation voice information;

[0185] analyzing the dialogue voice information to determine a state attribute of the first dialogue text, wherein the state attribute comprises an intermediate state or a final state, the intermediate state indicating that a sentence of dialogue text is incomplete, and the final state indicating that a sentence of dialogue text is complete;

[0186] generating an instant message containing the speaker identity, the first dialogue text, and the state attribute of the first dialogue text.

[0187] It should be noted that the embodiment is a service end corresponding to the embodiment shown in FIG. 2, and the specific implementation can refer to the related description in the embodiment shown in FIG. 2. To avoid repetition, details are not repeated here.

[0188] The multi-person dialogue voice transcription method of the embodiment of the application, the service end receives dialogue voice information collected in a multi-person dialogue scene sent by a first client; converts the dialogue voice information into an instant message, wherein the instant message contains a speaker identity, a first dialogue text corresponding to the dialogue voice information, and a state attribute of the first dialogue text, and the state attribute is used to indicate the completeness of a sentence of dialogue text; and sends the instant message to a second client. In this way, the state attribute of the multi-person dialogue text is identified by the service end, and the dialogue text, the state attribute thereof, and the speaker identity are sent to the client. The client updates the corresponding dialogue data in the dialogue page view based on the received state attribute of the dialogue text and the speaker identity, which can ensure that the multi-person dialogue text is displayed in a reasonable and orderly manner without repetition according to the speaker and the speaking state, and improves the effect of multi-person dialogue transcription text.

[0189] The embodiment of the application also provides a multi-person dialogue voice transcription device. Referring to FIG. 6, FIG. 6 is a structural diagram of a multi-person dialogue voice transcription device provided by the embodiment of the application. Since the multi-person dialogue voice transcription device solves problems in a similar principle to the multi-person dialogue voice transcription method in the embodiment of the application, the implementation of the multi-person dialogue voice transcription device can refer to the implementation of the method, and repeated details are not repeated.

[0190] As shown in FIG. 6, the multi-person dialogue voice transcription device 600 includes:

[0191] The first receiving module 601 is configured to receive an instant message in a multi-person dialogue scene sent by a service end, wherein the instant message contains a speaker identity, a first dialogue text, and a state attribute of the first dialogue text, the first dialogue text is a dialogue text corresponding to dialogue voice information received by the service end, and the state attribute is used to indicate the completeness of a sentence of dialogue text.

[0192] The updating module 602 is configured to, in a case where the speaker identity identifier corresponds to the dialogue text in the multi-person dialogue data displayed in the dialogue page view, update the dialogue text corresponding to the speaker identity identifier in the multi-person dialogue data according to a state attribute of a second dialogue text pre-stored, and update the dialogue page view by using the first dialogue text to update the dialogue text corresponding to the speaker identity identifier in the multi-person dialogue data, wherein the second dialogue text is the last dialogue text corresponding to the speaker identity identifier in the multi-person dialogue data before the first dialogue text is received.

[0193] Optionally, the first receiving module 601 further includes:

[0194] The loading module is configured to, in a case where the speaker identity identifier does not correspond to the dialogue text in the multi-person dialogue data, load the speaker identity identifier and the first dialogue text as a new dialogue data into the multi-person dialogue data.

[0195] Optionally, the updating module 602 includes:

[0196] The replacing unit is configured to, in a case where the state attribute of the second dialogue text is an intermediate state, replace the second dialogue text with the first dialogue text, and record a state attribute of the first dialogue text, wherein the first dialogue text contains the second dialogue text, and the intermediate state indicates that a sentence of the dialogue text is incomplete.

[0197] The loading unit is configured to, in a case where the state attribute of the second dialogue text is a final state, load the speaker identity identifier and the first dialogue text as a new dialogue data into the multi-person dialogue data, and record a state attribute of the first dialogue text, wherein the final state indicates that a sentence of the dialogue text is complete.

[0198] Optionally, the multi-person dialogue speech transcription apparatus 600 further includes:

[0199] The creating module is configured to create a speaker set, a new dialogue list and a final dialogue list, wherein the new dialogue list is used to store new dialogue data, and the final dialogue list is used to store the multi-person dialogue data.

[0200] The emptying module is configured to, in a case where the instant message of the server is not received for the first time, empty the new dialogue list.

[0201] The updating module 602 is configured to, in a case where the speaker set is not empty and the speaker identity identifier is included in the speaker set, update the dialogue text corresponding to the speaker identity identifier in the final dialogue list by using the first dialogue text according to a state attribute of a second dialogue text associated with the speaker identity identifier stored in the speaker set.

[0202] Optionally, the loading module is configured to, in a case that the speaker set is empty, or in a case that the speaker set is not empty and the speaker identity is not included in the speaker set, store the speaker identity and the first dialogue text as a new dialogue data in the new dialogue list, and store the speaker identity and a state attribute of the first dialogue text in the speaker set, and add the new dialogue list to the final dialogue list.

[0203] Optionally, the replacing unit is configured to, in a case that the state attribute of the second dialogue text is an intermediate state, replace the second dialogue text with the first dialogue text according to a timing update interval corresponding to the intermediate state.

[0204] The loading unit is configured to, in a case that the state attribute of the second dialogue text is a final state, load the speaker identity and the first dialogue text as a new dialogue data to the multi-person dialogue data according to a timing update interval corresponding to the final state.

[0205] Optionally, the multi-person dialogue voice transcription apparatus 600 further includes:

[0206] The deleting module is configured to, in a case that a size of the multi-person dialogue data exceeds a preset threshold, delete a first N dialogue data in the multi-person dialogue data according to a dialogue time sequence of the dialogue data in the multi-person dialogue data, N being a positive integer.

[0207] Optionally, the multi-person dialogue voice transcription apparatus 600 further includes:

[0208] The data loading module is configured to, in a case that an operation of loading historical dialogue data of the dialogue page view is received, acquire a preset number of historical dialogue data before a time point of dialogue data at a top of the dialogue page view from the server based on the time point, and load and display the preset number of historical dialogue data in the dialogue page view.

[0209] The multi-person dialogue voice transcription apparatus 600 provided by the embodiments of the present application can execute the method embodiment shown in FIG. 2, and the implementation principle and technical effects are similar, and the present embodiment will not be described here.

[0210] The multi-person conversation voice transcription device 600 according to the embodiments of the present application receives an instant message sent by a server in a multi-person conversation scenario, wherein the instant message comprises a speaker identity identifier, a first conversation text and a state attribute of the first conversation text, the first conversation text is a conversation text corresponding to the conversation voice information received by the server, and the state attribute is used to indicate the sentence integrity of the conversation text; in a case where the multi-person conversation data displayed in a conversation page view includes the conversation text corresponding to the speaker identity identifier, the first conversation text is used to update the conversation text corresponding to the speaker identity identifier in the multi-person conversation data according to a pre-stored state attribute of a second conversation text, so as to update the conversation page view, wherein the second conversation text is the last conversation text corresponding to the speaker identity identifier in the multi-person conversation data before the first conversation text is received. In this way, the state attribute of the multi-person conversation text is identified by the server, and the conversation text, the state attribute and the speaker identity identifier are sent to the client, the client updates the corresponding conversation data in the conversation page view based on the received state attribute of the conversation text and the speaker identity identifier, so that the multi-person conversation text can be displayed in a reasonable and orderly manner without repetition according to the speaker and the speaking state, and the effect of the multi-person conversation transcription text is improved.

[0211] The embodiments of the present application also provide another multi-person conversation voice transcription device. Referring to FIG. 7, FIG. 7 is a structural diagram of a multi-person conversation voice transcription device according to the embodiments of the present application. Since the multi-person conversation voice transcription device solves problems in the same principle as the multi-person conversation voice transcription method according to the embodiments of the present application, the implementation of the multi-person conversation voice transcription device can be referred to the implementation of the method, and the repeated parts will not be described herein.

[0212] As shown in FIG. 7, the multi-person conversation voice transcription device 700 comprises:

[0213] The second receiving module 701 is configured to receive the conversation voice information collected in the multi-person conversation scenario and sent by the first client;

[0214] The conversion module 702 is configured to convert the conversation voice information into an instant message, wherein the instant message comprises a speaker identity identifier, a first conversation text corresponding to the conversation voice information, and a state attribute of the first conversation text, and the state attribute is used to indicate the sentence integrity of the conversation text;

[0215] The sending module 703 is configured to send the instant message to the second client.

[0216] Optionally, the conversion module 702 comprises:

[0217] determining unit, configured to convert the dialogue voice information into first dialogue text, and determine a speaker identity corresponding to the dialogue voice information;

[0218] an analyzing unit, configured to analyze the dialogue voice information to determine a state attribute of the first dialogue text, wherein the state attribute comprises an intermediate state or a final state, the intermediate state indicating that a sentence of dialogue text is incomplete, and the final state indicating that a sentence of dialogue text is complete;

[0219] a generating unit, configured to generate an instant message containing the speaker identity, the first dialogue text, and the state attribute of the first dialogue text.

[0220] The multi-person dialogue voice transcription apparatus 700 provided by the embodiment of the present application can execute the method embodiment shown in FIG. 7, and has similar implementation principles and technical effects, which will not be repeated here.

[0221] The multi-person dialogue voice transcription apparatus 700 provided by the embodiment of the present application receives dialogue voice information collected in a multi-person dialogue scene sent by a first client; converts the dialogue voice information into an instant message, wherein the instant message contains a speaker identity, first dialogue text corresponding to the dialogue voice information, and a state attribute of the first dialogue text, the state attribute being used to indicate the completeness of a sentence of dialogue text; and sends the instant message to a second client. In this way, by identifying the state attribute of multi-person dialogue text on the server, and sending the dialogue text, the state attribute thereof, and the speaker identity to the client, the client updates the corresponding dialogue data in the dialogue page view based on the received state attribute of the dialogue text and the speaker identity, which can ensure that the multi-person dialogue text is displayed in a reasonable and orderly manner without repetition according to the speaker and the speaking state, and improves the effect of multi-person dialogue transcription text.

[0222] The embodiment of the present application further provides an electronic device. Since the principle of solving the problem by the electronic device is similar to the multi-person dialogue voice transcription method in the embodiment of the present application, the implementation of the electronic device can be referred to the implementation of the method, and the repeated parts will not be repeated. As shown in FIG. 8, the electronic device of the embodiment of the present application comprises a processor 800, a transceiver 810, and a memory 820.

[0223] In an implementation manner, the electronic device is a client, and the processor 800 is configured to read a program in the memory 820, and execute the following process:

[0224] receive, by the transceiver 810, an instant message sent by a server in a multi-person conversation scenario, wherein the instant message contains a speaker identity, a first conversation text and a state attribute of the first conversation text, the first conversation text is a conversation text corresponding to conversation voice information received by the server, and the state attribute is used to indicate the sentence integrity of the conversation text;

[0225] in a case where the multi-person conversation data displayed in the conversation page view contains the conversation text corresponding to the speaker identity, update the conversation text corresponding to the speaker identity in the multi-person conversation data by using the first conversation text according to a pre-stored state attribute of a second conversation text, to update the conversation page view, wherein the second conversation text is the last conversation text corresponding to the speaker identity existing in the multi-person conversation data before the first conversation text is received.

[0226] The transceiver 810 is configured to receive and send data under the control of the processor 800.

[0227] In FIG. 8, the bus architecture can include any number of interconnected buses and bridges, which are collectively represented by the bus interface. The bus architecture can also link various other circuits such as peripheral devices, voltage stabilizers, and power management circuits, which are well known in the art and thus will not be described further herein. The bus interface provides an interface. The transceiver 810 can be a plurality of elements, i.e., including a transmitter and a receiver, which provide units for communicating with various other devices on a transmission medium. The processor 800 is responsible for managing the bus architecture and general processing, and the memory 820 can store data used by the processor 800 when performing operations.

[0228] Optionally, the processor 800 is further configured to read a program in the memory 820 and perform the following steps:

[0229] in a case where the multi-person conversation data displayed in the conversation page view contains the conversation text corresponding to the speaker identity, update the conversation text corresponding to the speaker identity in the multi-person conversation data by using the first conversation text according to a pre-stored state attribute of a second conversation text, to update the conversation page view, wherein the second conversation text is the last conversation text corresponding to the speaker identity existing in the multi-person conversation data before the first conversation text is received.

[0230] Optionally, the processor 800 is further configured to read a program in the memory 820 and perform the following steps:

[0231] in a case where the state attribute of the second conversation text is an intermediate state, replace the second conversation text with the first conversation text, and record the state attribute of the first conversation text, wherein the first conversation text contains the second conversation text, and the intermediate state indicates that the sentence of the conversation text is incomplete.

[0232] In a case where the state attribute of the second dialogue text is a final state, loading the speaker identity and the first dialogue text as a piece of new dialogue data into the multi-person dialogue data, and recording the state attribute of the first dialogue text, wherein the final state indicates that a dialogue text is complete.

[0233] Optionally, the processor 800 is further configured to read programs in the memory 820 and perform the following steps:

[0234] creating a speaker set, a new dialogue list and a final dialogue list, wherein the new dialogue list is used to store new dialogue data, and the final dialogue list is used to store the multi-person dialogue data;

[0235] In a case where the instant message of the server is not received for the first time, emptying the new dialogue list;

[0236] In a case where the speaker set is not empty and the speaker identity is included in the speaker set, updating dialogue text corresponding to the speaker identity in the final dialogue list by using the first dialogue text according to the state attribute of the second dialogue text associated with the speaker identity stored in the speaker set.

[0237] Optionally, the processor 800 is further configured to read programs in the memory 820 and perform the following steps:

[0238] In a case where the speaker set is empty or the speaker set is not empty and the speaker identity is not included in the speaker set, storing the speaker identity and the first dialogue text as a piece of new dialogue data in the new dialogue list, and storing the state attribute of the speaker identity and the first dialogue text in the speaker set, and adding the new dialogue list to the final dialogue list.

[0239] Optionally, the processor 800 is further configured to read programs in the memory 820 and perform the following steps:

[0240] In a case where the state attribute of the second dialogue text is an intermediate state, replacing the second dialogue text with the first dialogue text according to a timing update interval corresponding to the intermediate state;

[0241] In a case where the state attribute of the second dialogue text is a final state, loading the speaker identity and the first dialogue text as a piece of new dialogue data into the multi-person dialogue data according to a timing update interval corresponding to the final state.

[0242] Optionally, the processor 800 is further configured to read a program in the memory 820 and perform the following steps:

[0243] In a case where the size of the current multi-person conversation data exceeds a preset threshold, the first N pieces of conversation data in the multi-person conversation data are deleted in chronological order of the conversation data in the multi-person conversation data, and N is a positive integer.

[0244] Optionally, the processor 800 is further configured to read a program in the memory 820 and perform the following steps:

[0245] In a case where the loading history conversation data operation for the conversation page view is received, a preset number of history conversation data before a time point of the conversation data at the top of the conversation page view is obtained from the server based on the time point, and the preset number of history conversation data is loaded and displayed in the conversation page view.

[0246] In another embodiment, the electronic device is a server, and the processor 800 is configured to read a program in the memory 820 and perform the following processes:

[0247] The transceiver 810 receives the conversation voice information collected in the multi-person conversation scene sent by the first client;

[0248] The conversation voice information is converted into an instant message, wherein the instant message contains a speaker identity, a first conversation text corresponding to the conversation voice information, and a state attribute of the first conversation text, and the state attribute is used to indicate the sentence integrity of the conversation text;

[0249] The transceiver 810 sends the instant message to the second client.

[0250] Optionally, the processor 800 is further configured to read a program in the memory 820 and perform the following steps:

[0251] The conversation voice information is converted into a first conversation text, and a speaker identity corresponding to the conversation voice information is determined;

[0252] The conversation voice information is analyzed to determine a state attribute of the first conversation text, wherein the state attribute includes an intermediate state or a final state, the intermediate state indicates that the sentence of the conversation text is incomplete, and the final state indicates that the sentence of the conversation text is complete;

[0253] An instant message containing the speaker identity, the first conversation text, and the state attribute of the first conversation text is generated.

[0254] The electronic device provided in the embodiments of the present application can execute the method embodiments shown in FIG. 2 or FIG. 5, and the implementation principles and technical effects are similar, and the embodiments will not be described here again.

[0255] In addition, the computer readable storage medium in the embodiments of the present application is used for storing a computer program, and the computer program can be executed by a processor to implement each step in the method embodiments shown in FIG. 2 or FIG. 5.

[0256] In several embodiments provided in the present application, it should be understood that the disclosed method and device can be implemented by other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0257] In addition, each functional unit in the embodiments of the present application can be integrated in a processing unit, or each unit can be physically included separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware, or in the form of hardware plus software functional unit.

[0258] The integrated unit realized in the form of software functional unit can be stored in a computer readable storage medium. The software functional unit stored in the storage medium includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute part of the steps of the transceiving method described in the embodiments of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0259] The above is the preferred embodiment of the present application, and it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

A multi-person conversation voice transcription method is executed by a client, and the method comprises: receiving instant messages sent by a server in a multi-person conversation scenario, wherein the instant messages contain a speaker identity, a first conversation text, and a state attribute of the first conversation text, the first conversation text is a conversation text corresponding to conversation voice information received by the server, and the state attribute is used to indicate the integrity of a sentence of the conversation text; in a case where the multi-person conversation data displayed in a conversation page view contains conversation text corresponding to the speaker identity, updating the conversation text corresponding to the speaker identity in the multi-person conversation data by using the first conversation text according to a pre-stored state attribute of a second conversation text, to update the conversation page view, wherein the second conversation text is the last conversation text corresponding to the speaker identity existing in the multi-person conversation data before the first conversation text is received. The method of claim 1, wherein, After the instant messages sent by the server in the multi-person conversation scenario are received, the method further comprises: in a case where the multi-person conversation data does not contain conversation text corresponding to the speaker identity, loading the speaker identity and the first conversation text as a piece of new conversation data to the multi-person conversation data. The method of claim 1, wherein, The updating of the conversation text corresponding to the speaker identity in the multi-person conversation data by using the first conversation text according to the pre-stored state attribute of the second conversation text comprises: in a case where the state attribute of the second conversation text is an intermediate state, replacing the second conversation text with the first conversation text, and recording the state attribute of the first conversation text, wherein the first conversation text contains the second conversation text, and the intermediate state indicates that the sentence of the conversation text is incomplete; in a case where the state attribute of the second conversation text is a final state, loading the speaker identity and the first conversation text as a piece of new conversation data to the multi-person conversation data, and recording the state attribute of the first conversation text, wherein the final state indicates that the sentence of the conversation text is complete. The method of claim 1, wherein, Before the instant messages sent by the server in the multi-person conversation scenario are received, the method further comprises: creating a speaker set, a new conversation list, and a final conversation list, wherein the new conversation list is used to store new conversation data, and the final conversation list is used to store the multi-person conversation data; After the instant messages sent by the server in the multi-person conversation scenario are received, before the updating of the conversation text corresponding to the speaker identity in the multi-person conversation data by using the first conversation text according to the pre-stored state attribute of the second conversation text, the method further comprises: in a case where the instant messages of the server are not received for the first time, emptying the new conversation list; In a case where the speaker identity identifier corresponds to the conversation text in the multi-person conversation data displayed in the conversation page view, the conversation text corresponding to the speaker identity identifier in the multi-person conversation data is updated according to a state attribute of the second conversation text stored in advance, including: In a case where the speaker set is not empty and the speaker identity identifier is included in the speaker set, the conversation text corresponding to the speaker identity identifier in the final conversation list is updated according to a state attribute of the second conversation text associated with the speaker identity identifier stored in the speaker set. The method of claim 2, wherein, Before receiving the instant message in the multi-person conversation scenario sent by the server, the method further includes: Creating a speaker set, an added conversation list, and a final conversation list, wherein the added conversation list is used to store added conversation data, and the final conversation list is used to store the multi-person conversation data; After receiving the instant message in the multi-person conversation scenario sent by the server, before loading the speaker identity identifier and the first conversation text as a piece of added conversation data into the multi-person conversation data, the method further includes: In a case where the instant message of the server is not received for the first time, the added conversation list is emptied; In a case where the speaker identity identifier does not correspond to the conversation text in the multi-person conversation data, the speaker identity identifier and the first conversation text are loaded as a piece of added conversation data into the multi-person conversation data, including: In a case where the speaker set is empty or in a case where the speaker set is not empty and the speaker identity identifier is not included in the speaker set, the speaker identity identifier and the first conversation text are stored as a piece of added conversation data in the added conversation list, and the speaker identity identifier and the state attribute of the first conversation text are associated and stored in the speaker set, and the added conversation list is added to the final conversation list. A multi-person conversation voice transcription method is executed by a server, and the method includes: Receiving conversation voice information collected in a multi-person conversation scenario sent by a first client; Converting the conversation voice information into an instant message, wherein the instant message includes a speaker identity identifier, a first conversation text corresponding to the conversation voice information, and a state attribute of the first conversation text, and the state attribute is used to indicate the completeness of the sentence of the conversation text; Sending the instant message to a second client. The method of claim 6, wherein, The conversion of the conversation voice information into an instant message includes: Converting the conversation voice information into a first conversation text and determining a speaker identity identifier corresponding to the conversation voice information; Analyzing the conversation voice information to determine a state attribute of the first conversation text, wherein the state attribute includes an intermediate state or a final state, the intermediate state indicates that the sentence of the conversation text is incomplete, and the final state indicates that the sentence of the conversation text is complete; generating an instant message containing the speaker identity, the first dialogue text, and a state attribute of the first dialogue text. An electronic device, comprising: a transceiver, a memory, a processor, and a computer program stored on the memory and executable on the processor; the processor is configured to read a program in the memory to implement the steps in the multi-person dialogue speech transcription method according to any one of claims 1 to 5; or implement the steps in the multi-person dialogue speech transcription method according to any one of claims 6 to 7. A computer readable storage medium for storing a computer program, the computer program being executable by a processor to implement the steps in the multi-person dialogue speech transcription method according to any one of claims 1 to 5; or implement the steps in the multi-person dialogue speech transcription method according to any one of claims 6 to 7. A computer program product comprising computer instructions, the computer instructions being executable by a processor to implement the steps in the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Conference sound acquisition method and device, conference record method and device and conference record presentation method and device

    CN111739553A

  • Conference record generation method and device, server and storage medium

    CN116013306A

  • Judicial scene multi-person dialogue recording method and system based on target voice separation

    CN116665678A

  • Audio processing method and device, computer equipment and storage medium

    CN117711401A

  • Multi-person dialogue voice transcription method, device, equipment, medium and program product

    CN118366456A