Conference data processing method, device and system, computer device and storage medium

By recording and automatically playing the participants' voice messages in the online meeting system, the problem of inefficiency caused by frequent switching of speaking permissions is solved, achieving efficient speech management and avoiding the chaos caused by multiple people speaking at the same time.

CN115344231BActive Publication Date: 2025-12-09INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210993681.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-18
Publication Date
2025-12-09
Estimated Expiration
2042-08-18

AI Technical Summary

Technical Problem

In existing technologies, the method of frequently switching the speaking permissions of participants in online meetings is inefficient, affects the efficiency of the meeting, and can easily lead to chaos in the meeting.

Method used

By working together with the client and server, the system records the participants' speech and stores it in a queue of audio messages to be played. The messages are then played automatically after the current speaker finishes speaking, eliminating the need for manual switching of speaking permissions.

Benefits of technology

This improved the efficiency of the meeting, ensuring that only one participant spoke at a time and avoiding chaos in the meeting room.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115344231B_ABST
    Figure CN115344231B_ABST
Patent Text Reader

Abstract

The application relates to a conference data processing method, device and system, computer equipment, a storage medium and a computer program product, relates to the technical field of artificial intelligence, and can be applied to the field of financial technology or other fields. The method comprises the following steps: in response to a speaking instruction, a first speaking request message is sent to a server; in response to a speaking preemption failure message sent by the server, speaking voice of a local user is recorded; in response to a recording end instruction, a second speaking request message containing recording information is sent to the server; the second speaking request message is used for instructing the server to store the recording information in a to-be-played voice queue, and in the case that current speaking voice information playing is ended, first target recording information to be played is acquired from the to-be-played voice queue, a first voice playing instruction is sent to a client; and in response to the first voice playing instruction, the first target recording information is played. By adopting the method, the speaker's voice can be automatically played, and the speaking efficiency of a conference can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a conference data processing method, device, system, computer device, storage medium and computer program product. BACKGROUND

[0002] With the popularization of multi-site collaborative office, online conference software appears. Participants can log in to the online conference software through their own terminal devices to participate in online conferences. In order to avoid on-site chaos caused by multiple people speaking at the same time during the conference, the conference speaking process can be managed by controlling the speaking rights of the speakers. In related technologies, the conference initiator is generally given the speaking control right, and other participants can request to speak to the conference initiator when they need to speak, and then the conference initiator manually switches the speaking rights of other participants, thereby managing the conference speaking process.

[0003] However, the speaking rights of participants often need to be frequently switched during the conference, and the method of manually switching the speaking rights to play the speaker's voice is low in efficiency, which affects the conference speaking efficiency. SUMMARY

[0004] Therefore, it is necessary to provide a conference data processing method, device, system, computer device, computer readable storage medium and computer program product capable of improving conference speaking efficiency to solve the above technical problems.

[0005] In a first aspect, the present application provides a conference data processing method. The method is applied to a client, and the method comprises:

[0006] In response to a speaking instruction, a first speaking request message is sent to the server;

[0007] In response to the speaking preemption failure message sent by the server, the speaking voice of the local user is recorded;

[0008] In response to a recording end instruction, recording information containing the speaking voice of the local user is obtained, and a second speaking request message containing the recording information is sent to the server; the second speaking request message is used to instruct the server to store the recording information in a to-be-played voice queue, and in the case that the current speaking voice information is played, to obtain first target recording information to be played from the to-be-played voice queue, and to send a first voice playing instruction containing the first target recording information to the client;

[0009] In response to the first voice playing instruction, the first target recording information contained in the first voice playing instruction is played.

[0010] In one of the embodiments, the method further comprises:

[0011] in response to the server sending the speech preemption success message, sending real-time voice information of the local user to the server; the real-time voice information is used for the server to send a second voice playing instruction containing the real-time voice information to other clients in the current conference.

[0012] In one of the embodiments, the method further comprises:

[0013] receiving recording summary information of each recording information in the to-be-played voice queue sent by the server, and displaying the recording summary information on the current interface;

[0014] in response to a playing instruction for second target recording information, sending a recording playing request message to the server; the recording playing request message is used to instruct the server to obtain the second target recording information from the to-be-played voice queue and send it to the client;

[0015] receiving the second target recording information sent by the server, and playing the second target recording information.

[0016] In one of the embodiments, the method further comprises:

[0017] determining a display box corresponding to a current speaker in the current interface;

[0018] displaying the display box of the current speaker according to a preset display mode.

[0019] In one of the embodiments, the method further comprises:

[0020] determining a display box corresponding to a current speaker in the current interface;

[0021] determining a target volume corresponding to the current speaker according to position information of the display box in the current interface and a preset volume correction strategy;

[0022] playing the current speech voice information according to the target volume.

[0023] In a second aspect, the application further provides a conference data processing method, which is applied to a server, and the method comprises:

[0024] receiving a first speech request message sent by a client, and determining a current state of the server; the current state comprises a waiting preemption state and a preoccupied state; the preoccupied state indicates that there is currently playing speech voice;

[0025] in a case where the current state of the server is a preempted state, sending a speech preemption failure message to the client; the speech preemption failure message is used for the client to acquire recording information of speech of the local user and send a second speech request message containing the recording information to the server;

[0026] in a case where the second speech request message sent by the client is received, storing the recording information in a to-be-played voice queue;

[0027] in a case where the current speech information is played, acquiring first target recording information to be played from the to-be-played voice queue and sending a first voice playing instruction containing the first target recording information to each client in the current conference; the first voice playing instruction is used to instruct the clients to play the first target recording information.

[0028] In one of the embodiments, the method further comprises:

[0029] in a case where the current state of the server is a waiting preemption state, sending a speech preemption success message to the client and switching the current state of the server to a preempted state; the speech preemption success message is used to instruct the client to send real-time voice information of the local user to the server;

[0030] receiving the real-time voice information sent by the client and sending a second voice playing instruction containing the real-time voice information to other clients in the current conference; the second voice playing instruction is used to instruct the other clients to play the real-time voice information.

[0031] In one of the embodiments, the acquiring, in a case where the current speech information is played, first target recording information to be played from the to-be-played voice queue and sending a first voice playing instruction containing the first target recording information to each client in the current conference comprises:

[0032] in a case where the current speech information is played, switching the current state of the server to a waiting preemption state;

[0033] if the first speech request message is not received within a preset time period, acquiring first target recording information to be played from the to-be-played voice queue, sending a first voice playing instruction containing the first target recording information to each client in the current conference and switching the current state of the server to a preempted state;

[0034] if the first speech request message is received within the preset time period, sending a speech preemption success message to the client sending the first speech request message and switching the current state of the server to a preempted state.

[0035] In one of the embodiments, the method further comprises:

[0036] In the case that the current speech information is real-time speech information, and the current speech information exceeds a preset time length and / or the number of speech pauses exceeds a preset number, obtaining second target recording information to be played from the queue of speech to be played, and sending a third speech playing instruction containing the second target recording information to each client in the current conference; the third speech playing instruction is used to instruct the clients to play the second target recording information.

[0037] In a third aspect, the present application further provides a conference data processing apparatus. The apparatus comprises:

[0038] A first sending module, configured to send a first speech request message to the server in response to a speech instruction;

[0039] A recording module, configured to record speech of a local user in response to a speech preemption failure message sent by the server;

[0040] A second sending module, configured to obtain recording information containing the speech of the local user in response to a recording end instruction, and send a second speech request message containing the recording information to the server; the second speech request message is used to instruct the server to store the recording information in a queue of speech to be played, and obtain first target recording information to be played from the queue of speech to be played in the case that current speech information playing ends, and send a first speech playing instruction containing the first target recording information to the client;

[0041] A first playing module, configured to play the first target recording information contained in the first speech playing instruction in response to the first speech playing instruction.

[0042] In one of the embodiments, the apparatus further comprises a third sending module, configured to send real-time speech information of the local user to the server in response to a speech preemption success message sent by the server; the real-time speech information is used for the server to send a second speech playing instruction containing the real-time speech information to other clients in the current conference.

[0043] In one of the embodiments, the apparatus further comprises:

[0044] A first display module, configured to receive recording summary information of each recording information in the queue of speech to be played sent by the server, and display the recording summary information on a current interface;

[0045] A fourth sending module is configured to, in response to a playing instruction for second target recording information, send a recording playing request message to the server, where the recording playing request message is used to instruct the server to obtain the second target recording information from the voice queue to be played and send the second target recording information to the client.

[0046] A second playing module is configured to receive the second target recording information sent by the server and play the second target recording information.

[0047] In one of the embodiments, the apparatus further includes:

[0048] A first determining module is configured to determine a display box corresponding to a current speaker in a current interface.

[0049] A second display module is configured to display the display box of the current speaker according to a preset display mode.

[0050] In one of the embodiments, the apparatus further includes:

[0051] A second determining module is configured to determine a display box corresponding to a current speaker in a current interface.

[0052] A third determining module is configured to determine a target volume corresponding to the current speaker according to position information of the display box in the current interface and a preset volume correction strategy.

[0053] A third playing module is configured to play the current speech information according to the target volume.

[0054] In a fourth aspect, the application further provides a conference data processing apparatus, which includes:

[0055] A determining module is configured to receive a first speech request message sent by a client and determine a current state of the server, where the current state includes a waiting preemption state and a preoccupied state, and the preoccupied state indicates that there is currently playing speech.

[0056] A first sending module is configured to, in a case where the current state of the server is the preoccupied state, send a speech preemption failure message to the client, where the speech preemption failure message is used for the client to obtain recording information containing speech of a local user and send a second speech request message containing the recording information to the server.

[0057] A storage module is configured to, in a case where the second speech request message sent by the client is received, store the recording information in a voice queue to be played.

[0058] The second sending module is configured to, in a case where the current speech information is played, acquire first target recording information to be played from the queue of speech information to be played, and send a first speech playing instruction containing the first target recording information to each client in the current conference; the first speech playing instruction is configured to instruct the clients to play the first target recording information.

[0059] In one of the embodiments, the apparatus further includes:

[0060] The third sending module is configured to, in a case where the current state of the server is a waiting preemption state, send a speech preemption success message to the client, and switch the current state of the server to a preoccupied state; the speech preemption success message is configured to instruct the client to send real-time speech information of the local user to the server.

[0061] The fourth sending module is configured to receive the real-time speech information sent by the client, and send a second speech playing instruction containing the real-time speech information to other clients in the current conference; the second speech playing instruction is configured to instruct the other clients to play the real-time speech information.

[0062] In one of the embodiments, the second sending module is specifically configured to, in a case where the current speech information is played, switch the current state of the server to a waiting preemption state; if no first speech request message is received within a preset time period, acquire first target recording information to be played from the queue of speech information to be played, send a first speech playing instruction containing the first target recording information to each client in the current conference, and switch the current state of the server to a preoccupied state; if a first speech request message is received within the preset time period, send a speech preemption success message to the client sending the first speech request message, and switch the current state of the server to a preoccupied state.

[0063] In one of the embodiments, the apparatus further includes:

[0064] The fifth sending module is configured to, in a case where the current speech information is real-time speech information, and the current speech information exceeds a preset time length and / or a number of speech pauses exceeds a preset number, acquire second target recording information to be played from the queue of speech information to be played, and send a third speech playing instruction containing the second target recording information to each client in the current conference; the third speech playing instruction is configured to instruct the clients to play the second target recording information.

[0065] In a fifth aspect, the application further provides a conference data processing system, including a plurality of clients and a server, wherein:

[0066] The client is configured to send a first speech request message to the server in response to a speech instruction;

[0067] The server is configured to receive the first speech request message sent by the client, and determine a current state of the server; the current state includes a waiting preemption state and a preempted state; the preempted state indicates that there is currently a playing speech voice; and the server is configured to send a speech preemption failure message to the client in a case where the current state of the server is the preempted state;

[0068] The client is further configured to record a speech voice of a local user in response to the speech preemption failure message sent by the server, and obtain recording information containing the speech voice of the local user in response to an end-of-recording instruction, and send a second speech request message containing the recording information to the server;

[0069] The server is further configured to store the recording information in a to-be-played voice queue in a case where the second speech request message sent by the client is received, and obtain first target recording information to be played from the to-be-played voice queue in a case where a current speech voice information is played to an end, and send a first speech playing instruction containing the first target recording information to each client in the current conference;

[0070] The client is further configured to play the first target recording information contained in the first speech playing instruction in response to the first speech playing instruction.

[0071] In a sixth aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method in the first aspect or the second aspect when executing the computer program.

[0072] In a seventh aspect, the present application further provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program implements the steps of the method in the first aspect or the second aspect when executed by a processor.

[0073] In an eighth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and the computer program implements the steps of the method in the first aspect or the second aspect when executed by a processor.

[0074] The conference data processing method, device, system, computer device, storage medium and computer program product, when a participant (corresponding to a user) wants to speak, a speaking instruction can be triggered, a real-time speaking request (first speaking request message) is sent to the server through the client, if the current conference is playing the speaking voice of other users, the client will receive the speaking preemption failure message sent by the server, the client can record the speaking voice of the local user to obtain recording information, and then send the second speaking request message containing the recording information to the server, so that the server stores the recording information in the to-be-played voice queue. In the case that the current speaking voice playing is ended, the server can obtain target recording information (first target recording information, such as the first recording information in the queue) from the to-be-played voice queue, and construct a voice playing instruction (first voice playing instruction) containing the target recording information and send it to each client in the current conference, so that each client plays the target recording information. In the method, each participant can trigger a speaking instruction to apply for speaking, if the speaking voice of other participants is being played, the speaking voice of the participant applying for speaking can be recorded and stored in the to-be-played voice queue for queuing and playing, so that the speaker voice can be played automatically. Compared with manually switching the speaker permission, the method can improve the conference speaking efficiency. At the same time, the method can control the speaking voice of only one participant to be played at the same time, avoiding the chaos caused by multiple people speaking at the same time. BRIEF DESCRIPTION OF DRAWINGS

[0075] Figure 1 A structural schematic diagram of a conference data processing system in an example;

[0076] Figure 2 A flowchart of a conference data processing method in an example;

[0077] Figure 3 A flowchart of a conference data processing method in another example;

[0078] Figure 4 A flowchart of a conference data processing method in another example;

[0079] Figure 5 A flowchart of a conference data processing method in another example;

[0080] Figure 6 A flowchart of a conference data processing method in another example;

[0081] Figure 7 A flowchart of sending a first voice playing instruction in an example;

[0082] Figure 8 A structural block diagram of a conference data processing device in an example;

[0083] Figure 9 is a structural block diagram of a conference data processing apparatus in another embodiment;

[0084] Figure 10 is an internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0085] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0086] First, before specifically introducing the technical solutions of the embodiments of the present application, the technical background or technical evolution context based on which the embodiments of the present application are introduced. With the popularization of multi-site collaborative office, online conference software appears. Participants can log in to the online conference software through their own terminal devices to participate in online conferences. In order to avoid the scene chaos caused by multiple people speaking at the same time during the conference, the conference speaking process can be managed by controlling the speaking rights of the speakers. In related technologies, the conference initiator is generally given the speaking control right, and other participants need to request speaking from the conference initiator when they need to speak, and then the conference initiator manually switches the speaking rights of other participants, thereby managing the conference speaking process. However, the speaking rights of the participants often need to be frequently switched during the conference, and the method of manually switching the speaking rights is low in efficiency, which affects the conference speaking efficiency. Based on this background, the applicant proposes the conference data processing method of the present application through long-term research and development and experimental verification, which can automatically play the speaking voice of each participant, avoid manual switching, improve the conference speaking efficiency, and can control the speaking voice of only one participant at the same time to avoid the scene chaos of the conference. In addition, it should be noted that the applicant has made a lot of creative labor for the discovery of the technical problem of the present application and the technical solutions introduced in the following embodiments.

[0087] The conference data processing method provided by the embodiments of the present application can be applied to the conference data processing system 100 as shown in Figure 1 The conference data processing system 100 includes a plurality of clients 102 and a server 104, and each client 102 communicates with the server 104 through a network. The client 102 can be a conference software, an application program, etc. installed on a terminal, and the terminal can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0088] In an embodiment, as shown in Figure 2As shown, a conference data processing method is provided, and the method is applied to Figure 1 The client 102 in the conference system 100 is taken as an example for illustration, and the method comprises the following steps:

[0089] In step 201, a first speech request message is sent to the server in response to a speech instruction.

[0090] In implementation, a user can log in and participate in an online conference through a client. After entering the conference, the user can perform an operation triggering a speech instruction on the current interface. For example, the current interface of the client can display a microphone icon, and the user can click the microphone icon to trigger the speech instruction. The client can then send a first speech request message to the server in response to the speech instruction. The first speech request message is used to request the local user of the client to make a real-time speech (i.e., to request to grant the local user a speech right to play the real-time speech of the user in the current conference).

[0091] In step 202, the speech of the local user is recorded in response to a speech preemption failure message sent by the server.

[0092] In implementation, after receiving the first speech request message sent by the client, the server can determine whether the local user of the client can make a real-time speech at present. For example, if the speech of the user of another client is being played in the conference, i.e., the current state of the server is a preoccupied state, the local user of the client cannot make a real-time speech, and the server can return a speech preemption failure message to the client (the client sending the first speech request message). After receiving the speech preemption failure message, the client can start a recording function to record the speech of the local user.

[0093] It can be understood that, after receiving the speech preemption failure message, the client can directly start the recording function, or can display a prompt information (such as a prompt information “failed to grab the microphone, do you want to record and wait for playing” and a selection control to let the user select whether to record.

[0094] In step 203, recording information containing the speech of the local user is obtained in response to a recording end instruction, and a second speech request message containing the recording information is sent to the server.

[0095] In implementation, during recording the speech voice of the local user, the user can trigger an end recording instruction. For example, when the client records the speech voice of the local user, a recording interface and an end recording control can be displayed, and the recording interface can also display information such as recording duration. The user can click the end recording control to trigger the end recording instruction. It can be understood that a recording duration limit condition (such as limiting the maximum recording time to 5 minutes) can be set in advance, which can be set by the conference initiator or set by default by the conference software. If the recorded speech voice exceeds the limited recording duration, the end recording instruction can be triggered automatically.

[0096] The client can obtain recording information containing the speech voice of the local user in response to the end recording instruction, that is, recording information obtained by recording the speech voice of the local user. Then, the client can construct a second speech request message containing the recording information and send it to the server. The second speech request message is used to instruct the server to store the recording information in the second speech request message in the to-be-played voice queue. In the case that the current speech voice information is played to the end, the server can obtain the first target recording information to be played (the first recording information in the to-be-played voice queue can be taken as the first target recording information) from the to-be-played voice queue. Then, the server can construct a first voice playing instruction according to the first target recording information and send the first voice playing instruction to all clients in the current conference.

[0097] It can be understood that the server can confirm the clients in the current conference through the heartbeat information sent by the clients. For example, the client participating in the current conference can send a heartbeat packet to the server at a regular time, and the server can determine the client in the heartbeat packet period as the client in the current conference (that is, the client receives the heartbeat packet sent by the client in the heartbeat packet period). Moreover, the server can feed back a heartbeat response packet for each heartbeat packet, and if the client does not receive the heartbeat response packet in the heartbeat packet period, it can be considered that the client has been disconnected.

[0098] Step 204, in response to the first voice playing instruction, playing the first target recording information contained in the first voice playing instruction.

[0099] In implementation, after the client receives the first voice playing instruction sent by the server, the first target recording information contained in the first voice playing instruction can be played. During playing the first target recording information, the first target recording information is the current speech voice information.

[0100] In the conference data processing method, each participant (corresponding to a user) can trigger a speech instruction to apply for speech, if speech voice of other participant is currently being played, speech voice of the participant applying for speech can be recorded to obtain recording information, and the recording information is stored in a to-be-played voice queue for queuing and playing. After the current speech voice playing ends, target recording information can be obtained from the to-be-played voice queue for playing, so that automatic playing of speech voice of a speaker can be realized, manual switching of speaker permission for playing is avoided, and conference speech efficiency is improved. Meanwhile, the method can control speech voice of only one participant to be played at the same time, so that conference site chaos caused by simultaneous speech of multiple participants is avoided.

[0101] In one embodiment, the method further includes the steps of: in response to the speech preemption success message sent by the server, sending real-time voice information of the local user to the server; and the real-time voice information is used for the server to send a second speech playing instruction containing the real-time voice information to other clients in the current conference.

[0102] In implementation, after the server receives the first speech request message sent by the client, if it is judged that there is no speech voice being played (i.e., the current state of the server is a waiting preemption state), the server can send a speech preemption success message to the client. After the client receives the speech preemption success message sent by the server, the client can send real-time voice information of the local user to the server. The real-time voice information is used for the server to send a second speech playing instruction containing the real-time voice information to other clients in the current conference, so that the other clients play the real-time voice information according to the second speech playing instruction. That is, the local user of the client that has successfully preoccupied the microphone (receives the speech preemption success message) can make real-time speech, and the real-time speech voice of the user is played by other clients as current speech voice. It can be understood that, if a client has successfully preoccupied the microphone, for other clients that send the first speech request message after the client, since there is current speech voice being played (real-time voice information of the client that has successfully preoccupied the microphone) at this time, the other clients will receive a speech preemption failure message sent by the server, i.e., the process described in step 202 can be performed.

[0103] In the embodiment, if there is no speech voice being played at present, such as when a conference just starts or during a gap period after playing of the last speech voice ends, when a participant triggers a speech instruction through a client, the participant can successfully preoccupy the microphone and make real-time speech, so that automatic playing of speech voice of a speaker is realized and conference speech efficiency is improved. The clients that fail to preoccupy the microphone can record speech voice of the local user for queuing and playing, so that speech voice of only one participant is played at the same time, conference speech order is maintained, and conference site chaos is avoided.

[0104] In one embodiment, as Figure 3As shown, the method further includes the following steps:

[0105] In step 301, the client receives the recording summary information of each recording information in the voice queue to be played sent by the server, and displays the recording summary information on the current interface.

[0106] In implementation, the client can receive the recording summary information of each recording information in the voice queue to be played sent by the server. For example, if the recording information in the voice queue to be played is updated (new recording information is added or old recording information is deleted), the server can send the summary information of each recording information (such as recording user identity information, recording duration, recording time, recording queue order, etc.) to each client in the current conference, and send the recording summary information to each client. After receiving the recording summary information, the client can display the recording summary information on the current interface, such as displaying the recording summary information in the recording queue display area of the current interface, and can display the list according to the queue order.

[0107] In step 302, in response to the playing instruction for the second target recording information, a recording playing request message is sent to the server.

[0108] The recording playing request message is used to instruct the server to obtain the second target recording information from the voice queue to be played and send it to the client.

[0109] In implementation, the user can trigger the playing instruction operation for the second target recording information on the current interface. For example, the user can click a piece of recording summary information in the displayed recording summary information, that is, trigger the playing instruction for the recording information. Or the current interface displays a playing control corresponding to each recording summary information, and the user can click the playing control, that is, trigger the playing instruction for the recording information. The client can determine the recording summary information as the recording summary information of the second target recording information, and send the recording playing request message for the second target recording information to the server. The recording playing request message is used to instruct the server to obtain the detailed information of the second target recording information from the voice queue to be played and send it to the client (the client sending the recording playing request message).

[0110] In step 303, the second target recording information sent by the server is received, and the second target recording information is played.

[0111] In implementation, after receiving the second target recording information sent by the server, the client can play the second target recording information. It can be understood that if the current conference is playing the current speech voice, the client (the client receiving the second target recording information) can simultaneously play the current speech voice and the second target recording information. The client can adjust the playing volume of the current speech voice according to a preset volume adjustment strategy, and play the second target recording information according to a preset volume. For example, the client can reduce the left and right channel volumes of the current speech voice to 60% of the initial volume, and the initial volume refers to the playing volume of the current speech voice before playing the second target recording information. After the second target recording information is played, the playing volume of the current speech voice is restored to the initial volume. The preset volume for playing the second target recording information can be 10% volume for the left channel and 100% volume for the right channel.

[0112] In some other examples, during the playing of the second target recording information, the client can display a playing progress bar, and the user can drag the progress bar to select a playing time point.

[0113] In some other examples, the user can also trigger a cancel operation on the recording information of the user himself in the recording summary information displayed in the current interface, and the client can send a cancel instruction on the recording information to the server, so that the server deletes the recording information from the voice queue to be played and no longer queues and plays.

[0114] In the embodiment, the client can display the summary information of each queued recording information in the voice queue to be played, so as to facilitate the user to check the queued recording situation, and the user can select and play the recording information according to needs.

[0115] In one embodiment, the method further includes the following steps: determining the display box corresponding to the current speaker in the current interface; and displaying the display box of the current speaker according to a preset display mode.

[0116] In implementation, the current interface of the client can display the display boxes of each user (generally one client corresponds to one user) in the current conference. For example, the video image display box or the avatar display box of the user can be displayed, and one user corresponds to one display box. The client can determine the display box corresponding to the current speaker (i.e. the user corresponding to the current speech voice information, which can be the user speaking in real time or the user recording) in the current interface.

[0117] The display boxes of each user can be arranged and displayed according to a preset layout mode. For example, a square matrix layout mode of n users per row and n users per column (i.e. n*n) can be used, and the value of n is determined according to the number of current conference users. If the number of current users is x, n satisfies (n-1) 2 <x<n2 Meanwhile, an upper limit of the number of persons displayed in each row and each column is set, i.e., the maximum value N of n is set max If the number of conference users exceeds N max , the display is paginated. The size of each user display box is the current interface pixel size divided by n, for example, if the resolution of the total area of the user display box of the current interface is 800*600, and there are 22 participants, n is 5 (4 2 <22<5 2 ), and the resolution of the display box area of each user is 160*120. If the resolution ratio of the video image recorded by the user is not 4:3, the image is cropped.

[0118] After the client determines the display box of the current speaker, the display box of the current speaker can be displayed according to a preset display mode. For example, the display box of the current speaker can be displayed at a proportion of 100%, and the display boxes of other users can be displayed at a proportion of 80%. If the resolution of the display box area of each user is 160*120 in the aforementioned example, the proportion of 100% is 160*120, and the proportion of 80% is 144*96, and the display can be centered (the other 20% of the non-display box area in the display area can be filled with black).

[0119] In this embodiment, the display box of the current speaker can be distinguished from the display boxes of other users in the current interface of the client, so that the user can quickly locate the current speaker according to the display effect.

[0120] In one embodiment, as shown in Figure 4 , the method further includes the following steps:

[0121] Step 401: determining the display box corresponding to the current speaker in the current interface.

[0122] In implementation, the current interface of the client can display the display boxes of users in the current conference. The client can determine the display box corresponding to the current speaker in the current interface.

[0123] Step 402: determining the target volume corresponding to the current speaker according to the position information of the display box in the current interface and a preset volume correction strategy.

[0124] In implementation, the client can determine the target volume corresponding to the current speaker according to the position information of the display box corresponding to the current speaker and a preset volume correction strategy. For example, the display boxes of users can be arranged in an n*n square matrix, and the number of rows in which the display box is located in the square matrix can be taken as the horizontal coordinate (denoted as N i , i.e., the i-th row in the square matrix), and the number of columns can be taken as the vertical coordinate (denoted as N jThe position information of the display frame is the position information of the display frame in the jth row of the matrix.

[0125] The preset volume correction strategy can be to correct the volume to simulate the position of the speaker by voice. Specifically, the target volume can be calculated by using the following formula:

[0126]

[0127]

[0128]

[0129] N max is the maximum value of n of the preset n*n layout, N now is the value of n of the actual n*n layout of the current interface, L def is the preset maximum volume correction value, L now is the current volume correction unit value, W def is the preset default volume size, N i is the horizontal coordinate of the current speaker, N j is the vertical coordinate of the current speaker (if the display frame of the current speaker is located in the second row and the third column, N i = 3, N j = 2), W left is the corrected left channel output volume size, W right is the corrected right channel output volume size. The client can determine W left and W right as the target volume.

[0130] Step 403: playing the current speech voice information according to the target volume.

[0131] In implementation, the client can play the current speech voice information according to the target volume determined in step 402.

[0132] In this embodiment, the target volume of the current speech voice, including the left channel volume and the right channel volume, can be corrected by the position information of the display frame of the current speaker, so as to facilitate the voice simulation positioning of the current speaker. The user can locate the approximate position of the display frame of the current speaker according to the left and right channel volumes of the current speech voice, and then quickly locate the current speaker. In addition, the user can also notice the change of the speaker according to the change of the target volume, so as to improve the attention of the user participating in the conference.

[0133] In one embodiment, as Figure 5 shown, a conference data processing method applied to a server is also provided, and the method includes the following steps:

[0134] In step 501, a first speech request message sent by a client is received, and a current state of the server is determined.

[0135] In implementation, the server can determine the current state of the server upon receiving the first speech request message sent by the client. The current state includes a waiting preemption state and a preempted state. The preempted state indicates that there is currently a playing speech voice (including real-time speech voice and recording).

[0136] In step 502, a speech preemption failure message is sent to the client when the current state of the server is the preempted state.

[0137] In implementation, the server can send a speech preemption failure message to the client when the current state of the server is the preempted state, i.e., there is currently a playing speech voice. The speech preemption failure message indicates that the client fails to preoccupy the microphone. After receiving the preemption failure message, the client can obtain recording information containing the speech voice of the local user, and send a second speech request message containing the recording information to the server. The specific execution process of the client can be referred to the description in the foregoing embodiments, which will not be described here.

[0138] In step 503, the recording information is stored in a to-be-played voice queue when the second speech request message sent by the client is received.

[0139] In implementation, the server can store the recording information contained in the second speech request message in the to-be-played voice queue when the second speech request message sent by the client is received. The to-be-played voice queue is used to store the recording information sent by each client in the current conference, so as to queue and play.

[0140] In step 504, first target recording information to be played is obtained from the to-be-played voice queue when the current speech voice information playing ends, and a first voice playing instruction containing the first target recording information is sent to each client in the current conference.

[0141] In implementation, after the current speech voice information playing ends, the server can obtain the first target recording information to be played from the to-be-played voice queue (the first recording information in the to-be-played voice queue can be taken as the first target recording information), then construct a first voice playing instruction according to the first target recording information, and send the first voice playing instruction to all clients in the current conference. The server can determine each client in the current conference according to the heartbeat information sent by the client. The first voice playing instruction is used to instruct each client to play the first target recording information, i.e., to play the first target recording information as the current speech voice information.

[0142] In this embodiment, when the server receives the speech request sent by the client, the server can determine whether the client can speak in real time according to the current state. If the current state of the server is the preempted state, that is, there is currently a speech voice being played, the server sends a speech preemption failure message to the client, and then the client can record the speech voice of the local user and send the recorded voice information to the server to store in the to-be-played voice queue. When the current speech voice is played, the server can obtain the target voice information from the to-be-played voice queue and send the target voice information to each client for playing. In this way, the speaker voice can be automatically played, the manual switching of the speaker permission for playing can be avoided, the conference speech efficiency can be improved, and the speech voice of only one user at the same time can be played to avoid the conference site being in chaos due to multiple users speaking at the same time.

[0143] In one embodiment, as shown in FIG. 6, the method further includes the following steps: Figure 6

[0144] Step 601: In a case where the current state of the server is the waiting preemption state, a speech preemption success message is sent to the client, and the current state of the server is switched to the preempted state.

[0145] In implementation, after the server receives the first speech request message sent by the client, if it is determined that the current state of the server is the waiting preemption state, for example, when the conference just starts or the last speech voice has been played, there is no speech voice being played at this time, and the state of the server can be identified as the waiting preemption state, the server can send a speech preemption success message to the client (the client that sends the first speech request message when the state of the server is the waiting preemption state), that is, the client successfully preempts the microphone, and the local user of the client can speak in real time. After the client receives the speech preemption success message, the real-time voice information of the local user can be sent to the server.

[0146] Step 602: Real-time voice information sent by the client is received, and a second voice playing instruction containing the real-time voice information is sent to other clients in the current conference.

[0147] In implementation, after the server receives the real-time voice information sent by the client (that is, the client that successfully preempts the microphone), the server can send a second voice playing instruction containing the real-time voice information to other clients in the current conference (that is, other clients in the current conference except the client that successfully preempts the microphone). After the other clients receive the second voice playing instruction, the real-time voice information contained in the second voice playing instruction can be played. At this time, the real-time voice information is the current speech voice information, or the speech voice information being played at present.

[0148] ​In the embodiment, if the current state of the server is the waiting preemption state, the client sending the first speech request message can succeed in the microphone preemption, the local user of the client can make real-time speech, thereby realizing automatic playing of the speaker voice and improving the conference speech efficiency. The clients failing in the microphone preemption can record the speech voice of the local user and play the recorded voice in sequence, thereby controlling the playing of the speech voice of only one person at the same time, maintaining the order of the conference speech, and avoiding the confusion of the conference site.

[0149] In one embodiment, as shown in FIG. 5, the process of sending the first voice playing instruction in step 504 specifically includes the following steps. Figure 7

[0150] Step 701, in the case that the current speech voice information playing is ended, switching the current state of the server to the waiting preemption state.

[0151] In the implementation, the server can switch the current state of the server from the preemption state to the waiting preemption state after the current speech voice information playing is ended.

[0152] Step 702, if the first speech request message is not received within a preset time period, obtaining the first target recording information to be played from the voice queue to be played, sending the first voice playing instruction containing the first target recording information to each client in the current conference, and switching the current state of the server to the preemption state.

[0153] In the implementation, after the server switches the current state to the waiting preemption state, if the first speech request message sent by each client in the current conference is not received within a preset time period (which can be set according to the situation, for example, within 1 minute after the switching to the waiting preemption state), the first target recording information to be played is obtained from the voice queue to be played, then the first voice playing instruction containing the first target recording information is sent to all online clients in the current conference, and the current state of the server is switched to the preemption state.

[0154] Step 703, if the first speech request message is received within a preset time period, sending a speech preemption success message to the client sending the first speech request message, and switching the current state of the server to the preemption state.

[0155] ​In the implementation, after the server switches the current state to the waiting preemption state, if a first speech request message sent by any client in the current conference is received within a preset time period, the server can send a speech preemption success message to the client, and switch the current state of the server to a preoccupied state. The speech preemption success message is used to instruct the client to obtain real-time voice information of the local user, and send a real-time speech message containing the real-time voice information to the server. It can be understood that, for other clients that send the first speech request message after the client, since the server has switched the state to the preoccupied state at this time, the server will return a speech preemption failure message to the other clients that send the speech request. Then the server can receive the real-time speech message sent by the client, and send a second speech playing instruction containing the real-time voice information to other clients in the current conference. The specific execution process of the client is described in the foregoing embodiments, and will not be described here.

[0156] In the embodiment, after the current speech voice information playing is ended, the server can switch the current state to the waiting preemption state. If a client preoccupies the microphone (sends the first speech request message) within a preset time period, the client can be allowed to successfully preoccupy the microphone and make real-time speech. If no client preoccupies the microphone within the preset time period, target recording information is obtained from the voice queue to be played, and is played as the current speech voice information. That is, by setting the preoccupation time length (the preset time period after the state is switched to the waiting preemption state), a user who needs to make speech before playing the recording can make real-time speech, and the flexibility of conference speech is improved.

[0157] In one embodiment, the method further includes the following steps: in the case that the current speech voice information is real-time voice information, and the current speech voice information exceeds a preset time length and / or the number of speech pauses exceeds a preset number, obtaining second target recording information to be played from the voice queue to be played, and sending a third speech playing instruction containing the second target recording information to each client in the current conference; the third speech playing instruction is used to instruct the client to play the second target recording information.

[0158] In implementation, if the current speech information is real-time speech information, the server can determine the duration of the current speech information and the number of speech pauses. For example, the server can convert the current speech information into text by using a relevant algorithm (e.g., by using a text-independent voiceprint system), and if there is no text in the text corresponding to the current speech information within a preset duration (e.g., within 1 second), the number of speech pauses can be increased by 1. If the duration of the current speech information exceeds a preset duration (e.g., 5 minutes) and / or the number of speech pauses exceeds a preset number (e.g., 3 times), the server can obtain second target recording information (e.g., the first recording information in the current queue) to be played from the queue of speech to be played, construct a third speech playing instruction containing the second target recording information, and send the third speech playing instruction to each online client in the current conference to instruct the clients to play the second target recording information. At this time, the server will no longer send a playing instruction containing the aforementioned real-time speech information, or the server can send an instruction to the clients to stop playing the aforementioned real-time speech information, and at the same time, send the third speech playing instruction containing the second target recording information, so that the clients stop playing the aforementioned real-time speech information and play the second target recording information instead.

[0159] In the embodiment, if the current speech information is real-time speech information and the current speech information exceeds a preset duration and / or the number of speech pauses exceeds a preset number, the server can instruct the clients to play the recording information, thereby stopping the playing of the real-time speech and realizing the function of automatically cutting off the microphone to automatically maintain the order of conference speech.

[0160] It should be understood that although each step in the flowchart involved in each embodiment described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment described above can include multiple steps or stages, which are not necessarily executed at the same time but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential but can be executed in rotation or alternation with at least part of other steps or stages in other steps.

[0161] Based on the same inventive concept, the embodiments of the present application further provide a conference data processing system for implementing the conference data processing method as described above. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme described in the above method, so the specific limitations in one or more conference data processing system embodiments provided below can be referred to the limitations of the conference data processing method in the above, which will not be repeated here.

[0162] In one embodiment, a conference data processing system is provided, comprising a plurality of clients and a server, wherein:

[0163] The client is configured to send a first speech request message to the server in response to a speech instruction.

[0164] The server is configured to receive the first speech request message sent by the client, determine the current state of the server, and send a speech preemption failure message to the client when the current state of the server is a preoccupied state.

[0165] The client is further configured to record the speech of the local user in response to the speech preemption failure message sent by the server, obtain the recording information containing the speech of the local user in response to an end of recording instruction, and send a second speech request message containing the recording information to the server.

[0166] The server is further configured to store the recording information in a to-be-played voice queue when the second speech request message sent by the client is received, obtain first target recording information to be played from the to-be-played voice queue when the current speech information is played, and send a first speech playing instruction containing the first target recording information to each client in the current conference.

[0167] The client is further configured to play the first target recording information contained in the first speech playing instruction in response to the first speech playing instruction.

[0168] Based on the same inventive concept, the embodiments of the present application further provide a conference data processing system for implementing the conference data processing method as described above. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme described in the above method, so the specific limitations in one or more conference data processing system embodiments provided below can be referred to the limitations of the conference data processing method in the above, which will not be repeated here.

[0169] In one embodiment, as Figure 8As shown, a conference data processing apparatus 800 is provided, comprising a first sending module 801, a recording module 802, a second sending module 803 and a first playing module 804, wherein:

[0170] The first sending module 801 is configured to send a first speech request message to a server in response to a speech instruction.

[0171] The recording module 802 is configured to record speech of a local user in response to a speech preemption failure message sent by the server.

[0172] The second sending module 803 is configured to acquire recording information containing the speech of the local user and send a second speech request message containing the recording information to the server in response to a recording end instruction; the second speech request message is configured to instruct the server to store the recording information in a to-be-played speech queue, and acquire first target recording information to be played from the to-be-played speech queue and send a first speech playing instruction containing the first target recording information to the client in the case that the current speech information playing ends.

[0173] The first playing module 804 is configured to play the first target recording information contained in the first speech playing instruction in response to the first speech playing instruction.

[0174] In one of the embodiments, the apparatus further comprises a third sending module configured to send real-time speech information of the local user to the server in response to a speech preemption success message sent by the server; the real-time speech information is configured to instruct the server to send a second speech playing instruction containing the real-time speech information to other clients in the current conference.

[0175] In one of the embodiments, the apparatus further comprises a first display module, a fourth sending module and a second playing module, wherein:

[0176] The first display module is configured to receive recording abstract information of each recording information in the to-be-played speech queue sent by the server, and display the recording abstract information on a current interface.

[0177] The fourth sending module is configured to send a recording playing request message to the server in response to a playing instruction for the second target recording information; the recording playing request message is configured to instruct the server to acquire the second target recording information from the to-be-played speech queue and send to the client.

[0178] The second playing module is configured to receive the second target recording information sent by the server and play the second target recording information.

[0179] In one of the embodiments, the apparatus further comprises a first determining module and a second display module, wherein:

[0180] The first determining module is configured to determine a display box corresponding to the current speaker in the current interface.

[0181] The second displaying module is configured to display the display box of the current speaker according to a preset display mode.

[0182] In one of the embodiments, the apparatus further includes a second determining module, a third determining module and a third playing module, wherein:

[0183] The second determining module is configured to determine a display box corresponding to the current speaker in the current interface.

[0184] The third determining module is configured to determine a target volume corresponding to the current speaker according to the position information of the display box in the current interface and a preset volume correction strategy.

[0185] The third playing module is configured to play the current speech information according to the target volume.

[0186] In one of the embodiments, as shown in Figure 9 The apparatus 900 further includes a determining module 901, a first sending module 902, a storage module 903 and a second sending module 904, wherein:

[0187] The determining module 901 is configured to receive a first speech request message sent by a client and determine a current state of a server; the current state includes a waiting occupation state and an occupied state; the occupied state indicates that there is currently a speech voice being played.

[0188] The first sending module 902 is configured to send a speech occupation failure message to the client in a case where the current state of the server is the occupied state; the speech occupation failure message is used for the client to obtain recording information of the speech voice of the local user and send a second speech request message containing the recording information to the server.

[0189] The storage module 903 is configured to store the recording information in a to-be-played voice queue in a case where the second speech request message sent by the client is received.

[0190] The second sending module 904 is configured to obtain first target recording information to be played from the to-be-played voice queue and send a first voice playing instruction containing the first target recording information to each client in the current conference in a case where the current speech voice information is played to an end; the first voice playing instruction is used for instructing each client to play the first target recording information.

[0191] In one of the embodiments, the apparatus further includes a third sending module and a fourth sending module, wherein:

[0192] The third sending module is configured to send a speech preemption success message to the client and switch the current state of the server to a preoccupied state, if the current state of the server is a waiting preemption state. The speech preemption success message is configured to instruct the client to send the real-time voice information of the local user to the server.

[0193] The fourth sending module is configured to receive the real-time voice information sent by the client and send a second voice playing instruction containing the real-time voice information to other clients in the current conference. The second voice playing instruction is configured to instruct the other clients to play the real-time voice information.

[0194] In one of the embodiments, the second sending module 904 is specifically configured to: switch the current state of the server to a waiting preemption state, if the current speech voice information playing ends; acquire first target recording information to be played from the voice queue to be played, send a first voice playing instruction containing the first target recording information to each client in the current conference, and switch the current state of the server to a preoccupied state, if the first speech request message is not received within a preset time period; and send a speech preemption success message to the client sending the first speech request message and switch the current state of the server to a preoccupied state, if the first speech request message is received within the preset time period.

[0195] In one of the embodiments, the apparatus further includes a fifth sending module configured to acquire second target recording information to be played from the voice queue to be played, send a third voice playing instruction containing the second target recording information to each client in the current conference, if the current speech voice information is real-time voice information and the current speech voice information exceeds a preset time length and / or the number of speech pauses exceeds a preset number. The third voice playing instruction is configured to instruct each client to play the second target recording information.

[0196] The modules in the conference data processing apparatus described above can be all or partially implemented by software, hardware, or a combination thereof. The modules described above can be embedded in or independent of a processor in a computer device in a hardware form, or stored in a memory in a computer device in a software form, so as to be called and executed by a processor to perform the operations corresponding to the modules.

[0197] In one embodiment, a computer device is provided, which can be a terminal. An internal structure diagram of the computer device can be as shown in Figure 10As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The communication interface of the computer device is used for wired or wireless communication with external terminals. Wireless communication can be achieved through WIFI, mobile cellular network, NFC (Near Field Communication) or other technologies. The computer program is executed by the processor to implement a conference data processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0198] Those skilled in the art can understand that Figure 10 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. A specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0199] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in the above method embodiments.

[0200] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in the above method embodiments.

[0201] In one embodiment, a computer program product is provided, including a computer program, and the computer program is executed by a processor to implement the steps in the above method embodiments.

[0202] The conference data processing method, device, system, computer device, storage medium and computer program product provided by the present application relate to the field of artificial intelligence technology, and can be used in the field of financial technology or other related fields. The application does not limit the application field of the conference data processing method, device, system, computer device, storage medium and computer program product.

[0203] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.

[0204] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments. Any reference to memory, database or other medium used in the embodiments provided by the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0205] The technical features of the above embodiments can be combined in any way. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictory, it should be considered as the scope of the present application.

[0206] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. A conference data processing method characterized by comprising: The method is applied to a client, and the method comprises: In response to a speaking instruction, a first speaking request message is sent to a server; In response to a speaking preemption failure message sent by the server, speaking voice of a local user is recorded; In response to a recording end instruction, recording information containing the speaking voice of the local user is obtained, and a second speaking request message containing the recording information is sent to the server; the second speaking request message is used to instruct the server to store the recording information in a to-be-played voice queue, and in the case that current speaking voice information is played to an end, to obtain first target recording information to be played from the to-be-played voice queue, and to send a first voice playing instruction containing the first target recording information to the client; In response to the first voice playing instruction, the first target recording information contained in the first voice playing instruction is played; The playing of the first target recording information contained in the first voice playing instruction comprises: In a current interface, a display box corresponding to a current speaker is determined; According to position information of the display box in the current interface and a preset volume correction strategy, a target volume corresponding to the current speaker is determined; The first target recording information is played according to the target volume.

2. The method of claim 1, wherein, The method further comprises: In response to a speaking preemption success message sent by the server, real-time voice information of the local user is sent to the server; the real-time voice information is used to instruct the server to send a second voice playing instruction containing the real-time voice information to other clients in a current conference.

3. The method of claim 1, wherein, The method further comprises: Receiving recording summary information of each recording information in the to-be-played voice queue sent by the server, and displaying the recording summary information in a current interface; In response to a playing instruction for second target recording information, a recording playing request message is sent to the server; the recording playing request message is used to instruct the server to obtain the second target recording information from the to-be-played voice queue and send it to the client; Receiving second target recording information sent by the server, and playing the second target recording information.

4. The method of claim 1, wherein, The method further comprises: In a current interface, a display box corresponding to a current speaker is determined; According to a preset display mode, the display box of the current speaker is displayed.

5. A conference data processing method characterized by comprising: The method is applied to a server, and the method comprises: A first speaking request message sent by a client is received, and a current state of the server is determined; the current state comprises a waiting preemption state and a preoccupied state; the preoccupied state indicates that there is currently playing speaking voice information; In the case that the current state of the server is the preoccupied state, a speaking preemption failure message is sent to the client; the speaking preemption failure message is used to instruct the client to obtain recording information containing speaking voice of a local user, and to send a second speaking request message containing the recording information to the server; In the case that the second speaking request message sent by the client is received, the recording information is stored in a to-be-played voice queue; In a case where the current speech information is played, first target recording information to be played is obtained from the to-be-played speech queue, and a first speech playing instruction containing the first target recording information is sent to each client in the current conference; the first speech playing instruction is used to instruct the clients to play the first target recording information. The first speech playing instruction is also used to instruct the clients to perform the following steps: In the current interface, a display box corresponding to the current speaker is determined; According to position information of the display box in the current interface and a preset volume correction strategy, a target volume corresponding to the current speaker is determined; The first target recording information is played according to the target volume.

6. The method of claim 5, wherein, The method further comprises: In a case where the current state of the server is a waiting preemption state, a speech preemption success message is sent to the client, and the current state of the server is switched to a preoccupied state; the speech preemption success message is used to instruct the client to send real-time speech information of the local user to the server; The real-time speech information sent by the client is received, and a second speech playing instruction containing the real-time speech information is sent to other clients in the current conference; the second speech playing instruction is used to instruct the other clients to play the real-time speech information.

7. The method of claim 5, wherein, The method further comprises: In a case where the current speech information is played, first target recording information to be played is obtained from the to-be-played speech queue, and a first speech playing instruction containing the first target recording information is sent to each client in the current conference; the first speech playing instruction is used to instruct the clients to play the first target recording information. In a case where the current speech information is played, the current state of the server is switched to a waiting preemption state; If no first speech request message is received within a preset time period, first target recording information to be played is obtained from the to-be-played speech queue, a first speech playing instruction containing the first target recording information is sent to each client in the current conference, and the current state of the server is switched to a preoccupied state; 8. The method of claim 5, wherein, If a first speech request message is received within the preset time period, a speech preemption success message is sent to the client sending the first speech request message, and the current state of the server is switched to a preoccupied state. The method further comprises:

9. A conference data processing system characterized by comprising: In a case where the current speech information is real-time speech information, and the current speech information exceeds a preset time length and / or a speech pause number exceeds a preset number, second target recording information to be played is obtained from the to-be-played speech queue, and a third speech playing instruction containing the second target recording information is sent to each client in the current conference; the third speech playing instruction is used to instruct the clients to play the second target recording information. The conference data processing system comprises a plurality of clients and a server, wherein: The client is used to send a first speech request message to the server in response to a speech instruction; The server is used for receiving the first speech request message sent by the client, determining a current state of the server; the current state comprises a waiting preemption state and a preempted state; the preempted state indicates that there is currently playing speech voice; in a case where the current state of the server is the preempted state, a speech preemption failure message is sent to the client; The client is further used for, in response to the speech preemption failure message sent by the server, recording speech voice of a local user; in response to an end-of-recording instruction, obtaining recording information containing the speech voice of the local user, and sending a second speech request message containing the recording information to the server; The server is further used for, in a case where the second speech request message sent by the client is received, storing the recording information in a to-be-played voice queue; in a case where current speech voice information playing ends, obtaining first target recording information to be played from the to-be-played voice queue, and sending a first voice playing instruction containing the first target recording information to each client in the current conference; The client is further used for, in response to the first voice playing instruction, playing the first target recording information contained in the first voice playing instruction; The client is further used for, in response to the first voice playing instruction, determining a display box corresponding to a current speaker in a current interface; According to position information of the display box in the current interface and a preset volume correction strategy, a target volume corresponding to the current speaker is determined; The first target recording information is played according to the target volume.

10. A conference data processing apparatus characterized by comprising: The device for executing the conference data processing method according to any one of claims 1-4 comprises: A first sending module is configured to send a first speech request message to a server in response to a speech instruction; A recording module is configured to record speech voice of a local user in response to a speech preemption failure message sent by the server; A second sending module is configured to obtain recording information containing the speech voice of the local user in response to an end-of-recording instruction, and send a second speech request message containing the recording information to the server; the second speech request message is used to instruct the server to store the recording information in a to-be-played voice queue, and in a case where current speech voice information playing ends, obtain first target recording information to be played from the to-be-played voice queue, and send a first voice playing instruction containing the first target recording information to a client; A first playing module is configured to play the first target recording information contained in the first voice playing instruction in response to the first voice playing instruction.

11. A conference data processing apparatus characterized by comprising: The device for executing the conference data processing method according to any one of claims 5-8 comprises: A determining module is configured to receive a first speech request message sent by a client, and determine a current state of a server; the current state comprises a waiting preemption state and a preempted state; the preempted state indicates that there is currently playing speech voice; The first sending module is configured to send a speech preemption failure message to the client when the current state of the server is the preemption state; the speech preemption failure message is used by the client to obtain recording information containing speech voice of a local user and send a second speech request message containing the recording information to the server; The storage module is configured to store the recording information in a to-be-played voice queue when the second speech request message sent by the client is received. The second sending module is configured to obtain first target recording information to be played from the to-be-played voice queue when current speech voice information playing is ended, and send a first voice playing instruction containing the first target recording information to each client in the current conference; the first voice playing instruction is used to instruct the clients to play the first target recording information.

12. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 4 or 5 to 8.

13. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 4 or 5 to 8.

14. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 4 or 5 to 8.

Citation Information

Patent Citations

  • Mobile video conference system and implementing method thereof

    CN105376516A

  • Voice interaction method, device and equipment for online class and storage medium

    CN114760274A