Voice information processing method and device, electronic equipment and readable storage medium

By selecting the nearest electronic device for recording and amplifying the volume in multi-person meetings, the problem of not being able to hear clearly in offline or on-site meetings is solved, and the participants are able to obtain complete information.

CN116506552BActive Publication Date: 2026-01-09VIVO MOBILE COMM CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310475550.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-27
Publication Date
2026-01-09
Estimated Expiration
2043-04-27

AI Technical Summary

Technical Problem

In offline or in-person meetings with multiple participants, the noise of the meeting environment, the speaker's soft voice, or the participant's hearing problems may prevent them from hearing the content clearly, thus affecting their ability to obtain complete meeting information.

Method used

By selecting the electronic device with the smallest distance from the first participant among N electronic devices to collect the speaker's audio, and when the volume is insufficient, determining the target electronic device to amplify the volume of the speaker's voice so that the volume is greater than or equal to the participant's volume threshold.

Benefits of technology

This ensures that all participants can clearly hear the speeches and obtain complete meeting information, thus resolving the problem of not being able to hear the speeches clearly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116506552B_ABST
    Figure CN116506552B_ABST
Patent Text Reader

Abstract

The application discloses a speech information processing method and device, electronic equipment and a readable storage medium, and belongs to the technical field of speech processing. In the case that N electronic devices are connected to a first conference system, for a first conference object in the first conference system except a speaking object, a first electronic device with the smallest distance to the first conference object in the N electronic devices collects a first audio of the speaking object; in the case that the volume of the first audio is less than a first volume threshold corresponding to the first conference object, a target electronic device corresponding to the first conference object is determined from the N electronic devices; the second audio of the speaking object is obtained based on the target electronic device; and the first audio is played based on the target electronic device according to an audio parameter, and the audio parameter includes a volume value, and the volume value is greater than or equal to the first volume threshold.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of voice processing, and particularly relates to a voice information processing method and device, an electronic device and a readable storage medium. BACKGROUND

[0002] For a multi-person offline meeting or live meeting, due to various factors such as a noisy meeting scene, a small voice of a speaker, or hearing problems of a meeting participant, the meeting participant may not be able to clearly hear the speech content of the speaker, so that the meeting participant cannot obtain complete meeting information in the meeting. SUMMARY

[0003] Embodiments of the present application provide a voice information processing method, device, electronic device and readable storage medium, which can solve the problem that a meeting participant cannot clearly hear the speech content of a speaker in an offline meeting or live meeting.

[0004] In a first aspect, embodiments of the present application provide a voice information processing method, which comprises:

[0005] In the case where N electronic devices are connected to the first conference system, for a first meeting participant other than the speaker in the first conference system, a first audio of the speaker is collected based on a first electronic device with the smallest distance to the first meeting participant among the N electronic devices, where N is a positive integer greater than 1;

[0006] In the case where the volume of the first audio is less than a first volume threshold corresponding to the first meeting participant, a target electronic device corresponding to the first meeting participant is determined from the N electronic devices;

[0007] The second audio of the speaker is obtained based on the target electronic device;

[0008] The second audio is played based on the target electronic device according to an audio parameter, and the audio parameter includes a volume value, and the volume value is greater than or equal to the first volume threshold.

[0009] In a second aspect, embodiments of the present application provide a voice information processing device, which comprises:

[0010] The first sound acquisition module is configured to, in the case where N electronic devices are connected to the first conference system, collect a first audio of a speaker based on a first electronic device with the smallest distance to a first meeting participant among the N electronic devices for the first meeting participant other than the speaker in the first conference system, where N is a positive integer greater than 1;

[0011] The target device determining module is configured to determine, from the N electronic devices, a target electronic device corresponding to the first conference participant, in a case where the volume of the first audio is less than a first volume threshold corresponding to the first conference participant.

[0012] The second sound obtaining module is configured to obtain, based on the target electronic device, second audio of the speaking object.

[0013] The playing module is configured to play, based on the target electronic device, the second audio according to an audio parameter, the audio parameter including a volume value, the volume value being greater than or equal to the first volume threshold.

[0014] In a third aspect, an electronic device is provided, and the device includes a processor and a memory, the memory storing programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the voice information processing method according to the first aspect.

[0015] In a fourth aspect, a readable storage medium is provided, and the readable storage medium stores programs or instructions, and the programs or instructions, when executed by a processor, implement the steps of the method according to the first aspect.

[0016] In a fifth aspect, a chip is provided, and the chip includes a processor and a communication interface, the communication interface being coupled to the processor, and the processor being configured to run programs or instructions to implement the method according to the first aspect.

[0017] In a sixth aspect, a computer program product is provided, and the program product is stored in a storage medium, and the program product is executed by at least one processor to implement the method according to the first aspect.

[0018] In the embodiment of the present application, in the case that N electronic devices are connected to the first conference system, for the first conference object in the first conference system except the speaking object, the first audio of the speaking object is collected based on the first electronic device in the N electronic devices which is closest to the first conference object, and in the case that the volume of the first audio is less than the first volume threshold corresponding to the first conference object, the target electronic device corresponding to the first conference object is determined from the N electronic devices, the second audio of the speaking object is obtained based on the target electronic device, and the second audio is played based on the target electronic device according to the volume, which is greater than or equal to the first volume threshold. According to the embodiment, in the case that the first conference object cannot hear the speaking sound of the speaking object, the speaking sound of the speaking object can be played based on the target electronic device at a volume that the first conference object can hear, so that the first conference object can obtain the speaking content of the speaking object based on the sound played by the target electronic device, thereby solving the problem that the conference object cannot hear the speaking content of the speaking object, and further ensuring that the conference object can obtain complete conference information in the conference. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiments of the present application will be briefly introduced as follows. Those skilled in the art can also obtain other drawings according to these drawings without creating any creative labor.

[0020] Figure 1 is a scene schematic diagram of an offline conference provided by the embodiments of the present application;

[0021] Figure 2 is a display interface schematic diagram of a virtual conference room provided by the embodiments of the present application;

[0022] Figure 3 is a flow schematic diagram of a voice information processing method provided by the embodiments of the present application;

[0023] Figure 4 is a flow schematic diagram of a method for determining the first volume threshold corresponding to the first conference object provided by the embodiments of the present application;

[0024] Figure 5 is a flow schematic diagram of a method for determining the target electronic device provided by the embodiments of the present application;

[0025] Figure 6 is a flow schematic diagram of a method for determining the target volume provided by the embodiments of the present application;

[0026] Figure 7 is a structural schematic diagram of a voice information processing device provided by the embodiments of the present application;

[0027] Figure 8 is a structural schematic diagram of an electronic device provided by an embodiment of the present application.

[0028] Figure 9 is a hardware structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0029] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0030] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second", and the like are generally of a kind and are not limited in number, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.

[0031] The voice information processing method, device, electronic device and readable storage medium provided by the embodiments of the present application will be described in detail below in combination with the drawings and specific embodiments and their application scenarios.

[0032] First, the voice information processing method provided by the embodiments of the present application will be introduced.

[0033] With the development of science and technology, the functions of electronic devices are becoming more and more perfect, so that electronic devices can gradually be applied in offline meetings or live meetings to improve the intelligentization and diversification of meeting scenarios. A meeting usually includes multiple electronic devices, and the multiple electronic devices can be connected to the same meeting system (hereinafter referred to as a first meeting system) through a network or other connection methods, so as to form a virtual meeting site on the network.

[0034] For example Figure 1As shown in FIG. 1, a schematic diagram of a scenario of an offline meeting is shown, which includes five meeting participants, namely, user 1, user 2, user 3, user 4 and user 5. Each of the meeting participants is provided with an electronic device, namely, device 1, device 2, device 3, device 4 and device 5. The offline meeting further includes a first conference system 100. The devices 1-5 are connected to the first conference system 100 through a network. After the connection, a virtual conference room is formed on the network. The virtual conference room can be displayed on the electronic devices 1-5. The display interface is shown in FIG. 1. The first conference system 100 can include a central control device of multiple electronic devices, which is used to control the multiple electronic devices. Figure 2

[0035] In the meeting, any meeting participant can participate in the speech. The meeting participant participating in the speech is the speech participant. For example, as shown in FIG. 2, in the meeting, the user 1, the user 2 and the user 5 participate in the speech. Therefore, the user 1, the user 2 and the user 5 are the speech participants. Figure 1

[0036] The electronic device can include a mobile phone, a computer, a smart watch and the like, which has a sound receiving and playing function. The electronic device can be provided in the meeting site in advance or can be provided by the meeting participant himself / herself, for example, the mobile phone of the meeting participant.

[0037] The speech information processing method provided by the embodiments of the present application can be applied to the offline meeting or the live meeting scenario including multiple electronic devices and the first conference system 100 as shown in FIG. 1, to solve the technical problem that the meeting participant cannot clearly hear the speech content of the speech participant. Figure 1

[0038] Referring to FIG. 3, a flowchart of a speech information processing method provided by the embodiments of the present application is shown. As shown in FIG. 3, the method can include the following steps S31-S34. Figure 3 Figure 3

[0039] S31. In the case that N electronic devices are connected to the first conference system, for a first meeting participant in the first conference system except the speech participant, a speech sound of the speech participant is collected based on a first electronic device in the N electronic devices, which is closest to the first meeting participant, wherein N is a positive integer greater than 1.

[0040] The first meeting participant is any meeting participant except the speech participant among all the meeting participants participating in the meeting.

[0041] The distances between different electronic devices in the N electronic devices and the first meeting participant are usually different. The electronic device closest to the first meeting participant is taken as the first electronic device. It can be seen that the first electronic device is included in the N electronic devices.​​​​​

[0042] When the first audio of the speaking object is collected based on the first electronic device, the speaking voice of the speaking object can be live broadcasted based on the radio function of the first electronic device, so as to obtain the first audio. Because the first electronic device is closest to the first participant, the volume of the first audio collected by the first electronic device is closest to the volume of the speaking voice of the speaking object heard by the first participant in the conference site. Therefore, whether the first participant can hear the speaking voice of the speaking object can be determined based on the volume of the first audio collected by the first electronic device.

[0043] S32. In a case where the volume of the first audio is less than a first volume threshold corresponding to the first participant, a target electronic device corresponding to the first participant is determined from the N electronic devices.

[0044] The first volume threshold corresponding to the first participant is a threshold value preset to distinguish whether the volume is a volume that the first participant can hear. Because different participants have different definitions of hearing, the first volume threshold corresponding to different participants is different.

[0045] The first participant can hear the sound greater than or equal to the first volume threshold corresponding to the first participant, and cannot hear the sound less than the first volume threshold corresponding to the first participant. Based on this, the volume of the first audio collected by the first electronic device can be compared with the first volume threshold corresponding to the first participant. In a case where the volume of the first audio is less than the first volume threshold, it is determined that the first participant cannot hear the speaking voice of the speaking object, and in a case where the volume of the first audio is greater than or equal to the first volume threshold, it is determined that the first participant can hear the speaking voice of the speaking object.

[0046] In a case where it is determined that the first participant cannot hear the speaking voice of the speaking object, in order to make the first participant hear the speaking voice of the speaking object, further, a target electronic device is determined from the N electronic devices, so that the speaking voice of the speaking object can be processed through the target electronic device, thereby enabling the first participant to hear the speaking voice of the speaking object. The target electronic device can be the first electronic device, or can be an electronic device other than the first electronic device in the N electronic devices.

[0047] In a case where it is determined that the first participant can hear the speaking voice of the speaking object, the speaking voice of the speaking object for the first participant can not be processed, that is, a target electronic device corresponding to the first participant does not need to be determined, and S33 and S34 do not need to be executed again.

[0048] S33. The second audio of the speaking object is obtained based on the target electronic device.

[0049] In the process of the target electronic device processing the speech sound of the speech object, the speech sound of the speech object is first acquired based on the target electronic device, and the acquired speech sound of the speech object is determined as the second audio.

[0050] S34. The target electronic device plays the second audio according to the audio parameter, and the audio parameter includes a volume value, and the volume value is greater than or equal to the first volume threshold.

[0051] After the second audio is acquired, the target electronic device can adjust and play the acquired second audio based on the audio parameter, wherein the audio parameter includes a volume value greater than or equal to the first volume threshold, so that the first participant can hear clearly when the adjusted second audio is played. Because the second audio is actually the speech sound of the speech object, the first participant can hear clearly the speech sound of the speech object through this way, thereby solving the problem of not being able to obtain complete conference information due to not being able to hear clearly.

[0052] The voice information processing method provided by the embodiments of the present application, in the case that N electronic devices are connected to the first conference system, for the first participant in the first conference system except the speech object, the first electronic device with the smallest distance from the first participant is used to collect the first audio of the speech object in the N electronic devices, and in the case that the volume of the first audio is less than the first volume threshold corresponding to the first participant, the target electronic device corresponding to the first participant is determined from the N electronic devices, the second audio of the speech object is acquired based on the target electronic device, and the second audio is played according to the volume value based on the target electronic device, and the volume value is greater than or equal to the first volume threshold. According to the present embodiment, in the case that the first participant cannot hear clearly the speech sound of the speech object, the target electronic device can play the speech sound of the speech object at a target volume that the first participant can hear clearly, so that the first participant can obtain the speech content of the speech object based on the sound played by the target electronic device, and solve the problem that the participant cannot hear clearly the speech content of the speech object, and further ensure that the participant can obtain complete conference information in the conference.

[0053] In some embodiments, after the participants enter the conference site, the electronic device with the smallest distance from each participant can be determined before step S31 is performed.

[0054] As an example, in the first conference system, each electronic device can be associated with the seat closest to it in the conference room. After the participants arrive at the conference room, their seats are determined, and each participant is associated with their seat in the first conference system. Thus, for any participant, when determining the electronic device closest to them, the association between the participant and the seat, and between the seat and the electronic device, can be used to identify the electronic device associated with the seat associated with the participant as the closest electronic device to that participant.

[0055] by Figure 1 For example, the electronic device placed in front of each participant is the one that is closest to the participant's seat. Thus, the electronic device placed in front of each participant can be determined as the one that is furthest from that participant. That is, the electronic device that is closest to user 1 is device 1, the one that is closest to user 2 is device 2, the one that is closest to user 3 is device 3, the one that is closest to user 4 is device 4, and the one that is closest to user 5 is device 5.

[0056] As another example, when electronic devices have distance measurement capabilities, for any participant, the electronic device with the smallest distance to that participant can be determined based on its distance measurement function. Specifically, after the participants have entered the meeting room, the electronic devices can use their distance measurement function to measure the distance between themselves and each participant in the vicinity. The measured distances are then correlated with the corresponding participants and uploaded to the first meeting system. The first meeting system, based on the distances uploaded by each electronic device and the correlation between the participants, determines the electronic device with the smallest distance to each participant. In this system, when electronic devices measure the distance to a participant based on their ranging function, they can collect the corresponding participant's voiceprint information. When associating the measured distance with the corresponding participant, the system can also associate the measured distance with the corresponding participant's voiceprint information. The first conference system can record the association between the voiceprint information of all participants and the participants. Thus, the first conference system can determine the participants corresponding to each distance based on the association between the distance and voiceprint information uploaded by the electronic devices and the association between the participants and voiceprint information recorded, and then determine the electronic device with the smallest distance to each participant.

[0057] As a further example, in a case where the electronic device is provided by the participant, the electronic device can be associated with its provider in the first conference system when the electronic device is connected to the first conference system. The participant usually carries his own electronic device, and thus it can be considered that the participant is closest to the electronic device provided by himself. Therefore, for any participant, when the electronic device closest to the participant is determined, the electronic device provided by the participant, i.e., the electronic device associated with the participant, can be determined as the electronic device closest to the participant.

[0058] After the electronic device closest to each participant is determined, S31 can be performed for the first participant.

[0059] In some embodiments, in S31, when the first audio of the speaking object is collected based on the first electronic device, in order to ensure that the first audio of the speaking object can be accurately collected, the voiceprint information of the speaking object can be pre-recorded in the first conference system. The first conference system can send the voiceprint information of the speaking object to the first electronic device. In this way, the first electronic device can collect the audio and take the collected audio matching the voiceprint information of the speaking object as the first audio of the speaking object.

[0060] Because the speaking object in the conference is usually not fixed, in order to ensure that the speaking sound of different speaking objects can be successfully and accurately collected, the voiceprint information corresponding to all participants can be recorded in the first conference system. During the conference, the first conference system can send the voiceprint information of the current speaking object to the first electronic device in real time.

[0061] After the first audio is obtained, whether the first participant can hear the speaking sound of the speaking object can be determined based on the first audio and the first volume threshold corresponding to the first participant through the above step S32.

[0062] In some embodiments, before the above step S32 is performed, the first volume threshold corresponding to the first participant is determined. As described above, when the first volume threshold corresponding to the first participant is determined, the following steps S41-S44 can be included. Figure 4

[0063] S41. Hearing test is performed on the first participant, and the first volume corresponding to the first participant is determined based on the test result.

[0064] ​In actual application, the hearing test can be performed on the first meeting participant based on the electronic device at the meeting site, for example, the hearing test can be performed on the first meeting participant based on the first electronic device. Taking the hearing test performed on the first meeting participant based on the first electronic device as an example, during the hearing test, the first electronic device can play speech sounds with different speech intensity, volume and / or speed, and let the first meeting participant select whether the speech sounds can be clearly heard, so as to determine the first volume suitable for the first meeting participant to listen to according to the selection of the first meeting participant.

[0065] For example, when the hearing test is performed on the first meeting participant based on the first electronic device, the speech intensity of the speech sounds played by the first electronic device generally can fluctuate between 250-1000 Hz in frequency, which can be determined according to actual conditions, the volume generally can be between 40-60, and the speed can be between 120-200 words per minute. The first volume that the first meeting participant can clearly hear is tested through the combination of speech intensity, volume and / or speed. For example, the speech recognition degree of the first meeting participant is 90% in a noise-free environment with a volume of 40 decibels and a speed of 150 words per minute; when the volume is reduced to 30 decibels, the speed is increased to 180 words per minute, and there is a 20-decibel speech noise, the speech recognition degree is reduced to 65, then it can be determined that the speech recognition degree of the first meeting participant needs to be improved by increasing the non-noise decibel or reducing the speed or filtering out invalid materials (such as wind noise, noise, etc.), so that the played speech sounds can be further adjusted on the basis of a volume of 40 decibels and a speed of 150 words per minute in a noise-free environment, so that the speech recognition degree of the first meeting participant reaches a preset value (such as 100%), and the volume that makes the speech recognition degree of the first meeting participant reach the preset value is taken as the first volume.

[0066] S42. Collect scene information of the meeting site where the first meeting participant is located.

[0067] Because there are usually some environmental noises in the meeting site, these environmental noises usually also affect the hearing effect of the first meeting participant on the speech sounds of the speaker. In view of this, in order to ensure that the volume threshold value can be accurately determined whether the first meeting participant can clearly hear the speech sounds of the speaker, in addition to the hearing test on the first meeting participant, the environmental noise of the meeting site can also be considered when determining the volume threshold value.

[0068] The environmental noise of the meeting site is usually related to the scene information of the meeting site, so the scene information of the meeting site can be collected.

[0069] S43. Determine the to-be-enhanced volume corresponding to the first meeting participant according to the scene information.

[0070] The scene information can include at least one of the following: the number of people in the meeting scene, and information of basic equipment in the meeting scene, wherein the information of the basic equipment can include: a sound volume generated by the basic equipment when in operation, and a distance between the first meeting participant and the basic equipment. The basic equipment can include common equipment in the meeting scene, such as an air conditioner, a projector, and the like.

[0071] In the determination of the to-be-enhanced volume corresponding to the first meeting participant according to the scene information, the following can be included:

[0072] determining the first interference volume according to the number of people in the meeting scene;

[0073] determining a second interference volume of the basic equipment to the first meeting participant according to the distance between the first meeting participant and the basic equipment and the sound volume generated by the basic equipment when in operation;

[0074] determining the to-be-enhanced volume corresponding to the first meeting participant according to at least one of the first interference volume and the second interference volume.

[0075] In the determination of the to-be-enhanced volume corresponding to the first meeting participant according to at least one of the first interference volume and the second interference volume, the following can be included: determining the first interference volume as the to-be-enhanced volume corresponding to the first meeting participant, determining the second interference volume as the to-be-enhanced volume corresponding to the first meeting participant, or determining a sum of the first interference volume and the second interference volume as the to-be-enhanced volume corresponding to the first meeting participant.

[0076] S44. determining a first volume threshold corresponding to the first meeting participant according to the first volume and the to-be-enhanced volume.

[0077] For example, a sum of the first volume and the to-be-enhanced volume can be determined as the first volume threshold corresponding to the first meeting participant, or a volume value obtained by weighted summation calculation of the first volume and the to-be-enhanced volume can be determined as the first volume threshold corresponding to the first meeting participant, wherein the weight values corresponding to the first volume and the to-be-enhanced volume respectively in the weighted summation calculation can be determined according to actual conditions.

[0078] In this way, in the determination of the first volume threshold corresponding to the first meeting participant, in addition to considering the hearing factor of the first meeting participant itself, the environmental noise of the meeting scene is also considered, and compared with only considering the hearing factor, the first volume threshold determined according to this kind of way can more accurately judge whether the first meeting participant can hear the speaking sound of the speaking object. In some embodiments, in the step S33, in the determination of the target electronic device corresponding to the first meeting participant, the first electronic device closest to the first meeting participant can be directly determined as the target electronic device.

[0079] Because the first electronic device is closest to the first participant, in general, the sound propagation path between the first electronic device and the first participant is shorter than that between the other electronic devices and the first participant, the transmission consumes less time, and the volume loss in the process is also less. If the N electronic devices play sound at the same volume, the volume of the sound played by the first electronic device heard by the first participant will generally be the largest, and it will be easier to hear clearly. Therefore, taking the first electronic device as the target electronic device can make the first participant hear the speech sound of the speech object played by the target electronic device faster and more clearly.

[0080] In some embodiments, in the above step S33, considering the pose of the first participant in the meeting and the relative position relationship between the first participant and the speech object, it will generally also affect the hearing effect of the first participant, so when determining the target electronic device corresponding to the first participant, the pose information of the first participant and the layout information of the meeting site can also be used to determine the target electronic device corresponding to the first participant, as shown in Figure 5 The way of determining the target electronic device can include the following steps S51-S53.

[0081] S51. Obtain the pose information of the first participant and the layout information of the meeting site where the first participant is located.

[0082] The layout information of the meeting site can include the seat distribution of each participant and the position distribution between the seats in the meeting site and the meeting content display device (such as a conference screen, etc.). Because the layout information of the meeting site generally does not change during the meeting, it can be pre-recorded into the first conference system, so that it can be directly called when used.

[0083] The pose of the first participant may change during the meeting, so the pose information of the first participant can be obtained in real time.

[0084] As an example, in the case where an image acquisition device that can communicate with the first conference system is provided in the meeting site, the position information of the first participant can be obtained based on the image acquisition device.

[0085] For example, when the pose information of the first participant needs to be obtained, the image acquisition device can be used to acquire the image of the first participant, and the image can be transmitted to the first conference system. The first conference system analyzes the image based on image analysis technology to determine the pose information of the first participant.

[0086] S52. Determine the direction of the speech object or the meeting content display device relative to the first participant based on the pose information and the layout information.

[0087] Taking the direction of the conference content display device as an example, according to the layout information, the relative position relationship between the conference content display device and the seat where the first meeting participant is located can be determined, such as the conference content display device being located in front of the seat where the first meeting participant is located. At this time, if it is determined according to the pose information of the first meeting participant that the pose of the first meeting participant is directly facing the front, it can be determined that the target direction of the conference content display device relative to the first meeting participant is in front of the first meeting participant.

[0088] S53. Determine the electronic device located in the direction and having the smallest distance from the first meeting participant as the target electronic device corresponding to the first meeting participant.

[0089] Taking the target direction of the target object relative to the first meeting participant as the front of the first meeting participant as an example, when determining the target electronic device, the electronic device located in front of the first meeting participant can be determined, and the electronic device having the smallest distance from the first meeting participant is determined as the target electronic device corresponding to the first meeting participant.

[0090] The reason for selecting the electronic device located in the direction is to ensure that the sound played by the target electronic device heard by the first meeting participant and the speech sound of the speech object heard on site come from the same direction, and to improve the authenticity. The reason for selecting the electronic device having the smallest distance from the first meeting participant is to enable the first meeting participant to hear the sound played by the target electronic device more quickly and clearly.

[0091] By determining the target electronic device in this way, the authenticity of the heard sound can be improved while ensuring that the first meeting participant can hear the speech sound of the speech object played by the target electronic device more quickly and clearly.

[0092] Further, in order to avoid the distance between the target electronic device determined by the above-mentioned method and the first meeting participant being too large, when performing S53, the electronic device located in the direction and having a distance from the first meeting participant less than a preset distance threshold can be determined, and the determined electronic device is selected as a candidate electronic device. The electronic device having the smallest distance from the first meeting participant is determined as the target electronic device corresponding to the first meeting participant. The distance threshold can be set according to actual conditions.

[0093] By the above-mentioned method, the distance between the finally determined target electronic device and the first meeting participant can be ensured not to be too large, thereby on the one hand, problems such as long sound transmission time and serious sound attenuation caused by too large distance can be avoided, and on the other hand, disturbance to other meeting participants can be avoided to some extent.

[0094] Further, if no electronic device located in the direction and the distance between the first participant and the electronic device is less than the preset distance threshold, that is, the distance between the electronic device in the direction and the first participant is large, an electronic device with the smallest distance to the first participant within the distance threshold can be selected from other directions as the target electronic device corresponding to the first participant, so that the voice information processing based on the target electronic device can be finally realized.

[0095] After the target electronic device is determined, the second audio of the speaking object can be obtained based on the target electronic device through S34.

[0096] In some embodiments, the target electronic device can obtain the second audio of the speaking object in a live sound collecting manner in S34. Similar to collecting the first audio based on the first electronic device, the target electronic device can collect sound in real time according to the voiceprint information of the speaking object, and determine the speaking sound matching the voiceprint information as the second audio of the speaking object.

[0097] Considering that the conference site is usually noisy, live sound collection may result in unclear second audio, so in some embodiments, the second audio of the speaking object collected by the second electronic device can be received by the target electronic device in S34, and the second electronic device is the electronic device with the smallest distance to the speaking object among the N electronic devices.

[0098] Because the second electronic device has the smallest distance to the speaking object, the second electronic device can collect the audio of the speaking object more clearly. The second electronic device transmits the collected second audio of the speaking object to the target electronic device through network transmission, so that the target electronic device can obtain the clear second audio of the speaking object, and the finally played audio can also be clear.

[0099] In some embodiments, before performing the above step S35, the audio parameter including the volume value can be determined.

[0100] As shown in FIG. 6, the method for determining the volume value included in the audio parameter can include the following steps S61-S64. Figure 6

[0101] S61. Obtain the distance between the target electronic device and the first participant.

[0102] The distance between the target electronic device and the first participant can be measured based on the ranging function of the target electronic device.

[0103] S62. Determine the second volume threshold played by the target electronic device according to the distance between the target electronic device and the first participant and the first volume threshold. ​

[0104] Because the distance between the target electronic device and the first participant is usually not 0, if the first volume threshold is directly used as the playing volume of the target electronic device to play the speech sound, the volume transmitted to the ear of the first participant may be less than the first volume threshold due to the attenuation of the sound in the propagation process, and thus the first participant cannot clearly hear the second audio played by the target electronic device. In view of this, in order to ensure that the first participant can clearly hear the second audio played by the target electronic device, a first volume threshold at which the target electronic device plays can be determined according to the distance between the target electronic device and the first participant and the volume threshold.

[0105] In determining the second volume threshold at which the target electronic device plays, the corresponding volume attenuation amount can be calculated according to the distance between the target electronic device and the first participant and the existing volume attenuation formula, and then the sum of the volume attenuation amount and the volume threshold can be determined as the second volume threshold.

[0106] S63. Determine a volume interval according to the second volume threshold, and a lower limit value of the volume interval is greater than or equal to the second volume threshold.

[0107] The upper limit value of the volume interval can be determined according to actual conditions, as long as the target electronic device plays at a volume within the volume interval without making the first participant uncomfortable.

[0108] As an example, taking the second volume threshold determined through S62 as 40 decibels and the best volume at which the first participant feels comfortable as 50 decibels, the volume interval can be set as [40, 50].

[0109] S64. Determine a volume value from the volume interval as the volume value included in the audio parameter.

[0110] According to the setting of the volume interval as described above, when the target electronic device plays audio at a volume within the volume interval, the first participant can definitely hear it without feeling uncomfortable, so a volume value can be randomly selected from the volume interval as the volume value included in the audio parameter.

[0111] In some embodiments, considering that in addition to the first participant and the speaking participant, there can be other participants in the conference, in order to avoid affecting the other participants when the target electronic device plays the second audio, S64 can use the following method when determining a volume value from the volume interval as the volume value included in the audio parameter.

[0112] Determine a second participant in the conference site that is closest to the first participant;

[0113] determine a probability value of the second conference participant being disturbed by the target electronic device playing the sound at the candidate volume value according to a distance between the target electronic device and the second conference participant;

[0114] In a case where the probability value of the second conference participant being disturbed by the target electronic device playing the sound at the candidate volume value is less than a preset probability threshold, the candidate volume value is determined as the target volume.

[0115] The probability threshold can be set according to actual conditions.

[0116] The manner of determining the probability value of the second conference participant being disturbed by the target electronic device playing the sound at the candidate volume value can include:

[0117] The distance between the target electronic device and the second conference participant is determined, an attenuation amount of the candidate volume value transmitted by the target electronic device to the second conference participant is determined based on the distance, the actual volume that the second conference participant is expected to hear is obtained by subtracting the attenuation amount from the candidate volume value, and the probability value of the second conference participant being disturbed by the target electronic device playing the sound at the candidate volume value is determined according to the actual volume.

[0118] For example, when determining the probability value of the second conference participant being disturbed by the target electronic device playing the sound at the candidate volume value according to the actual volume, the actual volume can be compared with a preset influence volume. If the actual volume is less than the influence volume, the probability value of the second conference participant being disturbed is determined as a first probability value, and if the actual volume is greater than or equal to the influence volume, the probability value of the second conference participant being disturbed is determined as a second probability value. The first probability value is less than the probability threshold, the second probability value is greater than or equal to the probability threshold, and the influence volume can be a volume that is determined according to prior knowledge or experiments to have an impact on a person.

[0119] Further, in a case where the probability value of the second conference participant being disturbed by the target electronic device playing the sound at the candidate volume value is greater than or equal to the probability threshold, the candidate volume value can be reduced, wherein the reduced candidate volume value is still within the volume range, until the probability value of the second conference participant being disturbed by the target electronic device playing the sound at the candidate volume value is less than the probability threshold, and the candidate volume value is determined as the volume value included in the audio parameter.

[0120] In some embodiments, because reasons other than a small volume can also cause a user to be unable to hear clearly, such as a fast speech speed, the audio parameter can include a speech speed parameter in addition to the volume.

[0121] As an example, the speech speed in the audio parameter can be set according to the actual needs of the first conference participant, and generally can be a lower speech speed, because generally the lower the speech speed, the easier it is to hear clearly. In this way, when the target electronic device adjusts and plays the second audio based on the speech speed, if the speech speed of the second audio is fast, the speech speed of the second audio can be reduced, so that the first conference participant can more easily hear the second audio played by the target electronic device, thereby enhancing the hearing experience effect of the first conference participant.

[0122] In some embodiments, when the first conference participant and the speaking object are in a back-to relationship, the first conference participant is more difficult to hear the speaking sound of the speaking object than when they are in a face-to relationship. In view of this, before determining whether the first conference participant can hear the speaking sound of the speaking object according to the volume of the first audio, the following steps can be performed first:

[0123] Determine whether the first conference participant and the speaking object are in a back-to relationship, and in the case that the first conference participant and the speaking object are in a back-to relationship, increase the first volume threshold corresponding to the first conference participant. In this way, by comparing the volume of the first audio with the increased volume threshold, it can be determined whether the first conference participant can hear the speaking sound of the speaking object. In the case that the volume of the first audio is less than the increased volume threshold, it is determined that the first conference participant cannot hear the speaking sound of the speaking object.

[0124] Correspondingly, in the above step S32, in the case that the volume of the first audio is less than the increased volume threshold, the target electronic device corresponding to the first conference participant is determined from the N electronic devices.

[0125] For example, in the case that an image acquisition device is provided at the conference site, an image containing the first conference participant and the speaking object can be acquired based on the image acquisition device, and the image is sent to the first conference system. The first conference system analyzes the image based on image analysis technology to determine whether the first conference participant and the speaking object are in a back-to relationship.

[0126] For example, in the case that it is determined that the first conference participant and the speaking object are in a back-to relationship, a preset volume value can be obtained, and the increase in the volume threshold can be realized by adding the first volume threshold corresponding to the first conference participant to the preset volume value. The added volume value is used as the increased volume threshold. The preset volume value can be set according to actual conditions, for example, it can be 5 decibels.

[0127] When the first conference participant and the speaking object are in a back-to relationship, the volume threshold is increased, and based on the increased volume threshold, it can be more accurately determined whether the first conference participant can hear the speaking sound of the speaking object.

[0128] In some embodiments, it is considered that there can be a problem of insufficient accuracy in setting the volume threshold value, so that the first volume threshold value corresponding to the first conference participant cannot accurately identify whether the first conference participant can hear the speech sound of the speaker. For example, due to the small first volume threshold value, in the case that the first conference participant cannot hear the speech sound, it is determined that the first conference participant can hear based on the volume threshold value. In view of this, in the case that the first volume threshold value cannot accurately determine whether the first conference participant can hear the speech sound of the speaker, the setting of the first volume threshold value can be re-performed in order to improve the accuracy of identification.

[0129] For example, if the conference participant thinks that he or she cannot hear the speaker, but the electronic device closest to him or her does not identify that he or she cannot hear, the conference participant can choose to adjust the first volume threshold value corresponding to him or her. When adjusting the first volume threshold value, the conference participant can be provided with a selection of different volume sounds played by the electronic device closest to him or her until the volume value that the conference participant can hear is selected, and the volume value is updated as the latest first volume threshold value corresponding to the conference participant, thereby realizing the updating of the first volume threshold value.

[0130] In some embodiments, it is considered that there can be a plurality of speakers in the conference, but the first conference participant may only want to understand the speech content of a specific speaker. In view of this, before collecting the first audio of the speaker based on the first electronic device closest to the first conference participant among the N electronic devices in step S31, the following steps can be performed first:

[0131] In the case that there are M speakers in the conference, the identification information corresponding to the M speakers is displayed in the first electronic device closest to the first conference participant, where M is a positive integer greater than 1;

[0132] Receiving a first input of the identification information corresponding to the first speaker, the first speaker being any one of the M speakers;

[0133] In response to the first input, the first speaker is taken as the target speaker.

[0134] Correspondingly, in S31, the sound of the target speaker is collected based on the first electronic device closest to the first conference participant among the N electronic devices.

[0135] Correspondingly, in S33, the second audio of the target speaker is obtained based on the target electronic device.

[0136] In this way, when M speaking objects speak at the same time, the first conference participant can select a speaking object of the greatest interest to the first conference participant as a target speaking object in the first electronic device. Thus, the first electronic device only needs to determine whether the first conference participant can hear the speech of the target speaking object, and only processes the speech sound of the target speaking object when the first conference participant cannot hear the speech, so that targeted voice processing is achieved.

[0137] The speech information processing method provided in the embodiments of the present application can be executed by a speech information processing device. The speech information processing device provided in the embodiments of the present application is described below with reference to the speech information processing method executed by the speech information processing device.

[0138] Referring to Figure 7 The structure of the speech information processing device provided in the embodiments of the present application is shown in Figure 7 The device can include the following modules:

[0139] The first sound acquisition module 701 is configured to, in a case where N electronic devices are connected to the first conference system, acquire, for a first conference participant other than a speaking object in the first conference system, first audio of the speaking object based on a first electronic device that is closest to the first conference participant among the N electronic devices.

[0140] The target device determination module 702 is configured to, in a case where a volume of the first audio is less than a first volume threshold corresponding to the first conference participant, determine a target electronic device corresponding to the first conference participant from the N electronic devices.

[0141] The second sound acquisition module 703 is configured to acquire second audio of the speaking object based on the target electronic device.

[0142] The playing module 704 is configured to play the second audio according to an audio parameter based on the target electronic device, the audio parameter including a volume value, and the volume value being greater than or equal to the first volume threshold.

[0143] In some embodiments, the device can further include a threshold setting module (not shown in the figure).

[0144] The threshold setting module can include:

[0145] The hearing test submodule is configured to, in a case where a volume of the speech sound is less than a first volume threshold corresponding to the first conference participant, perform a hearing test on the first conference participant before determining the target electronic device corresponding to the first conference participant from the N electronic devices, and determine the first volume corresponding to the first conference participant based on a test result.

[0146] The scene information acquisition submodule is configured to acquire scene information of a conference site where the first conference participant is located.

[0147] The to-be-enhanced volume determination sub-module is configured to determine the to-be-enhanced volume corresponding to the first conference participant according to the scene information.

[0148] The threshold determination sub-module is configured to determine the first volume threshold corresponding to the first conference participant according to the first volume and the to-be-enhanced volume.

[0149] The scene information includes at least one of the following: the number of people in the conference site, and information of a basic device in the conference site, wherein the information of the basic device includes a sound volume generated by the basic device when the basic device is running, and a distance between the first conference participant and the basic device.

[0150] The to-be-enhanced volume determination sub-module is configured to:

[0151] determine the first interference volume according to the number of people in the conference site;

[0152] determine a second interference volume of the basic device to the first conference participant according to the distance between the first conference participant and the basic device and the sound volume generated by the basic device when the basic device is running; and

[0153] determine the to-be-enhanced volume corresponding to the first conference participant according to at least one of the first interference volume and the second interference volume.

[0154] The target device determination module 702 is configured to:

[0155] obtain pose information of the first conference participant and layout information of a conference site where the first conference participant is located;

[0156] determine a target direction relative to the first conference participant according to the pose information and the layout information, wherein the target direction includes a direction of a speaking participant or a conference content display device;

[0157] determine, from N electronic devices, an electronic device located in the target direction and closest to the first conference participant as a target electronic device corresponding to the first conference participant.

[0158] The second sound acquisition module 703 is configured to:

[0159] acquire, based on the target electronic device, a second audio of the speaking participant collected by a second electronic device, the second electronic device being a second electronic device closest to the speaking participant among the N electronic devices.

[0160] The apparatus can further include a target volume setting module (not shown in the figure);

[0161] The target volume setting module can include:

[0162] The distance determining sub-module is configured to acquire a distance between the target electronic device and the first conference participant before the target electronic device plays the second audio according to the audio parameter;

[0163] The minimum volume determining sub-module is configured to determine a second volume threshold of the target electronic device according to the distance between the target electronic device and the first conference participant and the volume threshold;

[0164] The volume interval determining sub-module is configured to determine a volume interval according to the second volume threshold, wherein a lower limit value of the volume interval is greater than or equal to the second volume threshold;

[0165] The target volume determining sub-module is configured to determine a volume value from the volume interval as a volume value included in the audio parameter.

[0166] The target volume determining sub-module is configured to:

[0167] determine a second conference participant who is closest to the first conference participant in the conference site;

[0168] determine a probability value of interfering the second conference participant when the target electronic device plays sound according to the candidate volume value, and determine a candidate volume value as an alternative volume value in the volume interval;

[0169] determine the candidate volume value as the target volume when the probability value of interfering the second conference participant when the target electronic device plays sound according to the candidate volume value is less than a preset probability threshold;

[0170] decrease the candidate volume value when the probability value of interfering the second conference participant when the target electronic device plays sound according to the candidate volume value is greater than or equal to the probability threshold, wherein the decreased candidate volume value is in the volume interval until the probability value of interfering the second conference participant when the target electronic device plays sound according to the candidate volume value is less than the probability threshold.

[0171] The apparatus further includes a threshold adjusting module (not shown in the figure) configured to:

[0172] determine whether the first conference participant and the speaker are in a back-to-back relationship before determining the target electronic device corresponding to the first conference participant from the N electronic devices when the volume of the first audio is less than the first volume threshold corresponding to the first conference participant;

[0173] increase the first volume threshold corresponding to the first conference participant when the first conference participant and the speaker are in the back-to-back relationship;

[0174] Correspondingly, the target device determining module 702 is configured to:

[0175] In a case where the volume of the first audio is less than the first volume threshold corresponding to the first meeting participant after being increased, a target electronic device corresponding to the first meeting participant is determined from the N electronic devices.

[0176] The apparatus further includes a target speaker determination module (not shown in the figure) configured to:

[0177] In a case where M speakers exist in the conference simultaneously before the first electronic device closest to the first meeting participant in the N electronic devices collects the first audio of the speaker, the identification information corresponding to the M speakers is displayed in the first electronic device closest to the first meeting participant, where M is a positive integer greater than 1.

[0178] A first input to the identification information corresponding to the first speaker is received, the first speaker being any one of the M speakers.

[0179] In response to the first input, the first speaker is taken as a target speaker.

[0180] The first sound acquisition module is configured to:

[0181] Based on the first electronic device closest to the first meeting participant in the N electronic devices, the sound of the target speaker is collected.

[0182] The voice information processing apparatus provided by the embodiments of the present application, in a case where N electronic devices are connected to a first conference system, for a first meeting participant other than a speaker in the first conference system, based on the first electronic device closest to the first meeting participant in the N electronic devices, the first audio of the speaker is collected, and in a case where the volume of the first audio is less than the first volume threshold corresponding to the first meeting participant after being increased, a target electronic device corresponding to the first meeting participant is determined from the N electronic devices, the second audio of the speaker is obtained based on the target electronic device, and the second audio is played based on the target electronic device according to the audio parameter, the audio parameter including the volume value, the volume value being greater than or equal to the first volume threshold. According to the embodiments, in a case where the first meeting participant cannot hear the speaking sound of the speaker clearly, the speaking sound of the speaker can be played based on the target electronic device at a target volume that the first meeting participant can hear clearly, so that the first meeting participant can obtain the speaking content of the speaker based on the sound played by the target electronic device, and the problem that the meeting participant cannot hear the speaking content of the speaker is solved, and the meeting participant can further obtain complete conference information in the conference.

[0183] The voice information processing apparatus in the embodiments of the present application can be an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The electronic device can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a cash register, or a self-service machine, etc. The embodiments of the present application are not limited in this regard.

[0184] The voice information processing apparatus in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an ios operating system, or other possible operating systems. The embodiments of the present application are not limited in this regard.

[0185] The voice information processing apparatus provided in the embodiments of the present application can implement the method embodiments, and each process of the method embodiments is not repeated here. Figures 3 to 6 The voice information processing apparatus provided in the embodiments of the present application can implement the method embodiments, and each process of the method embodiments is not repeated here.

[0186] Optionally, as shown in Figure 8 The embodiments of the present application also provide an electronic device 800, which includes a processor 801 and a memory 802. The memory 802 stores programs or instructions that can run on the processor 801. When the programs or instructions are executed by the processor 801, each step of the voice information processing method embodiments described above is implemented, and the same technical effects are achieved. Each step of the voice information processing method embodiments described above is not repeated here.

[0187] It should be noted that the electronic device in the embodiments of the present application includes the mobile electronic device and the non-mobile electronic device described above.

[0188] Figure 9 A hardware structure diagram of an electronic device is provided to implement the embodiments of the present application.

[0189] The electronic device 900 includes, but is not limited to, a radio frequency unit 901, a network module 902, an audio output unit 903, an input unit 904, a sensor 905, a display unit 906, a user input unit 907, an interface unit 908, a memory 909, and a processor 910, etc.

[0190] Those skilled in the art can understand that the electronic device 900 can also include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 910 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management through the power management system. Figure 9 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than the figure, or combine certain components, or different component arrangements, which are not described here.

[0191] The network module 902 is configured to be connected to other electronic devices.

[0192] The processor 910 is configured to, for a first conference participant other than a speaking participant, collect first audio of the speaking participant based on a first electronic device closest to the first conference participant among N electronic devices connected to the first conference participant, determine a target electronic device corresponding to the first conference participant from the N electronic devices in a case where a volume of the first audio is less than a first volume threshold corresponding to the first conference participant, acquire second audio of the speaking participant based on the target electronic device, and play the second audio based on the target electronic device according to an audio parameter, the audio parameter including a volume value, the volume value being greater than or equal to the first volume threshold.

[0193] In a case where the first conference participant cannot hear the speaking sound of the speaking participant, the target electronic device plays the speaking sound of the speaking participant at a target volume that the first conference participant can hear, so that the first conference participant can obtain the speaking content of the speaking participant based on the second audio played by the target electronic device, thereby solving the problem that the conference participant cannot hear the speaking content of the speaking participant, and further ensuring that the conference participant can obtain complete conference information in the conference.

[0194] It should be understood that in the embodiments of the present application, the input unit 904 can include a graphics processor (GPU) 9041 and a microphone 9042. The graphics processor 9041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 906 can include a display panel 9061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 907 includes at least one of a touch panel 9071 and other input devices 9072. The touch panel 9071 is also referred to as a touch screen. The touch panel 9071 can include two parts of a touch detection device and a touch controller. The other input devices 9072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, and the like), a trackball, a mouse, a joystick, and the like, which will not be described here.

[0195] The memory 909 can be used to store software programs and various data. The memory 909 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, and the like), and the like. In addition, the memory 909 can include a volatile memory or a non-volatile memory, or the memory 909 can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PRO), an erasable programmable read-only memory (EPRO), an electrically erasable programmable read-only memory (EEPRO), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRA), a dynamic random access memory (DRA), a synchronous dynamic random access memory (SDRA), a double data rate synchronous dynamic random access memory (DDRSDRA), an enhanced synchronous dynamic random access memory (ESDRA), a synch link dynamic random access memory (SLDRA), and a direct memory bus random access memory (DRRA). The memory 909 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0196] The processor 910 can include one or more processing units; optionally, the processor 910 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes a wireless communication signal, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 910.

[0197] The embodiment of the present application further provides a readable storage medium, and the readable storage medium stores a program or instructions, the program or instructions are executed by a processor to realize each process of the above-mentioned voice information processing method embodiment, and the same technical effects can be achieved, to avoid repetition, which will not be described here.

[0198] The processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes a computer readable storage medium, such as a computer readable memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0199] The embodiment of the present application further provides a chip, and the chip includes a processor and a communication interface, the communication interface is coupled with the processor, and the processor is used to run a program or instructions to realize each process of the above-mentioned voice information processing method embodiment, and the same technical effects can be achieved, to avoid repetition, which will not be described here.

[0200] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system-level chip, a system chip, a chip system, or a system-on-chip chip, etc.

[0201] The embodiment of the present application provides a computer program product, and the program product is stored in a storage medium, and the program product is executed by at least one processor to realize each process of the above-mentioned voice information processing method embodiment, and the same technical effects can be achieved, to avoid repetition, which will not be described here.

[0202] It should be noted that, in the present document, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without further constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element. Furthermore, it is to be understood that the method and apparatus of the present embodiments can be carried out by more than one process, method, article, or apparatus either simultaneously, concurrently, or in reverse order. In other words, the steps described herein can be carried out in any order as would be appreciated by those skilled in the art, such as carrying out the described methods in a different order, removing, or adding steps, or combining steps, and the like. In addition, features described in relation to one example can be combined in a combination of examples.

[0203] From the above description of the embodiments, it is apparent that the above-described method of the embodiments can be implemented by means of software and the necessary universal hardware platform, of course, can also be implemented by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, an optical disk), and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the method described in the various embodiments of the present application.

[0204] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative, rather than limiting, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.

Claims

1. A voice information processing method characterized by comprising: The method comprises the following steps: In the case that N electronic devices are connected to the first conference system, for a first conference participant in the first conference system except for a speaking object, a first audio of the speaking object is collected based on a first electronic device in the N electronic devices that is closest to the first conference participant, wherein N is a positive integer greater than 1; In the case that the volume of the first audio is less than a first volume threshold corresponding to the first conference participant, a target electronic device corresponding to the first conference participant is determined from the N electronic devices; Second audio of the speaking object is obtained based on the target electronic device; The second audio is played based on the target electronic device according to an audio parameter, the audio parameter comprising a volume value, the volume value being greater than or equal to the first volume threshold; The method further comprises the following steps before the step of determining the target electronic device corresponding to the first conference participant from the N electronic devices in the case that the volume of the first audio is less than the first volume threshold corresponding to the first conference participant: A hearing test is performed on the first conference participant, and a first volume corresponding to the first conference participant is determined based on a test result; Scene information of a conference site where the first conference participant is located is collected; A to-be-enhanced volume corresponding to the first conference participant is determined according to the scene information; 2. The method of claim 1, wherein, A first volume threshold corresponding to the first conference participant is determined according to the first volume and the to-be-enhanced volume. The scene information comprises at least one of the following: a number of people in the conference site, and information of a basic device in the conference site, wherein the information of the basic device comprises a sound volume generated by the basic device when the basic device is running, and a distance between the first conference participant and the basic device; the step of determining the to-be-enhanced volume corresponding to the first conference participant according to the scene information comprises the following steps: A first interference volume is determined according to the number of people in the conference site; A second interference volume of the first conference participant caused by the basic device is determined according to the distance between the first conference participant and the basic device and the sound volume generated by the basic device when the basic device is running; The to-be-enhanced volume corresponding to the first conference participant is determined according to at least one of the first interference volume and the second interference volume.

3. The method of claim 2, wherein, The step of obtaining the second audio of the speaking object based on the target electronic device comprises the following step: The second audio of the speaking object collected by a second electronic device is received based on the target electronic device, the second electronic device being an electronic device in the N electronic devices that is closest to the speaking object. ​ ​ 4. The method according to any one of claims 1 to 3, characterized in that, ​ ​ 5. The method according to any one of claims 1 to 3, characterized in that, The method further includes, before the target electronic device plays the second audio according to the audio parameter: obtaining a distance between the target electronic device and the first conference participant; determining a second volume threshold value of the target electronic device according to the distance between the target electronic device and the first conference participant and the volume threshold value; determining a volume interval according to the second volume threshold value, a lower limit value of the volume interval being greater than or equal to the second volume threshold value; determining a volume value from the volume interval as a volume value included in the audio parameter.

6. The method of claim 5, wherein, The method further includes: determining a second conference participant having a minimum distance to the first conference participant in the conference site; determining a probability value of interfering the second conference participant when the target electronic device plays sound according to a candidate volume value in the volume interval according to a distance between the target electronic device and the second conference participant, the candidate volume value being any volume value in the volume interval; determining the candidate volume value as a target volume value when the probability value of interfering the second conference participant when the target electronic device plays sound according to the candidate volume value is less than a preset probability threshold value; decreasing the candidate volume value until the probability value of interfering the second conference participant when the target electronic device plays sound according to the candidate volume value is less than the probability threshold value, when the probability value of interfering the second conference participant when the target electronic device plays sound according to the candidate volume value is greater than or equal to the probability threshold value.

7. The method according to any one of claims 1 to 3, characterized in that, The method further includes, before determining the target electronic device corresponding to the first conference participant from the N electronic devices when the volume of the first audio is less than the first volume threshold value corresponding to the first conference participant: determining whether the first conference participant and the speaker are in a back-to-back relationship; increasing the first volume threshold value corresponding to the first conference participant when the first conference participant and the speaker are in a back-to-back relationship. The method further includes, before determining the target electronic device corresponding to the first conference participant from the N electronic devices when the volume of the first audio is less than the first volume threshold value corresponding to the first conference participant: determining the target electronic device corresponding to the first conference participant from the N electronic devices when the volume of the first audio is less than the increased first volume threshold value corresponding to the first conference participant.

8. The method according to any one of claims 1 to 3, characterized in that, The method further includes, before the first electronic device having a minimum distance to the first conference participant in the N electronic devices collects the first audio of the speaker: displaying identification information corresponding to M speakers in the first electronic device having a minimum distance to the first conference participant when M speakers exist in the conference simultaneously, M being a positive integer greater than 1. receive a first input of identification information corresponding to a first speaking object, the first speaking object being any one of the M speaking objects; determine the first speaking object as a target speaking object in response to the first input; the first audio of the speaking object is collected by a first electronic device with the smallest distance to the first participant among N electronic devices, including: the first audio of the target speaking object is collected by a first electronic device with the smallest distance to the first participant among N electronic devices.

9. A voice information processing apparatus characterized by comprising: including: The first sound acquisition module is used to collect the first audio of the speaking object based on the first electronic device with the smallest distance to the first participant among N electronic devices for the first participant in the first conference system except the speaking object in the case of N electronic devices connected to the first conference system, wherein N is a positive integer greater than 1; The target device determination module is used to determine a target electronic device corresponding to the first participant from N electronic devices in the case that the volume of the first audio is less than the first volume threshold corresponding to the first participant; The second sound acquisition module is used to acquire the second audio of the speaking object based on the target electronic device; The playing module is used to play the second audio according to the audio parameters based on the target electronic device, the audio parameters including the volume value, and the volume value being greater than or equal to the first volume threshold; The target device determination module is used to: acquire the pose information of the first participant and the layout information of the conference site where the first participant is located; determine the direction of the target object relative to the first participant according to the pose information and the layout information, wherein the target object includes the speaking object or the conference content display device; determine the electronic device with the smallest distance to the first participant among the N electronic devices located in the direction as the target electronic device corresponding to the first participant.

10. An electronic device, comprising: The electronic device includes a processor and a memory, the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the voice information processing method in any one of claims 1-8.

11. A readable storage medium, characterized by, The readable storage medium stores computer program instructions, and the computer program instructions are executed by the processor to implement the voice information processing method in any one of claims 1-8.

Citation Information

Patent Citations

  • Remote multiparty conference volume adjusting system and method

    CN103546109A

  • Video conference volume adjusting method and device, terminal device and storage medium

    CN112511786A

  • Conference earphone intelligent regulation and control management system based on scene analysis management and control

    CN114422916A