Voice processing method and device, equipment and storage medium
By moving the target object identifier in the voice communication interface to adjust the sound volume, the problem of inflexible sound volume adjustment in the prior art is solved, and information acquisition efficiency and voice communication experience are improved.
Patent Information
- Application Number
- CN202410095199.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-25
AI Technical Summary
When multiple objects communicate with voice, the flexibility of sound volume adjustment in the prior art is low, making it difficult for the object to clearly hear the required content in a noisy environment, affecting the efficiency of information acquisition and communication experience.
By moving the target object identifier in the voice communication interface, adjusting the voice playback volume of the target object based on the movement of the object identifier, realizing personalized sound volume control.
It improves the flexibility of sound volume adjustment and information acquisition efficiency, helps the subjects clearly hear the required content in noisy environments, and improves the voice communication experience.
Smart Images

Figure CN120378534A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technologies, specifically to the field of communication technologies, and particularly to a voice processing method, apparatus, device, and storage medium. Background Art
[0002] With the development of Internet technologies and communication technologies, more and more applications support voice communication; the so-called voice communication can also be referred to as a voice call, which is a communication method that enables objects (users) to communicate with each other by speaking. During the process of voice communication among multiple objects, each of these multiple objects can hear in real time the sounds (i.e., voice audio data) emitted by other objects.
[0003] Currently, during the process of voice communication among multiple objects, only any one object is supported to uniformly adjust the voice playback volume of multiple objects through the physical volume control keys on the terminal (such as a mobile phone) it uses. The flexibility of volume adjustment is relatively low, and in this way, when any one object faces a noisy environment where multiple other objects are speaking simultaneously, it is difficult to clearly hear the content it wants to listen to, thereby affecting the efficiency of the object to obtain information and resulting in a poor voice communication experience for the object. Summary of the Invention
[0004] Embodiments of this application provide a voice processing method, apparatus, device, and storage medium, which can improve the flexibility of volume adjustment and the efficiency of an object to obtain information, thereby enhancing the voice communication experience of the object.
[0005] On the one hand, embodiments of this application provide a voice processing method, and the method includes:
[0006] Display a voice communication interface, where the voice communication interface includes: object identifiers of multiple objects participating in the voice communication; the multiple objects include: a first object and at least one second object;
[0007] In response to a moving operation on a target object identifier, move the target object identifier in the voice communication interface; the target object identifier refers to the object identifier of a target second object, and the target second object is any one of the second objects participating in the voice communication;
[0008] Based on the moving situation of the target object identifier relative to the object identifier of the first object, adjust the voice playback volume of the target second object.
[0009] On the other hand, embodiments of this application provide a voice processing apparatus, and the apparatus includes:
[0010] A display unit for displaying a voice communication interface, where the voice communication interface includes: object identifiers of multiple objects participating in the voice communication; the multiple objects include: a first object and at least one second object;
[0011] The display unit is further configured to, in response to a moving operation on a target object identifier, move the target object identifier in the voice communication interface; the target object identifier refers to the object identifier of a target second object, and the target second object is any one of the second objects participating in the voice communication;
[0012] A processing unit for adjusting the voice playback volume of the target second object based on the moving condition of the target object identifier relative to the object identifier of the first object.
[0013] On the other hand, an embodiment of the present application provides a computer device, which includes an input interface and an output interface, and the computer device further includes:
[0014] A processor and a computer storage medium;
[0015] Wherein, the processor is adapted to implement one or more instructions, and the computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the above-mentioned voice processing method.
[0016] On the other hand, an embodiment of the present application provides a computer storage medium, which stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by a processor to perform the above-mentioned voice processing method.
[0017] On the other hand, an embodiment of the present application provides a computer program product, which includes one or more instructions; when the one or more instructions in the computer program product are executed by a processor, the above-mentioned voice processing method is implemented.
[0018] After the embodiment of the present application displays the voice communication interface, it can respond to a moving operation on the object identifier of any second object in the voice communication interface, move the corresponding object identifier in the voice communication interface, and thus adjust the voice playback volume of the corresponding second object based on the moving situation of the corresponding object identifier relative to the object identifier of the first object. It can be seen that the embodiment of the present application can support the first object to independently adjust the voice playback volume of a single second object by moving the object identifier of the single second object during the voice communication between the first object and at least one second object; this can not only improve the flexibility of volume adjustment, but also enable the first object to separately control the voice playback volume of each second object according to its actual demands in a noisy environment of a multi-person conversation, which is conducive to the first object hearing clearly the content it wants to hear, improving the information acquisition efficiency of the first object, and further enhancing the voice communication experience of the first object. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1a It is a system architecture diagram of a voice communication system provided by an embodiment of the present application;
[0021] Figure 1b It is another system architecture diagram of a voice communication system provided by an embodiment of the present application;
[0022] Figure 1c It is a schematic flowchart of a voice processing solution provided by an embodiment of the present application;
[0023] Figure 1d It is another schematic flowchart of a voice processing solution provided by an embodiment of the present application;
[0024] Figure 2 It is a schematic flowchart of a voice processing method provided by an embodiment of the present application;
[0025] Figure 3a It is a distribution schematic diagram of the object identifiers of each second object provided by an embodiment of the present application;
[0026] Figure 3b It is a state switching schematic diagram of a voice switch component provided by an embodiment of the present application;
[0027] Figure 3c It is a schematic diagram of moving a target object identifier provided by an embodiment of the present application;
[0028] Figure 3d It is a schematic diagram of another mobile target object identifier provided by an embodiment of the present application;
[0029] Figure 3e It is a logical schematic diagram of adjusting the voice playback volume of a target second object through a mobile target object identifier provided by an embodiment of the present application;
[0030] Figure 3f It is a schematic flowchart of a first object triggering a server to perform proportional calculation through a moving avatar provided by an embodiment of the present application;
[0031] Figure 4 It is a schematic flowchart of a voice processing method provided by another embodiment of the present application;
[0032] Figure 5a It is a schematic flowchart of a first object participating in a voice communication provided by an embodiment of the present application;
[0033] Figure 5b It is a schematic diagram of the relationship of multiple audio channels provided by an embodiment of the present application;
[0034] Figure 5c It is a schematic diagram of a first object dragging the avatar of a second object to adjust the volume provided by an embodiment of the present application;
[0035] Figure 5d It is a schematic diagram of displaying a voice private chat area by triggering a voice private chat prompt provided by an embodiment of the present application;
[0036] Figure 5e It is another schematic diagram of displaying a voice private chat area provided by an embodiment of the present application;
[0037] Figure 5f It is a schematic diagram of entering a voice private chat by dragging an avatar provided by an embodiment of the present application;
[0038] Figure 5g It is a schematic diagram of the interaction logic between a terminal and a server provided by an embodiment of the present application;
[0039] Figure 6 It is a schematic diagram of the structure of a voice processing device provided by an embodiment of the present application;
[0040] Figure 7 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0041] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0042] An embodiment of the present application proposes a voice processing solution to improve the flexibility of volume adjustment and the efficiency of obtaining information for objects during voice communication, thereby enhancing the voice communication experience of the objects. Among them, the general principle of this voice processing solution is as follows: During the voice communication process of multiple objects using terminals, the terminal of any object (let's assume object A) can obtain the voice audio data of the corresponding object in real time, as well as the voice audio data of other objects from other terminals. In addition, the terminal of any object (i.e., object A) can also display a voice communication interface containing the object identifiers of multiple objects, so that any object (i.e., object A) can, according to its own needs, move the object identifiers of at least one other object to trigger the terminal used by any object (i.e., object A) to adjust the playback volume of the voice audio data of the object corresponding to the moved object identifier (i.e., the voice playback volume) based on the movement of the moved object identifier relative to the object identifier of any object (i.e., object A), and then perform a mixing process based on the voice audio data with adjusted volume and the obtained voice audio data of other objects whose playback volumes are not adjusted, and render and play the mixed audio data. Among them, volume can also be referred to as sound volume, which refers to the magnitude of sound.
[0043] It should be noted that the embodiment of the present application does not limit the data interaction method between the terminals used by multiple objects. For example, the terminals used by multiple objects can establish a voice communication connection based on a short-range communication technology (such as Bluetooth technology, NFC (Near Field Communication) technology, etc.) and perform data interaction based on this short-range communication technology; in this case, the terminals used by multiple objects participating in voice communication can form a voice communication system as shown in Figure 1a Figure []. Another example is that the terminals used by multiple objects can establish a voice communication connection through a server and perform data interaction through the server; in this case, the terminals used by multiple objects participating in voice communication and the server can form a voice communication system as shown in Figure 1b Figure []. For the convenience of explanation, hereinafter, the voice communication system shown in Figure 1b Figure [] will be used as an example for illustration.
[0044] Among them, the above-mentioned terminal refers to a device with voice communication capabilities, such as a smart phone, a computer (such as a tablet computer, a notebook computer, a desktop computer, etc.), a smart wearable device (such as a smart watch, smart glasses), a smart voice interaction device, a smart home appliance (such as a smart TV), a vehicle-mounted terminal, or an aircraft, etc.; an application with voice communication functions can be installed and run in the terminal, such as a voice live broadcast application, a social application, a conference application, a multimedia playback application (such as a video playback application, a music playback application, etc.), etc. The above-mentioned server refers to a device that can establish a voice communication connection between various terminals and provide multiple services such as data storage and data forwarding for each terminal; specifically, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc., and so on.
[0045] Further, when implementing the voice processing solution proposed in the embodiments of the present application based on the Figure 1b shown voice communication system, the implementation solution of the voice processing solution can be any of the following:
[0046] Implementation solution one: The terminal performs audio processing; that is, in this implementation solution, the audio processing is mainly executed on the terminal (such as a mobile phone, a tablet, etc.) used by the object. Refer to Figure 1c shown, this implementation solution one can generally include the following steps:
[0047] 1. The terminal performs audio collection: During the voice communication of multiple objects, the terminal used by each object (which can be called the sender terminal 1, the sender terminal n, etc. at this time, n is a positive integer) can collect the voice of the corresponding object to collect the voice audio signal of the corresponding object. Taking the sender terminal 1 as an example, it can use the microphone array configured in the corresponding terminal to achieve high-quality collection of the voice of the corresponding object; and according to the needs of the application scenario, parameters such as the sampling rate, bit depth, and number of channels can be adjusted to meet different sound quality requirements.
[0048] 2. The terminal performs audio encoding: For example, after the sender terminal 1 collects the voice audio signal of the corresponding object, it can use advanced encoding algorithms such as AAC (Advanced Audio Coding), Opus (a voice coding format), etc. to perform audio encoding on the collected voice audio signal to achieve efficient compression of the digitized voice audio signal and obtain the encoded voice audio data (i.e., the voice audio data of the corresponding object). Optionally, the sender terminal 1 can also dynamically adjust the encoding bit rate according to the network conditions to adapt to different network environments.
[0049] 3. The terminal performs audio transmission: For example, after the sender terminal 1 obtains the encoded voice audio data through audio encoding, it can send the voice audio data to the server through the network. Specifically, the sender terminal 1 can use the Real-time Transport Protocol (RTP) or other real-time communication technologies (Web Real-Time Communications, WebRTC) to transmit the voice audio data to ensure low-latency transmission of the voice audio data. Moreover, the sender terminal 1 can also use technologies such as packet loss retransmission and forward error correction to improve the stability of audio transmission.
[0050] 4. The server receives and forwards the audio: The server can receive the voice audio data transmitted by multiple terminals such as the sender terminal 1 and the sender terminal n. One terminal corresponds to one audio channel, and the voice audio data transmitted by multiple terminals can form multi-channel audio data. The server can forward (i.e., backhaul) the multi-channel audio data to each terminal. It can be understood that the server can use the same transmission technology as in step 3 to achieve low-latency backhaul of the audio data.
[0051] 5. The terminal performs audio decoding: Any terminal (which can be called the receiver terminal at this time) can decode the multi-channel audio data after receiving it from the server to obtain the voice audio signals of multiple audio channels. Specifically, the receiver terminal can use a high-quality decoder for decoding to achieve high-fidelity restoration of the voice audio signal.
[0052] 6. The terminal obtains command words: The receiving terminal can obtain the volume gain control parameters of each object participating in the voice communication, and add the volume gain control parameters of each object to the command words of the object on the receiving terminal side. Among them, the volume gain control parameter is a parameter used to adjust the voice playback volume of the object. The process of obtaining the volume gain control parameters of each object may include: displaying a voice communication interface containing the object identifiers of multiple objects. For any other object (set as object B) except the object on the receiving terminal side (set as object I) among the multiple objects, if it is detected that the object identifier of object B is moved, the volume gain control parameter of object B can be determined according to the movement situation of the movement identifier of object B relative to the object identifier of object I, so that when the voice playback volume of object B is adjusted based on the volume gain control parameter of object B subsequently, the voice playback volume of object B can change; if it is not detected that the object identifier of object B is moved, the zero value can be used as the volume gain control parameter of object B, so that when the voice playback volume of object B is adjusted based on the volume gain control parameter of object B subsequently, the voice playback volume of object B can be prevented from changing. Optionally, the command word may further include the mixing parameters of each object. The mixing parameter is a parameter used to make the voice audio data of the object have a sense of space (azimuth) during playback; the mixing parameter of any object can be determined according to the positional relationship (or azimuth relationship) between the object identifier of the corresponding object and the object identifier of object I on the receiving terminal side.
[0053] 7. The terminal performs audio mixing and gain control: The receiving terminal can perform gain control (i.e., adjustment of the voice playback volume) and mixing processing (i.e., processing of mixing the voice audio signals of multiple channels into one audio data) on the voice audio signals of multiple channels using a linear mixing algorithm according to the command words of the corresponding objects, and obtain the mixed independent audio signal. It can be seen that the embodiments of the present application can adjust parameters such as the playback volume of the voice audio signals of different audio channels according to the object settings to provide personalized audio performance.
[0054] 8. The terminal performs audio rendering and playback: After obtaining the mixed independent audio signal, any terminal can render the independent audio signal and then play the rendered audio data through the speaker. Specifically, any terminal can use the audio rendering engine of the device system to perform the rendering operation to realize the restoration of the audio signal.
[0055] Implementation solution two: The server performs audio processing; that is, in this implementation solution, the audio processing is mainly completed on the server side. The terminal used by the object can interact with the server through the middle platform SDK (Software Development Kit) or the background service. As shown in Figure 1d shown, this implementation solution two may generally include the following steps:
[0056] 1. The terminal performs audio acquisition: The same as step 1 in the aforementioned implementation solution 1, the terminal used by each object can use the configured microphone array to perform audio acquisition.
[0057] 2. The terminal performs audio encoding: The same as step 2 in the aforementioned implementation solution 1, an advanced encoding algorithm can be used for audio compression.
[0058] 3. The terminal performs audio transmission and command word sending: Any terminal (such as the sending terminal 1) can send the encoded voice audio data and the command words of the corresponding object to the server through the network. That is, in addition to sending voice audio data, the terminal also needs to send command words related to audio processing, such as mixing parameters, volume gain control parameters, etc.; these command words can be interacted between the terminal and the server through the middle platform SDK or the background service.
[0059] 4. The server performs audio reception and command word reception, as well as audio mixing and gain control processing: After receiving the multi-channel audio data and command words, the server performs gain control and mixing processing on each voice audio data in the multi-channel audio data according to the command words according to the linear mixing algorithm to generate an independent audio stream. That is, an audio processing engine can be used on the server side to adjust the playback volume of each voice audio data according to the command words; and the server can generate independent audio streams for different terminals according to the command words sent by different terminals to meet personalized needs.
[0060] 5. The server performs audio forwarding: After generating the independent audio stream, the server can send the independent audio stream to the specified terminal (i.e., the terminal that sent the command word for generating the independent audio stream) through the network. Specifically, the server can use the Real-Time Transport Protocol (RTP) or other real-time communication technologies to send the customized independent audio stream to the corresponding terminal.
[0061] 6. The terminal performs audio decoding, rendering and playing: The receiving terminal (i.e., the terminal that receives the independent audio stream) can decode and render the received independent audio stream, and then play the rendered audio data through the speaker. The same as step 8 in the aforementioned implementation solution 1, the receiving terminal can use the device decoder and the audio rendering engine of the device system to restore the audio signal.
[0062] It should be noted that the above only exemplarily elaborates two implementation solutions of the audio processing solution proposed in the embodiments of the present application, and is not an exhaustive list. Moreover, in practical applications, appropriate implementation solutions can be selected according to different scenarios and requirements to implement the audio processing solution proposed in the embodiments of the present application. For example, in scenarios with high requirements for audio processing performance, implementation solution two can be adopted to implement the audio processing solution. In this case, the audio processing task can be shared to the server side, thereby reducing the burden on the terminal on the object side; while in scenarios with high requirements for real-time performance, implementation solution one can be adopted to implement the audio processing solution. In this case, the audio processing task can be concentrated on the terminal on the object side to reduce the delay caused by network transmission.
[0063] Based on the above relevant description of the audio processing solution, it can be known that the audio processing solution proposed in the embodiments of the present application can support any object to adjust the voice playback volume of other objects specifically by moving the object identifiers of other objects during the process of voice communication among multiple objects; this can not only improve the flexibility of volume adjustment, but also enable any object to control the voice playback volumes of other individual objects according to its actual demands in a noisy environment of a multi-person conversation, thereby facilitating any object to clearly hear the content it wants to hear, improving the efficiency of the object to obtain information, and further enhancing the voice communication experience of the object.
[0064] It is also worth emphasizing that: in the embodiments of the present application, when related data such as user information (i.e., object information) (such as any operation performed by the object, the voice audio data of the object, etc.) is involved, when any method embodiment proposed in the embodiments of the present application is applied to a specific product or technology, these related data are collected with the permission or consent of the user, and the collection, use, and processing of the related data comply with the relevant laws, regulations, and standards of the relevant region.
[0065] Based on the above relevant description of the audio processing solution, the embodiments of the present application propose a voice processing method. This voice processing method can be executed by the target terminal (i.e., any terminal) in the above-mentioned voice communication system, or by an application with voice communication functions in the target terminal; for the convenience of elaboration, the embodiments of the present application take the target terminal executing this voice processing method as an example for illustration. Please refer to Figure 2 As shown, this voice processing method may include the following steps S201 - S203:
[0066] S201, display a voice communication interface.
[0067] In the embodiments of the present application, the voice communication scenario can be any one of the following: a voice live broadcast scenario established based on a voice live broadcast application or a multimedia playback application with a voice live broadcast function, a voice call scenario established based on a social application, a voice conference scenario established based on a conference application, and so on. Correspondingly, the voice communication interface can be any one of the following: the live broadcast interface of a voice live broadcast room, the call interface of a voice call, the conference interface of a voice conference, and so on. Among them, a voice live broadcast room can be understood as a virtual room where multiple objects can conduct voice chats through the Internet, which can include but is not limited to: a topic chat room based on a certain discussion topic, a voice dating live broadcast room, and so on; the so-called voice dating live broadcast room refers to a virtual room where objects can pair up during a voice chat to make friends.
[0068] Among them, the voice communication interface may include: the object identifiers of multiple objects participating in the voice communication; the object identifier of any object may include but is not limited to: the network nickname used by the corresponding object (i.e., the name used to identify the object in the network), the avatar used by the corresponding object (i.e., the image used to identify the object in the network). It can be understood that the voice communication in the embodiments of the present application refers to a communication method in which multiple objects communicate with each other based on voice. During the voice communication process, the real faces of each object may or may not be presented in the voice communication interface, and the embodiments of the present application do not make any limitations in this regard.
[0069] The multiple objects participating in the voice communication may include: a first object and at least one second object; the so-called first object refers to the object using the target terminal, and the second object refers to the objects other than the first object among the multiple objects. For example, if a total of 5 objects (object 1 - object 5) participate in the voice communication, and object 1 uses the target terminal for voice communication, then object 1 is the first object, and the other 4 objects (i.e., object 2 - object 5) except object 1 are all second objects; if object 2 uses the target terminal for voice communication, then object 2 is the first object, and the other 4 objects (i.e., object 1 and object 3 - object 5) except object 2 are all second objects, and so on.
[0070] In a specific implementation, the target terminal can determine the personal field of the first object in the voice communication interface, and thus display the object identifiers of each second object in the personal field of the first object. Herein, the personal field of the first object refers to the interface range (i.e., the display area) constructed with the object identifier of the first object as the center; the first object can perform interaction operations within the interface range and generate corresponding feedback, such as the first object can move the object identifier of any second object within the interface range. It should be noted that the personal field of the first object can be located at any position in the voice communication interface, such as the top, bottom, or the center of the interface, etc.; for the convenience of description, hereinafter, it will be described by taking the personal field of the first object located at the bottom of the voice communication interface as an example.
[0071] The object identifiers of each second object can be distributed and displayed around the object identifier of the first object, and the present application embodiment does not limit the distribution manner of the object identifiers of each second object. For example, using a solid circle to represent the object identifier of the first object and a dashed circle to represent the object identifier of the second object, then referring to Figure 3a as shown: the object identifiers of each second object can be randomly distributed around the object identifier of the first object, or the object identifiers of each second object can be evenly distributed on the circumference determined with the object identifier of the first object as the center of the circle, or alternatively, the object identifiers of each second object can be symmetrically distributed on both sides of the object identifier of the first object. Another example is that the object identifiers of each second object can be distributed around the object identifier of the first object according to the rule that the distance between the object identifier of any second object and the object identifier of the first object needs to be negatively correlated with the voice playback volume of the corresponding second object.
[0072] In addition, when the target terminal outputs and displays the voice communication interface, each object participating in the voice communication can have an initial voice playback volume. The initial voice playback volumes of each object can be the same or different, and this is not limited. For example, the initial voice playback volumes of each object can be set to the default volume. Another example is that the initial voice playback volume of the corresponding object can be determined according to the historical voice frequency of each object during the voice communication process, and the initial voice playback volume of any object can be positively correlated with the historical voice frequency of the corresponding object; the so-called historical voice frequency refers to: before the target terminal displays the voice communication interface, the frequency of the object making a sound in the voice communication (the ratio of the number of times of making a sound to the communication duration). For example, suppose there are 3 objects (i.e., the first object and 2 second objects) participating in the voice communication. Before the target terminal outputs the voice communication interface, the historical voice frequency of the second object a is 1 time per 10 minutes, while the historical voice frequency of the second object b is 1 time per 5 minutes. Then the initial voice playback volume of the second object a can be 50 decibels, and the initial voice playback volume of the second object b can be 80 decibels.
[0073] Optionally, the voice communication interface may further include a voice switch component 31, and when the voice communication interface is displayed, the voice switch component 31 is in a closed state; in this case, if the first object wants to have a voice conversation with other objects during the voice communication process, it needs to perform a triggering operation on the voice switch component to turn on the voice switch component 31. Correspondingly, the target terminal can respond to the triggering operation on the voice switch component 31, switch the state of the voice switch component from the closed state to the open state, and when the voice switch component is in the open state, collect the voice of the first object, so as to realize sending the voice audio data of the first object to other second objects for playing. Among them, the triggering operation on the voice switch component 31 may include any one of the following: a click operation, a pressing operation, an operation of inputting a gesture or a voice command for indicating to turn on the voice switch component; taking the triggering operation as a click operation as an example, the schematic diagram of the target terminal responding to the triggering operation and switching the state of the voice switch component 31 from the closed state to the open state can be seen in Figure 3b as shown.
[0074] S202. In response to a moving operation on the target object identifier, move the target object identifier in the voice communication interface.
[0075] Among them, the target object identifier refers to the object identifier of the target second object, and the target second object is any second object participating in the voice communication, and specifically, it can be the second object selected by the first object according to its own needs. For example, in practical applications, if the first object wants to focus on listening to the voice audio data of the second object a, then the first object can perform a moving operation on the object identifier of the second object a in order to increase the voice playback volume of the second object a; in this case, the target object identifier is the object identifier of the second object a, and the target second object is the second object a. If the first object wants to decrease the voice playback volume of the second object b, then the first object can perform a moving operation on the object identifier of the second object b; in this case, the target object identifier is the object identifier of the second object b, and the target second object is the second object b.
[0076] In a specific implementation, the first object can continuously press and drag the target object identifier with a finger or an external device (such as a mouse) to move the target object identifier; then in this specific implementation, the moving operation on the target object identifier can be: an operation of continuously pressing and dragging the target object identifier. In this case, the specific implementation method for the target terminal to move the target object identifier in the voice communication interface can be: in the voice communication interface, move the target object identifier along the trajectory of the first object dragging the target object identifier, such as Figure 3c as shown.
[0077] In another specific implementation, the target terminal can provide one or more quick movement buttons 32 for the first object on the voice communication interface, and different quick movement buttons 32 correspond to different movement directions; and for any quick movement button 32, each time a triggering operation is performed on this quick movement button 32, the object identifier can move a preset distance in the corresponding movement direction. Based on this, the first object can first select the target object identifier by finger or an external device (such as a mouse), and then perform one or more triggering operations on the movement button according to actual needs (that is, the triggering operation on the quick movement button 32 (such as a click operation, a press operation)); then in this specific implementation, the movement operation for the target object identifier can be: one or more triggering operations on the movement button, and the quick movement button 32 operated by each movement button triggering operation can be the same or different, and this is not limited. In this case, the specific implementation manner for the target terminal to move the target object identifier in the voice communication interface can be: each time a movement button triggering operation is detected, the target object identifier is controlled to move a preset distance in the movement direction corresponding to the currently operated quick movement button 32. Exemplarily, the operation process of the first object moving the target object identifier by selecting the target object identifier and through at least one quick movement button 32 can be as Figure 3d shown.
[0078] S203, adjust the voice playback volume of the target second object based on the movement of the target object identifier relative to the object identifier of the first object.
[0079] In a specific implementation, as long as the target terminal detects that the target object identifier is moved, it can trigger the execution of step S203, so that during the movement of the target object identifier, the voice playback volume of the target second object can be dynamically and real-time adjusted. Or, the target terminal can trigger the execution of step S203 after detecting that the target object identifier stops moving, so that during the movement of the target object identifier, the voice playback volume of the target second object remains unchanged, and after the target object identifier stops moving, the voice playback volume of the target second object changes.
[0080] Among them, the movement of the target object identifier relative to the object identifier of the first object may include at least one of the following: movement direction and target distance. Specifically, the movement direction here may include: the direction away from the object identifier of the first object, or the direction towards the object identifier of the first object; the target distance here refers to: after moving the target object identifier, the distance between the target object identifier and the object identifier of the first object. Exemplarily, the distance between the target object identifier and the object identifier of the first object may refer to: the length of the line connecting the center points of the object identifier of the first object and the target object identifier; and the calculation method of the length of this line may be: establish a coordinate system with the center point of the object identifier of the first object or a certain point in the voice communication interface as the coordinate origin, and determine the position coordinates of the center point of the target object identifier in this coordinate system, as well as the position coordinates of the center point of the object identifier of the first object in this coordinate system, so as to calculate the length of the line based on the two determined position coordinates.
[0081] The embodiments of the present application can adjust the voice playback volume of the target second object by constructing a realistic interaction scenario and mapping the principle that the sound volume is larger when closer and smaller when farther in the real scene, so as to meet the more personalized sound adjustment needs of the first object, create a realistic interaction presence, and enhance the communication and interaction experience of voice communication. Based on this, when the movement situation includes the movement direction, if the movement direction is the direction away from the object identifier of the first object, it can be determined that the moved target object identifier is farther away from the object identifier of the first object. In this case, the voice playback volume of the target second object is decreased; if the movement direction is the direction towards the object identifier of the first object, it can be determined that the moved target object identifier is closer to the object identifier of the first object. In this case, the voice playback volume of the target second object is increased. When the movement situation includes the target distance, assuming that before moving the target object identifier, the distance between the target object identifier and the object identifier of the first object is the initial distance, then: if the target distance is greater than the initial distance, it can be determined that the moved target object identifier is farther away from the object identifier of the first object. In this case, the voice playback volume of the target second object is decreased; if the target distance is less than the initial distance, it can be determined that the moved target object identifier is closer to the object identifier of the first object. In this case, the voice playback volume of the target second object is increased.
[0082] Further, the voice playback volume of the target second object can be specifically adjusted based on the volume gain control parameter; in this case, the determination method of the volume gain control parameter for adjusting the voice playback volume of the target second object can include: determining the distance difference between the initial distance and the target distance. Specifically, the larger value of the initial distance and the target distance can be subtracted from the smaller value to obtain the distance difference. After obtaining the distance difference, the volume gain control parameter can be determined according to the ratio between the distance difference and the initial distance; specifically, the ratio between the distance difference and the initial distance can be directly used as the volume gain control parameter, or a preset linear parameter can be used to perform a linear transformation on the ratio between the distance difference and the initial distance to obtain the volume gain control parameter.
[0083] Based on the above description, for the logic of adjusting the voice playback volume of the target second object by moving the target object identifier, an exemplary reference can be made to Figure 3e as shown in the following: Taking the line connecting the center points of the object identifier of the first object and the target object identifier, the initial length x is taken as the initial distance. If the line is shortened after moving the target object identifier, the length y of the line is less than x, and the length y is taken as the target distance. In this case, the audio channel corresponding to the target second object can be associated (determined), and the volume of the audio channel corresponding to the target second object can be controlled by the indentation ratio to reduce the voice playback volume of the target second object, and the reduction ratio of the volume is (x - y) / x. It can be understood that the method for adjusting the voice playback volume when the line is lengthened is the same and will not be elaborated here; generally speaking, in the embodiments of the present application, the ratio obtained by dividing the increased or decreased distance by the initial distance is used to reduce or increase the volume of the audio channel corresponding to the target second object, so as to adjust the voice playback volume of the target second object.
[0084] Optionally, considering that the first object adjusts the voice playback volume of the target second object by moving the target object identifier, in order to enable the first object to timely know the adjustment situation of the voice playback volume of the target second object, if the target object identifier is moved, the target terminal can also display a volume control bar around the target object identifier (such as Figure 3eThe element identified by 33 is adopted), and this volume control bar is used to dynamically prompt the current voice playback volume of the target second object (i.e., the latest voice playback volume). Specifically, the volume control bar can have multiple display styles, and different display styles are used to indicate different voice playback volumes. The target terminal can update the display style of the volume control bar in real time according to the current voice playback volume of the target second object. Further optionally, after the target object identifier stops moving, the target terminal can cancel the display of this volume control bar; specifically, the target terminal can cancel the display of this volume control bar immediately after the target object identifier stops moving, or cancel the display of this volume control bar after waiting for a preset duration (such as 3 seconds), and no limitation is made on this.
[0085] It can be understood that the ratio mentioned in the above Figure 3e related descriptions is essentially the volume gain control parameter mentioned above. And in practical applications, this ratio can be calculated by the target terminal or by the server, and no limitation is made on this. Taking the object identifier of any object including the avatar of the corresponding object as an example, Figure 3f exemplarily shows a process in which the first object triggers the server to calculate the ratio by moving the avatar. As Figure 3f shown: The target terminal can randomly arrange the avatars of each second object around the avatar of the first object, and in response to the movement operation of the first object on the avatar of the target second object (assuming object A), move the avatar of object A, and notify the server to open and associate the audio channel of object A. At this time, a volume control bar appears next to the avatar of object A. After the avatar of object A is moved a certain distance, the target terminal can notify the server to calculate the ratio of the sound increase or decrease according to the movement ratio (i.e., the ratio between the distance difference generated by the movement and the initial distance).
[0086] After the embodiment of the present application displays the voice communication interface, it can respond to the movement operation of the object identifier of any second object in the voice communication interface, move the corresponding object identifier in the voice communication interface, and thus adjust the voice playback volume of the corresponding second object based on the movement situation of the corresponding object identifier relative to the object identifier of the first object. It can be seen that the embodiment of the present application can support the first object to independently adjust the voice playback volume of a single second object by moving the object identifier of the single second object during the voice communication process between the first object and at least one second object; this can not only improve the flexibility of volume adjustment, but also enable the first object to respectively control the voice playback volumes of each second object according to its actual demands in a noisy environment of a multi-person conversation, which is beneficial for the first object to clearly hear the content it wants to hear, improve the efficiency of the first object to obtain information, and further improve the voice communication experience of the first object.
[0087] Based on the above Figure 2Based on the related description of the embodiments of the voice processing method shown, the embodiments of the present application propose another voice processing method; in the embodiments of the present application, it is still described by taking the target terminal executing the voice processing method as an example. Please refer to Figure 4 As shown, the voice processing method may include the following steps S401 - S407:
[0088] S401, play multimedia data, and determine a discussion topic associated with the multimedia data.
[0089] Among them, multimedia refers to the integration of multiple media, which may include various media forms such as text, sound, and images; correspondingly, the multimedia data mentioned in the embodiments of the present application may be music (songs) or videos. Further, the videos mentioned here may be short videos (videos with a playing duration less than the duration threshold), film and television drama videos, game videos (such as videos obtained by recording the game process), or vlogs (a type of video that records and shares personal life or themes), etc., and there is no limitation on this.
[0090] The discussion topic associated with the multimedia data may be preset by the publisher or operator of the multimedia data for the multimedia data. The embodiments of the present application do not limit the specific content of the discussion topic. For example, if the multimedia data is a video of singer A singing song B at a certain scene, the discussion topic associated with the multimedia data may be "How is singer A's live singing skills?", or the discussion topic may be "How is the melody of song B?", etc.
[0091] S402, in the playback interface of the multimedia data, display the room entrance of the voice conversation room established based on the discussion topic.
[0092] Among them, the voice conversation room can be understood as a virtual room where multiple objects can conduct voice chats through the Internet. It can be, for example, a voice live broadcast room (such as a certain topic chat room) or a multi - person voice call room in a social scenario, etc. The room entrance of the voice conversation room can be displayed at any position of the playback interface, such as the top, bottom, the upper left or upper right corner of the interface, etc. Optionally, the room entrance may include at least one of the following: the discussion topic, the object identifiers of all or part of the objects that have entered the voice conversation room. It can be understood that when the room entrance of the voice conversation room is displayed, the first object has not entered the voice conversation room, that is, the first object has not participated in voice communication. At this time, each object that has entered the voice conversation room is a second object that is participating in voice communication.
[0093] S403, if the room entrance is triggered, control the first object to enter the voice conversation room to participate in voice communication, and trigger the display of the voice communication interface.
[0094] In a specific implementation, the first object can trigger the room entrance by performing a click operation, a press operation, or other operations on the room entrance; correspondingly, if the target terminal detects that the room entrance is triggered, it can control the first object to enter the voice conversation room to participate in voice communication and trigger the display of the voice communication interface. Among them, the voice communication interface is the conversation interface of the voice conversation room; when the voice conversation room is a voice live broadcast room, the conversation interface of the voice conversation room can be referred to as the live broadcast interface. Further, the voice communication interface may include: the object identifier of the first object, the object identifiers of at least one second object (i.e., other objects in the voice conversation room except the first object), a voice switch component in a closed state, a discussion topic, and background media.
[0095] Among them, the so-called background media refers to: media presented in the form of a display background. In one implementation, the background media can be a dynamic media, which can include the multimedia data. In this case, in order not to interrupt the multimedia playback experience when the first object enters the voice conversation room from the multimedia playback scene, the target terminal can support the multimedia data to start playing based on the target media frame in the voice communication interface; the target multimedia frame mentioned here refers to: the media frame being played in the multimedia data when the room entrance is triggered. For example, if the multimedia data is a video, and when the room entrance is triggered, if the 6th video frame in the video is being played, then the target multimedia frame is the 6th video frame, and the target terminal can start playing the video in the form of a display background based on the 6th video frame in the voice communication interface. Optionally, in other implementations, the target terminal can support starting to play the multimedia data based on the first media frame of the multimedia data in the voice communication interface. Or, in other implementations, the background media can also be a static media, which can be the target media frame in the multimedia data or the cover image of the multimedia data, etc.
[0096] Optionally, during a voice communication process, when there is a speaking object among multiple objects, the target terminal may visually highlight the object identifier of the speaking object during the speaking process of the corresponding object. The object identifier of the speaking object may include a first image, which may be the avatar used by the target object; correspondingly, the visual highlighting may include at least one of the following: magnifying the display of the image, adding a bright light of a preset color to the corresponding image, and replacing the first image with a second image. The second image is an image obtained by changing the pose of the object in the first image, and the object in the second image is in a speaking pose; in a specific implementation, a neural network model built based on AI (Artificial Intelligence) technology may be called to change the pose of the object in the first image to obtain the second image. It can be seen that the embodiments of the present application support the following voice interaction modes: the avatar (i.e., the first image) of the object making a sound in the voice conversation room becomes larger and is accompanied by a bright light of a preset color; or, the avatar (i.e., the first image) of the object making a sound in the voice conversation room is replaced with a second image, and the second image becomes larger and is accompanied by a bright light of a preset color, and so on.
[0097] Based on the above description, taking the multimedia data as a short video and the object identifier of any object including an avatar as an example, an implementation process of the first object participating in voice communication can be exemplarily referred to Figure 5a as shown: The target terminal first exposes the room entrance 51 of the voice conversation room (taking the topic chat room as an example) established based on the discussion topic 50 in the short video playback interface. The first object can click on the room entrance 51 to enter the voice conversation room from the short video playback scene. At this time, the target terminal may display a voice communication interface (i.e., the conversation interface of the voice conversation room). The basic form of this voice communication interface may be a background media (which may not be set), plus the avatar 52 of the first object participating in the voice communication at the bottom, the avatars 53 of each second object, and the voice switch component 54. The first object can click on the voice switch component 54 to turn on the voice switch component 54 (i.e., go on the air). When the first object makes a sound by speaking, the avatar 52 (i.e., the first image) of the first object becomes larger and is replaced with a second image 55.
[0098] It should be noted that in the embodiments of the present application, as Figure 5bAs shown: Any object entering the voice conversation room will establish an independent audio channel, and the audio channels of each object are included in the room's public audio channel (this channel can be used to play background media), so that when the first object adjusts the volume of the voice conversation room through the physical volume control keys or volume adjustment gestures of the terminal, etc., it can affect the voice playback volume of each object in the voice conversation room. For example, if the first object adjusts the volume of the voice conversation room from the initial 30 decibels to 20 decibels, then the maximum voice playback volume of each object in the voice conversation room is adjusted to 20 decibels, that is, when adjusting the voice playback volume of a single object, the voice playback volume of this single object cannot exceed the volume of the voice conversation room; that is to say, after the volume of the voice conversation room is adjusted from 30 decibels to 20 decibels, the voice playback volume of any object can be adjusted between 0 - 20 decibels, and the maximum can be adjusted to 20 decibels.
[0099] S404, in response to a movement operation for the target object identifier, move the target object identifier in the voice communication interface, where the target object identifier refers to the object identifier of the target second object.
[0100] S405, based on the movement situation of the target object identifier relative to the object identifier of the first object, adjust the voice playback volume of the target second object.
[0101] It can be understood that the specific implementation manners of steps S404 - S405 can refer to the relevant step descriptions in the foregoing Figure 2 shown method embodiments, and will not be elaborated here. Refer to Figure 5c As shown, when the object identifier is an avatar, through steps S404 - S405, the first object can support dragging the avatar of any second object in the voice conversation room. The closer the dragged avatar of the second object is to its own avatar, the louder the voice of the other party (i.e., the greater the voice playback volume), and the farther the dragged avatar of the second object is from its own avatar, the smaller the voice of the other party (i.e., the smaller the voice playback volume); and, after the first object performs the dragging action, a volume control bar can appear beside the dragged avatar to reflect the voice size (i.e., the voice playback volume) of the corresponding object, so as to facilitate the first object to control the volume.
[0102] It should be emphasized that in the embodiment of the present application, when supporting the first object to adjust the voice playback volume of the corresponding second object by moving the object identifier of any second object, the adjustment control of this volume only takes effect on the first object, that is, only the first object hears the adjusted voice playback volume of the corresponding second object, and other second objects still hear the voice playback volume before adjustment of the corresponding second object, thus not affecting the public sound field of the voice conversation room.
[0103] In addition, the embodiments of the present application can also provide a voice private chat mode to support the establishment of a voice private chat room between the first object and any target guest state in the public domain scenario for two-party communication and interaction. For the convenience of description, the embodiments of the present application take the target second object as an example to describe the implementation logic of the voice private chat mode, and the specific implementation logic can be referred to the following steps S406 - S407.
[0104] S406. In response to a voice private chat trigger operation for the target second object, send a voice private chat prompt to the target second object.
[0105] Among them, the voice private chat trigger operation may include: an operation of moving the target object identifier of the target second object to the target display position where the object identifier of the first object is located (that is, an operation of moving the target object identifier so that the target object identifier coincides with the object identifier of the first object). Or, the voice private chat trigger operation may include: an operation of continuously pressing the target object identifier of the target second object so that the pressing duration is greater than the duration threshold.
[0106] In a specific implementation, the target terminal can send a voice private chat prompt to the terminal used by the target second object through the server, so that the corresponding terminal outputs the voice private chat prompt in the voice communication interface on the side of the target second object, so that after seeing the voice private chat prompt, the target second object can decide whether to have a voice private chat with the first object. Exemplarily, Figure 5d The left figure below presents a display form of the voice private chat prompt 56, but does not limit the display form of the voice private chat prompt; for example, in other embodiments, the voice private chat prompt can also be displayed in a pop-up window and float on the voice communication interface on the side of the target second object. Further, if the target second object agrees to have a voice private chat with the first object after seeing the voice private chat prompt, a confirmation operation (such as a click operation, a press operation, etc.) can be performed on the voice private chat prompt. At this time, the terminal used by the target second object can display a voice private chat area 57 in the voice communication interface on the side of the target second object, so that the target second object can have a voice private chat with the first object through the voice private chat area, as Figure 5d shown.
[0107] S407. After the target second object agrees to have a voice private chat with the first object, display a voice private chat area so that the first object can have a voice private chat with the target second object through the voice private chat area.
[0108] In a specific implementation, after the target second object agrees to have a voice private chat with the first object, the target terminal may request the server to establish a private room (i.e., a private voice room) between the target second object and the first object on the basis of the existing voice conversation room, and the established private room does not deviate from the original scenario. The private room can be connected to the original scenario through sound and background media, so that when the first object and the target second object are having a voice private chat, the voice audio data and background media playback of other objects in the existing voice conversation room can be maintained. After the server establishes the private room, the target terminal can display the voice private chat area; specifically, the target terminal can display the voice private chat area 57 at the target display position (i.e., the display position where the object identifier of the first object is located) in the voice communication interface, as Figure 5e shown. Alternatively, the voice private chat area can also be displayed at other positions in the voice communication interface; or, a new interface can also be output, and the voice private chat area can be displayed in the new interface, which is not limited herein. The voice private chat area includes: the target object identifier of the target second object and the object identifier of the first object.
[0109] Optionally, during the process of the first object having a voice private chat with the target second object through the voice private chat area, in order to ensure that the first object can clearly hear the voice audio data of the target second object, the target terminal can also maintain the playback of other media data involved in the voice communication interface and reduce the playback volume of other media data, so that the reduced playback volume of other media data is less than the voice playback volume of the first object and less than the voice playback volume of the second object. Among them, other media data includes at least one of the following: the background media in the voice communication interface, and the voice audio data emitted by other second objects except the target second object. Further, if the voice private chat area is closed, the target terminal can also increase the playback volume of other media data to the historical playback volume; where the historical playback volume refers to: the playback volume that other media data had before the first object and the target second object had a voice private chat.
[0110] It can be seen that when the object identifier includes an avatar, the embodiment of the present application can support dragging the avatar of other objects to its own avatar (as Figure 5f shown) to enter the private chat mode (i.e., the whisper mode), establish an independent and private voice room on the basis of the public voice field, and conduct voice chat interactions. At this time, the background media and the voices of other users are still played, but the volume is reduced. In this case, the interaction logic between the target terminal and the server can be referred to together Figure 5gAs shown: The first object can drag the avatar of object A to overlap with its own avatar, thereby triggering the target terminal to send a private chat request to the terminal used by object A through the server, causing the terminal used by object A to display a voice private chat prompt based on this private chat request. After object A receives this voice private chat prompt and agrees to have a voice private chat with the first object, object A and the first object can enter the private chat mode. At this time, the server can establish a room between these two objects, namely the first object and object A, associate the audio channels of these two objects, and notify the target terminal and the terminal used by object A to control the background media and the playback volume of other objects to become smaller. Practice has shown that by supporting the first object to slide the avatars of other objects, it is possible to enter a voice private chat room in the current scenario. This not only avoids the first object from leaving and interrupting the current voice communication scenario but also eliminates the need for the first object to switch scenarios, resulting in a smoother operation experience.
[0111] Optionally, if the first object wishes to exit the private chat session with the target second object, the target terminal can also respond to the close operation on the voice private chat area, close the voice private chat area at the target display position; and display the object identifier of the first object at the target display position, and move the target object identifier of the target second object from the target display position to the historical display position of the target object identifier. The historical display position of the target object identifier refers to the display position where the target object identifier is located when a voice private chat trigger operation is detected. During the process of displaying the voice private chat area in the voice communication interface, the voice communication interface also includes: a close component of the voice private chat area (such as Figure 5e the component identified by 58 in the figure); in this case, the close operation can include: a trigger operation on the close component. It can be understood that only one implementation method of the close operation is mentioned here, not an exhaustive list; for example, the close operation can also be an operation of inputting a gesture or voice command for indicating the closure of the voice private chat area, or an operation of clicking on the blank area of the screen, etc.
[0112] In the embodiments of the present application, in a voice conversation room, avatars of other objects in the live broadcast room can be aggregated based on personal avatars. Dragging the avatars of other objects can control the voice channels (i.e., audio channels) of these objects in the live broadcast room from the main perspective (i.e., oneself). By controlling the distance between other objects and oneself, the voice volume of each other object can be adjusted. Moreover, it is supported that the first object can enter the private chat mode by dragging the avatars of other objects to its own avatar, and a private chat room can be established in the public domain scenario for two-way communication and interaction. The above points enable the objects in the voice conversation room to control the voice volume of other objects according to their own demands in a noisy scenario of multi-person conversation, which is more conducive to the efficiency of obtaining information for themselves. And it constructs a realistic interaction scenario, mapping the principle of objects being larger when closer and smaller when farther in the real field, vividly and intuitively displaying the UI (User Interface) interface for easy understanding and operation by objects. And it enables faster private chat and thus precipitates the acquaintance relationship chain. The above points can all promote the interaction and communication in the voice conversation room.
[0113] Based on the description of the embodiments of the above voice processing method, embodiments of the present application also disclose a voice processing device; the voice processing device can be a computer program (including one or more instructions) running on a computer device, and this voice processing device can execute Figure 2 or Figure 4 each step in the method flow shown. Please refer to Figure 6 , the voice processing device can run the following units:
[0114] A display unit 601, configured to display a voice communication interface, where the voice communication interface includes: object identifiers of multiple objects participating in the voice communication; the multiple objects include: a first object and at least one second object;
[0115] The display unit 601 is further configured to, in response to a moving operation on a target object identifier, move the target object identifier in the voice communication interface; the target object identifier refers to the object identifier of a target second object, and the target second object is any second object participating in the voice communication;
[0116] A processing unit 602, configured to adjust the voice playback volume of the target second object based on the moving condition of the target object identifier relative to the object identifier of the first object.
[0117] In an implementation manner, before displaying the voice communication interface, the processing unit 602 can also be used to:
[0118] Play multimedia data and determine a discussion topic associated with the multimedia data;
[0119] In the playback interface of the multimedia data, display an entrance to a voice conversation room established based on the discussion topic;
[0120] If the room entrance is triggered, control the first object to enter the voice conversation room to participate in voice communication, and trigger the display of the voice communication interface; wherein, the voice communication interface is the conversation interface of the voice conversation room.
[0121] In another implementation, the voice communication interface further includes a voice switch component, and when the voice communication interface is displayed, the voice switch component is in a closed state; the processing unit 602 can also be used for:
[0122] In response to a trigger operation on the voice switch component, switch the state of the voice switch component from the closed state to the open state;
[0123] When the voice switch component is in the open state, perform voice collection on the first object.
[0124] In another implementation, the processing unit 602 can also be used for:
[0125] If the target object identifier is moved, display a volume control bar around the target object identifier, and the volume control bar is used to dynamically prompt the current voice playback volume of the target second object;
[0126] After the target object identifier stops moving, cancel the display of the volume control bar.
[0127] In another implementation, the processing unit 602 can also be used for:
[0128] During voice communication, when there is a speaking object among the multiple objects, during the speaking process of the speaking object, visually highlight the object identifier of the speaking object;
[0129] Wherein, the object identifier of the speaking object includes a first image; the visual highlighting includes at least one of the following: magnifying the display of the image, adding a bright light of a preset color to the corresponding image, and replacing the first image with a second image;
[0130] Wherein, the second image is an image obtained by changing the posture of the object in the first image, and the object in the second image is in a speaking posture.
[0131] In another implementation, the processing unit 602 can also be used for:
[0132] In response to a voice private chat trigger operation on the target second object, send a voice private chat prompt to the target second object; wherein, the voice private chat trigger operation includes: an operation of moving the target object identifier of the target second object to the target display position where the object identifier of the first object is located.
[0133] After the target second object agrees to have a voice private chat with the first object, a voice private chat area is displayed, so that the first object can have a voice private chat with the target second object through the voice private chat area.
[0134] In another implementation, the voice private chat area is displayed at a target display position in the voice communication interface, and the voice private chat area includes: the target object identifier of the target second object and the object identifier of the first object;
[0135] The processing unit 602 can also be used for:
[0136] In response to a close operation on the voice private chat area, close the voice private chat area at the target display position;
[0137] Display the object identifier of the first object at the target display position, and move the target object identifier of the target second object from the target display position to the historical display position of the target object identifier;
[0138] Wherein, the historical display position of the target object identifier refers to: the display position where the target object identifier is located when the voice private chat trigger operation is detected.
[0139] In another implementation, the voice private chat area is displayed in the voice communication interface; the processing unit 602 can also be used for:
[0140] During the process that the first object has a voice private chat with the target second object through the voice private chat area, keep playing the other media data involved in the voice communication interface, and reduce the playing volume of the other media data;
[0141] Wherein, the other media data includes at least one of the following: the background media in the voice communication interface, and the voice audio data sent by other second objects except the target second object.
[0142] In another implementation, the processing unit 602 can also be used for:
[0143] If the voice private chat area is closed, increase the playing volume of the other media data to the historical playing volume;
[0144] Wherein, the historical playing volume refers to: the playing volume that the other media data has before the first object and the target second object have a voice private chat.
[0145] According to another embodiment of the present application, Figure 6Each unit in the voice processing device shown can be separately or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units with more specific functions to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, based on the voice processing device, other units can also be included. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units.
[0146] According to another embodiment of this application, it is possible to construct the voice processing device as shown in Figure 6 by running a computer program (including one or more instructions) that can execute each step involved in any of the above method embodiments on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), and to implement the voice processing method of the embodiments of this application. The computer program can be recorded on, for example, a computer-readable storage medium, loaded into the above computing device through the computer-readable storage medium, and run therein.
[0147] It is worth noting that in the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can include a part of the overall module or unit with the function of that module or unit.
[0148] After the embodiment of the present application displays the voice communication interface, it can respond to a moving operation on the object identifier of any second object in the voice communication interface, move the corresponding object identifier in the voice communication interface, and thus adjust the voice playback volume of the corresponding second object based on the movement of the corresponding object identifier relative to the object identifier of the first object. It can be seen that the embodiment of the present application can support the first object to independently adjust the voice playback volume of a single second object by moving the object identifier of the single second object during the voice communication between the first object and at least one second object; this can not only improve the flexibility of volume adjustment, but also enable the first object to control the voice playback volume of each second object according to its actual demands in a noisy environment of a multi-person conversation, which is beneficial for the first object to clearly hear the content it wants to hear, improve the efficiency of the first object to obtain information, and further improve the voice communication experience of the first object.
[0149] Based on the descriptions of the above method embodiments and apparatus embodiments, the embodiments of the present application further provide a computer device. Please refer to Figure 7 , the computer device at least includes a processor 701, an input interface 702, an output interface 703, and a computer storage medium 704. Among them, the processor 701, the input interface 702, the output interface 703, and the computer storage medium 704 in the computer device can be connected through a bus or other means. The computer storage medium 704 can be stored in the memory of the computer device. The computer storage medium 704 is used to store a computer program, the computer program includes one or more instructions, and the processor 701 is used to execute one or more instructions stored in the computer storage medium 704. The processor 701 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, and it is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method process or corresponding function.
[0150] In one embodiment, the processor 701 described in the embodiments of the present application can be used to perform a series of voice communication processes, specifically including: displaying a voice communication interface, where the voice communication interface includes: object identifiers of multiple objects participating in the voice communication; the multiple objects include: a first object and at least one second object; responding to a moving operation on a target object identifier, moving the target object identifier in the voice communication interface; the target object identifier refers to the object identifier of a target second object, and the target second object is any second object participating in the voice communication; adjusting the voice playback volume of the target second object based on the movement of the target object identifier relative to the object identifier of the first object, and so on.
[0151] An embodiment of the present application further provides a computer storage medium (Memory). The computer storage medium is a memory device in a computer device and is used to store computer programs and data. It can be understood that the computer storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer storage medium provides a storage space, and the operating system of the computer device is stored in this storage space. Moreover, a computer program is also stored in this storage space. The computer program includes one or more instructions suitable for being loaded and executed by a processor 701, and these instructions can be one or more program codes. It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer storage medium located far from the aforementioned processor.
[0152] In one embodiment, one or more instructions stored in the computer storage medium can be loaded and executed by the processor to implement the corresponding steps in the method embodiment as described above Figure 2 or Figure 4 shown; specifically, in implementation, one or more instructions in the computer storage medium can be loaded and executed by the processor to perform the following steps:
[0153] Display a voice communication interface, where the voice communication interface includes: object identifiers of multiple objects participating in the voice communication; the multiple objects include: a first object and at least one second object;
[0154] In response to a moving operation on a target object identifier, move the target object identifier in the voice communication interface; the target object identifier refers to the object identifier of a target second object, and the target second object is any one of the second objects participating in the voice communication;
[0155] Based on the moving situation of the target object identifier relative to the object identifier of the first object, adjust the voice playback volume of the target second object.
[0156] In one implementation manner, before displaying the voice communication interface, the one or more instructions can be loaded and specifically executed by the processor:
[0157] Play multimedia data and determine a discussion topic associated with the multimedia data;
[0158] In the playback interface of the multimedia data, display an entrance to a voice session room established based on the discussion topic;
[0159] If the entrance of the room is triggered, control the first object to enter the voice conversation room to participate in voice communication, and trigger the display of a voice communication interface; wherein, the voice communication interface is the conversation interface of the voice conversation room.
[0160] In another embodiment, the voice communication interface further includes a voice switch component, and when the voice communication interface is displayed, the voice switch component is in a closed state; the one or more instructions can be loaded and specifically executed by a processor:
[0161] In response to a trigger operation on the voice switch component, switch the state of the voice switch component from the closed state to the open state;
[0162] When the voice switch component is in the open state, perform voice collection on the first object.
[0163] In another embodiment, the one or more instructions can be loaded and specifically executed by a processor:
[0164] If the target object identifier is moved, display a volume control bar around the target object identifier, and the volume control bar is used to dynamically prompt the current voice playback volume of the target second object;
[0165] After the target object identifier stops moving, cancel the display of the volume control bar.
[0166] In another embodiment, the one or more instructions can be loaded and specifically executed by a processor:
[0167] During voice communication, when there is a speaking object among the multiple objects, during the speaking process of the speaking object, visually highlight the object identifier of the speaking object;
[0168] Wherein, the object identifier of the speaking object includes a first image; the visual highlighting includes at least one of the following: magnifying the display of the image, adding a bright light of a preset color to the corresponding image, and replacing the first image with a second image;
[0169] Wherein, the second image is an image obtained by changing the posture of the object in the first image, and the object in the second image is in a speaking posture.
[0170] In another embodiment, the one or more instructions can be loaded and specifically executed by a processor:
[0171] In response to a voice private chat trigger operation for the target second object, send a voice private chat prompt to the target second object; wherein, the voice private chat trigger operation includes: an operation of moving the target object identifier of the target second object to the target display position where the object identifier of the first object is located.
[0172] After the target second object agrees to have a voice private chat with the first object, display a voice private chat area so that the first object can have a voice private chat with the target second object through the voice private chat area.
[0173] In another implementation, the voice private chat area is displayed at a target display position in the voice communication interface, and the voice private chat area includes: the target object identifier of the target second object and the object identifier of the first object.
[0174] The one or more instructions can be loaded and specifically executed by a processor:
[0175] In response to a close operation for the voice private chat area, close the voice private chat area at the target display position.
[0176] Display the object identifier of the first object at the target display position, and move the target object identifier of the target second object from the target display position to the historical display position of the target object identifier.
[0177] Wherein, the historical display position of the target object identifier refers to: the display position where the target object identifier is located when the voice private chat trigger operation is detected.
[0178] In another implementation, the voice private chat area is displayed in the voice communication interface; the one or more instructions can be loaded and specifically executed by a processor:
[0179] During the process that the first object has a voice private chat with the target second object through the voice private chat area, keep playing the other media data involved in the voice communication interface and reduce the playing volume of the other media data.
[0180] Wherein, the other media data includes at least one of the following: the background media in the voice communication interface, and the voice audio data sent by other second objects except the target second object.
[0181] In another implementation, the one or more instructions can be loaded and specifically executed by a processor:
[0182] If the voice private chat area is closed, increase the playing volume of the other media data to the historical playing volume.
[0183] Wherein, the historical playback volume refers to the playback volume of the other media data before the first object and the target second object conduct a voice private chat.
[0184] After the voice communication interface is displayed in an embodiment of the present application, in response to a moving operation on the object identifier of any second object in the voice communication interface, the corresponding object identifier can be moved in the voice communication interface. Thus, based on the moving situation of the corresponding object identifier relative to the object identifier of the first object, the voice playback volume of the corresponding second object can be adjusted. It can be seen that in an embodiment of the present application, during the voice communication between the first object and at least one second object, the first object is supported to independently adjust the voice playback volume of a single second object by moving the object identifier of the single second object. This can not only improve the flexibility of volume adjustment, but also enable the first object to respectively control the voice playback volumes of each second object according to its actual demands in a noisy environment of a multi-person conversation, thereby facilitating the first object to clearly hear the content it wants to hear, improving the efficiency of the first object in obtaining information, and further enhancing the voice communication experience of the first object.
[0185] It should be noted that according to one aspect of the present application, a computer program product or a computer program is further provided. The computer program product or the computer program includes one or more instructions, and the one or more instructions are stored in a computer storage medium. The processor of the computer device reads the one or more instructions from the computer storage medium, and the processor executes the one or more instructions, so that the computer device executes the methods provided in various optional manners in the above-mentioned aspects of the voice processing method embodiment. It should be understood that the above-disclosed is only a preferred embodiment of the present application, and of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A voice processing method, characterized in that Including: Display a voice communication interface, which includes: object identifiers of multiple objects participating in the voice communication; the multiple objects include: a first object and at least one second object; In response to a moving operation on the target object identifier, move the target object identifier in the voice communication interface; the target object identifier refers to the object identifier of the target second object, and the target second object is any one of the second objects participating in the voice communication; Adjust the voice playback volume of the target second object based on the moving situation of the target object identifier relative to the object identifier of the first object.
2. The method according to claim 1, characterized in that Before displaying the voice communication interface, the method further includes: Play multimedia data and determine a discussion topic associated with the multimedia data; In the playback interface of the multimedia data, display an entrance to a voice session room established based on the discussion topic; If the room entrance is triggered, control the first object to enter the voice session room to participate in the voice communication and trigger the display of the voice communication interface; wherein, the voice communication interface is the session interface of the voice session room.
3. The method according to claim 2, wherein The voice communication interface further includes background media, and the background media includes the multimedia data; Wherein, in the voice communication interface, the multimedia data starts playing based on a target media frame; the target media frame refers to: the media frame being played in the multimedia data when the room entrance is triggered.
4. The method according to claim 1, characterized in that, The voice communication interface further includes a voice switch component, and when the voice communication interface is displayed, the voice switch component is in a closed state; the method further includes: In response to a triggering operation on the voice switch component, switch the state of the voice switch component from the closed state to the open state; When the voice switch component is in the open state, perform voice collection on the first object.
5. The method according to claim 1, characterized in that, The method further includes: If the target object identifier is moved, display a volume control bar around the target object identifier, and the volume control bar is used to dynamically prompt the current voice playback volume of the target second object; After the target object identifier stops moving, cancel the display of the volume control bar.
6. The method according to claim 1, characterized in that The method further includes: During the voice communication process, when there is a speaking object among the multiple objects, visually highlight the object identifier of the speaking object during the speaking process of the speaking object; Wherein, the object identifier of the speaking object includes a first image; the visual highlighting includes at least one of the following: magnifying the display of the image, adding a bright light of a preset color to the corresponding image, replacing the first image with a second image; Wherein, the second image is an image obtained by changing the posture of the object in the first image, and the object in the second image is in a speaking posture.
7. The method according to claim 1, characterized in that The moving situation includes a target distance, and the target distance refers to: the distance between the target object identifier and the object identifier of the first object after moving the target object identifier; the distance between the target object identifier and the object identifier of the first object before moving the target object identifier is the initial distance; Wherein, if the target distance is greater than the initial distance, the voice playback volume of the target second object is decreased; if the target distance is less than the initial distance, the voice playback volume of the target second object is increased.
8. The method according to claim 7, wherein The voice playback volume of the target second object is adjusted based on a volume gain control parameter; the determination method of the voice gain control parameter includes: Determine the distance difference between the initial distance and the target distance; Determine the volume gain control parameter according to the ratio between the distance difference and the initial distance.
9. The method according to claim 1, wherein The method further includes: In response to a voice private chat trigger operation for the target second object, send a voice private chat prompt to the target second object; wherein, the voice private chat trigger operation includes: an operation of moving the target object identifier of the target second object to the target display position where the object identifier of the first object is located; After the target second object agrees to have a voice private chat with the first object, display a voice private chat area, so that the first object can have a voice private chat with the target second object through the voice private chat area.
10. The method according to claim 9, wherein The voice private chat area is displayed at a target display position in the voice communication interface, and the voice private chat area includes: the target object identifier of the target second object and the object identifier of the first object; The method further includes: In response to a close operation for the voice private chat area, close the voice private chat area at the target display position; Display the object identifier of the first object at the target display position, and move the target object identifier of the target second object from the target display position to the historical display position of the target object identifier; Wherein, the historical display position of the target object identifier refers to: the display position where the target object identifier is located when the voice private chat trigger operation is detected.
11. The method according to claim 10, wherein During the process of displaying the voice private chat area in the voice communication interface, the voice communication interface further includes: a close component of the voice private chat area; Wherein, the close operation includes: a trigger operation for the close component.
12. The method according to claim 9, wherein The voice private chat area is displayed in the voice communication interface; the method further includes: During the process of the first object having a voice private chat with the target second object through the voice private chat area, keep playing other media data involved in the voice communication interface, and decrease the playback volume of the other media data; Wherein, the other media data includes at least one of the following: the background media in the voice communication interface, and the voice audio data sent by other second objects except the target second object.
13. The method according to claim 12, wherein The method further includes: If the voice private chat area is closed, increase the playback volume of the other media data to the historical playback volume; Wherein, the historical playback volume refers to: the playback volume that the other media data has before the first object and the target second object have a voice private chat.
14. A voice processing device, characterized in that, Including: A display unit for displaying a voice communication interface, where the voice communication interface includes: object identifiers of multiple objects participating in the voice communication; the multiple objects include: a first object and at least one second object; The display unit is further configured to move the target object identifier in the voice communication interface in response to a moving operation on the target object identifier; the target object identifier refers to the object identifier of a target second object, and the target second object is any one of the second objects participating in the voice communication; A processing unit for adjusting the voice playback volume of the target second object based on the movement of the target object identifier relative to the object identifier of the first object.
15. A computer device includes an input interface and an output interface, characterized in that, It further includes: A processor and a computer storage medium; Wherein, the processor is adapted to implement one or more instructions, the computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the voice processing method according to any one of claims 1-13.
16. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the voice processing method according to any one of claims 1-13.
17. A computer program product, characterized in that, The computer program product includes one or more instructions; when the one or more instructions in the computer program are executed by the processor, the voice processing method according to any one of claims 1-13 is implemented.