Sound signal processing method and device

By using an automatic level control device in a multimedia device to adjust the volume according to the echo signal, the problem of manual control of volume is solved, and automatic adjustment of volume and appropriate volume during the meeting is realized.

CN112737535BActive Publication Date: 2025-05-09ALIBABA GROUP HOLDING LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201911032579.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-10-28
Publication Date
2025-05-09
Estimated Expiration
2039-10-28

AI Technical Summary

Technical Problem

The volume of the multimedia device used in the conference scene in the prior art requires manual control by the user, which can easily lead to misoperation and affect the conference process.

Method used

By picking up the current sound signal, including the echo signal, the automatic level control device adjusts the playback volume of the playback device according to the echo signal, and automatically adjusts the volume.

Benefits of technology

It solves the problem that users can easily misoperate by manually controlling the volume, realizes automatic adjustment of the volume, ensures appropriate volume during the meeting, and reduces human errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112737535B_ABST
    Figure CN112737535B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for processing sound signals. The method comprises: picking up a current sound signal, wherein the sound signal comprises: an echo signal generated when a current playback device plays sound; and controlling the playback volume of the playback device at least according to the echo signal. The present invention solves the technical problem in the prior art that the volume of multimedia devices used in conference scenarios needs to be manually controlled by the user, which is prone to misoperation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sound processing, and in particular to a method and device for processing sound signals. Background Art

[0002] At present, audio and video conferencing systems are usually used in conference rooms to assist meetings. Specific functions may include amplifying the speaker's speech through speakers. For conference rooms of different sizes or conference rooms with different number of participants, the required volume may be different. In this case, the audio and video conferencing system usually provides a manual volume control solution, specifically a remote control for controlling the volume of the audio and video conferencing system.

[0003] However, the problem is that the participants are not necessarily familiar with the use of the remote control of the audio and video conferencing system, so they may make mistakes, such as shutting down the audio and video conferencing system, switching the mode of the audio and video conferencing system (remote video, etc.), etc., which requires the audio and video conferencing system to be readjusted, thus affecting the progress of the meeting.

[0004] In view of the problem that the volume of multimedia devices used in conference scenarios in the prior art requires manual control by users, which may lead to easy misoperation, no effective solution has been proposed so far. Summary of the invention

[0005] The embodiments of the present invention provide a method and device for processing a sound signal, so as to at least solve the technical problem in the prior art that the volume of a multimedia device used in a conference scene needs to be manually controlled by a user, which may lead to easy misoperation.

[0006] According to one aspect of an embodiment of the present invention, a method for processing a sound signal is provided, comprising: picking up a current sound signal, wherein the sound signal comprises: an echo signal generated when a current playback device plays sound; and controlling the playback volume of the playback device at least according to the echo signal.

[0007] Furthermore, the sound signal is input to an automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume of the playback device.

[0008] Furthermore, the automatic level control device also obtains the amplitude of the echo signal and compares the amplitude of the echo signal with a preset amplitude. If the amplitude of the echo signal is greater than the preset amplitude, the gain is reduced; if the amplitude of the echo signal is less than the preset amplitude, the gain is increased.

[0009] Furthermore, the method further includes: acquiring an echo signal, selecting a sound pickup device with the largest signal amplitude among multiple sound pickup devices as a target sound pickup device; and extracting the echo signal from the sound signal picked up by the target sound pickup device.

[0010] Furthermore, the above method also includes: collecting image information of the current environment; and controlling the playback volume of the playback device according to the image information.

[0011] Furthermore, the number of subjects included in the environment and the distance between the subjects and the playback device are determined according to the image information; and the playback volume of the playback device is adjusted according to the number and the distance.

[0012] Furthermore, the current playback volume is scored according to the quantity and the distance, wherein the score is used to indicate the degree of matching between the current playback volume and the current environment; and the playback volume is adjusted according to the score.

[0013] Furthermore, obtain a first score for the current playback volume with respect to quantity and a second score for the current playback volume with respect to distance; obtain a first weight corresponding to the first score and a second weight corresponding to the second score; use the first weight and the second weight to weight the first score and the second score to obtain a scoring result for quantity and distance with respect to the current playback volume.

[0014] According to one aspect of an embodiment of the present invention, a method for processing a sound signal is provided, comprising: picking up a current sound signal, wherein the sound signal comprises: a sound signal generated by a current sound source and an echo signal generated when a current playback device plays the sound; and playing the sound information generated by the sound source, wherein the playback device controls the playback volume of the playback device at least according to the echo signal.

[0015] Furthermore, the sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound are input into an automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume when the playback device plays the sound signal generated by the current sound source.

[0016] According to one aspect of an embodiment of the present invention, a method for processing a sound signal is provided, comprising: picking up a current sound signal, wherein the sound signal comprises: a sound signal generated by a current sound source and an echo signal generated when a current playback device plays the sound; and controlling the playback volume of the playback device at least according to the echo signal, wherein the playback device is used to play the sound information generated by the sound source.

[0017] Furthermore, the sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound are input into the automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume of the playback device.

[0018] According to one aspect of an embodiment of the present invention, a sound signal processing device is provided, comprising: a pickup module for picking up a current sound signal, wherein the sound signal comprises: an echo signal generated when a current playback device plays sound; and a control module for controlling the playback volume of the playback device at least according to the echo signal.

[0019] According to one aspect of an embodiment of the present invention, a storage medium is provided, the storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to perform the following steps: picking up a current sound signal, wherein the sound signal includes: an echo signal generated when the current playback device plays sound; and controlling the playback volume of the playback device at least according to the echo signal.

[0020] According to one aspect of an embodiment of the present invention, a processor is provided, and the processor is used to run a program, wherein the following steps are performed when the program is running: picking up a current sound signal, wherein the sound signal includes: an echo signal generated when a current playback device plays sound; and controlling the playback volume of the playback device at least according to the echo signal.

[0021] In an embodiment of the present invention, the echo signal is used as an input signal for volume adjustment, and the volume of the playback device is adjusted according to the echo signal, thereby realizing a closed-loop link. The closed-loop link includes a power amplifier and a speaker. By maintaining the echo signal at a fixed preset value, the sound finally heard by the human ear is not affected by the reverberation of the room, thereby solving the technical problem in the prior art that the volume of multimedia devices used in conference scenarios needs to be manually controlled by the user, which is prone to misoperation, and achieves the effect of automatic loudness adjustment. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0023] Figure 1 A hardware structure block diagram of a computing device (or mobile device) for implementing a method for processing a sound signal is shown;

[0024] Figure 2 is a flowchart of a method for processing a sound signal according to Embodiment 1 of the present application;

[0025] Figure 3 It is a schematic diagram of processing sound signals in an audio and video conferencing system;

[0026] Figure 4 is a schematic diagram of an audio and video conferencing system processing sound signals according to an embodiment of the present application;

[0027] Figure 5 is a schematic diagram of an audio and video conferencing system;

[0028] Figure 6 is a flowchart of a method for processing a sound signal according to Embodiment 2 of the present application;

[0029] Figure 7 is a flowchart of a method for processing a sound signal according to Embodiment 3 of the present application;

[0030] Figure 8 is a schematic diagram of a sound signal processing device according to Embodiment 4 of the present application;

[0031] Fig. 9 is a schematic diagram of a sound signal processing device according to Embodiment 5 of the present application;

[0032] Fig.10 is a schematic diagram of a sound signal processing device according to Embodiment 6 of the present application; and

[0033] Fig.11 is a structural block diagram of a computing device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0036] Example 1

[0037] According to an embodiment of the present invention, an embodiment of a method for processing a sound signal is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0038] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computing device or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computing device (or mobile device) for implementing a method for processing a sound signal. Figure 1 As shown, the computing device 10 (or mobile device 10) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is for illustration only and does not limit the structure of the electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0039] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computing device 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0040] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for processing sound signals in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, to implement the vulnerability detection method of the above-mentioned application program. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computing device 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0041] The transmission module 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computing device 10. In one example, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0042] The display may be, for example, a touch screen liquid crystal display (LCD) that may enable a user to interact with a user interface of computing device 10 (or mobile device).

[0043] It should be noted that, in some optional embodiments, the above Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. It should be noted that Figure 1This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the above-described computer device (or mobile device).

[0044] Under the above operating environment, this application provides Figure 2 The method for processing the sound signal shown. Figure 2 This is a flowchart of a method for processing a sound signal according to Example 1 of the present application.

[0045] Step S21, picking up the current sound signal, wherein the sound signal includes: an echo signal generated when the current playing device plays the sound.

[0046] Specifically, the playback device may be a speaker, and the speaker may be a speaker in an audio and video device. Taking the audio and video device used in a conference scene as an example, the audio and video device includes a sound pickup device and a speaker. The sound pickup device collects information sent by the conference speaker, and the speaker amplifies the sound and plays it. The echo signal refers to the signal after the sound signal emitted by the playback device is reflected by an object (wall, etc.) in the environment when the sound signal is played.

[0047] In the above steps, when the device is collecting sound, it not only collects the sound emitted by the sound source, but also collects the echo signal generated by the sound playing device. In an optional embodiment, taking a conference scene as an example, when the speaker of the conference is speaking, the speaker is used as a sound source and the sound signal is collected by the sound pickup device. At the same time, the audio and video equipment in the conference scene will also play the sound signal collected by the sound pickup device through the speaker. Therefore, the played sound signal propagates in the conference room and forms an echo signal after being reflected by objects such as walls. In the above scheme, the audio and video equipment collects the sound signal generated by the speaker while collecting the echo signal.

[0048] Step S23, controlling the playback volume of the playback device at least according to the echo signal.

[0049] In the above steps, the device controls the playback volume according to the echo signal, which may be achieved by controlling the amplitude of the echo signal to a preset value or within a preset range through a power amplifier. The power amplifier is used to amplify the sound information generated by the sound source.

[0050] The above scheme uses the echo signal as the input signal for volume adjustment, and adjusts the volume of the playback device according to the echo signal, thereby realizing a closed-loop link. The closed-loop link includes a power amplifier and a speaker. By keeping the echo signal at a fixed preset value, the sound finally heard by the human ear is not affected by the room reverberation, thereby solving the technical problem in the prior art that the volume of multimedia devices used in conference scenarios requires manual control by the user, which is prone to misoperation, and achieves the effect of automatic loudness adjustment.

[0051] As an optional embodiment, controlling the playback volume of the playback device at least according to the echo signal includes: inputting the sound signal into an automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume of the playback device.

[0052] Specifically, the automatic level control device is ALC (Automatic Level Control), and the control purpose of the automatic level control device is to keep the amplitude of the output signal unchanged even if the amplitude of the input signal changes. The above steps input the echo signal to the automatic level control device, so that the amplitude of the echo signal can be controlled to a preset value through the automatic level control device.

[0053] Figure 3 It is a schematic diagram of audio and video conferencing system processing sound signals, combined with Figure 3 As shown in the figure, the microphone (mic) collects the sound signal while also collecting the echo signal from the speaker. After the acoustic echo canceller (AcousticEchoCanceller) eliminates the echo information in the sound signal collected by the microphone, the echo signal (echo estimate) output by the acoustic echo canceller is similar to the sound from the far end "heard" by the microphone, and then input into the audio encoder (Audio Encoder) through the digital automatic gain control (digital AGC), and the audio encoder encodes the audio data into a bit stream.

[0054] The audio decoder obtains the bit stream from the audio encoder, decodes the bit stream and restores it to the original audio data. Before playing the sound with the speaker, a manual volume control device is used for the user to adjust the volume. The manual volume control device controls the gain of the power amplifier (power amp) according to the user's adjustment, thereby controlling the output volume of the speaker (louspeaker).

[0055] Figure 4 is a schematic diagram of an audio and video conferencing system processing sound signals according to an embodiment of the present application, combined with Figure 4 As shown, the solution of this application introduces an ALC to replace Figure 3 The manual volume control device in the ALC. Unlike ordinary ALC (ordinary ALC has only one input signal), the ALC here takes echoestimate as the second input while inputting the sound signal generated by the current sound source. The goal of ALC gain adjustment is to make the amplitude of echo estimate reach a preset value. This preset value can be similar to the signal amplitude of the sound signal generated when the microphone receives the near-end speech. It can be seen that this is a closed-loop feedback system. The use of the echoestimate signal means that the closed loop contains factors such as speaker sensitivity, power amplifier, speaker, room reverberation, etc. Therefore, the ultimate goal of ALC gain adjustment is the loudness "heard" by the microphone, which is equivalent to the loudness of the audio signal heard by the ears of people in the conference room.

[0056] As an optional embodiment, the automatic level control device also obtains the amplitude of the echo signal and compares the amplitude of the echo signal with a preset amplitude. If the amplitude of the echo signal is greater than the preset amplitude, the gain is reduced; if the amplitude of the echo signal is less than the preset amplitude, the gain is increased.

[0057] Specifically, the automatic level control device may adjust the gain by comparing the amplitude of the echo signal with a preset amplitude.

[0058] In an optional embodiment, the automatic level control device obtains a preset amplitude and a gain adjustment step. After obtaining the echo signal, the automatic level device compares the amplitude of the echo signal with the preset amplitude. If the amplitude of the echo signal is greater than the preset amplitude, the gain is reduced according to the preset gain adjustment step, and the amplitude of the echo signal is continuously compared with the preset amplitude until the amplitude of the echo signal is the same as the preset amplitude. Similarly, if the amplitude of the echo signal is less than the preset amplitude, the gain is increased according to the preset gain adjustment step, and the amplitude of the echo signal is continuously compared with the preset amplitude until the amplitude of the echo signal is the same as the preset amplitude.

[0059] As an optional embodiment, the method further includes: acquiring an echo signal, wherein acquiring the echo signal includes: selecting a sound pickup device with the largest signal amplitude among multiple sound pickup devices as a target sound pickup device; and extracting the echo signal from the sound signal picked up by the target sound pickup device.

[0060] It should be noted that in some scenarios, multiple sound pickup devices may be included. For example, in a conference scenario, a microphone array is used to pick up sound information. Since the echo signals collected by each sound pickup device may be different, it is necessary to determine which sound pickup device's echo signal is selected for volume adjustment.

[0061] In the above embodiment, a sound pickup device with the largest signal amplitude among the multiple sound pickup devices is used as the target sound pickup device, and the volume of the playback device is controlled by the echo signal collected by the target sound pickup device.

[0062] Figure 5 This is a schematic diagram of an audio and video conferencing system. In this example, the sound pickup system uses a total of 10 directional microphones, including a main microphone (containing 4 directional microphone units) and one left and right extension microphone (containing 3 directional microphone units and 3 virtual microphones constructed by these 3 physical microphones). In this case, the echo signal can select the one with the largest signal amplitude from the 16 microphone signals.

[0063] As an optional embodiment, the method further includes: collecting image information of the current environment; and controlling the playback volume of the playback device according to the image information.

[0064] Specifically, the above image information can be obtained by an image acquisition device set at a preset position. The image acquisition device can be located in the same room or in the same environment as the above sound pickup device. The above image acquisition device can be an ordinary camera or a 3D camera.

[0065] In an optional embodiment, the image acquisition device can be controlled to acquire image information in the current environment according to a preset period, and analyze the acquired image information to obtain information about people in the current environment, thereby adjusting the playback volume of the playback device according to the information about people in the current environment.

[0066] It should be noted that the above-mentioned solution of controlling the playback volume of the playback device through image information can be carried out while controlling the playback volume based on the echo signal. The difference is that the image information is only used to fine-tune the playback volume, that is, the adjustment step is smaller each time.

[0067] Still in the conference scenario, participants sitting in front of the conference table (closer to the speaker) and sitting behind the conference table (farther from the speaker) obviously have different requirements for the loudness of the speaker. Therefore, the above solution also combines visual information to adjust the volume of the playback device.

[0068] As an optional embodiment, controlling the playback volume of the playback device according to the image information includes: determining the number of subjects included in the environment and the distance between the subjects and the playback device according to the image information; and adjusting the playback volume of the playback device according to the number and distance.

[0069] Specifically, the subject in the above environment can be a person in the environment. Taking a conference scene as an example, it can be a parameter. The distance between the subject and the playback device can be the distance between the participant and the audio and video conferencing system. When multiple participants are involved, the distance can be the average distance between each participant and the audio and video conferencing system.

[0070] Since the requirements for playback volume are obviously different when the number of subjects and the distance between the subjects and the playback device are different, the above steps introduce the parameters of the number of people and distance through visual information to adjust the playback volume to achieve a better adjustment result.

[0071] As an optional embodiment, adjusting the playback volume of the playback device according to the quantity and distance includes: scoring the current playback volume according to the quantity and distance, wherein the score is used to indicate the degree of matching between the current playback volume and the current environment; and adjusting the playback volume according to the score.

[0072] In the above scheme, the current playback volume is scored by the number of subjects and the distance between the subjects and the playback device, so that the playback volume is adjusted according to the scoring result. In an optional embodiment, a preset score threshold can be set, and if the scoring result is less than the score threshold, the playback volume of the playback device is adjusted. The adjustment method can be to search for the corresponding volume according to the current number in the corresponding relationship between the number and the volume, and search for the corresponding volume according to the current distance in the corresponding relationship between the distance and the volume, and then take the average of the two volumes found to obtain the target volume.

[0073] As an optional embodiment, scoring the current playback volume according to quantity and distance includes: obtaining a first score of the current playback volume for quantity and a second score of the current playback volume for distance; obtaining a first weight corresponding to the first score and a second weight corresponding to the second score; weighting the first score and the second score using the first weight and the second weight to obtain a scoring result of the quantity and distance for the current playback volume.

[0074] In an optional embodiment, the playback volume has different scores for different numbers of subjects, and similarly, the playback volume has different scores for different distances. According to the current playback volume, the first score corresponding to the current number of subjects and the second score corresponding to the current playback volume at the current distance are searched, and then the first score and the second score are weighted based on the pre-acquired first weight and second weight, so as to obtain the final scoring result.

[0075] As an optional embodiment, controlling the playback volume of the playback device according to the image information includes: determining the relative position between the subject contained in the environment and the playback device according to the image information; and adjusting the playback volume of the playback device according to the relative position.

[0076] Specifically, the above relative position can be expressed in the form of coordinates. For example, the coordinates of the playback device can be set as the origin (0, 0), the direction in which the playback device is set can be used as one of the coordinate axes, and the direction perpendicular to the coordinate axis can be used as the other coordinate axis to form a coordinate system with the playback device as the origin, thereby determining the position of the subject in the image information in the above coordinate system, that is, the relative position of the subject and the playback device.

[0077] The coordinate system can be divided into multiple areas, each area corresponds to a different adjustment coefficient, after determining the relative position between the main body and the playback device, the area to which the main body belongs is determined according to the coordinates representing the relative position, so that the playback volume is adjusted according to the adjustment coefficient corresponding to the area. The adjustment coefficient can be (-1, 1). If the determined adjustment coefficient belongs to (-1, 0), it means reducing the playback volume. If the determined adjustment coefficient belongs to (0, 1), it means increasing the playback volume.

[0078] When a room includes a plurality of subjects at different positions, the playback volume is adjusted according to the adjustment coefficient corresponding to the area including the largest number of subjects.

[0079] As an optional embodiment, controlling the playback volume of the playback device according to the image information includes: acquiring a positional relationship between a sound-emitting subject and a sound pickup device; and adjusting the playback volume of the playback device according to the positional relationship.

[0080] The sound pickup device can be of various types, such as single-directional, bi-directional or omni-directional. The single-directional playback device can include the following three types: cardioid, supercardioid and omni-directional. If it is an omni-directional sound pickup device, there is no attenuation in all directions. If it is a single-directional or bi-directional sound pickup device, the sensitivity of sound pickup is different due to the different positions of the sound source. Therefore, the above solution adjusts the playback volume of the playback device in combination with the position of the sound source and the sound pickup device, so that the playback effect is not affected by the position of the speaker.

[0081] In an optional embodiment, after determining the type of the sound pickup device, the directional range of the sound pickup device can be determined, and whether the sound-emitting subject is within the directional range of the sound pickup device can be determined based on the image information. If the sound pickup subject is within the directional range of the sound pickup device, the playback volume of the playback device can be maintained or reduced. If the sound pickup subject does not fall within the directional range of the sound pickup device, the playback volume of the playback device can be increased.

[0082] In another optional embodiment, the positional relationship between the above-mentioned sound-emitting subject and the sound pickup device can also be expressed in the form of coordinates, and the coordinates can be polar coordinates. A polar coordinate system can be constructed based on the directivity of the sound pickup device, and the polar coordinate system includes a range for representing the directivity of the sound pickup device. The coordinates of the sound-emitting subject can be determined in the constructed polar coordinate system, and it can be determined whether the coordinates of the sound-emitting subject are within the range for representing the directivity of the sound pickup device. If the coordinates of the sound pickup subject are within the above range, the playback volume of the playback device can be maintained or reduced, and if the coordinates of the sound pickup subject are not within the above range, the playback volume of the playback device can be increased.

[0083] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0084] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of various embodiments of the present invention.

[0085] Example 2

[0086] According to an embodiment of the present invention, an embodiment of a method for processing a sound signal is also provided. Figure 6 is a flowchart of a method for processing a sound signal according to Embodiment 2 of the present application, combined with Figure 6 As shown, the method includes:

[0087] Step S61, picking up the current sound signal, wherein the sound signal includes: the sound signal generated by the current sound source and the echo signal generated when the current playing device plays the sound.

[0088] Specifically, the current sound source may be a participant who is currently speaking, and the playback device may be a speaker, which may be a speaker in an audio and video device. Taking the audio and video device used in a conference scene as an example, the audio and video device includes a sound pickup device and a speaker. The sound pickup device collects information emitted by the conference speaker, and the speaker amplifies the sound and plays it. The echo signal refers to the signal emitted by the playback device after the sound signal is reflected by an object (wall, etc.) in the environment when the sound signal is played.

[0089] In the above steps, when the device is collecting sound, it not only collects the sound emitted by the sound source, but also collects the echo signal generated by the sound playing device. In an optional embodiment, taking a conference scene as an example, when the speaker of the conference is speaking, the speaker is used as a sound source and the sound signal is collected by the sound pickup device. At the same time, the audio and video equipment in the conference scene will also play the sound signal collected by the sound pickup device through the speaker. Therefore, the played sound signal propagates in the conference room and forms an echo signal after being reflected by objects such as walls. In the above scheme, the audio and video equipment collects the sound signal generated by the speaker while collecting the echo signal.

[0090] Step S63: play the sound information generated by the sound source, wherein the playing device controls the playing volume of the playing device at least according to the echo signal.

[0091] In the above steps, the device controls the playback volume according to the echo signal, which may be achieved by controlling the amplitude of the echo signal to a preset value or within a preset range through a power amplifier. The power amplifier is used to amplify the sound information generated by the sound source.

[0092] The above scheme uses the echo signal as the input signal for volume adjustment, and adjusts the volume of the playback device according to the echo signal, thereby realizing a closed-loop link. The closed-loop link includes a power amplifier and a speaker. By keeping the echo signal at a fixed preset value, the sound finally heard by the human ear is not affected by the room reverberation, thereby solving the technical problem in the prior art that the volume of multimedia devices used in conference scenarios requires manual control by the user, which is prone to misoperation, and achieves the effect of automatic loudness adjustment.

[0093] As an optional embodiment, the playback device controls the playback volume of the playback device at least according to the echo signal, including: inputting the sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound into an automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume when the playback device plays the sound signal generated by the current sound source.

[0094] Specifically, the automatic level control device is ALC (Automatic Level Control), and the control purpose of the automatic level control device is to keep the amplitude of the output signal unchanged even if the amplitude of the input signal changes. The above steps input the echo signal to the automatic level control device, so that the amplitude of the echo signal can be controlled to a preset value through the automatic level control device.

[0095] Example 3

[0096] According to an embodiment of the present invention, an embodiment of a method for processing a sound signal is also provided. Figure 7 is a flowchart of a method for processing a sound signal according to Embodiment 3 of the present application, combined with Figure 7 As shown, the method includes:

[0097] Step S71, picking up the current sound signal, wherein the sound signal includes: the sound signal generated by the current sound source and the echo signal generated when the current playing device plays the sound.

[0098] Specifically, the current sound source may be a participant who is currently speaking, and the playback device may be a speaker, which may be a speaker in an audio and video device. Taking the audio and video device used in a conference scene as an example, the audio and video device includes a sound pickup device and a speaker. The sound pickup device collects information emitted by the conference speaker, and the speaker amplifies the sound and plays it. The echo signal refers to the signal emitted by the playback device after the sound signal is reflected by an object (wall, etc.) in the environment when the sound signal is played.

[0099] In the above steps, when the device is collecting sound, it not only collects the sound emitted by the sound source, but also collects the echo signal generated by the sound playing device. In an optional embodiment, taking a conference scene as an example, when the speaker of the conference is speaking, the speaker is used as a sound source and the sound signal is collected by the sound pickup device. At the same time, the audio and video equipment in the conference scene will also play the sound signal collected by the sound pickup device through the speaker. Therefore, the played sound signal propagates in the conference room and forms an echo signal after being reflected by objects such as walls. In the above scheme, the audio and video equipment collects the sound signal generated by the speaker while collecting the echo signal.

[0100] Step S73, controlling the playback volume of a playback device at least according to the echo signal, wherein the playback device is used to play the sound information generated by the sound source.

[0101] In the above steps, the device controls the playback volume according to the echo signal, which may be achieved by controlling the amplitude of the echo signal to a preset value or within a preset range through a power amplifier. The power amplifier is used to amplify the sound information generated by the sound source.

[0102] The above scheme uses the echo signal as the input signal for volume adjustment, and adjusts the volume of the playback device according to the echo signal, thereby realizing a closed-loop link. The closed-loop link includes a power amplifier and a speaker. By keeping the echo signal at a fixed preset value, the sound finally heard by the human ear is not affected by the room reverberation, thereby solving the technical problem in the prior art that the volume of multimedia devices used in conference scenarios requires manual control by the user, which is prone to misoperation, and achieves the effect of automatic loudness adjustment.

[0103] As an optional embodiment, controlling the playback volume of the playback device at least according to the echo signal includes: inputting the sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound into an automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume of the playback device.

[0104] Specifically, the automatic level control device is ALC (Automatic Level Control), and the control purpose of the automatic level control device is to keep the amplitude of the output signal unchanged even if the amplitude of the input signal changes. The above steps input the echo signal to the automatic level control device, so that the amplitude of the echo signal can be controlled to a preset value through the automatic level control device.

[0105] Example 4

[0106] According to an embodiment of the present invention, there is also provided a sound signal processing device for implementing the sound signal processing method of embodiment 1. Figure 8 is a schematic diagram of a sound signal processing device according to Embodiment 4 of the present application, such as Figure 8 As shown, the device 800 includes:

[0107] The pickup module 802 is used to pick up the current sound signal, wherein the sound signal includes: an echo signal generated when the current playback device plays the sound.

[0108] The control module 804 is used to control the playback volume of the playback device at least according to the echo signal.

[0109] It should be noted that the above-mentioned picking module 802 and control module 804 correspond to steps S21 to S23 in Embodiment 1, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned modules as part of the apparatus can be run in the computing device 10 provided in Embodiment 1.

[0110] As an optional embodiment, the control module includes: an input submodule, used to input the sound signal to the automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume of the playback device.

[0111] As an optional embodiment, the automatic level control device also obtains the amplitude of the echo signal and compares the amplitude of the echo signal with a preset amplitude. If the amplitude of the echo signal is greater than the preset amplitude, the gain is reduced; if the amplitude of the echo signal is less than the preset amplitude, the gain is increased.

[0112] As an optional embodiment, the above-mentioned device also includes: an acquisition module for acquiring an echo signal, wherein the acquisition module includes: a selection submodule for selecting a sound pickup device with the largest signal amplitude among multiple sound pickup devices as a target sound pickup device; and a lifting submodule for extracting the echo signal from the sound signal picked up by the target sound pickup device.

[0113] As an optional embodiment, the above-mentioned device further includes: a collection module, which is used to collect image information of the current environment; and an adjustment module, which is used to control the playback volume of the playback device according to the image information.

[0114] As an optional embodiment, the adjustment module includes: a first determination submodule, used to determine the number of subjects contained in the environment and the distance between the subjects and the playback device based on the image information; a first adjustment submodule, used to adjust the playback volume of the playback device based on the number and distance.

[0115] As an optional embodiment, the first adjustment submodule includes: a scoring unit, used to score the current playback volume according to quantity and distance, wherein the score is used to indicate the degree of matching between the current playback volume and the current environment; and an adjustment unit, used to adjust the playback volume according to the score.

[0116] As an optional embodiment, the scoring unit includes: a first acquisition subunit, used to obtain a first score of the current playback volume for quantity and a second score of the current playback volume for distance; a second acquisition subunit, used to obtain a first weight corresponding to the first score and a second weight corresponding to the second score; a weighting unit, used to weight the first score and the second score using the first weight and the second weight to obtain a scoring result of quantity and distance for the current playback volume.

[0117] As an optional embodiment, the adjustment module includes: a second determination submodule, used to determine the relative position between the subject contained in the environment and the playback device according to the image information; and a second adjustment submodule, used to adjust the playback volume of the playback device according to the relative position.

[0118] As an optional embodiment, the adjustment module includes: a third acquisition subunit, used to obtain the position relationship between the sound-emitting subject and the sound pickup device; and a third adjustment submodule, used to adjust the playback volume of the playback device according to the position relationship.

[0119] Example 5

[0120] According to an embodiment of the present invention, there is also provided a sound signal processing device for implementing the sound signal processing method of embodiment 2. Fig. 9 is a schematic diagram of a sound signal processing device according to Embodiment 5 of the present application, such as Fig. 9 As shown, the device 900 includes:

[0121] The pickup module 902 is used to pick up the current sound signal, wherein the sound signal includes: the sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound.

[0122] The playing module 904 is used to play the sound information generated by the sound source, wherein the playing device controls the playing volume of the playing device at least according to the echo signal.

[0123] It should be noted that the above-mentioned picking module 9022 and playing module 904 correspond to steps S61 to S63 in Example 2, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules as part of the device can be run in the computing device 10 provided in Example 1.

[0124] As an optional embodiment, the sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound are input into an automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume when the playback device plays the sound signal generated by the current sound source.

[0125] Example 6

[0126] According to an embodiment of the present invention, there is also provided a sound signal processing device for implementing the sound signal processing method of embodiment 3. Fig.10 is a schematic diagram of a sound signal processing device according to Embodiment 6 of the present application, such as Fig.10 As shown, the device 1000 includes:

[0127] The pickup module 1002 is used to pick up the current sound signal, wherein the sound signal includes: the sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound.

[0128] The control module 1004 is used to control the playback volume of the playback device at least according to the echo signal, wherein the playback device is used to play the sound information generated by the sound source.

[0129] It should be noted that the above-mentioned picking module 1002 and control module 1004 correspond to steps S71 to S73 in Example 3, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules as part of the device can be run in the computing device 10 provided in Example 1.

[0130] As an optional embodiment, the sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound are input into an automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume when the playback device plays the sound signal generated by the current sound source.

[0131] Example 7

[0132] The embodiment of the present invention may provide a computing device, which may be any computing device in a computing device group. Optionally, in this embodiment, the computing device may also be replaced by a terminal device such as a mobile terminal.

[0133] Optionally, in this embodiment, the computing device may be located in at least one network device among a plurality of network devices of a computer network.

[0134] In this embodiment, the above-mentioned computing device can execute the program code of the following steps in the vulnerability detection method of the application: picking up the current sound signal, wherein the sound signal includes: the echo signal generated when the current playback device plays the sound; and controlling the playback volume of the playback device at least according to the echo signal.

[0135] Optionally, Fig.11 is a structural block diagram of a computing device according to an embodiment of the present invention. Fig.11 As shown, the computing device A may include: one or more (only one is shown in the figure) processors 1102 , a memory 1104 , and a peripheral interface 1106 .

[0136] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the security vulnerability detection method and device in the embodiment of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, the detection method of the above-mentioned system vulnerability attack is realized. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0137] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: pick up the current sound signal, wherein the sound signal includes: the echo signal generated when the current playback device plays the sound; control the playback volume of the playback device at least according to the echo signal.

[0138] Optionally, the processor may also execute program code of the following steps: inputting a sound signal to an automatic level control device, wherein the automatic level control device adjusts a gain of the input signal according to the echo signal, and the gain is used to determine a playback volume of a playback device.

[0139] Optionally, the automatic level control device also obtains the amplitude of the echo signal and compares the amplitude of the echo signal with a preset amplitude. If the amplitude of the echo signal is greater than the preset amplitude, the gain is reduced; if the amplitude of the echo signal is less than the preset amplitude, the gain is increased.

[0140] Optionally, the processor may also execute program codes of the following steps: selecting a sound pickup device with the largest signal amplitude among multiple sound pickup devices as a target sound pickup device; and extracting an echo signal from the sound signal picked up by the target sound pickup device.

[0141] Optionally, the processor may also execute program codes of the following steps: collecting image information of the current environment; and controlling the playback volume of the playback device according to the image information.

[0142] Optionally, the processor may also execute program codes of the following steps: determining the number of subjects included in the environment and the distance between the subjects and the playback device according to the image information; and adjusting the playback volume of the playback device according to the number and the distance.

[0143] Optionally, the processor may also execute program code of the following steps: scoring the current playback volume according to the quantity and distance, wherein the score is used to indicate the degree of matching between the current playback volume and the current environment; and adjusting the playback volume according to the score.

[0144] Optionally, the processor may also execute the program code of the following steps: obtaining a first score for the current playback volume with respect to quantity and a second score for the current playback volume with respect to distance; obtaining a first weight corresponding to the first score and a second weight corresponding to the second score; and weighting the first score and the second score using the first weight and the second weight to obtain a scoring result for quantity and distance with respect to the current playback volume.

[0145] Optionally, the processor may also execute program codes of the following steps: determining the relative position between the subject contained in the environment and the playback device according to the image information; and adjusting the playback volume of the playback device according to the relative position.

[0146] Optionally, the processor may also execute program codes of the following steps: obtaining the positional relationship between the sound-emitting subject and the sound pickup device; and adjusting the playback volume of the playback device according to the positional relationship.

[0147] According to an embodiment of the present invention, a method for processing a sound signal is provided. An echo signal is used as an input signal for volume adjustment, and the volume of a playback device is adjusted according to the echo signal, thereby realizing a closed-loop link, which includes a power amplifier and a speaker. By keeping the echo signal at a fixed preset value, the sound finally heard by the human ear is not affected by the reverberation of the room, thereby solving the technical problem in the prior art that the volume of multimedia devices used in conference scenes needs to be manually controlled by the user, which is easy to cause misoperation, and achieving the effect of automatic loudness adjustment.

[0148] Those skilled in the art will appreciate that the structure shown in the figure is for illustration only, and the computing device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (MID), a PAD, or other terminal devices. Fig.10 The structure of the electronic device is not limited. For example, the computing device 10 may also include Fig.10 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Fig.10 Different configurations are shown.

[0149] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0150] Example 4

[0151] The embodiment of the present invention further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the sound signal processing method provided in the first embodiment.

[0152] Optionally, in this embodiment, the above storage medium may be located in any one of the computing devices in the computing device group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0153] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: picking up a current sound signal, wherein the sound signal includes: an echo signal generated when the current playback device plays sound; and controlling the playback volume of the playback device at least based on the echo signal.

[0154] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0155] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0156] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0157] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0158] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0159] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0160] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for processing a sound signal, characterized in that: include: Picking up a current sound signal, wherein the sound signal includes: an echo signal generated when the current playback device plays the sound; controlling the playback volume of the playback device at least according to the echo signal; Wherein, picking up the current sound signal comprises: picking up the current sound signal by using a sound pickup device; The method also includes: determining the type of the sound pickup device; determining the directionality range of the sound pickup device based on the type; determining the positional relationship between the sound-emitting subject and the directionality range based on image information of the current environment; maintaining the playback volume or reducing the playback volume in response to the positional relationship that the sound-emitting subject is within the directionality range; and increasing the playback volume in response to the positional relationship that the sound-emitting subject is outside the directionality range.

2. The method according to claim 1, characterized in that Controlling the playback volume of the playback device at least according to the echo signal includes: The sound signal is input to an automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume of the playback device.

3. The method according to claim 2, characterized in that The automatic level control device also obtains the amplitude of the echo signal and compares the amplitude of the echo signal with a preset amplitude. If the amplitude of the echo signal is greater than the preset amplitude, the gain is reduced; if the amplitude of the echo signal is less than the preset amplitude, the gain is increased.

4. The method according to claim 1, characterized in that: The method further includes: acquiring the echo signal, wherein acquiring the echo signal includes: Among the multiple sound pickup devices, a sound pickup device with the largest signal amplitude is selected as the target sound pickup device; The echo signal is extracted from the sound signal picked up by the target sound pickup device.

5. The method according to claim 1, characterized in that The method further comprises: Determine the number of subjects included in the environment and the distance between the subjects and the playback device according to the image information; The playback volume of the playback device is adjusted according to the number and the distance.

6. The method according to claim 5, characterized in that Adjusting the playback volume of the playback device according to the number and the distance includes: Scoring the current playback volume according to the number and the distance, wherein the score is used to indicate the degree of matching between the current playback volume and the current environment; The playback volume is adjusted according to the score.

7. The method according to claim 6, characterized in that Scoring the current playback volume according to the quantity and the distance includes: Obtaining a first score of the current playback volume with respect to the quantity and a second score of the current playback volume with respect to the distance; Obtaining a first weight corresponding to the first score and a second weight corresponding to the second score; The first score and the second score are weighted using the first weight and the second weight to obtain a scoring result of the quantity and the distance for the current playback volume.

8. The method according to claim 1, characterized in that The method further comprises: Determine the relative position between the subject contained in the environment and the playback device according to the image information; The playback volume of the playback device is adjusted according to the relative position.

9. A method for processing a sound signal, characterized in that: include: Picking up the current sound signal, wherein the sound signal includes: the sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound; Playing the sound information generated by the sound source, wherein the playing device controls the playing volume of the playing device at least according to the echo signal; Wherein, picking up the current sound signal comprises: picking up the current sound signal by using a sound pickup device; The method also includes: determining the type of the sound pickup device; determining the directionality range of the sound pickup device based on the type; determining the positional relationship between the sound-emitting subject and the directionality range based on image information of the current environment; maintaining the playback volume or reducing the playback volume in response to the positional relationship that the sound-emitting subject is within the directionality range; and increasing the playback volume in response to the positional relationship that the sound-emitting subject is outside the directionality range.

10. The method according to claim 9, characterized in that The playback device controls the playback volume of the playback device at least according to the echo signal, including: The sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound are input into an automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume when the playback device plays the sound signal generated by the current sound source.

11. A method for processing a sound signal, characterized in that: include: Picking up the current sound signal, wherein the sound signal includes: the sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound; controlling the playback volume of the playback device at least according to the echo signal, wherein the playback device is used to play the sound information generated by the sound source; Wherein, picking up the current sound signal comprises: picking up the current sound signal by using a sound pickup device; The method also includes: determining the type of the sound pickup device; determining the directionality range of the sound pickup device based on the type; determining the positional relationship between the sound-emitting subject and the directionality range based on image information of the current environment; maintaining the playback volume or reducing the playback volume in response to the positional relationship that the sound-emitting subject is within the directionality range; and increasing the playback volume in response to the positional relationship that the sound-emitting subject is outside the directionality range.

12. The method according to claim 11, characterized in that Controlling the playback volume of the playback device at least according to the echo signal includes: The sound signal generated by the current sound source and the echo signal generated when the current playback device plays the sound are input into an automatic level control device, wherein the automatic level control device adjusts the gain of the input signal according to the echo signal, and the gain is used to determine the playback volume of the playback device.

13. A sound signal processing device, characterized in that: include: A pickup module, used to pick up the current sound signal, wherein the sound signal includes: an echo signal generated when the current playback device plays the sound; A control module, used for controlling the playback volume of the playback device at least according to the echo signal; The pickup module is further used to: pick up the current sound signal using a sound pickup device; The device is also used to: determine the type of the sound pickup device; determine the directionality range of the sound pickup device based on the type; determine the positional relationship between the sound-emitting subject and the directionality range based on image information of the current environment; in response to the positional relationship being that the sound-emitting subject is within the directionality range, maintain the playback volume or reduce the playback volume; in response to the positional relationship being that the sound-emitting subject is outside the directionality range, increase the playback volume.

14. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to perform the following steps: Picking up a current sound signal, wherein the sound signal includes: an echo signal generated when the current playback device plays the sound; controlling the playback volume of the playback device at least according to the echo signal; When the program is running, the device where the storage medium is located is also controlled to perform the following steps: using a sound pickup device to pick up the current sound signal; Determine the type of the sound pickup device; based on the type, determine the directional range of the sound pickup device; determine the positional relationship between the sound-emitting subject and the directional range according to image information of the current environment; in response to the positional relationship that the sound-emitting subject is within the directional range, maintain the playback volume or reduce the playback volume; in response to the positional relationship that the sound-emitting subject is outside the directional range, increase the playback volume.

15. A processor, characterized in that: The processor is used to run a program, wherein the following steps are performed when the program is run: Picking up a current sound signal, wherein the sound signal includes: an echo signal generated when the current playback device plays the sound; controlling the playback volume of the playback device at least according to the echo signal; When the program is running, the following steps are also performed: using a sound pickup device to pick up the current sound signal; Determine the type of the sound pickup device; based on the type, determine the directional range of the sound pickup device; determine the positional relationship between the sound-emitting subject and the directional range according to image information of the current environment; in response to the positional relationship that the sound-emitting subject is within the directional range, maintain the playback volume or reduce the playback volume; in response to the positional relationship that the sound-emitting subject is outside the directional range, increase the playback volume.

Citation Information

Patent Citations

  • Controller, control method, and program

    CN104662874A

  • Method and device of controlling smart home

    CN105487396A

  • Volume adjusting method and apparatus and smart terminal

    CN105979358A

  • Multimedia equipment capable of automatically adjusting volume and multimedia playing system

    CN202976756U