Volume adjustment method, device, apparatus and medium

By acquiring the environmental characteristics of the first party in an audio/video call and adjusting the volume of the second party in the call before the call, the problem of inappropriate volume in audio/video connections is solved, and volume adjustment adapted to the environment is achieved, improving call efficiency and privacy protection.

CN122496580APending Publication Date: 2026-07-31INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2025-08-15
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In audio and video calls, if the customer service representative's volume is too high, it will leak customer privacy; if the volume is too low, it will affect the customer's experience and lead to low call efficiency.

Method used

Before a call, the audio characteristics of the environment in which the first caller is located are obtained, and the volume of the second caller is adjusted according to these characteristics to ensure that the volume is appropriate for the ambient noise level, avoid privacy leaks, and ensure clear calls.

Benefits of technology

By acquiring environmental characteristics before a call and adjusting the volume in real time, the problem of unclear hearing or privacy leaks caused by inappropriate volume is solved, improving call efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496580A_ABST
    Figure CN122496580A_ABST
Patent Text Reader

Abstract

This invention discloses a volume adjustment method, apparatus, device, and medium, relating to the audio and video field and applicable to the fintech sector. The method includes: before a business call between a first party and a second party, acquiring a first audio feature of the environment in which the first party is located; during the call between the first party and the second party, acquiring the audio of the second party; and adjusting the call volume of the second party's audio based on the first audio feature. Embodiments of this invention enable users to clearly hear the other party's voice, thereby improving call efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio and video, and can be used in the field of financial technology, particularly to a volume adjustment method, apparatus, device, and medium. Background Technology

[0002] With the development of mobile finance, many functions that previously required offline staff intervention have been transferred to mobile clients and used remotely via audio and video connections.

[0003] However, if the volume of the customer service representative is too high during audio and video calls, it will leak the customer's privacy; if the volume is too low, it will affect the customer's use. Summary of the Invention

[0004] This invention provides a volume adjustment method, apparatus, device, and medium that enables users to clearly hear the other party's voice, thereby improving call efficiency.

[0005] According to one aspect of the present invention, an embodiment of the present invention provides a volume adjustment method, the method comprising:

[0006] Before a business call between the first party and the second party, the first audio features of the environment in which the first party is located are obtained.

[0007] During the call between the first party and the second party, the audio of the second party is acquired;

[0008] The call volume of the second party's audio is adjusted based on the first audio feature.

[0009] According to another aspect of the present invention, embodiments of the present invention also provide a volume adjustment device, the device comprising:

[0010] The feature acquisition module is used to acquire the first audio features of the environment in which the first caller is located before the business call between the first caller and the second caller.

[0011] An audio acquisition module is used to acquire the audio of the second party during a call between the first party and the second party.

[0012] The volume adjustment module is used to adjust the call volume of the second party's audio based on the first audio characteristics.

[0013] According to another aspect of the present invention, embodiments of the present invention also provide a volume adjustment device, the volume adjustment device comprising:

[0014] At least one processor; and

[0015] A memory that is communicatively connected to at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the volume adjustment method of any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the volume adjustment method of any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the volume adjustment method described in any embodiment of the present invention.

[0019] The technical solution of this invention obtains the first audio features of the environment of the first caller in advance before the call, and adjusts the volume of the audio of the second caller obtained during the call based on the first audio features. It can flexibly adjust the output audio volume based on the call environment, solving the problem in the prior art that the customer service volume is inappropriate, resulting in unclear hearing or privacy leakage. It can output audio with an environment-appropriate volume, so that the user can clearly hear the other party's voice, thereby improving call efficiency. Furthermore, by obtaining the audio features of the environment before the call and adjusting the output volume in real time during the call, the real-time performance of volume adjustment is improved.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a volume adjustment method provided according to an embodiment of the present invention;

[0023] Figure 2 This is a flowchart of a volume adjustment method provided according to an embodiment of the present invention;

[0024] Figure 3 This is a structural diagram of a volume adjustment device provided according to an embodiment of the present invention;

[0025] Figure 4 This is a schematic diagram of the structure of a volume adjustment device provided in an embodiment of the present invention. Detailed Implementation

[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] The acquisition, storage, and application of environmental audio, audio of both parties in a call, user characteristics, and user information involved in the technical solutions of this invention comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0029] Figure 1 This is a flowchart illustrating a volume adjustment method provided in an embodiment of the present invention. This embodiment is applicable to situations where a user adjusts the volume of a staff member during an audio / video call with the staff member via a client. The method can be executed by a volume adjustment device, which can be implemented in hardware and / or software.

[0030] See Figure 1 The volume adjustment methods shown include:

[0031] S101. Before the business call between the first caller and the second caller, obtain the first audio features of the environment in which the first caller is located.

[0032] In this context, the first party in the call is the user conducting the audio / video call using the client. The second party is the staff member conducting the audio / video call with the first party. The audio / video call can include either audio or video calls. A business call can refer to a call related to the client's functions. The environment of the first party can refer to their actual surroundings. The first audio feature can refer to the characteristics of the audio from the environment that can be captured. The first audio feature can be obtained by performing convolution or encoding on the captured environmental audio. Alternatively, it can be obtained by extracting features from the captured environmental audio in the frequency and / or time domains. The first audio feature is typically obtained with the authorization of the first party before the business call begins.

[0033] In some embodiments, "before the business call" can refer to the steps in a fixed business processing flow that precede the steps corresponding to the business call.

[0034] In some embodiments, ambient audio is acquired by capturing the environment of the first caller, and features are extracted to obtain first audio features. For example, a high-sensitivity and low-noise microphone can be selected to ensure accurate capture of ambient audio. An appropriate sampling frequency (e.g., 44.1 kHz) is set to ensure sufficient audio detail is captured. The acquired audio data is converted into a digital signal, and ambient audio is acquired in real time through the application programming interface provided by the mobile phone operating system.

[0035] In some implementations, the acquired ambient audio is preprocessed, for example, by using bandpass filters to remove high-frequency and low-frequency noise, preserving sound within the audible range (e.g., 20Hz-20kHz). Noise suppression algorithms, such as spectral subtraction or adaptive filters, are applied to reduce the impact of background noise. Redundant signal analysis can be eliminated, and the output volume can be adjusted only based on noise within the audible range, thus improving the accuracy of volume adjustment.

[0036] S102. During the call between the first party and the second party, obtain the audio of the second party.

[0037] The call process can refer to the client collecting the voice audio of the first party in the call, transmitting it to the terminal device of the second party in the call, and playing it, or the terminal device collecting the voice audio of the second party in the call and transmitting it to the client of the first party in the call for playback.

[0038] S103. Adjust the call volume of the second party's audio based on the first audio feature.

[0039] The first audio feature describes the volume of noise in the first party's environment. The call volume can refer to the volume of the second party's audio output in the environment of the first party. The call volume can be determined based on the first audio feature to ensure that the call volume is greater than the noise level in the first party's environment, and that the call volume is not too high to avoid privacy breaches. In some embodiments, a lower volume limit can be determined based on the first audio feature, and the call volume can be adjusted according to the lower volume limit to ensure that the call volume is greater than the lower volume limit. For example, the call volume can be adjusted to a volume slightly higher than the lower volume limit.

[0040] The technical solution of this invention obtains the first audio features of the environment of the first caller in advance before the call, and adjusts the volume of the audio of the second caller obtained during the call based on the first audio features. It can flexibly adjust the output audio volume based on the call environment, solving the problem in the prior art that the customer service volume is inappropriate, resulting in unclear hearing or privacy leakage. It can output audio with an environment-appropriate volume, so that the user can clearly hear the other party's voice, thereby improving call efficiency. Furthermore, by obtaining the audio features of the environment before the call and adjusting the output volume in real time during the call, the real-time performance of volume adjustment is improved.

[0041] In an optional embodiment, obtaining the first audio feature of the environment in which the first caller is located includes: obtaining the environmental audio of the first caller; segmenting the environmental audio into unit signals; determining the amplitude of the environmental audio based on the amplitude of each unit signal; converting the environmental audio into a frequency domain signal, and obtaining spectral features from the converted frequency domain signal; determining the noise frequency of at least one noise source based on the spectral features; and determining the first audio feature of the first caller based on the noise frequency of each noise source and the amplitude of the environmental audio.

[0042] Ambient audio refers to audio collected from the environment in which the first party in the call is located. Ambient audio is a continuous time-domain signal and can be divided into short frames (e.g., 20ms per frame), with each frame representing a unit signal. This segmentation of unit signals facilitates subsequent processing. The amplitude of each unit signal can be calculated; amplitude is also the energy value. The average amplitude of all unit signal amplitudes is then taken as the amplitude of the ambient audio. The amplitude of the ambient audio reflects the intensity of the sound. The average amplitude of multiple consecutive unit signals is calculated to smooth out fluctuations.

[0043] To perform frequency domain analysis on a time-domain signal, it's essential to first convert it to a frequency-domain signal. Fourier transform can be used to convert ambient audio into a frequency-domain signal. This frequency-domain signal can be a superposition of at least one frequency component. Analysis and processing can then be performed on the frequency-domain signal to obtain at least one frequency component. The Fast Fourier Transform (FFT) can be used to convert the time-domain signal back to a frequency-domain signal, and the intensity of each frequency component can be analyzed. Each frequency component and its intensity (peak frequency) are then defined as spectral characteristics. Based on the frequency components and their intensities, noise sources are identified, and the frequencies corresponding to these noise sources in the spectral characteristics are determined as the noise frequencies of the noise source signal.

[0044] The amplitude of the ambient audio, the noise source, and the noise audio are identified as the first audio feature.

[0045] Signals with higher frequencies and / or higher amplitudes typically interfere with the first party in a call hearing the other party's audio. Signals with higher frequencies and higher amplitudes in the ambient audio can be identified as noise sources.

[0046] It is evident that by performing time-domain analysis on ambient audio to obtain amplitude, and frequency-domain analysis on ambient audio to obtain noise source components and noise frequencies, the amplitude and frequency domain can be determined as the first audio feature, which can enrich the content dimension of the audio feature. Adjusting the volume based on the first audio feature can improve the accuracy of the volume adjustment.

[0047] In an optional embodiment, determining the first audio feature of the first caller based on the noise frequency of each of the noise sources includes: classifying the ambient audio for noise levels based on the noise frequency of each of the noise sources and the amplitude of the ambient audio to obtain a classification result; and determining the first audio feature of the first caller based on the classification result, the noise frequency of each of the noise sources, and the amplitude of the ambient audio.

[0048] This can be achieved by setting at least one threshold range corresponding to a specific type. The threshold range into which the ambient audio falls is determined based on the noise frequency and amplitude, and the classification result corresponding to that threshold range is identified as the noise level classification result for the ambient audio. This classification result can be added to the first audio feature.

[0049] In some embodiments, multiple threshold ranges are set according to a preset noise level standard (e.g., 0-100 dB), such as 0-20 dB, 21-40 dB, 41-60 dB, 61-80 dB, and 81-100 dB. The calculated amplitude and frequency can be mapped to the corresponding noise level. The noise level classification can be dynamically adjusted based on real-time user feedback to adapt to changes in different environments.

[0050] It is evident that classifying the noise level of ambient audio by amplitude and frequency, obtaining the classification result, and using the classification result as the first audio feature can enrich the content of the first audio feature and further improve the accuracy of volume adjustment.

[0051] In an optional embodiment, obtaining the first audio feature of the environment of the first caller before the business call between the first caller and the second caller includes: obtaining the environmental audio of the first caller when the first caller triggers the target step of the security verification process, wherein the second caller is the business user of the security verification process; the security verification process includes at least the target step and the call step, wherein the target step is before the call step, and the call step is the step of the business call between the first caller and the second caller.

[0052] The security verification process typically includes a call step. This call step is used for manual verification, specifically through a business call. The target step can be a step in the security verification process that precedes the call step. The target step is usually a prerequisite for the call step. For example, the target step could be an information input step. The business user then performs manual verification on the first party in the call.

[0053] In some embodiments, when a user logs into a business system account using a client, and a high-security verification operation fails, such as consecutive incorrect password entries or failed face verification, the client is triggered to execute a security verification process.

[0054] In one example, a user initiates a transaction request using a client. Before executing the transaction request, a security verification process is required. This security verification process consists of four steps, with the last step being the call step. The first step executed can be identified as the target step. The client can trigger an authorization confirmation message and provide it to the user. When the user confirms authorization, ambient audio is collected starting from the target step.

[0055] It is evident that by limiting the state before the business call between the first and second parties to the state when the target step in the security verification process is executed before the call steps, the ambient audio can be quickly acquired and the first audio feature extracted during the security verification process. The volume can be adjusted at the beginning of the business call, thereby improving the call efficiency of security verification and thus improving the execution efficiency of security verification operations.

[0056] Figure 2This is a flowchart illustrating a volume adjustment method provided in an embodiment of the present invention. Based on the above embodiments, the volume adjustment method is optimized as follows: during a call between the first and second parties, a second audio feature is acquired. The call volume of the second party's audio is adjusted according to the first audio feature, specifically: the call volume of the second party's audio is adjusted based on both the first and second audio features.

[0057] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments.

[0058] See Figure 2 The volume adjustment methods shown include:

[0059] S201. Before the business call between the first caller and the second caller, obtain the first audio features of the environment in which the first caller is located.

[0060] S202. During the call between the first party and the second party, obtain the audio of the second party.

[0061] S203. During the call between the first party and the second party, acquire the second audio feature.

[0062] The second audio feature can refer to features extracted from audio during a call. In some embodiments, the second audio feature includes at least one of the following: audio features of the first caller, audio features of the second caller, historical volume adjustment information of the first caller, and ambient audio during the call. The audio features of the first caller can be obtained by extracting features based on attribute information provided with user authorization. The audio features of the second caller may include features extracted from the second caller's ambient audio and / or features extracted from the second caller's voice audio.

[0063] In an optional embodiment, the second audio feature includes: the first caller's historical volume adjustment information and / or the second caller's voice audio features.

[0064] The historical volume adjustment information can refer to the output volume adjusted by the first party at a historical time. The speech audio features can refer to the features extracted from the speech audio of the second party. In some embodiments, time-domain and frequency-domain features can be extracted from the speech audio of the second party to obtain speech audio features. For example, the speech audio can be segmented into unit signals; the amplitude of the speech audio can be determined based on the amplitude of each unit signal; the speech audio can be converted into a frequency-domain signal, and spectral features can be obtained from the converted frequency-domain signal; the frequency of the speech audio can be determined based on the spectral features; and the speech audio features can be determined based on the frequency and amplitude of the speech audio. Alternatively, the speech audio can be input into a convolutional network or encoder for feature extraction processing to obtain the output speech audio features.

[0065] It is evident that by limiting the second audio feature to the historical volume adjustment information of the first caller and / or the voice audio features of the second caller, the call volume can be adjusted by combining the volume preference information of the first caller and the real-time audio of the second caller. This increases the deterministic factors for adjusting the call volume and thus improves the accuracy of volume adjustment.

[0066] S204. Adjust the call volume of the second caller's audio based on the first audio feature and the second audio feature.

[0067] In this method, the first audio feature and the second audio feature can be fused together, and the call volume of the second party's audio can be determined based on the fusion result. In some embodiments, the fusion result of the first audio feature and the second audio feature can be decoded to obtain the call volume. For example, a first volume can be determined based on the first audio feature, a second volume can be determined based on the second audio feature, and a fusion result can be determined based on the first volume and the second volume, with the fused volume being determined as the call volume. The fusion result can be the average or weighted average of the first and second volumes, etc.

[0068] This invention, by acquiring a second audio feature during a call and combining it with a first audio feature and a second audio feature to adjust the call volume, can enrich the content of the input data for adjusting the volume, adapt to various call scenarios, and flexibly and accurately adjust the output volume.

[0069] In an optional embodiment, adjusting the call volume of the second party's audio based on the first audio feature and the call audio feature includes: obtaining a first weight of the first audio feature and a second weight of the second audio feature; determining a target volume based on the first audio feature, the first weight, the second audio feature, and the second weight; and adjusting the call volume of the second party's audio using the target volume.

[0070] In this system, the first weight is greater than the second weight, which can be determined experimentally. The first volume can be determined based on the first audio feature, and the second volume based on the second audio feature. The first volume determines the lower limit of the call volume. The second volume determines the average call volume. A weighted average is calculated based on the first volume, the first weight, the second volume, and the second weight to obtain the target volume. The call volume is greater than or equal to the target volume. The target volume can be defined as the call volume, or a volume slightly larger than the target volume can be defined as the call volume.

[0071] As can be seen, by configuring the first weight of the first audio feature and the second weight of the second audio feature, and by combining the first weight with the second weight to determine the call volume, the weight of each audio feature affecting the call volume can be flexibly adjusted according to the business scenario, thereby providing targeted and accurate call volume for different scenarios.

[0072] The volume adjustment method implemented in this invention can improve user experience: by automatically adjusting the volume, users can obtain the best auditory experience in any environment; it can provide personalized services: taking into account the specific needs of different user groups, it provides more personalized services; and it can improve efficiency: reducing the number of times users manually adjust the volume and improving the efficiency of remote audio and video calls.

[0073] Figure 3 This is a schematic diagram of a volume adjustment device provided in an embodiment of the present invention. This embodiment of the present invention is applicable to situations where the volume of code files prepared for submission in a code submission pipeline is adjusted. The device can execute a volume adjustment method and can be implemented in hardware and / or software.

[0074] See Figure 3 The volume adjustment device shown includes:

[0075] The feature acquisition module 301 is used to acquire the first audio features of the environment in which the first caller is located before the business call between the first caller and the second caller.

[0076] The audio acquisition module 302 is used to acquire the audio of the second party during the call between the first party and the second party.

[0077] The volume adjustment module 303 is used to adjust the call volume of the second party's audio based on the first audio feature.

[0078] The technical solution of this invention obtains the first audio features of the environment of the first caller in advance before the call, and adjusts the volume of the audio of the second caller obtained during the call based on the first audio features. It can flexibly adjust the output audio volume based on the call environment, solving the problem in the prior art that the customer service volume is inappropriate, resulting in unclear hearing or privacy leakage. It can output audio with an environment-appropriate volume, so that the user can clearly hear the other party's voice, thereby improving call efficiency. Furthermore, by obtaining the audio features of the environment before the call and adjusting the output volume in real time during the call, the real-time performance of volume adjustment is improved.

[0079] Optionally, the feature acquisition module 301 is specifically used for:

[0080] Obtain the ambient audio of the first party in the call;

[0081] The ambient audio is converted into a frequency domain signal, and the spectral characteristics of the converted frequency domain signal are obtained.

[0082] The noise frequency of at least one noise source is determined based on the spectral characteristics.

[0083] The first audio feature of the first caller is determined based on the noise frequency of each of the noise sources.

[0084] Optionally, the feature acquisition module 301 is specifically used for:

[0085] Based on the noise frequency of each noise source, the noise levels of each noise source are classified to obtain the classification results;

[0086] The first audio feature of the first caller is determined based on the classification results and the noise frequencies of each noise source.

[0087] Optionally, the volume adjustment device also includes:

[0088] The second feature acquisition module is used to acquire a second audio feature during the call between the first party and the second party.

[0089] Volume adjustment module 303 is specifically used for:

[0090] The call volume of the second party's audio is adjusted based on the first audio feature and the second audio feature.

[0091] Optional, the volume adjustment module 303 is specifically used for:

[0092] Obtain the first weight of the first audio feature and the second weight of the second audio feature;

[0093] The target volume is determined based on the first audio feature, the first weight, the second audio feature, and the second weight;

[0094] The target volume is used to adjust the audio volume of the second party in the call.

[0095] Optionally, the second audio feature includes: the first party's historical volume adjustment information and / or the second party's voice audio features.

[0096] Optionally, the feature acquisition module 301 is specifically used for:

[0097] When the first caller triggers the target step of the security verification process, the environmental audio of the first caller is obtained, and the second caller is the business user of the security verification process; the security verification process includes at least the target step and the call step, the target step is before the call step, and the call step is the step of the business call between the first caller and the second caller.

[0098] The volume adjustment device provided in the embodiments of the present invention can execute the volume adjustment method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the volume adjustment method.

[0099] Figure 4 A schematic diagram of the structure of a volume adjustment device 400 that can be used to implement an embodiment of the present invention is shown.

[0100] like Figure 4 As shown, the volume adjustment device 400 includes at least one processor 401 and a memory, such as a read-only memory 402 or a random access memory 403, communicatively connected to the at least one processor 401. The memory stores computer programs executable by the at least one processor. The processor 401 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 402 or loaded from storage unit 408 into the random access memory 403. The random access memory 403 may also store various programs and data required for the operation of the volume adjustment device 400. The processor 401, read-only memory 402, and random access memory 403 are interconnected via a bus 404. An input / output interface 405 is also connected to the bus 404.

[0101] Multiple components in the volume adjustment device 400 are connected to the input / output interface 405, including: an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, optical disk, etc.; and a communication unit 409, such as a network card, modem, wireless transceiver, etc. The communication unit 409 allows the volume adjustment device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0102] Processor 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 401 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 401 performs the various methods and processes described above, such as volume adjustment methods.

[0103] In some embodiments, the volume adjustment method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the volume adjustment device 400 via read-only memory 402 and / or communication unit 409. When the computer program is loaded into random access memory 403 and executed by processor 401, one or more steps of the volume adjustment method described above may be performed. Alternatively, in other embodiments, processor 401 may be configured to perform the volume adjustment method by any other suitable means (e.g., by means of firmware).

[0104] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), complex programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0105] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0106] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0107] To provide user interaction, the systems and techniques described herein can be implemented on an operational detection device, which includes: a display device (e.g., a cathode ray tube or liquid crystal monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the volume adjustment device. Other types of devices can also be used to provide user interaction; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0108] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0109] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product within the cloud computing service system. This addresses the shortcomings of traditional physical hosts and virtual private servers, such as high management difficulty and weak business scalability.

[0110] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0111] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method of adjusting volume, characterized by, The method includes: Before a business call between the first party and the second party, the first audio features of the environment in which the first party is located are obtained. During the call between the first party and the second party, the audio of the second party is acquired; The call volume of the second party's audio is adjusted based on the first audio feature.

2. The method of claim 1, wherein, The step of obtaining the first audio feature of the environment in which the first caller is located includes: Obtain the ambient audio of the first party in the call; The ambient audio is segmented into unit signals; The amplitude of the ambient audio is determined based on the amplitude of each unit signal. The ambient audio is converted into a frequency domain signal, and the spectral characteristics of the converted frequency domain signal are obtained. The noise frequency of at least one noise source is determined based on the spectral characteristics. The first audio feature of the first caller is determined based on the noise frequency of each noise source and the amplitude of the ambient audio.

3. The method of claim 2, wherein, Determining the first audio feature of the first caller based on the noise frequency of each noise source and the amplitude of the ambient audio includes: Based on the noise frequency of each noise source and the amplitude of the ambient audio, the ambient audio is classified according to its noise level to obtain the classification result; Based on the classification results, the noise frequencies of each noise source, and the amplitude of the ambient audio, the first audio feature of the first caller is determined.

4. The method of claim 1, wherein, Also includes: During the call between the first party and the second party, the second audio feature is acquired; The step of adjusting the call volume of the second party's audio based on the first audio feature includes: The call volume of the second party's audio is adjusted based on the first audio feature and the second audio feature.

5. The method of claim 4, wherein, The step of adjusting the call volume of the second party's audio based on the first audio feature and the call audio feature includes: Obtain the first weight of the first audio feature and the second weight of the second audio feature; The target volume is determined based on the first audio feature, the first weight, the second audio feature, and the second weight; The target volume is used to adjust the audio volume of the second party in the call.

6. The method of claim 4, wherein, The second audio feature includes: the historical volume adjustment information of the first caller and / or the voice audio features of the second caller.

7. The method of claim 1, wherein, Prior to the business call between the first party and the second party, obtaining the first audio features of the environment in which the first party is located includes: When the first caller triggers the target step of the security verification process, the environmental audio of the first caller is obtained, and the second caller is the business user of the security verification process; the security verification process includes at least the target step and the call step, the target step is before the call step, and the call step is the step of the business call between the first caller and the second caller.

8. A volume adjusting apparatus characterized by comprising: The device includes: The feature acquisition module is used to acquire the first audio features of the environment in which the first caller is located before the business call between the first caller and the second caller. The audio acquisition module is used to acquire the audio of the second party during the call between the first party and the second party. The volume adjustment module is used to adjust the call volume of the second party's audio based on the first audio characteristics.

9. A volume adjusting apparatus characterized by comprising: The volume adjustment device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the volume adjustment method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the volume adjustment method according to any one of claims 1-7.