Private call method

CN122799792APending Publication Date: 2026-09-22MIGU MUSIC CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610842139.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-22

AI Technical Summary

Benefits of technology

[0024] Therefore, this solution addresses the shortcomings of traditional audio noise cancellation technology, such as high cost, poor adaptability, and unstable effect, by using environmental detection and reinforcement learning-driven adaptive cancellation. This reduces noise cancellation costs and improves environmental adaptability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799792A_ABST
    Figure CN122799792A_ABST
Patent Text Reader

Abstract

This disclosure proposes a privacy-preserving call method, relating to the field of voice call technology. The method includes: in response to a privacy-preserving call request, acquiring the current ambient audio, and generating a second audio corresponding to the first audio based on the first audio, wherein the first audio is the audio being played by the earpiece device of the terminal, and the second audio is an inverse audio with the same amplitude but opposite phase to the first audio; calculating a similarity value between the ambient audio and the first audio; in response to the calculation that the similarity value between the ambient audio and the first audio is greater than a similarity threshold, adjusting the second audio based on a Q-Learning algorithm to generate a third audio; and sending the third audio to the terminal's speaker device for playback. Therefore, this solution, through environmental detection and reinforcement learning-driven adaptive cancellation, solves the shortcomings of traditional audio cancellation techniques, such as high cost, poor adaptability, and unstable effects, reducing cancellation costs and improving environmental adaptability and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of voice call technology, and more particularly to a method for making private calls. Background Technology

[0002] In applications such as mobile communication, intelligent voice interaction, and privacy protection, preventing accidental leakage of call audio through device speakers or microphones (e.g., in hands-free calls, voice assistant responses) has become an important security requirement. To suppress such audio leakage, existing technologies often employ the principle of Active Noise Cancellation (ANC). This involves generating an anti-phase sound wave with the same amplitude but opposite phase as the original audio signal and superimposing it on the original audio in the leakage path, thereby achieving acoustic cancellation and a "silencing" effect. Summary of the Invention

[0003] This disclosure aims to at least partially address one of the technical problems in the related art.

[0004] Therefore, one purpose of this disclosure is to propose a method for private calling.

[0005] The second objective of this disclosure is to propose a privacy call device.

[0006] The third objective of this disclosure is to propose an electronic device.

[0007] The fourth objective of this disclosure is to provide a non-transitory computer-readable storage medium.

[0008] The fifth objective of this disclosure is to provide a computer program product.

[0009] To achieve the above objectives, a first aspect of this disclosure provides a privacy call method applied to a terminal, comprising: in response to a privacy call request, acquiring current ambient audio, and generating a second audio corresponding to the first audio based on the first audio, wherein the first audio is audio being played by the terminal's earpiece device, and the second audio is an out-of-phase audio with the same amplitude but opposite phase to the first audio; calculating a similarity value between the ambient audio and the first audio; in response to the similarity value between the ambient audio and the first audio being greater than a similarity threshold, adjusting the second audio based on a Q-Learning algorithm to generate a third audio; and sending the third audio to the terminal's speaker device for playback.

[0010] According to one embodiment of this disclosure, adjusting the second audio based on the Q-Learning algorithm to generate a third audio includes: selecting a target candidate action with the largest expected value from a target action space, the target action space including multiple candidate actions and the expected value corresponding to each candidate action; adjusting the second audio by executing the target candidate action; playing the adjusted second audio through the speaker device; calculating the similarity value between the adjusted first audio and the ambient audio; repeating the above steps of selecting the target candidate action with the largest expected value from the target action space and subsequent steps until the similarity value between the adjusted first audio and the ambient audio is less than or equal to the similarity threshold; and outputting the finally adjusted second audio as the third audio.

[0011] According to one embodiment of this disclosure, the method further includes: in response to playing the adjusted second audio and the similarity value between the first audio and the ambient audio being greater than the similarity threshold, updating the expected values ​​of all candidate actions in the target action space.

[0012] According to one embodiment of this disclosure, updating the expected value of all candidate actions in the target action space includes: for any candidate action in the target action space, calculating a first reward value for adjusting the second audio by executing the candidate action; after adjusting the second audio by executing the candidate action, selecting the next candidate action with the largest expected value from the target action space based on the candidate action, and calculating a second reward value for adjusting the second audio by executing the next candidate action; and calculating the expected value of the candidate action based on the first reward value and the second reward value.

[0013] According to one embodiment of this disclosure, calculating a target reward value, wherein the target reward value is either the first reward value or the second reward value, includes: after playing the adjusted second audio, calculating the similarity value between the first audio and the ambient audio; and using the reciprocal of the similarity value as the target reward value.

[0014] According to one embodiment of this disclosure, constructing the target action space includes: selecting target candidate actions from a candidate action space based on an ε-greedy exploration strategy, wherein the candidate action space includes multiple candidate actions and the expected value corresponding to each candidate action; after adjusting a fourth audio by executing the target candidate action, selecting the target candidate action with the largest expected value from the candidate action space, wherein the fourth audio is a test audio for adjusting the candidate action space; after adjusting the fourth audio by executing the target candidate action and playing the adjusted fourth audio through the speaker device, calculating the similarity value between the first audio and the test environment audio, wherein the test environment audio is the environmental audio for adjusting the candidate action space; in response to the similarity value between the first audio and the test environment audio being greater than the similarity threshold, updating the expected values ​​of all candidate actions in the candidate action space; repeating the above process of selecting target candidate actions and their subsequent actions from the candidate action space based on the ε-greedy exploration strategy until, after playing the adjusted fourth audio, the similarity value between the first audio and the test environment audio is less than or equal to the similarity threshold, and outputting the adjusted candidate action space as the target action space.

[0015] According to one embodiment of this disclosure, the step of selecting a target candidate action from the candidate action space based on the ε-greedy exploration strategy includes: selecting a target candidate action from the candidate action space based on a first selection strategy or a second selection strategy, wherein the probability of executing the first selection strategy is ε, and the probability of executing the second selection strategy is 1-ε.

[0016] According to one embodiment of this disclosure, the first selection strategy is to randomly select an action from the candidate action space, and the second selection strategy is to select the action with the highest expected value from the candidate action space.

[0017] According to one embodiment of this disclosure, calculating the similarity value between the ambient audio and the first audio includes: obtaining a first audio feature and a second audio feature of the ambient audio; and calculating the similarity value between the ambient audio and the first audio based on the first audio feature and the second audio feature.

[0018] According to one embodiment of this disclosure, calculating the similarity value between the ambient audio and the first audio based on the first audio feature and the second audio feature includes: calculating the feature cosine similarity value between the first audio feature and the second audio feature, and calculating the text cosine similarity value between the ambient audio and the first audio; and calculating the similarity value between the ambient audio and the first audio based on the feature cosine similarity value and the text cosine similarity value.

[0019] According to one embodiment of this disclosure, calculating the similarity value between the ambient audio and the first audio based on the first audio feature and the second audio feature includes: calculating the feature cosine similarity value between the first audio feature and the second audio feature, and calculating the text cosine similarity value between the ambient audio and the first audio; and calculating the similarity value between the ambient audio and the first audio based on the feature cosine similarity value and the text cosine similarity value.

[0020] To achieve the above objectives, a second aspect of this disclosure provides a privacy call device, comprising: a acquisition module, configured to acquire current ambient audio in response to a privacy call request, and generate a second audio corresponding to the first audio based on the first audio, wherein the first audio is audio being played by the earpiece device of a terminal, and the second audio is an out-of-phase audio with the same amplitude but opposite phase to the first audio; a calculation module, configured to calculate a similarity value between the ambient audio and the first audio; a generation module, configured to adjust the second audio based on a Q-Learning algorithm to generate a third audio in response to the calculation that the similarity value between the ambient audio and the first audio is greater than a similarity threshold; and a playback module, configured to send the third audio to the speaker device of the terminal for playback.

[0021] To achieve the above objectives, a third aspect of this disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to implement the privacy call method as described in the first aspect of this disclosure.

[0022] To achieve the above objectives, a fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to implement the privacy call method as described in the first aspect of this disclosure.

[0023] To achieve the above objectives, a fifth aspect of this disclosure provides a computer program product including a computer program that, when executed by a processor, implements the privacy call method as described in the first aspect of this disclosure.

[0024] Therefore, this solution addresses the shortcomings of traditional audio noise cancellation technology, such as high cost, poor adaptability, and unstable effect, by using environmental detection and reinforcement learning-driven adaptive cancellation. This reduces noise cancellation costs and improves environmental adaptability and user experience. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of a privacy call method according to one embodiment of the present disclosure;

[0026] Figure 2 This is a schematic diagram of another privacy call method according to one embodiment of this disclosure; Figure 3 This is a schematic diagram of another privacy call method according to one embodiment of this disclosure; Figure 4 This is a schematic diagram of another privacy call method according to one embodiment of this disclosure; Figure 5 This is a schematic diagram of another privacy call method according to one embodiment of this disclosure; Figure 6 This is a schematic diagram of a privacy call device according to one embodiment of the present disclosure; Figure 7 This is a schematic diagram of an electronic device according to one embodiment of the present disclosure. Detailed Implementation

[0027] Embodiments of this disclosure are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.

[0028] The acquisition, storage, use, and processing of data in this disclosed technical solution all comply with the relevant provisions of relevant laws and regulations.

[0029] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0030] Current noise cancellation solutions typically rely on a phase-locked loop (PLL) circuit in conjunction with a power amplifier to generate a high-precision inverted audio signal in real time. This PLL circuit tracks the frequency and phase of the original audio signal and drives the power amplifier to output a strictly inverted and amplitude-matched cancellation signal. However, this analog hardware-based noise cancellation architecture has significant technical limitations: First, it is highly hardware-dependent. PLL circuits are application-specific analog integrated circuits, which are complex to design, costly, and require close integration with audio codecs and power amplifier modules. This not only increases the hardware BOM cost of terminal devices (such as smartphones, smart speakers, and in-vehicle terminals) but also places higher demands on PCB layout, power supply noise suppression, and signal integrity, limiting the widespread application of this technology in resource-constrained or low-cost devices.

[0031] Secondly, it has poor environmental adaptability. In actual call scenarios, the acoustic environment is highly dynamic—including changes in device placement, differences in surrounding reflective surfaces, fluctuations in background noise, and changes in the user's distance from the microphone / speaker. These factors can cause amplitude attenuation, phase shift, or multipath reverberation in the original audio signal during the leakage path. Traditional phase-locked loop circuits can only generate a fixed inverted signal based on the source signal and cannot adaptively compensate for distortions in the propagation path. Therefore, the cancellation signal and the actual leaked signal cannot always maintain an ideal "equal amplitude and inverted phase" relationship, resulting in a significant decrease in the noise reduction effect, and even an enhancement effect in some frequency bands, further exacerbating the risk of privacy leakage.

[0032] To address the aforementioned issues, this disclosure proposes a method for private calling. Figure 1 This is a schematic diagram of a privacy call method according to one embodiment of the present disclosure, such as... Figure 1 As shown, this private call method includes the following steps: S101, in response to the privacy call request, obtain the current ambient audio, and generate a second audio corresponding to the first audio based on the first audio, wherein the first audio is the audio being played by the terminal's earpiece device, and the second audio is an inverse audio with the same amplitude but opposite phase to the first audio.

[0033] The privacy call method of this application embodiment can be applied to privacy call scenarios. The execution subject of the privacy call in this application embodiment can be the privacy call device of this application embodiment, which can be installed on an electronic device.

[0034] It should be noted that privacy call requests can be triggered manually by the user or automatically identified by the system based on context such as keywords in the call content, geographical location (e.g., public places), and device status (e.g., hands-free enabled).

[0035] In this embodiment of the disclosure, when the terminal initiates or receives a private call request (e.g., when "private mode" or "anti-eavesdropping mode" is enabled or when entering a sensitive voice interaction scenario), the system automatically activates the environmental monitoring module.

[0036] The system collects ambient audio signals in the current space in real time through the microphone array or high-sensitivity pickup unit built into the terminal device. Ambient audio contains acoustic components such as the original playback content that may be leaked, background noise, reverberation, and multipath reflections, and is a key basis for determining whether audio leakage exists.

[0037] It should be noted that the second audio refers to the inverse (or adversarial) audio signal generated according to this scheme and used to cancel the original audio signal leaked in the environment in privacy protection or active noise cancellation scenarios. Its core purpose is to weaken or even eliminate the audibility of sensitive speech in physical space through the principle of sound wave superposition, thereby preventing eavesdropping or recording by others.

[0038] The second audio signal is an artificially synthesized audio signal, typically designed to be highly correlated with the original playback audio (the first audio signal) in terms of frequency components and time alignment, but with specific adjustments to phase and / or amplitude to achieve acoustic cancellation.

[0039] S102, calculate the similarity value between the ambient audio and the first audio.

[0040] It should be noted that the first audio refers to the original audio signal currently being output by the terminal playback device (such as the other party's voice, voice assistant reply, etc.). This signal can be directly obtained from the audio driver layer or the system mixer (without relying on external sound pickup, avoiding delay and distortion).

[0041] There are various methods for calculating the similarity value between the ambient audio and the first audio, and no limitation is made here.

[0042] In this embodiment of the disclosure, the ambient audio and the first audio can be cross-domain aligned and feature compared. For example, acoustic features (such as Mel spectrum, MFCC, or deep embedding vector) can be extracted from the two audio streams respectively. Dynamic time warping (DTW) or cross-correlation function can be used to calculate their time-frequency similarity, or a normalized similarity value can be output through a pre-trained audio similarity model (such as a contrastive learning model based on Siamese network).

[0043] S103, in response to the fact that the similarity value between the ambient audio and the first audio is greater than the similarity threshold, the second audio is adjusted based on the Q-Learning algorithm to generate the third audio.

[0044] Traditional solutions typically generate fixed-phase inverted audio based on ideal acoustic models. This is insufficient to handle variations in leakage paths caused by differences in devices, wearing methods, environmental reverberation, or background noise in real-world use, often resulting in insufficient cancellation or overcompensation, leading to distortion or even feedback. Our solution, however, actively collects ambient audio and calculates its similarity to the first audio (the content played through the earpiece) after generating an initially inverted second audio. Once the similarity exceeds a threshold (indicating identifiable voice leakage), a dynamic adjustment mechanism based on the Q-Learning algorithm is triggered. This mechanism performs multi-dimensional optimization of the inverted audio (such as gain, phase, and frequency response) to generate a third audio that better matches the current acoustic environment. This closed-loop feedback mechanism can detect the degree of leakage online and intelligently correct the cancellation strategy. It not only effectively suppresses the risk of eavesdropping but also avoids audio distortion or power waste caused by blindly increasing the inverted volume, thus achieving a highly robust, low-distortion, and privacy-protected call experience in complex real-world scenarios.

[0045] To enhance privacy protection, it is necessary to further dynamically adjust the second audio (i.e., the antiphase sound wave currently used to cancel the leakage sound) based on the Q-Learning reinforcement learning algorithm to generate an optimized third audio, so as to achieve more accurate acoustic cancellation.

[0046] In this embodiment, when the similarity value exceeds a preset similarity threshold, it indicates the presence of significant original audio leakage signals in the environment, posing a privacy risk. At this time, the system triggers an intelligent noise reduction mechanism and executes a preset Q-Learning algorithm to adjust the second audio and generate a third audio. S104, the third audio is sent to the playback device of the terminal for playback.

[0047] In this embodiment, in response to a privacy call request, the current ambient audio is first obtained, and a second audio corresponding to the first audio is generated based on the first audio. The first audio is the audio being played by the terminal's earpiece, and the second audio is an out-of-phase audio with the same amplitude but opposite phase to the first audio. Then, the similarity value between the ambient audio and the first audio is calculated. Next, in response to the similarity value being greater than a similarity threshold, the second audio is adjusted based on the Q-Learning algorithm to generate a third audio. Finally, the third audio is sent to the terminal's speaker for playback. Therefore, this solution, through environmental detection and reinforcement learning-driven adaptive cancellation, solves the shortcomings of traditional audio cancellation techniques, such as high cost, poor adaptability, and unstable effects, reducing cancellation costs and improving environmental adaptability and user experience.

[0048] In one possible implementation, the current ambient audio can be acquired by using an ambient sound acquisition device, which is installed on the terminal and spaced at a preset distance from the playback device.

[0049] For example, the first audio signal is emitted from the earpiece of the mobile communication device, while the ambient sound acquisition device is located outside the earpiece. The ambient sound acquisition device should be positioned at the shortest distance from the first audio signal to meet call privacy requirements. For instance, if the call privacy requirement is that the target state can be achieved from a distance of 5cm from the earpiece, then the second microphone position would be 5cm from the earpiece.

[0050] By setting up an ambient sound acquisition device on the terminal and maintaining a certain distance from the playback device, it's possible to more accurately capture the sounds and background noise playing in the actual environment, rather than just the sound signal directly output from the playback device. This helps obtain more realistic environmental audio data, providing more precise input for subsequent processing. Simultaneously, the preset distance helps reduce the direct impact of sound emitted by the playback device on the ambient sound acquisition device, thereby reducing the direct echo or feedback volume in the acquired sound and improving the system's noise immunity. This is particularly important for applications requiring noise suppression, speech recognition, or audio separation.

[0051] In the above embodiments, the second audio is adjusted based on the Q-Learning algorithm to generate the third audio, and can also be achieved through... Figure 2 To further explain, the method includes: S201, Select the target candidate action with the largest expected value from the target action space. The target action space includes multiple candidate actions and the expected value of each candidate action.

[0052] It should be noted that the target action space is a set of operations that can be parametrically adjusted on the second audio, for example: Frequency band-level phase shift (e.g., applying a phase shift to the 500–2000 Hz human voice band) π phase shift); Amplitude scaling factor (e.g., 0.9×, 1.1×); Time-domain micro-delay (compensating for propagation delay from speaker to microphone); Spectrum shaping (enhancing / attenuating specific harmonic components to match leakage path characteristics).

[0053] Each candidate action is associated with an expected value (i.e., Q value), which represents the long-term cumulative reward (such as offsetting effect, perceptual distortion, and other comprehensive indicators) expected to be obtained by performing the action in the current state (s_t).

[0054] In one possible implementation, the target action space can be represented as shown in the table below:

[0055] The target action space is composed of the phase and amplitude of the second audio. The phase space includes: F represents advancing one unit, B represents retreating one unit, and - represents remaining unchanged. The amplitude space includes: E increases one unit, W decreases one unit, and - remains unchanged. The Q value is the expected value corresponding to each candidate action.

[0056] S202, perform the target candidate action to adjust the second audio, play the adjusted second audio through the speaker device, and calculate the similarity value between the adjusted first audio and the ambient audio.

[0057] The system applies the target candidate action to the second audio, generates an intermediate adjusted audio, and plays it briefly as a tentative cancellation signal or for simulation evaluation.

[0058] After execution, return the immediate reward (r_1) and update the system state to (s_1), which includes: Spectral characteristics of the adjusted audio; Superimposed signals collected by environmental microphones; Changes in acoustic path estimation, etc.

[0059] Based on the new state (s_1), the agent no longer uses ε-greedy, but directly selects the action that maximizes the Q value as the target candidate action, in order to quickly approach the local optimum in a single iteration and improve convergence efficiency.

[0060] After applying the second audio to the original first audio, the optimized audio for the current round is generated.

[0061] The system synchronously collects ambient audio. Calculate the similarity value between the adjusted first audio and the ambient audio. The specific calculation method can be referred to the content in the above embodiment, and will not be repeated here. Calculate the cross-correlation peak or DTW distance and normalize it to the [0,1] interval.

[0062] This similarity value reflects the degree of original audio leakage that still exists in the environment under the current cancellation strategy: the smaller the value, the better the cancellation effect.

[0063] S203, repeat the above steps of selecting the target candidate action with the largest expected value from the target action space and its subsequent steps until the similarity value between the adjusted first audio and the ambient audio is less than or equal to the similarity threshold, and output the finally adjusted second audio as the third audio.

[0064] If the current similarity value is greater than the similarity threshold, the system will treat the current optimized audio as a new starting point, return to S201, and start the next iteration. Update the current status; Again, a new target candidate action is selected based on the ε-greedy strategy; Perform a two-step action sequence (explore + exploit) to generate a new round of cancel audio; Recalculate the similarity and determine if the termination condition is met.

[0065] This process continues, forming a closed-loop optimization mechanism of "exploration-utilization-evaluation-re-exploration" until, at this point, it is determined that the audio leak has been reduced to a safe level.

[0066] In this embodiment of the disclosure, after the final selected target candidate action is obtained, the audio signal generated by the target candidate action (or its corresponding complete action sequence) that successfully achieves the similarity standard in the last round is determined as the final third audio and sent to the terminal playback device for formal playback, thereby realizing the active cancellation of environmentally leaked audio.

[0067] It should be noted that, in response to the playback of the adjusted second audio, the expected values ​​of all candidate actions in the target action space are updated. The update method can be as follows: Figure 3 As shown, the method includes: S301, for any candidate action in the target action space, calculate the first reward value for adjusting the second audio by executing the candidate action.

[0068] In this embodiment of the disclosure, when the system detects that the similarity value between the adjusted first audio and the ambient audio is still greater than the preset similarity threshold, it indicates that the current cancellation strategy has failed to effectively suppress audio leakage. The expected value (i.e., Q value) of each candidate action in the target action space needs to be dynamically updated to guide the agent to learn a better cancellation strategy.

[0069] For the target action space, the system first simulates or actually performs the action (a) to adjust the second audio, generating an intermediate cancellation signal.

[0070] Subsequently, a first reward value is obtained based on a predefined reward function. This reward function takes into account the following factors: Cancellation effect: The residual energy after the adjusted signal is superimposed on the ambient audio, for example: Perceptual distortion penalty: Avoid generating harsh or abnormal sounds, for example, by imposing a negative penalty based on PESQ or STOI metrics; Action complexity cost: Apply a regularization term to high gain or extreme phase shifts to prevent overfitting.

[0071] S302, after adjusting the second audio by executing the candidate action, select the next candidate action with the largest expected value from the target action space based on the candidate action, and calculate the second reward value for adjusting the second audio by executing the next candidate action.

[0072] After action (a) is performed, the system state transitions from (s) to a new state (s'). This new state (s') includes: Acoustic characteristics of the adjusted audio; Environmental feedback (such as superimposed signals collected by microphones); Contextual information such as channel estimation updates.

[0073] Based on the state (s'), the system uses a greedy policy to select the next candidate action (a') with the highest expected value from the target action space: Subsequently, the simulation executes (a') to further optimize the second audio (or the currently adjusted audio), generating a secondary adjustment signal and calculating the corresponding second reward value: S303, calculate the expected value of the candidate action based on the first reward value and the second reward value.

[0074] The system combines immediate rewards with expected future rewards and uses the Temporal Difference (TD) concept to update the expected value (Q(s,a)) of action (a).

[0075] Specifically, the target value is defined as:

[0076] in, Indicates taking an action in state s The reward value; Indicates the weight of future rewards; Indicates the current state Take action below The expected state was then obtained. In the expected state Take the best action The purpose of obtaining the maximum Q value is not only to calculate the reward of the action in the current state, but also to consider long-term reward factors. That is, in the training phase of the agent, in addition to considering the action that obtains the highest reward score in the existing target action space, it also considers that higher reward scores may be generated in the subsequent actions of other actions, and randomly selects another action to calculate the reward score in a comprehensive manner.

[0077] In this embodiment of the disclosure, the target action space can also be constructed by... Figure 4 To further explain, the method includes: S401, selects target candidate actions from the candidate action space based on the ε-greedy exploration strategy. The candidate action space includes multiple candidate actions and their respective expected values.

[0078] In this embodiment of the disclosure, the ε-greedy exploration strategy is pre-designed.

[0079] In one possible implementation, a target candidate action can be selected from the candidate action space based on a first selection strategy or a second selection strategy, wherein the probability of executing the first selection strategy is ε and the probability of executing the second selection strategy is 1-ε.

[0080] It should be noted that ε in the embodiments of this disclosure can be limited according to actual design needs. For example, ε can be 0.1.

[0081] The first selection strategy is to randomly select an action from the candidate action space, and the second selection strategy is to select the action with the highest expected value from the candidate action space.

[0082] It should be noted that in this embodiment, a second selection strategy is executed with probability, which selects an action randomly and uniformly, regardless of the current estimate of the value of each action. This strategy is used to proactively explore actions that have not yet been fully evaluated, preventing the agent from prematurely converging to a local optimum.

[0083] By employing the ε-greedy exploration strategy, the system introduces randomness with a controllable probability in each decision, which ensures the stability of the strategy while retaining the possibility of discovering better actions, thereby enhancing the agent's adaptability and robustness in complex acoustic environments.

[0084] S402, after adjusting the fourth audio by executing the target candidate action, select the third candidate action with the highest expected value from the candidate action space, and the fourth audio is the test audio for adjusting the candidate action space.

[0085] S403, execute the third candidate action to adjust the fourth audio, calculate the similarity value between the first audio and the test environment audio, where the test environment audio is the environment audio for which the candidate action space has been adjusted.

[0086] It should be noted that the specific implementation methods of S402 and S403 can be referred to the contents of the above embodiments, and will not be repeated here.

[0087] S404, in response to the similarity value between the adjusted first audio and the test environment audio being greater than the similarity threshold, updates the expected value of all candidate actions in the candidate action space.

[0088] S405, repeat the above-mentioned ε-greedy exploration strategy to select target candidate actions and their subsequent actions from the candidate action space until the adjusted fourth audio is played. If the similarity value between the first audio and the test environment audio is less than or equal to the similarity threshold, the adjusted candidate action space is output as the target action space.

[0089] In this embodiment of the disclosure, when the similarity meets the condition, the expected value of all actions in the candidate action space is updated incrementally, and the similarity is regarded as a "reward signal" to guide the system to converge toward highly adapted actions.

[0090] By combining the ε-greedy exploration mechanism in reinforcement learning with audio perception similarity feedback, a closed-loop self-optimizing audio action space adjustment framework was constructed. This not only solves the problem that traditional fixed strategies cannot adapt to dynamic acoustic environments, but also significantly improves the quality and personalization of the auditory experience in human-computer interaction through continuous updates of the expected value. The constructed target action space can provide a data foundation for subsequent action selection.

[0091] In this embodiment of the disclosure, the target reward value is calculated. The target reward value is either a first reward value or a second reward value. First, after playing the adjusted second audio, the similarity value between the first audio and the ambient audio is calculated, and then the reciprocal of the similarity value is used as the target reward value.

[0092] In this embodiment of the disclosure, the similarity value between the ambient audio and the first audio can be calculated by first obtaining the first audio feature and the second audio feature of the ambient audio, and then calculating the similarity value between the ambient audio and the first audio based on the first audio feature and the second audio feature.

[0093] In one possible implementation, a similarity value between the ambient audio and the first audio is calculated based on the first and second audio features. This can also be achieved through... Figure 5 To further explain, the method includes: S501, calculate the feature cosine similarity value of the first audio feature and the second audio feature, and calculate the text cosine similarity value between the ambient audio and the first audio.

[0094] S502, calculate the similarity value between the ambient audio and the first audio based on the feature cosine similarity value and the text cosine similarity value.

[0095] The similarity calculation method in this embodiment combines acoustic similarity and content similarity, evaluating the similarity between the environmental audio and the first audio by using a comprehensive similarity score. The calculation method is as follows: (1) Obtain and Mel-frequency cepstral coefficient matrix of segments within the same time interval.

[0096] (2) Calculate the feature cosine similarity value of the two using the obtained matrix. .

[0097] (3) Obtaining text through an audio-to-text model and Audio text , .

[0098] (4) Calculate using the fastText model and The vector representation is then used to calculate the text cosine similarity value. .

[0099] (5) Obtain the overall audio similarity ,in .

[0100] in, Indicates the first audio value. Indicates the second audio value. This represents the weight value of the feature cosine similarity value.

[0101] The reward score for the selected action is calculated using the following formula:

[0102] Calculate how to achieve state s by taking a phase-amplitude action a in the current state s. Then, find the action that yields the maximum Q value from the Q-table. To obtain its Q value .

[0103] Corresponding to the privacy call methods provided in the above embodiments, one embodiment of this disclosure also provides a privacy call device. Since the privacy call device provided in this disclosure corresponds to the privacy call methods provided in the above embodiments, the implementation methods of the above privacy call methods are also applicable to the privacy call device provided in this disclosure, and will not be described in detail in the following embodiments.

[0104] Figure 6 This is a schematic diagram of a privacy call device according to one embodiment of the present disclosure. As shown in FIG6, the privacy call device 600 includes: The acquisition module 610 is used to respond to a privacy call request, acquire the current ambient audio, and generate a second audio corresponding to the first audio based on the first audio. The first audio is the audio being played by the earpiece device of the terminal, and the second audio is an inverse audio with the same amplitude but opposite phase to the first audio.

[0105] The calculation module 620 is used to calculate the similarity value between the ambient audio and the first audio.

[0106] The generation module 630 is used to adjust the second audio based on the Q-Learning algorithm to generate the third audio in response to the similarity value between the ambient audio and the first audio being greater than the similarity threshold.

[0107] The playback module 640 is used to send the third audio to the speaker device of the terminal for playback.

[0108] According to one embodiment of this disclosure, adjusting a second audio based on the Q-Learning algorithm to generate a third audio includes: selecting a target candidate action with the largest expected value from a target action space, the target action space including multiple candidate actions and their respective expected values; adjusting the second audio by executing the target candidate action; playing the adjusted second audio through a speaker device after adjusting the second audio by executing the target candidate action; calculating the similarity value between the first audio and the ambient audio; repeating the above steps of selecting the target candidate action with the largest expected value from the target action space and subsequent steps until the similarity value between the first audio and the ambient audio is less than or equal to a similarity threshold, and outputting the finally adjusted second audio as the third audio.

[0109] According to one embodiment of this disclosure, the method further includes: in response to the playback of the adjusted second audio, if the similarity value between the first audio and the ambient audio is greater than a similarity threshold, updating the expected value of all candidate actions in the target action space.

[0110] According to one embodiment of this disclosure, updating the expected value of all candidate actions in the target action space includes: for any candidate action in the target action space, calculating a first reward value for adjusting the second audio by executing the candidate action; after adjusting the second audio by executing the candidate action, selecting the next candidate action with the largest expected value from the target action space based on the candidate action, and calculating a second reward value for adjusting the second audio by executing the next candidate action; and calculating the expected value of the candidate action based on the first reward value and the second reward value.

[0111] According to one embodiment of this disclosure, calculating a target reward value, which is either a first reward value or a second reward value, includes: after playing the adjusted second audio, calculating the similarity value between the first audio and the ambient audio; and using the reciprocal of the similarity value as the target reward value.

[0112] According to one embodiment of this disclosure, constructing a target action space includes: selecting target candidate actions from a candidate action space based on an ε-greedy exploration strategy, the candidate action space including multiple candidate actions and their respective expected values; after adjusting a fourth audio by executing the target candidate action, selecting a third candidate action with the largest expected value from the candidate action space, the fourth audio being a test audio for adjusting the candidate action space; adjusting the fourth audio by executing the third candidate action, playing the adjusted fourth audio through a speaker device, calculating the similarity value between a first audio and a test environment audio, the test environment audio being the environment audio for adjusting the candidate action space; updating the expected values ​​of all candidate actions in the candidate action space in response to the similarity value between the first audio and the test environment audio being greater than a similarity threshold; repeating the above selection of target candidate actions and their subsequent actions from the candidate action space based on the ε-greedy exploration strategy until, after playing the adjusted fourth audio, the similarity value between the first audio and the test environment audio is less than or equal to the similarity threshold, and outputting the adjusted candidate action space as the target action space.

[0113] According to one embodiment of the present disclosure, selecting a target candidate action from a candidate action space based on an ε-greedy exploration strategy includes: selecting a target candidate action from a candidate action space based on a first selection strategy or a second selection strategy, wherein the probability of executing the first selection strategy is ε, and the probability of executing the second selection strategy is 1-ε.

[0114] According to one embodiment of this disclosure, the first selection strategy is to randomly select an action from the candidate action space, and the second selection strategy is to select the action with the highest expected value from the candidate action space.

[0115] According to one embodiment of this disclosure, calculating the similarity value between ambient audio and a first audio includes: acquiring a first audio feature and a second audio feature of the ambient audio; and calculating the similarity value between the ambient audio and the first audio based on the first audio feature and the second audio feature.

[0116] According to one embodiment of this disclosure, calculating the similarity value between ambient audio and the first audio based on a first audio feature and a second audio feature includes: calculating the feature cosine similarity value of the first audio feature and the second audio feature, and calculating the text cosine similarity value between the ambient audio and the first audio; and calculating the similarity value between the ambient audio and the first audio based on the feature cosine similarity value and the text cosine similarity value.

[0117] According to one embodiment of this disclosure, calculating the similarity value between ambient audio and the first audio based on a first audio feature and a second audio feature includes: calculating the feature cosine similarity value of the first audio feature and the second audio feature, and calculating the text cosine similarity value between the ambient audio and the first audio; and calculating the similarity value between the ambient audio and the first audio based on the feature cosine similarity value and the text cosine similarity value.

[0118] Therefore, this solution addresses the shortcomings of traditional audio noise cancellation technology, such as high cost, poor adaptability, and unstable effect, by using environmental detection and reinforcement learning-driven adaptive cancellation. This reduces noise cancellation costs and improves environmental adaptability and user experience.

[0119] To implement the above embodiments, this disclosure also proposes an electronic device 700. Figure 7 This is a schematic diagram of an electronic device according to one embodiment of the present disclosure, such as... Figure 7 As shown, the electronic device 700 includes: a processor 701 and a memory 702 communicatively connected to the processor. The memory 702 stores instructions executable by at least one processor. The instructions are executed by at least one processor 701 to implement the functions described in this disclosure. Figures 1-5 The privacy call method in the embodiment.

[0120] To implement the above embodiments, this disclosure also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to implement the present disclosure. Figures 1-5 The privacy call method in the embodiment.

[0121] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program, which, when executed by a processor, implements the features of this disclosure. Figures 1-5 The privacy call method in the embodiment.

[0122] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0123] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.

[0124] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0125] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0126] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0127] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that contains, stores, communicates, propagates, or transmits programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0128] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0129] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.

[0130] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0131] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for making private calls, characterized in that, Applied to terminals, including: In response to a privacy call request, the system obtains the current ambient audio and generates a second audio corresponding to the first audio based on the first audio. The first audio is the audio being played by the terminal's earpiece device, and the second audio is an inverse audio with the same amplitude but opposite phase to the first audio. Calculate the similarity value between the ambient audio and the first audio; In response to the fact that the similarity value between the ambient audio and the first audio is greater than a similarity threshold, the second audio is adjusted based on the Q-Learning algorithm to generate a third audio. The third audio is sent to the speaker device of the terminal for playback.

2. The method according to claim 1, characterized in that, The adjustment of the second audio based on the Q-Learning algorithm to generate the third audio includes: Select the target candidate action with the largest expected value from the target action space, wherein the target action space includes multiple candidate actions and the expected value corresponding to each candidate action; The target candidate action is executed to adjust the second audio. After the adjusted second audio is played through the speaker device, the similarity value between the adjusted first audio and the ambient audio is calculated. Repeat the above steps of selecting the target candidate action with the highest expected value from the target action space and subsequent steps until the similarity value between the adjusted first audio and the ambient audio is less than or equal to the similarity threshold, and output the finally adjusted second audio as the third audio.

3. The method according to claim 2, characterized in that, The method further includes: In response to the playback of the adjusted second audio, if the similarity value between the first audio and the ambient audio is greater than the similarity threshold, the expected values ​​of all candidate actions in the target action space are updated.

4. The method according to claim 3, characterized in that, The step of updating the expected value of all candidate actions in the target action space includes: For any candidate action in the target action space, calculate a first reward value for adjusting the second audio by performing the candidate action; After adjusting the second audio by performing the candidate action, based on the candidate action, the next candidate action with the largest expected value is selected from the target action space, and a second reward value is calculated for adjusting the second audio by performing the next candidate action; The expected value of the candidate action is calculated based on the first reward value and the second reward value.

5. The method according to claim 4, characterized in that, Calculating the target reward value, wherein the target reward value is either the first reward value or the second reward value, includes: After calculating the adjusted second audio, calculate the similarity value between the first audio and the ambient audio. The reciprocal of the similarity value is used as the target reward value.

6. The method according to claim 2, characterized in that, Constructing the target action space includes: The target candidate action is selected from the candidate action space based on the ε-greedy exploration strategy. The candidate action space includes multiple candidate actions and the expected value corresponding to each candidate action. After the target candidate action is executed and the fourth audio is adjusted, the target candidate action with the largest expected value is selected from the candidate action space, and the fourth audio is the test audio for adjusting the candidate action space; The fourth audio is adjusted by performing the target candidate action. After the adjusted fourth audio is played through the speaker device, the similarity value between the first audio and the test environment audio is calculated. The test environment audio is the environment audio in which the candidate action space is adjusted. In response to the similarity value between the first audio and the test environment audio being greater than the similarity threshold, the expected values ​​of all candidate actions in the candidate action space are updated; Repeat the above ε-greedy exploration strategy to select target candidate actions and their subsequent actions from the candidate action space until the adjusted fourth audio is played. If the similarity value between the first audio and the test environment audio is less than or equal to the similarity threshold, the adjusted candidate action space is output as the target action space.

7. The method according to claim 6, characterized in that, The method of selecting target candidate actions from the candidate action space based on the ε-greedy exploration strategy includes: A target candidate action is selected from the candidate action space based on a first selection strategy or a second selection strategy, wherein the probability of executing the first selection strategy is ε, and the probability of executing the second selection strategy is 1-ε.

8. The method according to claim 7, characterized in that, The first selection strategy is to randomly select an action from the candidate action space, and the second selection strategy is to select the action with the highest expected value from the candidate action space.

9. The method according to any one of claims 1-8, characterized in that, The calculation of the similarity value between the environmental audio and the first audio includes: Obtain the first audio feature and the second audio feature of the ambient audio; Based on the first audio feature and the second audio feature, the similarity value between the ambient audio and the first audio is calculated.

10. The method according to claim 9, characterized in that, The step of calculating the similarity value between the environmental audio and the first audio based on the first audio feature and the second audio feature includes: Calculate the feature cosine similarity value between the first audio feature and the second audio feature, and calculate the text cosine similarity value between the environmental audio and the first audio. The similarity value between the environmental audio and the first audio is calculated based on the feature cosine similarity value and the text cosine similarity value.