Call control method and device
By obtaining and identifying audio content in voice calls on electronic devices, determining the conversation topic, and turning off the microphone when the topics do not match, the risk of privacy leakage in voice calls is solved, and effective protection of user privacy is achieved.
Patent Information
- Application Number
- CN202510375082.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-05-13
AI Technical Summary
During voice calls, users cannot effectively control the microphone, resulting in offline conversation contents that may be transmitted to online calls, and there is a risk of privacy leakage.
Semantic recognition is performed to determine the conversation topic by obtaining audio content of the audio output and input components on the electronic device, and if the online and offline conversation topics are not the same, the audio input component is turned off to prevent privacy leakage.
It effectively reduces the risk of privacy leakage and protects users' privacy, especially when users have online and offline conversations at the same time.
Smart Images

Figure CN119996566A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of electronic equipment, and specifically relates to a call control method and device. Background Art
[0002] With the rapid development of mobile communication technology, voice call technology exists in multiple scenarios such as daily communication and game connection. In the related technology, during the voice call, the user needs to manually turn on and off the microphone, which is more dependent on the user's operation. However, for some special scenarios, for example, during a voice call, user A and user B are talking online while user A is also talking offline with user C. Since user A cannot predict when to turn off the microphone, or his hands are occupied and he cannot operate the microphone, the content of the offline conversation is transmitted to the online call, which poses a risk of privacy leakage. Summary of the invention
[0003] The purpose of the embodiments of the present application is to provide a call control method and device that can reduce the risk of privacy leakage.
[0004] In a first aspect, an embodiment of the present application provides a call control method, which is applied to a first electronic device, and the method includes:
[0005] During a voice call between a first user using the first electronic device and a second user using a second electronic device, obtaining first audio through an audio output component of the first electronic device; and obtaining second audio through an audio input component of the first electronic device;
[0006] Performing semantic recognition on the first audio and the second audio to determine a first conversation topic and a second conversation topic; wherein the first conversation topic is a conversation topic of a voice call between the second user and the first user;
[0007] In case the first conversation topic is different from the second conversation topic, the audio input component is closed.
[0008] In a second aspect, an embodiment of the present application provides a call control device, applied to a first electronic device, the device comprising:
[0009] A first acquisition module is used to acquire first audio through an audio output component of the first electronic device during a voice call between a first user using the first electronic device and a second user using a second electronic device; and to acquire second audio through an audio input component of the first electronic device;
[0010] A second acquisition module, used for acquiring a second audio through an audio input component of the first electronic device;
[0011] a recognition module, configured to perform semantic recognition on the first audio and the second audio to determine a first conversation topic and a second conversation topic; wherein the first conversation topic is a conversation topic of a voice call between the second user and the first user;
[0012] A control module is used to close the audio input component when the first dialogue topic is different from the second dialogue topic.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the call control method described in the first aspect are implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the call control method described in the first aspect are implemented.
[0015] In a fifth aspect, an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the call control method as described in the first aspect.
[0016] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the call control method as described in the first aspect.
[0017] In an embodiment of the present application, during a voice call between a first user using a first electronic device and a second user using a second electronic device, a first audio is obtained through an audio output component of the first electronic device; and a second audio is obtained through an audio input component of the first electronic device; semantic recognition is performed on the first audio and the second audio to determine a first conversation topic and a second conversation topic; wherein the first conversation topic is a conversation topic for the voice call between the second user and the first user; and when the first conversation topic is different from the second conversation topic, the audio input component is turned off.
[0018] It can be seen that in the embodiment of the present application, considering that when the local user only has a voice call with the online user, the conversation topic of both parties is usually unchanged in a short period of time, and when the local user has a voice call with the online user and a conversation with the offline user at the same time, the conversation objects are different, and the conversation topics are usually different, and the previous voice content obtained by the audio output component and the audio input component can reflect the conversation topic. Therefore, the first conversation topic and the second conversation topic can be determined based on the previous voice content obtained by the audio output component and the audio input component. The first conversation topic is the conversation topic of the voice call between the online user and the local user. If the first conversation topic is different from the second conversation topic, it means that the local user is having a voice call with the online user while also having a conversation with the offline user. At this time, the audio input component is turned off to prevent the offline conversation content from being transmitted to the online voice call, thereby reducing the risk of privacy leakage, thereby achieving the purpose of protecting user privacy. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a flow chart of a call control method provided by some embodiments of the present application;
[0020] Figure 2 is one of the example diagrams of a call control interface provided by some embodiments of the present application;
[0021] Figure 3 This is a second example diagram of a call control interface provided by some embodiments of the present application;
[0022] Figure 4 is a flowchart of an implementation of step 102 provided in some embodiments of the present application;
[0023] Figure 5 is an example diagram of a call control method provided by some embodiments of the present application;
[0024] Figure 6 is a structural block diagram of a call control device provided in an embodiment of the present application;
[0025] Figure 7 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application;
[0026] Figure 8 It is a schematic diagram of the hardware structure of an electronic device implementing various embodiments of the present application. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all embodiments obtained by ordinary technicians in this field belong to the scope of protection of this application.
[0028] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0029] For ease of understanding, the call control method provided in the embodiment of the present application is described in detail below in conjunction with the accompanying drawings.
[0030] It should be noted that the call control method provided in the embodiment of the present application is applicable to electronic devices. In actual applications, the electronic devices may include: mobile terminals such as smart phones, tablet computers, smart watches, personal digital assistants, and may also include: computer devices such as laptop computers, desktop computers, and desktop computers. The embodiment of the present application does not limit this.
[0031] Figure 1 is a flow chart of a call control method provided by some embodiments of the present application, the call control method is applied to a first electronic device, such as Figure 1 As shown, the method may include the following steps: step 101, step 102 and step 103.
[0032] In step 101, during a voice call between a first user using a first electronic device and a second user using a second electronic device, first audio is acquired through an audio output component of the first electronic device; and second audio is acquired through an audio input component of the first electronic device.
[0033] In the embodiment of the present application, the above-mentioned voice call includes but is not limited to: voice call in video call, network voice call, telephone voice call, etc.
[0034] In the embodiments of the present application, the call control method can be applicable to a variety of scenarios, including but not limited to: online video communication scenarios, online conference scenarios, game connection scenarios, mobile phone call scenarios, etc.
[0035] In the embodiment of the present application, the first user is a local user, and the second user is an online user.
[0036] In the embodiment of the present application, the audio output component may include at least one of the following: an earpiece, a speaker.
[0037] It can be seen that in the embodiments of the present application, taking into account that in some voice call scenarios, some users like to use a receiver and some users like to use a speaker, for users who use a receiver, the first audio can be obtained through the receiver, and for users who use a speaker, the first audio can be obtained through the speaker. This can meet the voice call needs of different users, has high flexibility, and is applicable to a wide range of scenarios.
[0038] In the embodiment of the present application, the audio input component may include at least one of the following: a microphone, a microphone array.
[0039] It can be seen that in the embodiments of the present application, it is taken into consideration that the hardware configurations of different electronic devices may be different. For example, some electronic devices are only configured with one microphone, while some electronic devices are configured with a microphone array. Therefore, for an electronic device configured with only one microphone, the second audio can be obtained through this one microphone, and for an electronic device configured with a microphone array, the second audio can be obtained through the microphone array. This can meet the voice call scenarios of electronic devices with different hardware configurations and has a relatively high flexibility.
[0040] In an embodiment of the present application, the first audio is audio transmitted from the second electronic device, which is captured by the audio output component of the first electronic device. The audio records what the second user said during an online voice call between the first user and the second user.
[0041] In an embodiment of the present application, the second audio is the audio generated by a sound-making object in the environment where the first electronic device is located. The audio is captured by the audio input component of the first electronic device. If there is only one sound-making object, namely the first user, in the environment, then the audio records what the first user said. If there are not only one sound-making object, namely the first user, but also other sound-making objects, such as other users, in the environment, then the audio records what the first user said and what the other users said.
[0042] It can be understood that if the first electronic device is used as the dividing line, the first audio can be called "in-machine audio" or "online audio", that is, audio transmitted from the first electronic device; the second audio can be called "out-of-machine audio" or "offline audio", that is, audio input into the first electronic device.
[0043] In an embodiment of the present application, during a voice call between a first user and a second user, the first audio records what the second user said during the voice call, and the second audio records at least what the first user said during the voice call. Since the conversation subject is a topic of communication between the two parties, in order to determine through the conversation subject whether the first user is also having a conversation with other users offline during the voice call with the second user, it is necessary to obtain the first audio and the second audio at the same time, wherein the entire content of the first audio and part of the content of the second audio are used to determine the conversation subject of the voice call between the second user and the first user; and part of the content of the second audio is used to determine the conversation subject of the sound-emitting object in the environment where the first electronic device is located.
[0044] In an embodiment of the present application, the above-mentioned call control method can be integrated into an electronic device as a function, referred to as "microphone intelligent recognition intention function", and the function can be manually turned on by the user.
[0045] For example, Figure 2 As shown, a permission setting interface 22 of any application is displayed on the display screen 21 of the electronic device 20, and an "intelligent recognition of intent" function 23 is added under the microphone permission of the permission setting interface 22 for the user to manually open it.
[0046] In an embodiment of the present application, for the microphone intelligent recognition intention function, the user can also be reminded to manually turn it on when an application in the electronic device applies for microphone permission for the first time.
[0047] For example, Figure 3 As shown, when any application in the electronic device 20 applies for microphone permission for the first time, a reminder message 24 pops up on the display screen 21 to inform the user of the function, and the user decides whether to turn it on.
[0048] In an embodiment of the present application, after a user turns on the microphone intelligent recognition intention function for an application in an electronic device, the electronic device automatically executes the above-mentioned call control method during the process of running a voice call by the application.
[0049] In step 102, semantic recognition is performed on the first audio and the second audio to determine the first conversation topic and the second conversation topic; wherein the first conversation topic is the conversation topic of the voice call between the second user and the first user.
[0050] In the embodiment of the present application, considering that the conversation topic requires communication and answers between the two parties, and the first audio only records what the second user said, the first conversation topic cannot be determined based on the first audio alone. Therefore, it is also necessary to combine the second audio and perform semantic recognition on the first audio and the second audio to determine the first conversation topic and the second conversation topic.
[0051] In the embodiment of the present application, the first conversation topic refers to the conversation topic extracted based on the semantic content of the first audio and the semantic content of the second audio, which can also be called "on-machine conversation topic" or "online conversation topic". The second conversation topic refers to the conversation topic extracted based on the semantic content of the second audio, which can also be called "off-machine conversation topic" or "offline conversation topic".
[0052] In some embodiments, Figure 4 As shown, the above step 102 may include the following steps: step 1021, step 1022, step 1023, step 1024 and step 1025.
[0053] In step 1021, semantic recognition is performed on the first audio to obtain a first semantic recognition result.
[0054] In the embodiment of the present application, the first semantic recognition result includes the main content of what the second user said.
[0055] In step 1022, semantic recognition is performed on the second audio to obtain a second semantic recognition result.
[0056] In the embodiment of the present application, the second semantic recognition result at least includes the main content of the first user's speech. During the voice call between the first user and the second user, if the first user only has a conversation with the second user, the second semantic recognition result only includes the main content of the first user's speech; if the first user has an online conversation with the second user and an offline conversation with other users, the second semantic recognition result includes the main content of the first user's speech as well as the main content of the other users' speech.
[0057] In step 1023, the conversation topic corresponding to the second user is determined according to the first semantic recognition result and the second semantic recognition result.
[0058] In an embodiment of the present application, since the first semantic recognition result only includes the main content of what the second user said, and does not include the main content of what the first user said when having a voice call with the second user, and the second semantic recognition result includes the main content of what the first user said, it is necessary to determine the conversation topic corresponding to the second user based on the first semantic recognition result and the second semantic recognition result.
[0059] In step 1024, the conversation topic corresponding to the second user is determined as the first conversation topic.
[0060] In some embodiments, the conversation topic corresponding to the second user can be directly determined as the first conversation topic.
[0061] In some embodiments, in order to ensure the accuracy of the first conversation topic, some auxiliary conditions may be used to verify the correctness of the conversation topic corresponding to the second user. Accordingly, before the above step 1024, the following steps may also be included: determining a first application in the first electronic device; wherein the first application is an application running in the foreground;
[0062] Accordingly, the above step 1024 may include the following steps: step 10241;
[0063] In step 10241, when the semantic correlation between the program topic of the first application and the conversation topic corresponding to the second user is higher than a first threshold, the conversation topic corresponding to the second user is determined as the first conversation topic.
[0064] In the embodiment of the present application, considering that the application currently running in the foreground is usually the application that the user pays the most attention to and is also the application that is highly correlated with the current voice call, the correctness of the conversation topic corresponding to the second user can be verified through the program theme of the application currently running in the foreground.
[0065] In the embodiments of the present application, the program subject includes but is not limited to: the function of the program and the type of the program.
[0066] For example, for a video application, the program theme is "watch videos"; for a live broadcast application, the program theme is "watch live broadcasts"; for a game application, the program theme is "play games"; for a shopping application, the program theme is "buy things", etc.
[0067] Exemplarily, the first user and the second user have a voice call through an instant messaging program, and the application currently running on the first electronic device of the first user is a game application, then the program theme is playing games. If the corresponding conversation theme of the second user is shopping, it means that the semantic correlation between the program theme and the conversation theme corresponding to the second user is very low. If the corresponding conversation theme of the second user is games, it means that the semantic correlation between the program theme and the conversation theme corresponding to the second user is very high, and the conversation theme corresponding to the second user is determined as the first conversation theme.
[0068] It can be seen that in the embodiment of the present application, the correctness of the conversation topic corresponding to the second user can be verified by the program theme of the application currently running in the foreground. Since the application currently running in the foreground is usually the application with the highest user attention and is also the application with a relatively high correlation with the current voice call, the verification standard is relatively accurate and can ensure the accuracy of the verification result.
[0069] In some embodiments, in order to further ensure the accuracy of the first conversation topic, other auxiliary conditions can be used to verify the correctness of the conversation topic corresponding to the second user. Accordingly, the following steps can be included before the above step 10241: determining a second application in the first electronic device; wherein the second application is an application for performing voice calls; determining the communication topic based on the communication information between the first user and the second user recorded in the second application.
[0070] Accordingly, the above step 10241 may include the following steps: step 102411;
[0071] In step 102411, when the semantic correlation between the program topic and the conversation topic corresponding to the second user is higher than a first threshold, and the semantic correlation between the program topic of the first application and the communication topic is higher than a second threshold, the conversation topic corresponding to the second user is determined as the first conversation topic.
[0072] In the embodiment of the present application, considering that for the application currently running the voice call, there are usually two dimensions of communication between the two parties of the conversation, one is the text dimension communication (such as chat records in the instant messaging application, text comments in the short video application, etc.), and the other is the voice dimension communication, and the communication content in the voice dimension and the communication content in the text dimension usually overlap. For example, when two users use the instant messaging application to make a voice call, the text chat content and voice chat content of the two users are closely related. Therefore, the communication topic in the text dimension can be used to verify the correctness of the conversation topic corresponding to the second user in the voice dimension.
[0073] In the embodiment of the present application, communication in the text dimension includes but is not limited to: user notes, user previous chat records, text comments, text replies, etc.
[0074] Exemplarily, the first user and the second user make a voice call through an instant messaging program, and the application currently running on the first electronic device of the first user is a game application, then the program theme is playing games. The first user's note to the second user in the instant messaging program is: Wang Datou, Class 2, Grade 3. Combined with some chat content between the first user and the second user in the instant messaging program: class, after school, playing games, etc., it can be seen that the second user's label is a classmate, and the label is used as the communication theme. After comparison, it is found that the correlation between the communication theme and the program theme is very high. If the corresponding conversation theme of the second user is shopping, it means that the semantic correlation between the program theme and the conversation theme corresponding to the second user is very low. If the corresponding conversation theme of the second user is a game, it means that the semantic correlation between the program theme and the conversation theme corresponding to the second user is very high, and the correlation between the communication theme and the program theme is also very high, and the conversation theme corresponding to the second user is determined as the first conversation theme.
[0075] It can be seen that in the embodiment of the present application, the correctness of the conversation topic corresponding to the second user in the voice dimension can be verified through the communication topic in the text dimension. Since the application currently running the voice call is usually two-dimensional communication between the two parties in the conversation, one is the communication in the text dimension, and the other is the communication in the voice dimension. The communication content in the voice dimension and the communication content in the text dimension usually overlap, so the verification standard is relatively accurate and can ensure the accuracy of the verification result.
[0076] In step 1025, a second conversation topic is determined based on the second semantic recognition result.
[0077] In the embodiment of the present application, since the second semantic recognition result includes the main content of the words spoken by all people around the environment where the first electronic device is located, the second conversation topic can be determined based on the second semantic recognition result.
[0078] It can be seen that in the embodiment of the present application, semantic recognition can be performed on the first audio and the second audio respectively, and then the first conversation topic can be determined by combining the semantic recognition results of the two, and the second conversation topic can be determined according to the semantic recognition result of the second audio, and the determination result of the conversation topic is relatively accurate.
[0079] In step 103, when the first conversation topic is different from the second conversation topic, the audio input component is turned off.
[0080] In an embodiment of the present application, if the first conversation topic is different from the second conversation topic, it means that the local user is having a conversation with an offline user while having a voice call with an online user. At this time, the audio input component is turned off to prevent the offline conversation content from being transmitted to the online voice call to avoid privacy leakage.
[0081] In some embodiments, when the first conversation topic is different from the second conversation topic, the audio input component can be immediately closed.
[0082] In some embodiments, in order to prevent misjudgment, the user's facial orientation can be combined to further determine whether to turn off the audio input component. Accordingly, before the above step 103, the following steps can also be included: when the first conversation topic is different from the second conversation topic, detect whether the first user's face is facing the first electronic device.
[0083] Accordingly, the above step 103 may include the following steps: step 1031;
[0084] In step 1031, when the first conversation topic is different from the second conversation topic and the face of the first user is not facing the first electronic device, the audio input component is turned off.
[0085] In the embodiment of the present application, the purpose of detecting whether the first user's face is facing the first electronic device is to determine whether the first user has an obvious tendency to turn his head during a voice call with a second user. If the first user's face is facing the first electronic device, it means that the first user has no tendency to turn his head. In this case, the first user is usually concentrating on the voice call with the second user; if the first user's face is not facing the first electronic device, it means that the first user has a tendency to turn his head. In this case, the first user is likely to be having an offline conversation with other users nearby.
[0086] In the embodiment of the present application, if the first electronic device is equipped with a microphone array, the microphone array can be used to detect whether the first user's face is facing the first electronic device. Alternatively, a front camera can be used to detect whether the first user's face is facing the first electronic device, which is not limited in the embodiment of the present application.
[0087] In an embodiment of the present application, if the first conversation topic is different from the second conversation topic, and the face of the first user is not facing the first electronic device, it means that the local user is having a voice call with the online user while turning his head to have a conversation with the offline user. At this time, the audio input component is turned off to prevent the offline conversation content from being transmitted to the online voice call to avoid privacy leakage.
[0088] In some embodiments, the following situation may be considered: during a voice call between a first user and a second user, the first user starts to have an offline conversation with other users, but the offline conversation ends very quickly. At this time, in order to avoid immediately closing the audio input component and causing the first user and the second user to be unable to continue the voice call, and to ensure the voice call experience of both parties, the following steps may be included before the above step 103: when the first conversation topic is different from the second conversation topic, a third audio is obtained through the audio input component.
[0089] Accordingly, the above step 103 may include the following steps: step 1032;
[0090] In step 1032, when the first conversation topic is different from the second conversation topic and the third audio matches the second conversation topic, the audio input component is closed.
[0091] In the embodiment of the present application, the third audio is the audio subsequently captured by the audio input component. The third audio contains at least the words spoken by the first user. If the third audio matches the second conversation topic, it means that the third audio contains not only the words spoken by the first user, but also the words spoken by other users. At this time, in order to prevent the offline conversation content from being transmitted to the online voice call, the audio input component is turned off.
[0092] In order to facilitate the overall understanding of the call control method provided in the embodiment of the present application, Figure 5 The example diagram shown in FIG. 1 is used as an example, wherein the audio input component is a microphone and the audio output component is a receiver. Figure 5 As shown, the call control method may include the following steps:
[0093] Step 501 , step 502 , step 503 , step 504 , step 505 and step 506 .
[0094] In step 501, the microphone intelligent recognition intention function is turned on.
[0095] In step 502, a first application running in the foreground is identified, and a second application that uses a microphone for a voice call is identified.
[0096] In step 503, a first audio is acquired through an earpiece, and a second audio is acquired through a microphone.
[0097] In step 504, a program theme of the first application and a communication theme of the second application are determined.
[0098] In step 505, a first conversation topic and a second conversation topic are determined according to the first audio, the second audio, the program topic and the communication topic.
[0099] In step 506, when the first conversation topic is different from the second conversation topic, it is detected whether the local user is facing the local user, and if not, the microphone is turned off.
[0100] It can be seen that in the embodiment of the present application, a method for protecting user privacy by identifying user intention based on user facial orientation and voice content is proposed, and the final topic of the on-machine and off-machine conversation is formed by analyzing the application running the voice call, the application running in the foreground, and the audio content of the user's on-machine and off-machine conversations. On this basis, combined with the analysis of the user's facial orientation (i.e., whether the sound is directed toward the electronic device), the audio input component is adjusted in a timely and accurate manner when the user's hands are occupied, thereby bringing a good user experience to the user, and effectively preventing the off-machine conversation from leaking into the voice call through the turned-on audio input component, thereby achieving the purpose of isolating the on-machine and off-machine conversations.
[0101] As can be seen from the above embodiment, in this embodiment, during a voice call between a first user using a first electronic device and a second user using a second electronic device, a first audio is obtained through an audio output component of the first electronic device; and a second audio is obtained through an audio input component of the first electronic device; semantic recognition is performed on the first audio and the second audio to determine a first conversation topic and a second conversation topic; wherein the first conversation topic is a conversation topic for the voice call between the second user and the first user; and when the first conversation topic and the second conversation topic are different, the audio input component is closed.
[0102] It can be seen that in the embodiment of the present application, considering that when the local user only has a voice call with the online user, the conversation topic of both parties is usually unchanged in a short period of time, and when the local user has a voice call with the online user and a conversation with the offline user at the same time, the conversation objects are different, and the conversation topics are usually different, and the previous voice content obtained by the audio output component and the audio input component can reflect the conversation topic. Therefore, the first conversation topic and the second conversation topic can be determined based on the previous voice content obtained by the audio output component and the audio input component. The first conversation topic is the conversation topic of the voice call between the online user and the local user. If the first conversation topic is different from the second conversation topic, it means that the local user is having a voice call with the online user while also having a conversation with the offline user. At this time, the audio input component is turned off to prevent the offline conversation content from being transmitted to the online voice call, thereby reducing the risk of privacy leakage, thereby achieving the purpose of protecting user privacy.
[0103] The call control method provided in the embodiment of the present application can be executed by a call control device. In the embodiment of the present application, the call control device provided in the embodiment of the present application is described by taking the call control method executed by the call control device as an example.
[0104] Figure 6 is a structural block diagram of a call control device provided by some embodiments of the present application, which is applied to a first electronic device, such as Figure 6 As shown, the call control device 600 may include: a first acquisition module 601, a second acquisition module 602, an identification module 603 and a control module 604;
[0105] The first acquisition module 601 is used to acquire first audio through the audio output component of the first electronic device during a voice call between a first user using the first electronic device and a second user using the second electronic device;
[0106] The second acquisition module 602 is used to acquire the second audio through the audio input component of the first electronic device;
[0107] The recognition module 603 is used to perform semantic recognition on the first audio and the second audio to determine a first conversation topic and a second conversation topic; wherein the first conversation topic is a conversation topic of a voice call between the second user and the first user;
[0108] The control module 604 is used to close the audio input component when the first dialogue topic is different from the second dialogue topic.
[0109] As can be seen from the above embodiment, in this embodiment, during a voice call between a first user using a first electronic device and a second user using a second electronic device, a first audio is obtained through an audio output component of the first electronic device; and a second audio is obtained through an audio input component of the first electronic device; semantic recognition is performed on the first audio and the second audio to determine a first conversation topic and a second conversation topic; wherein the first conversation topic is a conversation topic for the voice call between the second user and the first user; and when the first conversation topic and the second conversation topic are different, the audio input component is closed.
[0110] It can be seen that in the embodiment of the present application, considering that when the local user only has a voice call with the online user, the conversation topic of both parties is usually unchanged in a short period of time, and when the local user has a voice call with the online user and a conversation with the offline user at the same time, the conversation objects are different, and the conversation topics are usually different, and the previous voice content obtained by the audio output component and the audio input component can reflect the conversation topic. Therefore, the first conversation topic and the second conversation topic can be determined based on the previous voice content obtained by the audio output component and the audio input component. The first conversation topic is the conversation topic of the voice call between the online user and the local user. If the first conversation topic is different from the second conversation topic, it means that the local user is having a voice call with the online user while also having a conversation with the offline user. At this time, the audio input component is turned off to prevent the offline conversation content from being transmitted to the online voice call, thereby reducing the risk of privacy leakage, thereby achieving the purpose of protecting user privacy.
[0111] Optionally, as an embodiment, the identification module 603 may include:
[0112] A first recognition submodule, configured to perform semantic recognition on the first audio to obtain a first semantic recognition result;
[0113] A second recognition submodule, used to perform semantic recognition on the second audio to obtain a second semantic recognition result;
[0114] A first determination submodule, configured to determine a conversation topic corresponding to the second user according to the first semantic recognition result and the second semantic recognition result;
[0115] A second determining submodule, configured to determine the conversation topic corresponding to the second user as the first conversation topic;
[0116] The third determination submodule is used to determine a second dialogue topic according to the second semantic recognition result.
[0117] Optionally, as an embodiment, the call control device 600 may further include:
[0118] A determination module, configured to determine a first application in the first electronic device; wherein the first application is an application running in the foreground;
[0119] The second determining submodule may include:
[0120] A determination unit is used to determine the conversation topic corresponding to the second user as the first conversation topic when the semantic correlation between the program topic of the first application and the conversation topic corresponding to the second user is higher than a first threshold.
[0121] Optionally, as an embodiment, the call control device 600 may further include:
[0122] A detection module, configured to detect whether the first user faces toward the first electronic device when the first conversation topic is different from the second conversation topic;
[0123] The control module 604 may include:
[0124] The first control submodule is used to turn off the audio input component when the first conversation topic is different from the second conversation topic and the face of the first user is not facing the first electronic device.
[0125] Optionally, as an embodiment, the call control device 600 may further include:
[0126] A third acquisition module, configured to acquire a third audio through the audio input component when the first conversation topic is different from the second conversation topic;
[0127] The control module 604 may include:
[0128] The second control submodule is used to close the audio input component when the first dialogue topic is different from the second dialogue topic and the third audio complies with the second dialogue topic.
[0129] The call control device in the embodiment of the present application can be an electronic device or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or a device other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (Augmented Reality, AR) / virtual reality (Virtual Reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (Ultra-Mobile Personal Computer, UMPC), a netbook, or a personal digital assistant (Personal Digital Assistant, PDA), etc. It can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (Personal Computer, PC), a television (Television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.
[0130] The call control device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or any other operating system, which is not specifically limited in the embodiment of the present application.
[0131] The call control device provided in the embodiment of the present application can implement each process implemented in the above method embodiment, and will not be described again here to avoid repetition.
[0132] Alternatively, if Figure 7 As shown, an embodiment of the present application further provides an electronic device 700, including a processor 701 and a memory 702, wherein the memory 702 stores programs or instructions that can be executed on the processor 701, and when the program or instructions are executed by the processor 701, the various steps of the above-mentioned call control method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, they are not described here.
[0133] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0134] Figure 8 It is a schematic diagram of the hardware structure of an electronic device implementing various embodiments of the present application.
[0135] The electronic device 800 includes but is not limited to: a radio frequency unit 801, a network module 802, an audio output unit 803, an input unit 804, a sensor 805, a display unit 806, a user input unit 807, an interface unit 808, a memory 809, a processor 810 and other components.
[0136] Those skilled in the art will appreciate that the electronic device 800 may also include a power source (such as a battery) for supplying power to each component, and the power source may be logically connected to the processor 810 through a power management system, thereby implementing functions such as managing charging, discharging, and power consumption management through the power management system. Figure 8 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be described in detail here.
[0137] Among them, when the electronic device is a first electronic device, the processor 810 is used to obtain the first audio through the audio output component of the first electronic device during a process in which a first user uses the first electronic device to have a voice call with a second user using the second electronic device; and obtain the second audio through the audio input component of the first electronic device; perform semantic recognition on the first audio and the second audio to determine a first conversation topic and a second conversation topic; wherein the first conversation topic is the conversation topic of the voice call between the second user and the first user; when the first conversation topic is different from the second conversation topic, close the audio input component.
[0138] It can be seen that in the embodiment of the present application, considering that when the local user only has a voice call with the online user, the conversation topic of both parties is usually unchanged in a short period of time, and when the local user has a voice call with the online user and a conversation with the offline user at the same time, the conversation objects are different, and the conversation topics are usually different, and the previous voice content obtained by the audio output component and the audio input component can reflect the conversation topic. Therefore, the first conversation topic and the second conversation topic can be determined based on the previous voice content obtained by the audio output component and the audio input component. The first conversation topic is the conversation topic of the voice call between the online user and the local user. If the first conversation topic is different from the second conversation topic, it means that the local user is having a voice call with the online user while also having a conversation with the offline user. At this time, the audio input component is turned off to prevent the offline conversation content from being transmitted to the online voice call, thereby reducing the risk of privacy leakage, thereby achieving the purpose of protecting user privacy.
[0139] Optionally, as an embodiment, the processor 810 is specifically used to perform semantic recognition on the first audio to obtain a first semantic recognition result; perform semantic recognition on the second audio to obtain a second semantic recognition result; determine a conversation topic corresponding to the second user based on the first semantic recognition result and the second semantic recognition result; determine the conversation topic corresponding to the second user as the first conversation topic; and determine the second conversation topic based on the second semantic recognition result.
[0140] Optionally, as an embodiment, the processor 810 is specifically used to determine a first application in the first electronic device; wherein the first application is an application running in the foreground; when the semantic correlation between the program topic of the first application and the conversation topic corresponding to the second user is higher than a first threshold, the conversation topic corresponding to the second user is determined as the first conversation topic.
[0141] Optionally, as an embodiment, the processor 810 is specifically used to detect whether the first user's face is facing the first electronic device when the first conversation topic is different from the second conversation topic; and to turn off the audio input component when the first conversation topic is different from the second conversation topic and the first user's face is not facing the first electronic device.
[0142] Optionally, as an embodiment, the processor 810 is specifically used to obtain a third audio through the audio input component when the first dialogue topic is different from the second dialogue topic; and to close the audio input component when the first dialogue topic is different from the second dialogue topic and the third audio conforms to the second dialogue topic.
[0143] It should be understood that in the embodiment of the present application, the input unit 804 may include a graphics processor (GPU) 8041 and a microphone 8042, and the graphics processor 8041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 806 may include a display panel 8061, and the display panel 8061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 807 includes at least one of a touch panel 8071 and an input device 8072. The touch panel 8071 is also called a touch screen. The touch panel 8071 may include two parts: a touch detection device and a touch controller. The input device 8072 may include, but is not limited to, a physical keyboard, a function key (such as a volume control key, a switch key, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
[0144] The memory 809 can be used to store software programs and various data. The memory 809 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 809 may include a volatile memory or a non-volatile memory, or the memory 809 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 809 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0145] The processor 810 may include one or more processing units; optionally, the processor 810 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor 810.
[0146] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned call control method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0147] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0148] An embodiment of the present application also provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned call control method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0149] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0150] The embodiment of the present application also provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned call control method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0151] It should be noted that, in this article, the terms "comprise", "include" or any variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes elements that are not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in the examples.
[0152] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for enabling a terminal (such as a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0153] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A call control method, applied to a first electronic device, characterized in that: The method comprises: During a voice call between a first user using the first electronic device and a second user using a second electronic device, obtaining first audio through an audio output component of the first electronic device; and obtaining second audio through an audio input component of the first electronic device; Performing semantic recognition on the first audio and the second audio to determine a first conversation topic and a second conversation topic; wherein the first conversation topic is a conversation topic of a voice call between the second user and the first user; In case the first conversation topic is different from the second conversation topic, the audio input component is closed.
2. The method according to claim 1, characterized in that The performing semantic recognition on the first audio and the second audio to determine the first conversation topic and the second conversation topic includes: Performing semantic recognition on the first audio to obtain a first semantic recognition result; Performing semantic recognition on the second audio to obtain a second semantic recognition result; Determining a conversation topic corresponding to the second user according to the first semantic recognition result and the second semantic recognition result; Determining the conversation topic corresponding to the second user as the first conversation topic; A second conversation topic is determined according to the second semantic recognition result.
3. The method according to claim 2, characterized in that Before the step of determining the conversation topic corresponding to the second user as the first conversation topic, the method further includes: Determine a first application in the first electronic device; wherein the first application is an application running in the foreground; The step of determining the conversation topic corresponding to the second user as the first conversation topic includes: When the semantic correlation between the program topic of the first application and the conversation topic corresponding to the second user is higher than a first threshold, the conversation topic corresponding to the second user is determined as the first conversation topic.
4. The method according to claim 1, characterized in that: Before the step of closing the audio input component, the method further includes: In a case where the first conversation topic is different from the second conversation topic, detecting whether the first user faces toward the first electronic device; When the first conversation topic is different from the second conversation topic, closing the audio input component includes: When the first conversation topic is different from the second conversation topic and the face of the first user is not facing the first electronic device, the audio input component is turned off.
5. The method according to claim 1, characterized in that Before the step of closing the audio input component, the method further includes: When the first conversation topic is different from the second conversation topic, obtaining a third audio through the audio input component; When the first conversation topic is different from the second conversation topic, closing the audio input component includes: When the first conversation topic is different from the second conversation topic and the third audio complies with the second conversation topic, the audio input component is closed.
6. A call control device, applied to a first electronic device, characterized in that: The device comprises: A first acquisition module is used to acquire first audio through an audio output component of the first electronic device during a voice call between a first user using the first electronic device and a second user using a second electronic device; A second acquisition module, used for acquiring a second audio through an audio input component of the first electronic device; a recognition module, configured to perform semantic recognition on the first audio and the second audio to determine a first conversation topic and a second conversation topic; wherein the first conversation topic is a conversation topic of a voice call between the second user and the first user; A control module is used to close the audio input component when the first dialogue topic is different from the second dialogue topic.
7. The device according to claim 6, characterized in that The identification module comprises: A first recognition submodule, configured to perform semantic recognition on the first audio to obtain a first semantic recognition result; A second recognition submodule, used to perform semantic recognition on the second audio to obtain a second semantic recognition result; A first determination submodule, configured to determine a conversation topic corresponding to the second user according to the first semantic recognition result and the second semantic recognition result; A second determining submodule, configured to determine the conversation topic corresponding to the second user as the first conversation topic; The third determination submodule is used to determine a second dialogue topic according to the second semantic recognition result.
8. The device according to claim 7, characterized in that The device also includes: A determination module, configured to determine a first application in the first electronic device; wherein the first application is an application running in the foreground; The second determining submodule includes: A determination unit is used to determine the conversation topic corresponding to the second user as the first conversation topic when the semantic correlation between the program topic of the first application and the conversation topic corresponding to the second user is higher than a first threshold.
9. The device according to claim 6, characterized in that The device also includes: A detection module, configured to detect whether the first user faces toward the first electronic device when the first conversation topic is different from the second conversation topic; The control module comprises: The first control submodule is used to turn off the audio input component when the first conversation topic is different from the second conversation topic and the face of the first user is not facing the first electronic device.
10. The device according to claim 6, characterized in that The device also includes: A third acquisition module, configured to acquire a third audio through the audio input component when the first conversation topic is different from the second conversation topic; The control module comprises: The second control submodule is used to close the audio input component when the first dialogue topic is different from the second dialogue topic and the third audio complies with the second dialogue topic.