Network adjustment method and device

By detecting voice content and network quality parameters in video calls and judging abnormal call risks, the problem of AI face swapping technology being used for video fraud is solved, and the security and detection speed of video calls are improved.

CN120151464APending Publication Date: 2025-06-13VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510285389.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

AI face-changing technology is easily used by criminals in video calls, resulting in video fraud, making it difficult for users to distinguish the authenticity, and the security of video calls is poor.

Method used

By detecting preset keywords in voice content during video calls, obtaining network quality parameters, and when abnormal conditions are met, the call risk abnormality is judged based on image and audio data, and a prompt message is displayed to warn the user.

Benefits of technology

It improves the speed and accuracy of detecting AI face-changing videos, improves the security of video calls, and reduces the risk of users falling into fraud traps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151464A_ABST
    Figure CN120151464A_ABST
Patent Text Reader

Abstract

The invention discloses a network adjustment method and device, and belongs to the technical field of electronics. The method comprises the following steps: in a video call process, if the voice content of the video call contains a preset keyword, acquiring a network quality parameter in the video call process; under the condition that the network quality parameter meets the network abnormal condition, determining whether the video call has a call risk abnormity or not based on at least one of image data and audio data in the video call process; and under the condition that the video call has the call risk abnormity, first prompt information is displayed, and the first prompt information is used for prompting that the video call has the call risk abnormity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic technology, and particularly relates to a network adjustment method and device. Background Art

[0002] With the rapid development of Artificial Intelligence (AI) technology, AI face-swapping technology has been widely applied in various fields. For example, during a video call, AI face-swapping technology can use deep learning algorithms to replace the face of a person in the video with the face of another person. By analyzing a large amount of facial image data, the AI model can generate highly realistic human faces and make them naturally present various expressions in the video.

[0003] However, AI face-swapping technology is easily exploited by criminals to carry out video fraud. When criminals have a video call with the victim, they often use AI face-swapping technology to forge the appearance of others and create highly realistic videos, making it difficult for the victim to distinguish the true from the false. Without the user's knowledge, it is easy to fall into a fraud trap, resulting in property losses and privacy leaks. Thus, the security of video calls through electronic devices is relatively poor. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a network adjustment method and device, which can improve the security of video calls through electronic devices.

[0005] In a first aspect, the embodiments of this application provide a network adjustment method, which includes: during a video call, if the voice content of the video call contains a preset keyword, obtain the network quality parameters during the video call; when the network quality parameters meet the network anomaly condition, determine whether there is a call risk anomaly in the video call based on at least one of the image data and audio data during the video call; when there is a call risk anomaly in the video call, display a first prompt message, which is used to prompt that there is a call risk anomaly in the video call.

[0006] In a second aspect, the embodiments of this application provide a network adjustment device, which includes: an acquisition module, a determination module, and a display module. The above acquisition module is used to obtain the network quality parameters during the video call if the voice content of the video call contains a preset keyword during the video call. The above determination module is used to determine whether there is a call risk anomaly in the video call based on at least one of the image data and audio data during the video call when the network quality parameters obtained by the acquisition module meet the network anomaly condition. The above display module is used to display a first prompt message, which is used to prompt that there is a call risk anomaly in the video call when the determination module determines that there is a call risk anomaly in the video call.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instructions that can run on the processor. When the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instructions are stored. When the program or instructions are executed by a processor, the steps of the method described in the first aspect are implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instructions to implement the method described in the first aspect.

[0010] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0011] In the embodiment of the present application, during a video call, if the voice content of the video call contains a preset keyword, network quality parameters during the video call are obtained. Then, when the network quality parameters meet the network anomaly condition, based on at least one of the image data and audio data during the video call, it is determined whether there is a call risk anomaly in the video call. When there is a call risk anomaly in the video call, a first prompt message is displayed, and the first prompt message is used to prompt that there is a call risk anomaly in the video call. In this solution, since the network quality parameters during the video call can be obtained, it is possible to determine whether the network anomaly condition is met according to the network quality parameters, that is, whether the call quality of the video call is poor. Then, when the network quality parameters meet the network anomaly condition, that is, the call quality is poor, based on at least one of the image data and audio data during the video call, it is determined whether there is a call risk anomaly in the video call, that is, the defects of the AI face-swapping video can be magnified through the poor call quality, and the anomalies of the AI face-swapping video can be exposed. Then, in this case, the video call can be analyzed and judged, thereby improving the speed and accuracy of detecting the AI face-swapping video. In this way, the security of video calls through the electronic device is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is one of the flowcharts of the network adjustment method provided by the embodiment of the present application;

[0013] Figure 2 is the second flowchart of the network adjustment method provided by the embodiment of the present application;

[0014] Figure 3 It is the third flowchart of the network adjustment method provided by the embodiments of the present application;

[0015] Figure 4 It is the fourth flowchart of the network adjustment method provided by the embodiments of the present application;

[0016] Figure 5 It is the fifth flowchart of the network adjustment method provided by the embodiments of the present application;

[0017] Figure 6 It is a schematic diagram of the video call interface provided by the embodiments of the present application;

[0018] Figure 7 It is the sixth flowchart of the network adjustment method provided by the embodiments of the present application;

[0019] Figure 8 It is the seventh flowchart of the network adjustment method provided by the embodiments of the present application;

[0020] Figure 9 It is a schematic diagram of the execution process of the network adjustment method provided by the embodiments of the present application;

[0021] Figure 10 It is one of the schematic diagrams of the network adjustment device provided by the embodiments of the present application;

[0022] Figure 11 It is another schematic diagram of the network adjustment device provided by the embodiments of the present application;

[0023] Figure 12 It is a schematic diagram of the structure of the electronic device provided by the embodiments of the present application;

[0024] Figure 13 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0025] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application fall within the protection scope of the present application.

[0026] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or more. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0027] The terms "at least one (item)", "at least one of", etc. in this application refer to any one, any two or more combinations of the objects it contains. For example, at least one (item) of a, b, and c can represent: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two (items)" means two or more, and its meaning is similar to that of "at least one (item)".

[0028] The network adjustment method and device provided by the embodiments of this application will be described in detail below in conjunction with the accompanying drawings, through specific embodiments and their application scenarios.

[0029] The embodiments of this application can be applied to the scenario where a user uses an electronic device for a video call.

[0030] Taking some specific scenarios of the embodiments of this application as examples, the network adjustment method provided by the embodiments of this application will be described exemplarily below.

[0031] Scenario 1: Suppose the user receives a video call from a scammer on a social software. The scammer uses AI face-swapping technology to forge the face of the user's relative and asks for money in the name of seeing a doctor. The electronic device can detect the voice content during the video call. When it monitors that the voice content contains the preset keyword "borrow money", the electronic device can detect whether the current network quality parameter meets the network anomaly condition, so as to judge and adjust the current call quality. In the case of poor call quality, it can further analyze the audio and video, such as checking whether there is any abnormality in the face of the other party, whether the voice is synchronized with the picture, etc., to determine whether there is an abnormal call risk. Once an abnormality is found, the electronic device can display a prompt message to prompt the user that they may have encountered fraud, thereby raising the user's vigilance and protecting the user's property safety.

[0032] Scenario 2: Suppose that when a user is participating in a video conference, a fraudster steals the account of another conference member and uses AI face-swapping technology to replace that member in the video conference. The electronic device can detect the voice content during the video call of the video conference. When it monitors that the voice content contains business preset keywords such as "contract" and "funds", the electronic device can detect whether the current network quality parameters meet the network anomaly conditions, so as to judge and adjust the current call quality. And when the call quality is poor, it further analyzes the audio and video, such as checking whether there are abnormalities in the faces of each conference member and whether the voice is synchronized with the picture, etc., to determine whether there is an abnormal call risk. Once it is found that the video call of a certain conference member is abnormal, the electronic device can display a prompt message to prompt the user that there is an abnormality in this conference member, who may be a fraudster, and can notify other participants, thereby raising the vigilance of the participants and protecting the privacy and property safety of the participants.

[0033] Scenario 3: Suppose that in a financial transaction scenario, a user makes a video call with a customer service staff through the bank's online video service to confirm identity and handle a large amount of fund transfer business. A fraudster uses AI face-swapping technology to impersonate the bank's customer service staff and induces the user to provide privacy information for improper operations. Then the electronic device can detect the voice content during the video call. When it monitors that the voice content contains financial-related preset keywords such as "transfer" and "account", the electronic device can detect whether the current network quality parameters meet the network anomaly conditions, so as to judge and adjust the current call quality. And when the call quality is poor, it further analyzes the audio and video, such as checking whether there are abnormalities in the face of the customer service staff and whether the voice is synchronized with the picture, etc., to determine whether there is an abnormal call risk. Once an abnormality is found, the electronic device can display a prompt message to prompt the user that they may have been scammed, thereby raising the user's vigilance and protecting the user's property safety and privacy.

[0034] It should be noted that the above Scenarios 1 to 3 only exemplarily list some scenarios where the embodiments of the present application may be applied. In actual implementation, the embodiments of the present application can also be applied to more scenarios where network adjustment is required for any possible needs, and the embodiments of the present application are not limited thereto.

[0035] An embodiment of the present application provides a network adjustment method and device. Since the electronic device can obtain network quality parameters during a video call, the electronic device can determine whether the network anomaly condition is satisfied according to the network quality parameters, that is, whether the call quality of the video call is poor. Then, when the network quality parameters satisfy the network anomaly condition, that is, the call quality is poor, the electronic device can determine whether there is a call risk anomaly in the video call based on at least one of the image data and audio data during the video call. That is, the electronic device can amplify the defects of the AI face-swapping video through the poor call quality and expose the anomalies of the AI face-swapping video. Then, in this case, the video call can be analyzed and judged, thereby improving the speed and accuracy of detecting the AI face-swapping video. In this way, the security of the electronic device during video calls is improved.

[0036] The execution subject of the network adjustment method provided by the embodiment of the present application can be a network adjustment device, and this network adjustment device can be an electronic device, or a functional module or functional entity in the electronic device. Hereinafter, taking the electronic device as an example, the technical solution provided by the embodiment of the present application will be described.

[0037] Figure 1 The flowchart of a network adjustment method provided by an embodiment of the present application is shown. As Figure 1 shown, the network adjustment method provided by the embodiment of the present application may include the following steps 201 to 203.

[0038] Step 201: During a video call, if the voice content of the video call contains a preset keyword, the electronic device obtains network quality parameters during the video call.

[0039] In the embodiment of the present application, the above-mentioned video call refers to the realization of real-time video and audio communication between two or more users through the video call function of the electronic device. For example, home video communication, internal company video conferencing, video business handling, video online education, video medical consultation, etc.

[0040] Optionally, in the embodiment of the present application, the above-mentioned video call can be realized through an application with a video call function. Specifically:

[0041] (1) Home video communication can be realized through the video call function of a social application.

[0042] (2) Internal company video conferencing can be realized through the conferencing function of an office application or a conferencing application.

[0043] (3) Video business handling can be realized through the video handling function of a banking application or a government affairs application.

[0044] (4) Video online education can be realized through the live course function of an online course application or an education application.

[0045] (5) Video medical consultation can be achieved through functions such as remote consultation and video consultation of medical applications.

[0046] Optionally, in the embodiments of the present application, the electronic device can determine whether the current interface is related to a video call by obtaining the package name. Specifically, the electronic device can obtain the package name of the current top-level application and the current foreground application interface through a specific command, such as the dumpsys activity top|grep ACTIVITY command, where dumpsysactivity top is a specific command for obtaining the foreground application interface with which the user is currently interacting, and grepACTIVITY is used to filter the output result to filter out the key information related to the activity, so that the specific activity of the current foreground application can be located more efficiently, avoiding searching in a large amount of irrelevant information.

[0047] Optionally, in the embodiments of the present application, the electronic device can take a screenshot or enable the accessibility service to monitor the changes of interface elements and analyze their content. When making a video call, the electronic device can detect image elements in the interface, such as the image display frame of the front camera, specific buttons (such as the "hang up" button), etc., to determine that the current interface is a video call interface.

[0048] Optionally, in the embodiments of the present application, the electronic device can determine whether it is in the process of a video call by detecting the camera and the phone status. Specifically, when making a video call, the camera of the electronic device is usually called to collect video images. The electronic device can obtain the usage status of the camera by accessing the system service related to the camera, such as checking whether the camera is turned on and whether it is collecting images, etc., for determining whether a video call is in progress. The electronic device can also obtain the phone status information by listening to the broadcasts related to the phone status, such as listening to the status changes of the phone being connected, hung up, ringing, etc. If it is detected that the phone is in the connected state, combined with the judgment results of other methods, it can be further confirmed that an audio call is in progress.

[0049] Optionally, in the embodiments of the present application, the electronic device can further determine the specific scenario of the current video call and determine whether to perform subsequent operations according to different scenarios. For example, the electronic device can continuously detect the call content in the case of an online meeting with a higher security level, or the electronic device can continuously detect the call content in the case of a video call with a stranger or a relative or friend, or the electronic device can default to detecting the call content of all video calls to ensure the communication security of the user.

[0050] In the embodiments of the present application, the above-mentioned preset keywords are some sensitive words set in advance, such as words related to money or privacy, such as "borrowing money", "contract", "funds", "transfer", etc. When the voice content of the video call contains these preset keywords, the electronic device can determine that the video call involves a risk topic, thereby triggering subsequent operations such as obtaining network quality parameters.

[0051] Optionally, in the embodiments of the present application, during the call, the electronic device can collect the user's voice through the microphone, convert it into a voice signal, and then the electronic device can use the speech recognition library to convert the voice signal into text content, and analyze and process the text content, which can include keyword matching, grammar analysis, and semantic analysis, etc. For example, the electronic device can use regular expressions to find out whether the text contains specific preset keywords.

[0052] Optionally, in the embodiments of the present application, during the call, the electronic device can receive the encoded voice data of the call partner, and then the electronic device can use the audio decoder to decode the voice data to obtain an audio signal, and then the electronic device can use the speech recognition library to convert the voice signal into text content, and analyze and process the text content. For the specific method, refer to the above description and will not be elaborated here.

[0053] In the embodiments of the present application, the above-mentioned network quality parameters are parameters used to measure the network condition, and these parameters can reflect the stability and transmission ability of the network, and have an important impact on the quality of the video call.

[0054] Optionally, in the embodiments of the present application, the above-mentioned network quality parameters may include at least one of the following: network parameters during the video call, video parameters during the video call.

[0055] Exemplarily, the above-mentioned network parameters may include but are not limited to at least one of the following: network delay, packet loss rate, bandwidth, etc. Among them:

[0056] (1) The network delay is the time elapsed from when the data is sent from the sending end to when it is received by the receiving end. For example, the network delay can be 50 milliseconds (ms), 100 ms, etc. In the case of a relatively high network delay, for example, when the network delay is greater than 1000 ms, there will be an obvious delay in the transmission of the picture and sound in the video call, resulting in an unsmooth interaction between the callers. For example, when one caller asks a question, the other caller needs to wait for a few seconds to hear the question before answering, and then it takes another few seconds to transmit so that the questioner can hear, seriously affecting the communication efficiency.

[0057] (2) The packet loss rate is the ratio of the number of lost data packets to the total number of data packets during data transmission, such as 1% to 3%. In the case of a relatively high packet loss rate, for example, when the packet loss rate is 10%

[0058] When the above occurs, problems such as video frame freezes, blurry images, and audio interruptions may occur during a video call. For example, when a call participant is explaining a solution, the video may suddenly freeze for a few seconds and then resume playing, or the image may become blurry, or the sound may suddenly cut out for a few seconds and then return, affecting the viewing experience of other call participants and reducing the call quality.

[0059] (3) Bandwidth is the amount of data that the network can transmit per unit of time, such as 1 megabit per second (Mbps)

[0060] to 2 Mbps. When the bandwidth is insufficient, for example, when it is less than 128 kilobits per second (Kbps)

[0061] the resolution and frame rate of the video call may decrease. The originally clear image may become blurry, the video may freeze, and the movement may not be smooth, affecting the viewing experience of the call participants.

[0062] Moreover, insufficient bandwidth can also limit group video calls. For example, in a 10-person video conference, due to low bandwidth, only the video and audio of 5 people can be normally displayed and transmitted, while the video and audio of others will experience freezing, latency, or loss, affecting the normal progress of the conference.

[0063] Exemplarily, the above video parameters may include, but are not limited to, at least one of the following: frame rate, bit rate, resolution, encoding method, etc. These parameters can reflect the network quality of the video call from the perspective of video quality. Specifically:

[0064] (1) The frame rate is the number of times the video image is updated per second, usually expressed in frames per second (FPS). The higher the frame rate, the smoother the video image. When the network condition is good, the video call can maintain a high frame rate; when the network condition is poor, in order to ensure the smoothness of the video call, the system may reduce the frame rate. For example, when the network bandwidth is insufficient, the frame rate of the video call may be reduced from 30 FPS to 15 FPS, resulting in video frame freezes.

[0065] (2) The bit rate is the transmission rate of video data per unit of time, usually expressed in Kbps or Mbps. The higher the bit rate, the better the details and quality of the video image. When the network condition is good, the video call can use a high bit rate; when the network condition is poor, in order to ensure the smoothness of the video call, the electronic device may reduce the bit rate of the video call. For example, when the network bandwidth is insufficient, the bit rate of the video call may be reduced from 1 Mbps to 500

[0066] Kbps, resulting in phenomena such as blurred images and loss of details.

[0067] (3) Resolution refers to the number of pixels in the video frame, usually represented by the number of pixels in width and height, such as 1920×1080. The higher the resolution, the clearer the video frame. When the network condition is good, a higher resolution can be used for video calls; while when the network condition is poor, in order to ensure the smoothness of video calls, the electronic device may reduce the resolution. For example, when the network bandwidth is insufficient, the resolution of the video call may be reduced from 1920×1080 to 1280×720, resulting in phenomena such as blurred images and loss of details.

[0068] (4) The encoding method, also known as the encoding format, refers to the way of converting video signals into digital signals and performing compression encoding. Common ones include H.264, H.265, MPEG-4, etc. Different video encoding formats have different requirements for network bandwidth. Formats with higher encoding efficiency can provide better video quality than those with lower encoding efficiency under the same bandwidth. If the network bandwidth is insufficient, using a format with lower encoding efficiency for video calls may result in stuttering, degraded video quality, etc. For example, in a poor network condition, a video call using the H.265 encoding format may be smoother than a video call using the H.264 encoding format because the H.265 encoding format can provide a higher compression ratio and better video quality under the same bandwidth.

[0069] In the embodiments of the present application, when it is detected that a video call is in progress, the electronic device can continuously monitor the voice content of the video call, convert this voice content into text, and compare it with preset keywords. When the voice content matches the preset keywords, it is determined that the video call involves a risk topic. When it is detected that the voice content of the video call contains preset keywords, the electronic device can obtain the network quality parameters during the video call to determine whether the current network condition is normal, providing a basis for subsequent abnormal call risk detection.

[0070] Step 202: When the network quality parameters meet the network anomaly conditions, the electronic device determines whether there is an abnormal call risk in the video call based on at least one of the image data and audio data during the video call.

[0071] In the embodiments of the present application, the above-mentioned image data is a series of image data received by the electronic device during the video call, including each frame of the video frame.

[0072] In the embodiments of the present application, the above-mentioned audio data is a series of audio data received by the electronic device during the video call, including all the voice content of the video call.

[0073] In the embodiments of the present application, the above-mentioned call risk anomaly refers to an abnormal situation where the identity information of the call object may not match the actual situation, posing a risk of identity theft, which may cause users to suffer property losses, privacy leaks, or other security threats.

[0074] Optionally, in the embodiments of the present application, when the network quality parameter meets the network anomaly condition, that is, when the first score is less than the first threshold, the electronic device can analyze the image data during the video call through the audio-visual acquisition and detection module to check whether there is an anomaly in the face of the other party. For example, the electronic device can detect the key points of the face through the audio-visual acquisition and detection module and determine whether the positions, proportions, movements, etc. of the facial features conform to the characteristics of normal people. The electronic device can also analyze information such as the texture and color of the face to determine whether there are traces of forgery. In addition, the electronic device can also check the clarity and stability of the video screen to determine whether there are abnormal image changes.

[0075] Optionally, when the network quality parameter meets the network anomaly condition, that is, when the first score is less than the first threshold, the electronic device can analyze the audio data during the video call to check whether there is an anomaly in the voice of the other party. For example, the electronic device can convert the voice into text through speech recognition technology and compare it with the lip movements in the video screen to determine whether there is a mismatch. The electronic device can also analyze features such as the timbre, intonation, and speech rate of the voice to determine whether there are abnormal voice changes.

[0076] Optionally, in the embodiments of the present application, in combination with Figure 1 , as Figure 2 shown, the above step 202 can be specifically implemented through the following step 202a and step 202b.

[0077] Step 202a: When the network quality parameter meets the network anomaly condition, the electronic device obtains facial feature information from the image data.

[0078] In the embodiments of the present application, the above-mentioned facial feature information is information related to the human face extracted by the electronic device from the image data, including the key points of the face, such as the positions of the eyes, nose, mouth, eyebrows, etc., and also includes the proportions of the facial features, facial contours, expressions, skin colors, etc.

[0079] Optionally, in the embodiments of the present application, after introducing a delay in the network adjustment module, the electronic device can display a banner prompt on the video call interface to prompt the user to pay attention to the fluctuations of the other person's face and the abnormalities of the voice. Because after the audio and video are stuck, the possibility of AI failure increases, and it is possible that a certain frame or a certain segment is not replaced by AI for face and voice, and thus returns to the original face and voice of the attacker, or there may be obvious video freezes, facial feature misalignments, synchronization errors, etc. Similarly, for AI voice conversion, there may be audio abnormalities, inconsistencies, etc., which are easily perceived by the user. If the call partner does not use AI face or voice conversion technology but uses a real camera, the video quality may decrease slightly, but the overall impact is not significant.

[0080] Step 202b: When the facial feature information meets the facial abnormality condition, the electronic device determines that there is a call risk abnormality in the video call.

[0081] Optionally, in the embodiments of the present application, the above facial abnormality condition includes at least one of the following:

[0082] There is an offset between the facial pixels and the background edge pixels;

[0083] There is a difference between the facial shadow and the background shadow;

[0084] Abnormal facial movements;

[0085] The pupil highlight points of the face do not conform to the scene light source;

[0086] The facial features do not match the face images stored locally.

[0087] In the embodiments of the present application, the specific determination method of the above facial abnormality condition is as follows:

[0088] (1) There is an offset between the facial pixels and the background edge pixels: The electronic device can identify the boundary line between the facial area and the background area, compare the positional relationship between the facial edge pixels and the background edge pixels, and check whether there is an abnormal displacement or misalignment. For example, in a normal face image, the facial edge pixels should gradually transition to the background pixels to form a smooth boundary, and the offset should be less than or equal to 1 pixel. If it is found that the facial edge pixels are suddenly interrupted or there is an obvious misalignment, such as an offset of more than 2 pixels between the face and the background edge, the electronic device can determine that there is a call risk abnormality in the video call.

[0089] (2) There are differences between facial shadows and background shadows: Since the formation of facial shadows is usually related to background shadows to a certain extent because they are affected by the same environmental light source. The electronic device can analyze the light and shadow distribution in the picture, detect the shadow intensity, direction, and distribution of the facial and background areas, and analyze through a grayscale histogram to determine whether there are inconsistent situations. For example, if the difference in the shadow direction or intensity between the face and the environment exceeds 20%, the electronic device can determine that there is an abnormal call risk in the video call.

[0090] (3) Abnormal facial movements: Normal facial expressions and movements have a certain degree of continuity and reasonableness. For example, actions such as blinking and opening the mouth have certain rules. The electronic device can analyze the dynamic data in the facial feature information and track the movement trajectories and expression changes of facial muscles. The electronic device can judge whether there are abnormal facial movements by calculating the change rate of facial expressions and the continuity of facial muscle movements. For example, the normal blinking frequency of a person is 15 - 20 times per minute, while the blinking frequency of a fake face with AI face swapping may be much higher or lower than this range, such as 5 times. When the electronic device detects sudden, stiff, or uncoordinated facial movements, such as abnormal blinking frequency or abnormal mouth opening and closing amplitude, the electronic device can determine that there is an abnormal call risk in the video call.

[0091] (4) The highlight points of the pupils on the face do not match the scene light source: The highlight points of the pupils are the reflection points formed when light shines on the pupils. Under normal circumstances, the position and size of the highlight points depend on the direction and intensity of the environmental light source. The electronic device can identify the light source position and intensity in the picture and compare the relationship between the highlight points of the pupils and the light source. If the highlight points of the pupils do not appear in the position where they should be, or their size and shape do not match the light source characteristics, this may be a sign of face synthesis. For example, in an indoor scene where the light source comes from above, but the highlight points of the pupils on the face appear at the lower part of the pupils, this indicates that the face may have undergone AI face swapping.

[0092] (5) Facial features do not match the face images stored locally: When there are face images of the contacts in the video call stored locally in the electronic device, the electronic device can use face recognition technology to compare the facial feature information in the current video with the facial image features in the photo of this object stored locally. The electronic device can calculate the similarity between the two facial features. If the similarity is lower than a certain threshold, the electronic device can judge that the facial features do not match the face images stored locally, and the face in the current video call may be forged.

[0093] In this way, the electronic device can accurately judge whether there is an abnormal call risk in the video call through facial abnormal conditions, so as to timely discover abnormal situations and remind users, improving the security of the video call.

[0094] Optionally, in the embodiments of the present application, in combination with Figure 1 , as Figure 3 shown, the above step 202 can be specifically implemented by the following steps 202c to 202e.

[0095] Step 202c: When the network quality parameter meets the network anomaly condition, the electronic device obtains voiceprint feature information from the audio data.

[0096] In the embodiments of the present application, the above voiceprint feature information is acoustic feature data extracted from the audio data that can characterize the unique identity of the speaker, and may include parameters such as the frequency, amplitude, speech rate, and pitch of the voice signal, for example.

[0097] Optionally, in the embodiments of the present application, the electronic device can intercept a segment of real-time audio signal from the audio data of the current video call, and use voiceprint recognition technology to extract features from the preprocessed audio signal to obtain voiceprint feature information. For example, the electronic device can perform spectral analysis on the audio signal, extract features such as fundamental frequency, formant, and short-time energy, and form a voiceprint feature vector.

[0098] Step 202d: Based on the voiceprint feature information, the electronic device determines whether there is voiceprint distortion in the audio data.

[0099] In the embodiments of the present application, the above voiceprint distortion means that there are obvious differences or abnormal changes between the voiceprint features in the audio data and the normal voiceprint features.

[0100] Optionally, in the embodiments of the present application, the electronic device can analyze the fundamental frequency (F0) fluctuation of the extracted voiceprint information. Generally, the fundamental frequency fluctuation of human speech is less than 20 Hertz (Hz). When it exceeds 50 Hz, the electronic device can determine that there is voiceprint distortion.

[0101] Optionally, in the embodiments of the present application, the electronic device can perform comparative analysis on the extracted voiceprint feature information and a preset normal voiceprint template. The normal voiceprint template can be established based on the previous normal voice data of the call object and contains stable voiceprint features. The electronic device can judge whether there is distortion by calculating the similarity between the voiceprint feature information and the normal voiceprint template. Common similarity calculation methods include Euclidean distance, cosine similarity, etc. If the similarity is lower than the preset threshold, the electronic device can determine that there is voiceprint distortion.

[0102] Optionally, in the embodiments of the present application, the electronic device may compare the extracted voiceprint feature information with the mouth movement of the face in the image data. For example, by analyzing features such as the frequency, amplitude, and shape of the mouth movement, it is matched with the voice features in the audio data. If the mouth movement is inconsistent with the voice features, such as the mouth movement lagging behind or leading the voice, or the mouth movement not matching the voice content, the electronic device may determine that there is voiceprint distortion.

[0103] Step 202e: When there is voiceprint distortion in the voiceprint, the electronic device determines that there is an abnormal call risk in the video call.

[0104] Optionally, in the embodiments of the present application, when it is determined that there is distortion in the voiceprint, the electronic device may combine other detection results, such as facial feature judgment, etc., to comprehensively determine whether there is an abnormal call risk in the video call. If the degree of voiceprint distortion is relatively high, the electronic device may determine that there is an abnormal call risk in the video call.

[0105] In this way, the electronic device can further determine whether there is an abnormal call risk in the video call, such as situations that may cause risks such as voice synthesis and voice tampering, thereby improving the security of the video call.

[0106] Step 203: When there is an abnormal call risk in the video call, the electronic device displays a first prompt message.

[0107] In the embodiments of the present application, the above first prompt message is used to prompt that there is an abnormal call risk in the video call.

[0108] Optionally, in the embodiments of the present application, the above first prompt message includes at least one of the following: a prompt message to terminate the call, a prompt message to record a video, and a prompt message for identity verification.

[0109] Exemplarily, the user uses a social application to make a video call with a friend. During the call, the electronic device analyzes and finds that the other party's blinking frequency is 5 times per minute, far lower than the normal range, and the fundamental frequency fluctuation in the voiceprint feature information exceeds 50 Hz. The electronic device determines that there is an abnormal call risk and displays a window in the video call interface, and displays the first prompt message in the window. The prompt content is as follows: "Call risk reminder: There may be an abnormal call risk in the current video call. It is suspected that the call is not from the person himself. Please be vigilant. It is recommended that you terminate the call to ensure safety. If you still need to continue the call, you can choose to record a video for subsequent verification, or request the other party to perform identity verification." And the electronic device displays three controls below the window, namely "Terminate call", "Record video", and "Continue call", and the user can select operations according to the actual situation to ensure communication security.

[0110] In the embodiments of the present application, the above-mentioned call termination prompt message is used to advise the user to end the current video call that may have risks, so as to prevent the user from continuing to be threatened by potential risks. The above-mentioned video recording prompt message is used to advise the user to record the current video call, so that the user can save the call record for subsequent investigation or verification. The above-mentioned identity verification prompt message is used to advise the user to verify the identity of the other party, and to prompt the user to verify the true identity of the other party through other means, such as phone calls, text messages, third-party verification platforms, etc., to ensure the security of the call.

[0111] The embodiments of the present application provide a network adjustment method and device. Since the electronic device can obtain the network quality parameters during the video call, the electronic device can determine whether the network anomaly condition is met according to the network quality parameters, that is, whether the call quality of the video call is poor. Then, when the network quality parameters meet the network anomaly condition, that is, when the call quality is poor, the electronic device can determine whether there is a call risk anomaly in the video call based on at least one of the image data and audio data during the video call. That is, the electronic device can magnify the defects of the AI face-swapping video through the poor call quality and expose the anomalies of the AI face-swapping video. Then, in this case, the video call can be analyzed and judged, thereby improving the speed and accuracy of detecting the AI face-swapping video. In this way, the security of the electronic device for video calls is improved.

[0112] Optionally, in the embodiments of the present application, in combination with Figure 1 , as Figure 4 shown, after the above step 201, the network adjustment method provided by the embodiments of the present application further includes the following steps 301 and 302.

[0113] Step 301: The electronic device determines a first score based on the network quality parameters.

[0114] In the embodiments of the present application, the above-mentioned first score is used to characterize the call quality of the video call.

[0115] Optionally, in the embodiments of the present application, the electronic device may set weights for parameters such as network latency, packet loss rate, bandwidth, frame rate, bit rate, resolution, and encoding format respectively, multiply each weight by the corresponding parameter value, and then add up all the results to obtain a first score. The setting of the weights depends on the degree of influence of each parameter on the video call quality. The electronic device may set higher weights for parameters such as network latency and video latency that have a greater impact on real-time performance. For the convenience of calculation, the electronic device may normalize these parameter values and convert them into values between 0 and 1. For example, if the maximum value of network latency is set to 200 ms and the minimum value is set to 0 ms, then when the network latency is 100 ms, the normalized value of network latency is (200 - 100) / (200 - 0) = 0.5. Similarly, other parameters can be normalized in a similar manner. In particular, for the encoding format, specific values can be set for specific encoding formats. For example, the value of H.265 is set to 0.9, and the value of H.264 is set to 0.7, etc. Specifically, it can be set according to the compression efficiency and video quality of the encoding format.

[0116] Exemplarily, assume that when calculating the first score, the electronic device uses seven parameters: network latency, packet loss rate, bandwidth, video latency, frame rate, bit rate, and resolution, and sets weights for them respectively. Assume that the weights of the seven parameters are: the weight of network latency is 0.2, the weight of packet loss rate is 0.15, the weight of bandwidth is 0.1, the weight of frame rate is 0.2, the weight of bit rate is 0.1, the weight of resolution is 0.1, and the weight of encoding format is 0.05. Since network latency and frame rate have a greater impact on real-time performance, their weights are relatively high; while bandwidth and bit rate have a relatively small impact on the video call quality, so their weights are relatively low. Assume that in a video call, the values of the seven parameters obtained by the electronic device are: network latency 100 milliseconds, packet loss rate 2%, bandwidth 5 Mbps, frame rate 25 FPS, bit rate 1 Mbps, resolution 1280×720, and encoding format H.265. Assume that after converting the seven parameters into normalized parameter values, they are network latency 0.5, packet loss rate 0.4, bandwidth 0.6, frame rate 0.8, bit rate 0.7, resolution 0.9, and encoding format 0.9 respectively. Then the first score is 0.2×0.5 + 0.15×0.4 + 0.1×0.6 + 0.2×0.8 + 0.1×0.7 + 0.1×0.9 + 0.05×0.9 = 0.685.

[0117] Step 302: The electronic device determines whether the network quality parameter meets the network anomaly condition based on the first score.

[0118] In this way, the electronic device can quantify multiple network quality parameters into a first score through weighted combination, so as to more comprehensively and intuitively compare the impacts of different network qualities on video calls.

[0119] Optionally, in the embodiments of the present application, in combination with Figure 4 , such as Figure 5 shown, the above step 302 can be specifically implemented by the following step 302a or 302b.

[0120] Step 302a: When the first score is less than the first threshold, the electronic device determines that the network quality parameter meets the network anomaly condition.

[0121] In the embodiments of the present application, the above first threshold is a preset threshold for determining whether the network quality parameter meets the network anomaly condition.

[0122] In the embodiments of the present application, the above network anomaly condition is a standard for determining whether the network is abnormal. When the network quality parameter meets the network anomaly condition, it indicates that the network may have an abnormal situation, which may affect the normal progress of the video call.

[0123] Optionally, in the embodiments of the present application, the above score threshold can be the default of the electronic device or pre-set by the user. For example, the above score threshold can be 0.6, 0.7, etc. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not limit it.

[0124] Optionally, in the embodiments of the present application, after calculating the first score, the electronic device can compare the first score with a score threshold. If the first score is greater than or equal to the score threshold, the electronic device can determine that the network quality parameter is in an abnormal state. If the first score is less than the score threshold, the electronic device can determine that the network quality parameter is not in an abnormal state.

[0125] Exemplarily, assuming that the first score is 0.685 and the score threshold is 0.7, the electronic device can determine that the network quality parameter is in a network abnormal state.

[0126] Step 302b: When the first score is greater than or equal to the first threshold, the electronic device adjusts at least one of the network delay parameter and the video codec delay parameter of the video call until the first score is less than the first threshold, and determines that the network quality parameter meets the network anomaly condition.

[0127] It should be noted that, under normal circumstances, the network latency of a video call is about 50 ms, and the range of 50 - 100 ms is acceptable. When the network bandwidth decreases and the network latency increases to more than 100 ms, obvious latency will occur. When it exceeds 200 ms, the experience will be significantly affected, resulting in video frame freezes, pixelation, blurring, etc., affecting the smoothness of AI face swapping, leading to loss of facial details and mismatched facial features, such as deviations in the positions of eyes and mouths. At the same time, stable high frame rates are required for AI face swapping to maintain the coherence of facial details, such as above 24 fps. Otherwise, unnatural facial movements will occur, such as sudden expression changes and blurred edges, and pixel inconsistencies will also occur, such as skin tone breaks. On the other hand, for AI voice conversion, stable voice transmission requires a network latency of less than 150 ms. If the latency exceeds 300 ms, the smoothness of the voice can be disrupted, resulting in disrupted voice rhythms, sudden pitch changes, inconsistent speech speeds, noise interference, etc. in AI voice conversion, and may also cause long-term voice latency and interruptions. In the embodiments of the present application, when the first score is greater than or equal to the first threshold, that is, when the call quality is good, the electronic device can achieve the effect of actively reducing the call quality by adjusting at least one of the network latency parameters and video codec latency parameters of the video call, thereby exposing the synthesis traces of AI face swapping and AI voice conversion.

[0128] In the embodiments of the present application, the above network latency parameters are various parameters affecting the network state. It can be understood that the larger the value of the network latency parameter, the worse the network quality, the slower the network speed, the worse the real-time performance, the more video call freezes, and the more voice call interruptions.

[0129] Optionally, in the embodiments of the present application, the above network latency parameters may include at least one of the following: network rate, network bandwidth, packet loss rate, network jitter.

[0130] In the embodiments of the present application, the above video codec latency parameters are parameters of the audio and video codec, which are used to adjust the video quality. The audio and video codec is a software or hardware device for encoding and decoding audio and video data. The main function of the audio and video codec is to compress and encode the original audio and video data to reduce the data volume for easy storage and transmission. At the same time, at the receiving end, the audio and video codec can also decode the compressed audio and video data and restore it to the original audio and video signal for the user to play and watch.

[0131] Optionally, in the embodiments of the present application, the above video codec latency parameters may include at least one of the following: resolution, frame rate, bit rate, audio bit rate, encoding method.

[0132] Optionally, in the embodiments of the present application, the electronic device can control the device-side network latency at the operating system level to reduce the first score. The specific methods include:

[0133] (1) By means of a traffic control mechanism, such as traffic shaping, limit the uplink / downlink bandwidth and set the traffic transmission rate. For example, an electronic device can use the TrafficStats class or a third-party library to limit the network rate of a specific application, thereby controlling the speed of data sending and receiving and increasing network latency.

[0134] (2) During the data transmission process, insert fixed or random delays using system-level application programming interfaces (APIs). For example, an electronic device can indirectly increase the end-to-end delay and make the data packets take longer to transmit in the network by modifying Transmission Control Protocol / Internet Protocol (TCP / IP) protocol stack parameters, such as increasing the retransmission timeout or reducing the window size.

[0135] (3) Control the number of data packets transmitted per second, reduce the data throughput, and dynamically adjust the packet sending interval. For example, an electronic device can introduce jitter by means of random delays, making the packet sending intervals uneven and increasing the uncertainty and latency of the network.

[0136] (4) Lower the transmission priority of audio and video data packets so that they are queued or discarded in case of network congestion. In this way, an electronic device can simulate the situation of network congestion, increase the transmission delay of data packets, and thus reduce the quality of video calls.

[0137] (5) Use packet loss injection tools, such as network simulation modules, to actively discard some data packets. In a network simulation environment, an electronic device can randomly discard data packets according to the configured packet loss rate through these tools, creating an environment similar to network congestion or poor signal, thereby increasing the unreliability and latency of the network.

[0138] Optionally, in the embodiments of the present application, an electronic device can reduce the resolution, frame rate, bit rate, and audio bit rate by adjusting the parameters of the audio and video codec, and force the use of an inefficient coding format. For example, an electronic device can disable hardware acceleration and only use the Central Processing Unit (CPU) for encoding and decoding to reduce the speed; disable dynamic bit rate adjustment and reduce the audio sampling rate, etc. These methods can increase the computational complexity and processing time during the video encoding and decoding process, thereby increasing the video encoding and decoding latency and reducing the real-time performance and smoothness of video calls.

[0139] Optionally, in the embodiments of the present application, when the first score is greater than or equal to the first threshold, it indicates that the network quality parameter has not reached an abnormal state, but there may be certain risks. The electronic device can adjust at least one of the network delay parameter and the video codec delay parameter of the video call to reduce the first score until the first score is less than the first threshold. When the first score is reduced to less than the first threshold, the electronic device can determine that the network quality parameter meets the network anomaly condition, that is, the network quality is poor and there may be network anomalies.

[0140] For example, the electronic device can increase the video call delay and reduce the call quality by reducing the bandwidth of the video call, inserting a fixed delay, etc., so as to reduce the first score.

[0141] For another example, the electronic device can adjust parameters such as resolution and frame rate to reduce the video image quality, or select a less efficient video coding format to increase the delay in the video codec process, so as to reduce the first score.

[0142] Optionally, in the embodiments of the present application, the electronic device can adjust the parameters randomly or according to a specific pattern, or can perform dynamic intelligent adjustment according to the actual network condition of the user. During the process of adjusting the parameters, the current network can be continuously tested through the network monitoring module, and the network delay policy can be continuously adjusted according to the actual situation to balance security and user experience. At the same time, it also makes the AI face swapping algorithm difficult to predict and increases the cost for attackers to counteract, thereby enhancing the security of the video call.

[0143] Optionally, in the embodiments of the present application, after adjusting at least one of the network delay parameter and the video codec delay parameter of the video call, the electronic device can re-determine the first score and determine whether the first score is greater than or equal to the first threshold. If the first score is less than the first threshold, the electronic device can determine that the network quality parameter meets the network anomaly condition. If the first score is greater than or equal to the first threshold, the electronic device can continue to adjust at least one of the network delay parameter and the video codec delay parameter of the video call, and re-determine the first score after the adjustment. And so on, the electronic device can continuously adjust the parameters until the first score is less than the first threshold.

[0144] Optionally, in the embodiments of the present application, when the first score is greater than or equal to the first threshold, the electronic device can display the first information, and the user can input the first information to trigger the electronic device to adjust at least one of the network delay parameter and the video codec delay parameter of the video call, and perform subsequent judgment on the abnormal call risk.

[0145] Exemplarily, taking the electronic device as a mobile phone as an example, as Figure 6As shown in (a) of [the figure], it is assumed that the victim and the scammer are having a video call, and the mobile phone displays the video call interface 10. When the scammer asks to borrow money from the user and the electronic device detects the preset keyword "borrow money", then as Figure 6 shown in (b) of [the figure], the electronic device can display the first information 11 in the form of a window in the video call interface 10. The user can input the first information to trigger the electronic device to adjust at least one of the network delay parameter and the video codec delay parameter of the video call, and perform subsequent judgment on the abnormal call risk.

[0146] In this way, the electronic device can judge the call quality of the video call by comparing the first score with the first threshold, and when the first score is greater than or equal to the first threshold, that is, when the call quality is good, actively adjust the parameters to reduce the call quality of the video call, so that the video call using AI face swapping or AI voice conversion exposes the synthesis traces, increasing the possibility of being recognized, thereby improving the security of the call.

[0147] Optionally, in the embodiments of the present application, in combination with Figure 1 , as Figure 7 shown, after the above step 203, the network adjustment method provided by the embodiments of the present application further includes the following step 401 and step 402.

[0148] Step 401: The electronic device determines at least one contact associated with the video call.

[0149] In the embodiments of the present application, the above at least one contact associated with the video call is also called an associated contact, which is a person related to the current video call, such as a participant in the video call or other person related to the call content.

[0150] Optionally, in the embodiments of the present application, the electronic device can convert the voice content into text through voice recognition technology, and then use natural language processing technology to analyze the personnel information in the text. The electronic device can determine at least one contact associated with the video call according to the call information and the analysis result of the call content. These contacts can be both parties of the call or other persons mentioned in the call. For example, if a certain colleague of the user is mentioned in the video call, the electronic device can determine this colleague as an associated contact.

[0151] Step 402: The electronic device sends a second prompt message to at least one contact.

[0152] In the embodiments of the present application, the above second prompt message is used to prompt that there is an abnormal call risk in the video call.

[0153] It can be understood that, different from the first prompt message, the second prompt message is a prompt message sent to associated contacts to prompt them to pay attention to the possible risks in the video call.

[0154] Optionally, in the embodiments of the present application, the electronic device may generate corresponding second prompt messages according to the abnormal call risks of the video call. The second prompt message may include content indicating that there are abnormal risks in the video call, and measures recommended for the contacts, such as verifying the identity of the call counterpart and increasing vigilance. The electronic device may obtain the contact information of the associated contacts, such as phone numbers, email addresses, social application accounts, etc., from the address book, social applications or other relevant applications, and send the second prompt message to at least one relevant contact through corresponding means, such as text messages, emails, messages, etc.

[0155] Optionally, in the embodiments of the present application, if the user discovers an abnormality during the call, the user may input to the electronic device to trigger the electronic device to send a second prompt message to the relevant contacts.

[0156] In this way, the electronic device can send the second prompt message to the relevant contacts in a timely manner to prompt them to pay attention to the possible risks in the video call, improve the prevention awareness of the relevant contacts, prevent the spread of risks, thereby protecting the property and privacy security of more people, and further enhancing the security of the video call.

[0157] Optionally, in the embodiments of the present application, the electronic device may combine multiple identity verification means, such as SMS verification codes, real-time behavior verification, etc., and encourage the use of a secure communication platform with end-to-end encryption to cope with the evolving AI fraud means.

[0158] Optionally, in the embodiments of the present application, such as Figure 8As shown in the figure, the network adjustment method provided by the embodiment of the present application may include a voice monitoring module, a network monitoring module, a network adjustment module, a real-time audio and video acquisition and detection module, and an early warning module. Specifically, during a video call, the electronic device can continuously detect the call content through the voice monitoring module. When a preset keyword is detected, the network quality parameters during the video call are obtained through the network monitoring module, and a first score is calculated based on the network quality parameters. The first score is compared with a first threshold. If the first score is greater than or equal to the first threshold, the electronic device can adjust at least one of the network delay parameter and the video codec delay parameter through the network adjustment module to reduce the first score until the first score is less than the first threshold. When the first score is less than the first threshold, the electronic device can obtain and detect video data and audio data through the real-time audio and video acquisition and detection module to determine whether there is an abnormal call risk in the video call. When there is an abnormal call risk in the video call, the electronic device can display a first prompt message through the early warning module and send a second prompt message to relevant contacts to prompt the possible risks in the video call to the user and relevant contacts.

[0159] Figure 9 It is a schematic diagram of the execution process of the network adjustment method provided by the embodiment of the present application. As Figure 9 shown, the memory allocation method provided by the embodiment of the present application may include the following steps 10 to 15.

[0160] Step 10: When it is detected that the user answers the video call, the electronic device monitors high-risk words through the voice monitoring module.

[0161] Step 11: When high-risk words are detected, the electronic device obtains parameters such as the network status through the network monitoring module, scores them, and obtains a first score.

[0162] Step 12: The electronic device determines whether it is necessary to adjust the network of the current video call according to the first score through the network adjustment module.

[0163] When the electronic device determines that it is necessary to adjust the network of the current video call, the electronic device executes the following step 13. When the electronic device determines that it is not necessary to adjust the network of the current video call, the electronic device executes the following step 14.

[0164] Step 13: The electronic device introduces a small amount of network delay or jitter to the current terminal device on the premise of not seriously affecting the current user experience.

[0165] Step 14: The electronic device analyzes the video call through the audio and video acquisition and detection module and prompts the user to pay attention to the video call status.

[0166] Step 15: When a risk call anomaly is detected, the electronic device executes corresponding measures through the warning module.

[0167] Optionally, in the embodiments of the present application, the electronic device can increase network transmission packet loss by reducing the network bandwidth. When the network bandwidth is insufficient, the picture in a video call will be pixelated, blurred, and stuck. The AI face swapping technology usually requires high-quality video and a stable frame rate to maintain the coherence of facial details. Therefore, under network latency, low-quality network may cause video detail loss, making the transmission of real-time video data of AI face swapping possibly interfered, so that the facial movements and changes of the deepfake face model are unnatural, with pixel inconsistencies, etc., resulting in the victim being able to detect abnormal facial expressions or signs of AI manipulation, exposing its forgery with obvious distortion, increasing the possibility of the scammer being identified, and thus improving the security of the call. Similarly, low-quality network may also cause delays, stuck, or audio loss of the voice signal, thus destroying the fluency of the voice. When the network quality is poor, the changes in the voice after AI voice conversion may become more obvious or unnatural, such as abruptly intermittent, blurred and misaligned, or the speaking rhythm is incorrect. Long-term voice delay or interruption can also alert the victim.

[0168] Each of the above method embodiments, or various possible implementation manners in each method embodiment, can be executed alone, or any two or more of them can be combined and executed. It can be specifically determined according to actual usage requirements. The embodiments of the present application do not limit this.

[0169] For the network adjustment method provided by the embodiments of the present application, the execution subject can be a network adjustment device. In the embodiments of the present application, taking the network adjustment device executing the network adjustment method as an example, the network adjustment device provided by the embodiments of the present application is described.

[0170] Figure 10 Shows a possible structural schematic diagram of a network adjustment device involved in some embodiments of the present application. As Figure 10 shown, the network adjustment device 70 may include: an acquisition module 71, a determination module 72, and a display module 73.

[0171] The above acquisition module 71 is used to obtain network quality parameters during a video call if the voice content of the video call contains a preset keyword.

[0172] The above determination module 72 is used to determine whether there is a call risk anomaly in the video call based on at least one of the image data and audio data during the video call when the network quality parameters obtained by the acquisition module 71 meet the network anomaly condition.

[0173] The above display module 73 is used to display a first prompt message when the determination module 72 determines that there is an abnormal call risk in the video call, and the first prompt message is used to prompt that there is an abnormal call risk in the video call.

[0174] In a possible implementation manner, the above determination module 72 is further configured to: after the acquisition module 71 acquires the network quality parameters during the video call, determine a first score based on the network quality parameters acquired by the acquisition module 71, where the first score is used to characterize the call quality of the video call; and, based on the first score, determine whether the network quality parameters meet the network anomaly condition.

[0175] In a possible implementation manner, the above determination module 72 is specifically configured to: when the first score is less than the first threshold, determine that the network quality parameters meet the network anomaly condition; or, when the first score is greater than or equal to the first threshold, adjust at least one of the network delay parameter and the video codec delay parameter of the video call until the first score is less than the first threshold, and determine that the network quality parameters meet the network anomaly condition.

[0176] In a possible implementation manner, the above determination module 72 is specifically configured to obtain facial feature information from the image data; and, when the facial feature information meets the facial anomaly condition, determine that there is an abnormal call risk in the video call.

[0177] In a possible implementation manner, the above determination module 72 is specifically configured to obtain voiceprint feature information from the audio data; and, based on the voiceprint feature information, determine whether there is voiceprint distortion in the voiceprint of the audio data; and, when there is voiceprint distortion in the voiceprint, determine that there is an abnormal call risk in the video call.

[0178] In a possible implementation manner, by way of example, in combination Figure 10 , such as Figure 11 shown, the network adjustment device provided by the embodiments of the present application further includes: a sending module 74. The above determination module 72 is further configured to determine at least one contact associated with the video call after the display module 73 displays the first prompt message. The sending module 74 is used to send a second prompt message to at least one contact determined by the determination module 72, and the second prompt message is used to prompt that there is an abnormal call risk in the video call.

[0179] In an embodiment of the present application, a network adjustment device is provided. Since the network adjustment device can obtain network quality parameters during a video call, the network adjustment device can determine whether the network anomaly condition is met based on the network quality parameters, that is, whether the call quality of the video call is poor. Then, when the network quality parameters meet the network anomaly condition, that is, when the call quality is poor, the network adjustment device can determine whether there is a call risk anomaly in the video call based on at least one of the image data and audio data during the video call. That is, the network adjustment device can magnify the defects of the AI face-swapping video through the poor call quality, expose the anomalies of the AI face-swapping video, and then analyze and judge the video call in this case, thereby improving the speed and accuracy of detecting the AI face-swapping video. In this way, the security of the network adjustment device for video calls is improved.

[0180] The network adjustment device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than the terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiment of the present application does not make specific limitations.

[0181] The network adjustment device in the embodiment of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiment of the present application does not make specific limitations.

[0182] The network adjustment device provided in the embodiment of the present application can implement each process implemented in the above method embodiment. To avoid repetition, it will not be elaborated here.

[0183] Optionally, as Figure 12As shown in the figure, an embodiment of the present application further provides an electronic device 1000, including a processor 1001 and a memory 1002. A program or instruction that can run on the processor 1001 is stored on the memory 1002. When the program or instruction is executed by the processor 1001, each step of the above-mentioned network adjustment method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0184] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0185] Figure 13 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.

[0186] The electronic device 100 includes but is not limited to: a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110, etc.

[0187] Those skilled in the art can understand that the electronic device 100 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 110 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 13 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements, which will not be elaborated here.

[0188] Among them, the above-mentioned processor 110 is used to obtain network quality parameters during a video call if the voice content of the video call contains a preset keyword.

[0189] The above-mentioned processor 110 is used to determine whether there is a call risk anomaly in the video call based on at least one of the image data and audio data during the video call when the network quality parameters meet the network anomaly conditions.

[0190] The above-mentioned display unit 106 is used to display a first prompt message when there is a call risk anomaly in the video call. The first prompt message is used to prompt that there is a call risk anomaly in the video call.

[0191] Optionally, the above-mentioned processor 110 is further used to: after the network quality parameters during the video call, determine a first score based on the network quality parameters. The first score is used to characterize the call quality of the video call; and, based on the first score, determine whether the network quality parameters meet the network anomaly conditions.

[0192] Optionally, the above-mentioned processor 110 is specifically configured to: when the first score is less than the first threshold, determine that the network quality parameter meets the network anomaly condition; or, when the first score is greater than or equal to the first threshold, adjust at least one of the network delay parameter and the video codec delay parameter of the video call until the first score is less than the first threshold, and determine that the network quality parameter meets the network anomaly condition.

[0193] Optionally, the above-mentioned processor 110 is specifically configured to obtain facial feature information from the image data; and, when the facial feature information meets the facial anomaly condition, determine that there is a call risk anomaly in the video call.

[0194] Optionally, the above-mentioned processor 110 is specifically configured to obtain voiceprint feature information from the audio data; and, based on the voiceprint feature information, determine whether there is voiceprint distortion in the voiceprint in the audio data; and, when there is voiceprint distortion in the voiceprint, determine that there is a call risk anomaly in the video call.

[0195] Optionally, the above-mentioned processor 110 is further configured to determine at least one contact associated with the video call after displaying the first prompt message. The above-mentioned radio frequency unit 101 is configured to send a second prompt message to at least one contact, and the second prompt message is used to prompt that there is a call risk anomaly in the video call.

[0196] The embodiment of the present application provides an electronic device. Since the electronic device can obtain the network quality parameter during the video call, the electronic device can determine whether it meets the network anomaly condition according to the network quality parameter, that is, whether the call quality of the video call is poor. Then, when the network quality parameter meets the network anomaly condition, that is, the call quality is poor, the electronic device can determine whether there is a call risk anomaly in the video call based on at least one of the image data and the audio data during the video call. That is, the electronic device can amplify the defects of the AI face-swapping video through the poor call quality, expose the anomalies of the AI face-swapping video, and then can analyze and judge the video call in this case, thereby improving the speed and accuracy of detecting the AI face-swapping video. In this way, the security of the electronic device for video calls is improved.

[0197] The electronic device provided by the embodiment of the present application can implement each process implemented by the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. The beneficial effects of various implementation manners in this embodiment can specifically refer to the beneficial effects of the corresponding implementation manners in the above method embodiment. To avoid repetition, it will not be elaborated here.

[0198] It should be understood that in the embodiments of the present application, the input unit 104 may include a Graphics Processing Unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. The other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0199] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory may be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 109 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0200] The processor 110 may include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor 110 either.

[0201] The embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the network adjustment method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0202] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk, or an optical disc, etc.

[0203] The embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned embodiment of the network adjustment method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0204] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0205] The embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above-mentioned embodiment of the network adjustment method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0206] It should be noted that, in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0207] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0208] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit and scope protected by the claims of the present application, can also make many forms, all of which fall within the protection scope of the present application.

Claims

1. A network adjustment method, characterized in that: include: During a video call, if the voice content of the video call contains a preset keyword, obtaining a network quality parameter during the video call; When the network quality parameter satisfies the network abnormality condition, determining whether there is a call risk abnormality in the video call based on at least one of the image data and the audio data during the video call; In the case that the video call has a call risk abnormality, a first prompt information is displayed, where the first prompt information is used to prompt that the video call has a call risk abnormality.

2. The method according to claim 1, characterized in that After obtaining the network quality parameters during the video call, the method further includes: Determine a first score based on the network quality parameter, where the first score is used to characterize the call quality of the video call; Based on the first score, it is determined whether the network quality parameter satisfies the network abnormality condition.

3. The method according to claim 2, characterized in that The determining, based on the first score, whether the network quality parameter satisfies the network abnormality condition includes: When the first score is less than a first threshold, determining that the network quality parameter meets the network abnormality condition; or, When the first score is greater than or equal to the first threshold, at least one of the network delay parameter and the video codec delay parameter of the video call is adjusted until the first score is less than the first threshold, and it is determined that the network quality parameter meets the network abnormality condition.

4. The method according to any one of claims 1 to 3, characterized in that: The determining whether there is a call risk abnormality in the video call based on at least one of the image data and the audio data during the video call includes: Acquire facial feature information from the image data; When the facial feature information satisfies a facial abnormality condition, it is determined that there is a call risk abnormality in the video call.

5. The method according to any one of claims 1 to 3, characterized in that: The determining whether there is a call risk abnormality in the video call based on at least one of the image data and the audio data during the video call includes: Acquire voiceprint feature information from the audio data; Based on the voiceprint feature information, determining whether the voiceprint in the audio data has voiceprint distortion; When the voiceprint is distorted, it is determined that there is a call risk abnormality in the video call.

6. The method according to claim 1, characterized in that After displaying the first prompt information, the method further includes: determining at least one contact associated with the video call; A second prompt message is sent to the at least one contact, where the second prompt message is used to prompt that there is a call risk abnormality in the video call.

7. A network adjustment device, characterized in that: include: Acquisition module, determination module and display module; The acquisition module is used to acquire the network quality parameters during the video call if the voice content of the video call contains preset keywords; The determining module is configured to determine whether there is a call risk abnormality in the video call based on at least one of the image data and the audio data during the video call when the network quality parameter obtained by the obtaining module meets the network abnormality condition; The display module is used to display first prompt information when the determination module determines that there is a call risk abnormality in the video call, and the first prompt information is used to prompt that there is a call risk abnormality in the video call.

8. The device according to claim 7, characterized in that The determining module is also used for: After the acquisition module acquires the network quality parameter during the video call, determining a first score based on the network quality parameter acquired by the acquisition module, wherein the first score is used to characterize the call quality of the video call; as well as, Based on the first score, it is determined whether the network quality parameter satisfies the network abnormality condition.

9. The device according to claim 8, characterized in that The determination module is specifically used for: When the first score is less than a first threshold, determining that the network quality parameter meets the network abnormality condition; or, When the first score is greater than or equal to the first threshold, at least one of the network delay parameter and the video codec delay parameter of the video call is adjusted until the first score is less than the first threshold, and it is determined that the network quality parameter meets the network abnormality condition.

10. The device according to any one of claims 7 to 9, characterized in that The determining module is specifically used to obtain facial feature information from the image data; and When the facial feature information acquired by the acquisition module meets a facial abnormality condition, it is determined that there is a call risk abnormality in the video call.

11. The device according to any one of claims 7 to 9, characterized in that The determining module is specifically configured to obtain voiceprint feature information from the audio data; as well as, Based on the voiceprint feature information acquired by the acquisition module, determining whether the voiceprint in the audio data has voiceprint distortion; and When the voiceprint is distorted, it is determined that there is a call risk abnormality in the video call.

12. The device according to claim 10, characterized in that The device also includes: a sending module; The determination module is further configured to determine at least one contact associated with the video call after the display module displays the first prompt information; The sending module is used to send a second prompt message to the at least one contact determined by the determining module, wherein the second prompt message is used to prompt that there is a call risk abnormality in the video call.