Video call method and mobile terminal
By employing active and passive detection mechanisms in video calls, mobile terminals can identify and alert users to the risk of face swapping, thus solving the problem of fake faces in video calls and improving user experience and security.
Patent Information
- Application Number
- PCT/CN2025/092213
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-29
- Filing Date
- 2025-04-29
- Publication Date
- 2026-01-02
AI Technical Summary
During video calls, there are instances where AI-powered face-swapping technology is used to forge the other party's face, leading to a loss of user trust, financial losses, and privacy breaches. Existing technologies struggle to effectively detect and alert users to such incidents.
During video calls, mobile terminals can identify face-swapping risks through active and passive detection methods, and trigger different prompt mechanisms under different conditions, including displaying detection controls and screen recording functions, to improve the targeting of detection and user experience.
It effectively identifies and alerts users to the risk of face swapping, reduces the impact on video calls, improves detection efficiency and accuracy, and protects user privacy and property security.
Smart Images

Figure CN2025092213_02012026_PF_FP_ABST
Abstract
Description
Video call method and mobile terminal
[0001] The present application claims priority to the Chinese Patent Application No. 202410844074.7, filed on June 26, 2024, entitled "Video call method and mobile terminal", and the Chinese Patent Application No. 202411380212.7, filed on September 29, 2024, entitled "Video call method and mobile terminal", the contents of which are incorporated herein by reference in their entirety. TECHNICAL FIELD
[0002] Embodiments of the present application relate to the technical field of mobile terminal, and in particular to a video call method and mobile terminal. BACKGROUND
[0003] Mobile terminals such as mobile phones and tablets usually provide a video call function. For example, a social chat application is installed in the mobile terminal, and the mobile terminal can provide a video call function through the social chat application.
[0004] During the video call, the two parties participating in the call can see each other's face, which makes it easier to confirm the identity of the other party and easily gain each other's trust. However, some people take advantage of this and use artificial intelligence (AI) face changing technology to change their face to that of a "familiar person" of the other party, such as a family member, friend, or leader, during the video call to gain the trust of the other party, and then make demands such as transferring money or helping with payment, which can easily result in property loss and a poor user experience. SUMMARY
[0005] Embodiments of the present application provide a video call method and mobile terminal, which can prompt the user in a timely manner in the case of face changing risk during the video call, thereby improving the call experience.
[0006] To achieve the above-mentioned purpose, embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, the present application provides a video call method applied to a mobile terminal, specifically: after establishing a video call, whether the face in the video call interface has a face changing risk can be detected. Wherein, under different conditions, face changing detection can be triggered in different ways. Specifically, under the condition that the first condition (the detection condition in the active detection below) is met, whether the first face in the video call interface has a face changing risk is detected, and after detecting that there is a face changing risk, a first prompt (prompt e below) is displayed. That is, if the first condition is met, active detection is performed to determine whether there is a face changing risk, and a prompt is given when there is a risk. It can be seen that, in the active detection mode, no impact on the video call will be generated except in the case of detecting a face changing risk.
[0008] Under the condition that the second condition (the push condition of prompt g or prompt h below) is met, a second prompt (prompt g or prompt h below) for detecting a face changing risk is displayed, and in response to a first operation (such as a trigger operation on a detection control) on the second prompt, whether there is a face changing risk is detected, and a detection process is prompted, and after obtaining a detection result, a third prompt (prompt j below) is displayed, and the detection result includes a face changing risk or no face changing risk. That is, if the second condition is met, a face changing risk can be prompted first, and then passive detection is triggered to determine whether there is a face changing risk after the user confirms. During passive detection, the detection process will be prompted, such as displaying a detection in progress animation, and a prompt will be given in the case of a risk or no risk. It can be seen that in the passive detection mode, the detection process and the detection result are visible, so that the user can understand the detection process and the detection result after confirming the detection.
[0009] The first condition and the second condition are different, and generally, the first condition has lower requirements and the second condition has higher requirements. In other words, the first condition is easier to meet and the second condition is more difficult to meet. In this way, in most cases, the active detection mode can be used to detect the face changing risk, reducing the impact on the video call, and in a few cases, the passive detection mode can be used to detect the face changing risk, so that the user can understand the detection process and the detection result.
[0010] In summary, by using the scheme of the present application, under different conditions, face changing detection can be triggered and prompted in different ways, so that different detection experiences can be provided, such as no perception in the active detection mode if there is no face changing risk, and full perception in the passive detection mode.
[0011] In a possible design of the first aspect, the first condition or the second condition includes at least one of the following:
[0012] The time difference between the moment when the first user is added as a friend and the start time of the video call is within a first time length (e.g., time length 5, time length 6, or time length 10), that is, the first user is a friend added for a short time. The first user is an account logged in on a device opposite to the mobile terminal in establishing the video call.
[0013] The first user is the first added friend. For example, the mobile terminal adds the first friend after installing a chat application.
[0014] The video call is the first video call with the first user.
[0015] The video call is a video call actively initiated by the first user.
[0016] The first user is in a detection list (e.g., list 1) of the mobile terminal, and the detection list records a first number (e.g., number 1) of user names newly added by the mobile terminal, that is, the first user is one of the first number of friends newly added by the mobile terminal.
[0017] The mobile terminal has a risk event, such as receiving a risk message, receiving an overseas call, downloading an unknown application, and performing an operation of inputting a bank account number, transferring money, or sharing a screen.
[0018] The second time when the mobile terminal is detected to have the risk event is within a second time length from the start time of the video call. That is, if the video call is received shortly after the mobile terminal is detected to have the risk event, face changing detection can be triggered for the video call. The second time length is the same as or different from the first time length.
[0019] The mobile terminal is determined to have a risk based on a plurality of associated risk behavior factors. For example, after receiving an overseas call, some unknown applications are downloaded for screen sharing within a short time. Therefore, face changing detection can be triggered for the video call in this case.
[0020] The third time is within a third time length from the start time of the video call, and the third time is the time when the mobile terminal is determined to have a risk based on a plurality of associated risk behavior factors. That is, if the video call is received shortly after the mobile terminal is determined to have a risk based on a series of associated risk behavior factors, face changing detection can be triggered for the video call. The third time length is the same as or different from the second time length or the first time length.
[0021] That is, the mobile terminal detects or prompts to detect face changing risks for some video calls that may have face changing risks, so as to improve the pertinence of face changing detection and avoid a large number of invalid detections.
[0022] In a possible design manner of the first aspect, the displaying the first prompt includes: in a case where the video call is not ended when the face-swap risk exists, displaying the first prompt on the video call interface, the first prompt including a re-detection control and / or a screen recording control; in a case where the video call is ended when the face-swap risk exists, displaying the first prompt on an interface after the video call is ended, the first prompt not including the re-detection control and the screen recording control; the re-detection control is used to trigger the mobile terminal to re-perform face-swap detection; the screen recording control is used to record a screen of the video call interface after being triggered.
[0023] In this way, the first prompt is displayed whenever the face-swap risk exists, so as to prompt the face-swap risk. Further, since the re-detection and screen recording cannot be performed after the video call is ended, the re-detection control and the screen recording control are not provided in the first prompt in a case where the video call is ended when the face-swap risk exists, so as to avoid providing invalid controls. In addition, the screen recording control is provided to record the screen of the video call interface, so as to save a screen recording of the video call with the face-swap risk, to serve as evidence when necessary, or to update a detection model for face-swap detection.
[0024] In a possible design manner of the first aspect, the detection result is that the face-swap risk does not exist. The third prompt is displayed after the detection result is obtained, including: in a case where the video call is not ended when the detection result is obtained, displaying the third prompt on the video call interface, the third prompt including a re-detection control and / or a screen recording control. In a case where the video call is ended when the detection result is obtained, displaying the third prompt on an interface after the video call is ended, the third prompt not including the re-detection control and the screen recording control. For details, refer to the foregoing description of the first prompt.
[0025] In a possible design manner of the first aspect, the prompt including the re-detection control (such as the first prompt or the third prompt described above) is displayed on the video call interface, including: in a case where a number of times of detecting whether the face-swap risk exists for the video call does not exceed a first number of times (such as a number of times 1 described below), the prompt including the re-detection control is displayed on the video call interface. In this way, malicious detection for the same video call can be avoided.
[0026] In a possible design manner of the first aspect, in a case where the number of times of detecting whether the face-swap risk exists for the video call exceeds the first number of times, a prompt not including the re-detection control is displayed on the video call interface.
[0027] In a possible design manner of the first aspect, the mobile terminal can display a prompt of undetected detection in a case where the second condition is met and the second face is not detected; or display the prompt of undetected detection in an interface of ending the video call in a case where the second condition is met and the video call is ended before the video frame of the video call is acquired. That is, in the passive detection manner, for an abnormal case that causes undetected detection, such as undetected face, undetected image frame, and the like, the prompt of undetected detection can be displayed, so that the user can be explicitly aware of the abnormality.
[0028] Conversely, in a case where the first condition is met and the first face is not detected, or in a case where the video call is ended before the video frame of the video call is acquired, the prompt of undetected detection is not displayed. That is, in the active detection manner, for an abnormal case that causes undetected detection, such as undetected face, undetected image frame, and the like, the prompt is not displayed, so as to avoid interference to the user.
[0029] In a possible design manner of the first aspect, the prompt detection process includes: displaying at least one of the following information: a prompt of starting detection (such as a prompt of "starting detection"), a prompt of detecting (such as a prompt of "detecting"), and a detection animation effect.
[0030] In a possible design manner of the first aspect, displaying the second prompt of detecting the face swap risk includes: in a case where the second prompt is displayed for the first to Nth times, displaying the second prompt in a card form, and the second prompt in the card form includes a function introduction of the face swap detection function and / or a detection confirmation control (such as a detection control, an opening and detection control, and the like in the following); and in a case where the second prompt is displayed for the N+1th time, displaying the second prompt in a capsule form, and the second prompt in the capsule form does not include the function introduction of the face swap detection function and the detection confirmation control. N is a natural number greater than 1, such as N=3.
[0031] In this way, the details of the prompt can be presented to the user in the first few times of displaying the second prompt, so that the user can be explicitly aware of the content of the prompt, and the capsule can be presented to the user in the subsequent display of the second prompt, so as to reduce the impact on the call.
[0032] In a possible design of the first aspect, the second condition includes a first sub-condition and a second sub-condition. The second prompt for detecting the face swapping risk is displayed in a case where the second condition is met, including: the second prompt for detecting the face swapping risk (as prompt g below) is displayed in a case where the first sub-condition (as a push condition of prompt g below) is met, and the first sub-condition includes that the face swapping detection switch of the mobile terminal is turned on. The second prompt for turning on the face swapping detection function and detection (as prompt h below) is displayed in a case where the second sub-condition (as a push condition of prompt h below) is met, and the second sub-condition includes that the face swapping detection switch of the mobile terminal is turned off.
[0033] In addition, the first sub-condition has a lower requirement, and the second sub-condition has a higher requirement. In other words, the first sub-condition is easier to meet, and the second sub-condition is more difficult to meet. In practice, the mobile terminal is by default turned on the risk protection, and the face swapping detection function is by default turned on in response. If the face swapping detection switch is turned off, that is, the face swapping detection function is not turned on, it is indicated that the face swapping detection function is most likely manually turned off by the user. In this case, the second sub-condition can be used to prompt the face swapping detection function to be turned on at a lower frequency, so as to avoid disturbing the user by frequently prompting the user about the function that is not needed by the user. After the face swapping detection function is turned on, the first sub-condition can be used to prompt the face swapping detection function to be used to detect the face swapping risk at a slightly higher frequency, so as to improve the use rate of the face swapping detection function.
[0034] In a possible design of the first aspect, the mobile terminal can intercept a video frame in the video call interface in a case where the first condition is met; perform face region detection on the intercepted video frame, identify a face region that meets the opposite-end face constraint condition from the detected face regions as a first face; and perform face swapping risk detection according to the first face in the video frame.
[0035] In the above scheme, the video frame stream of the video call is intercepted, and the opposite-end face region that needs to be detected for the face swapping risk can be accurately identified based on the face region detection, thereby improving the accuracy of subsequent face swapping risk detection.
[0036] In a possible design of the first aspect, the opposite-end face constraint condition includes a preset position constraint condition and / or a preset size constraint condition.
[0037] In the above scheme, the mobile terminal can accurately identify the opposite-end face region according to the preset position constraint condition and / or the preset size constraint condition.
[0038] In a possible design of the first aspect, before identifying the face region that meets the opposite-end face constraint condition in the detected face region as the first face, the mobile terminal can acquire a window switching number when the video frame is intercepted; the window switching number refers to a number of times of switching display windows of the two parties in the video call interface when the video frame is intercepted; and the opposite-end face constraint condition corresponding to the video frame is determined based on the window switching number.
[0039] In the above scheme, the window switching in the video call is considered, and the opposite-end face constraint condition for identifying the opposite-end face region is more accurately determined based on the statistical window switching number, so that the opposite-end face region can be accurately identified.
[0040] In a possible design of the first aspect, the intercepted video frame is a plurality of video frames; the mobile terminal can perform face swapping identification based on the first face in each of the intercepted video frames, to obtain a face forgery confidence corresponding to each video frame; the face forgery confidences corresponding to the video frames are fused to obtain a target confidence; and it is determined whether there is a face swapping risk according to the target confidence.
[0041] In the above scheme, the mobile terminal can more accurately identify the face swapping risk identification result of the current video call based on the fusion result of the face forgery confidences of the plurality of video frames.
[0042] In a possible design of the first aspect, the plurality of video frames are intercepted from the video call according to at least one of a preset interception time, an interception frequency, or an interception total time length.
[0043] In the above scheme, in the video call, the video frames are intercepted based on the preset interception parameter, so that the video frames required for face swapping detection can be obtained, and the accuracy of subsequent face swapping risk detection can be improved.
[0044] In a possible design of the first aspect, the mobile terminal can screen face forgery confidences corresponding to part of the video frames from the face forgery confidences corresponding to the video frames; the screened face forgery confidences are higher than the face forgery confidences that are not screened; and the screened face forgery confidences are weighted and fused to obtain a target confidence.
[0045] In the above scheme, the face swapping risk identification result can be obtained without detecting all frames, system resources are saved, and the detection efficiency is improved.
[0046] In a possible design manner of the first aspect, after the mobile terminal intercepts each video frame and calculates the face forgery confidence of the video frame, the mobile terminal updates and stores the face forgery confidence corresponding to the video frame into an array, and fuses the face forgery confidence currently stored in the array by weighting to obtain a target confidence. Further, the mobile terminal can identify a risk level of the face swapping risk according to the target confidence; if the risk level is a medium risk, the mobile terminal continues to intercept a next video frame; if the risk level is a low risk or a high risk, the mobile terminal stops intercepting the video frame, and in the case of the high risk, it is determined that the face swapping risk exists.
[0047] In the foregoing solution, in combination with the array, the comprehensive target confidence corresponding to the currently intercepted video frame can be calculated after each video frame is intercepted, to identify the risk level of the current video call. Only in the case of the medium risk, the video frame needs to be continuously intercepted. If it is the low risk or the high risk, the video frame does not need to be continuously intercepted, thereby saving system resources and enabling accurate identification of the face swapping risk.
[0048] In the second aspect, the present application further provides a mobile terminal. The mobile terminal includes a display screen, a memory, and one or more processors. The display screen, the memory, and the processor are coupled. The memory is configured to store computer program code. The computer program code includes computer instructions. When the computer instructions are executed by the processor, the mobile terminal performs the method in the first aspect and any possible design manner thereof.
[0049] In the third aspect, the present application provides a chip system. The chip system is applied to a mobile terminal including a display screen and a memory. The chip system includes one or more interface circuits and one or more processors. The interface circuit and the processor are interconnected through a line. The interface circuit is configured to receive a signal from the memory of the mobile terminal and send a signal to the processor. The signal includes computer instructions stored in the memory. When the processor executes the computer instructions, the mobile terminal performs the method in the first aspect and any possible design manner thereof.
[0050] In the fourth aspect, the present application provides a computer readable storage medium. The computer readable storage medium includes computer instructions. When the computer instructions are run on a mobile terminal, the mobile terminal performs the method in the first aspect and any possible design manner thereof.
[0051] In the fifth aspect, the present application provides a computer program product. When the computer program product is run on a computer, the computer performs the method in the first aspect and any possible design manner thereof.
[0052] It can be understood that the mobile terminal of the second aspect, the chip system of the third aspect, the computer readable storage medium of the fourth aspect, and the computer program product of the fifth aspect can achieve the beneficial effects as described in the first aspect and any possible design of the first aspect, and details are not repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0053] FIG. 1A is a schematic diagram of a scenario according to an embodiment of the present application;
[0054] FIG. 1B is a hardware structure diagram of a mobile phone according to an embodiment of the present application;
[0055] FIG. 1C is a software structure diagram of a mobile phone according to an embodiment of the present application and a flowchart of a video call method based on the software structure diagram;
[0056] FIG. 1D is a schematic diagram of an interface for starting face changing detection according to an embodiment of the present application;
[0057] FIG. 2 is a schematic diagram of another interface for starting face changing detection according to an embodiment of the present application;
[0058] FIG. 3 is a schematic diagram of an interface for starting a system manager service according to an embodiment of the present application;
[0059] FIG. 4 is a schematic diagram of an interface for understanding face changing detection details according to an embodiment of the present application;
[0060] FIG. 5 is a schematic diagram of an interface for viewing face changing detection records according to an embodiment of the present application;
[0061] FIG. 6 is a schematic diagram of an interface for a security setting item according to an embodiment of the present application;
[0062] FIG. 7 is a schematic diagram of a third interface for starting face changing detection according to an embodiment of the present application;
[0063] FIG. 8 is a schematic diagram of a prompt a and an interface response thereof according to an embodiment of the present application;
[0064] FIG. 9 is a schematic diagram of a disappearance process of the prompt a according to an embodiment of the present application;
[0065] FIG. 10 is a schematic diagram of a prompt d and an interface response thereof according to an embodiment of the present application;
[0066] FIG. 11 is a schematic diagram of a principle of a proactive detection scheme according to an embodiment of the present application;
[0067] FIG. 12 is a schematic diagram of an interface of the proactive detection scheme according to an embodiment of the present application;
[0068] FIG. 13 is a schematic diagram of an interface for adding a friend and starting a video call according to an embodiment of the present application;
[0069] FIG. 14 is an interface schematic diagram for prompting to use face swapping detection in a video call according to an embodiment of the present application;
[0070] FIG. 15A is an interface schematic diagram for prompting to start face swapping detection in a video call according to an embodiment of the present application;
[0071] FIG. 15B is another interface schematic diagram for prompting to start face swapping detection in a video call according to an embodiment of the present application;
[0072] FIG. 16 is a principle diagram of a passive detection scheme according to an embodiment of the present application;
[0073] FIG. 17 is an interface schematic diagram of a passive detection scheme according to an embodiment of the present application;
[0074] FIG. 18 is another interface schematic diagram of a passive detection scheme according to an embodiment of the present application;
[0075] FIG. 19 is a third interface schematic diagram of a passive detection scheme according to an embodiment of the present application;
[0076] FIG. 20 is a fourth interface schematic diagram of a passive detection scheme according to an embodiment of the present application;
[0077] FIG. 21 is a fifth interface schematic diagram of a passive detection scheme according to an embodiment of the present application;
[0078] FIG. 22 is a sixth interface schematic diagram of a passive detection scheme according to an embodiment of the present application;
[0079] FIG. 23 is an interface schematic diagram for a case where there is a face swapping risk according to an embodiment of the present application;
[0080] FIG. 24 is another interface schematic diagram for a case where there is a face swapping risk according to an embodiment of the present application;
[0081] FIG. 25A is a third interface schematic diagram for a case where there is a face swapping risk according to an embodiment of the present application;
[0082] FIG. 25B is a fourth interface schematic diagram for a case where there is a face swapping risk according to an embodiment of the present application;
[0083] FIG. 26 is a first interface schematic diagram for closing face swapping detection according to an embodiment of the present application;
[0084] FIG. 27 is a second interface schematic diagram for closing face swapping detection according to an embodiment of the present application;
[0085] FIG. 28 is a flowchart of a video call method according to an embodiment of the present application;
[0086] FIG. 29 is a schematic diagram of a video call interface according to an embodiment of the present application;
[0087] FIG. 30 is a schematic diagram of interface changes triggered by adding a friend to face changing risk detection according to an embodiment of the present application;
[0088] FIG. 31 is a schematic diagram of interface changes triggered by adding a friend to passive detection according to an embodiment of the present application;
[0089] FIG. 32 is a schematic diagram of a video call method based on a software structure according to an embodiment of the present application;
[0090] FIG. 33 is a timing diagram of a video call method according to an embodiment of the present application;
[0091] FIG. 34 is a timing diagram of face changing detection and disposal steps according to an embodiment of the present application;
[0092] FIG. 35 is a timing diagram of a video call method based on passive detection according to an embodiment of the present application;
[0093] FIG. 36 is a timing diagram of another video call method based on passive detection according to an embodiment of the present application. DETAILED DESCRIPTION
[0094] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing the specific embodiments and are not intended to be limiting on the present application. As used in the specification and the appended claims of the present application, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that "at least one" and "one or more" refer to one or two or more (including two) in the following embodiments of the present application. The term "and / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships; for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. In the description of the embodiments, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0095] Reference in the specification to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places in the specification are not necessarily all referring to the same embodiment, although it can. The terms "including," "comprising," "having" and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. The terms "connected," "coupled," and "pathway" are used broadly and encompass both direct and indirect connections, couplings and pathways, and are meant to include items connected indirectly together via another interconnected item. The terms "first," "second," and "third" and the like in the description and in the claims, are used for describing various elements, but do not mean to imply relative importance or a required ordering. Thus, a feature specified as "first" can in some embodiments be termed as "second" or "third." The use of the term "about" in conjunction with a quantity is intended to include the value of the quantity + / - 10% of the stated value.
[0096] The use of the terms "example," "such as," or "for instance" in the description is merely meant to illustrate an example, an instance, or a configuration and does not mean or imply that the example, instance or configuration is preferred or superior to other examples, instances or configurations. Thus, the use of the terms "example," "such as," or "for instance" is not meant to limit or confine the scope of the application in any way.
[0097] In the scenario of a video call, as shown in FIG. 1A, after adding Lisa (a user device user) as a friend, the user X of the PC end operates the face changing software installed on the PC end. The PC end uses the face spoofing AI algorithm based on the face changing software to perform face spoofing on the video stream collected by the camera, and uses the voice changer AI algorithm to perform voice spoofing on the voice stream collected by the microphone, and performs AI deep synthesis on the spoofed face and voice, thereby performing a video call as Lisa's friend David. For example, in FIG. 1A, the mouth, nose, eyebrows, eyes and other features of the face displayed in the video call interface of the PC end are different from those of X, and are close to the face of Lisa's friend David, which is achieved by the face spoofing AI algorithm. After the PC end spoofs the face to initiate a video call to Lisa, it will send the spoofed video stream to the user device based on the video coding standards such as H.265 or H.264 and the transmission protocols such as RTP or RTCP. The video call interface displayed by the user device will include the spoofed face. The user device user thinks that it is David who initiates the video call to him / her. In this way, the user device user (Lisa) is likely to suffer from adverse effects such as property loss or personal privacy leakage due to false cognition during the video call with the spoofed friend.
[0098] To avoid the above-mentioned adverse effects, the present application proposes a video call method applied to a mobile terminal, so that when the mobile terminal is used by a user (such as Lisa in FIG. 1A) for video call, the mobile terminal can detect the face-swapping risk in real time during the video call process and give a prompt. For example, as shown in FIG. 1A, the face-swapping risk detection mechanism can include extracting video stream images (i.e. video frames), detecting whether the opposite face in the video frame exists face-swapping according to the video frame, such as detecting face-swapping boundaries, eye contact, eyebrows, lips and other synthetic flaws in the video frame, and / or performing inter-frame coherence, screen flicker, jitter detection to identify face-swapping risks. It should be noted that this is only a simple illustrative description of face-swapping risk detection, and the specific means of detecting face-swapping risk are not limited in the present application. For example, the mobile terminal can detect whether there is a face-swapping risk based on an AI model. For another example, the mobile terminal can compare the face with the user's white list face library to detect whether there is a face-swapping risk. For another example, the mobile terminal can detect whether there is a face-swapping risk by analyzing the facial expressions and body movements of the face.
[0099] For example, the mobile terminal is installed with an application (application, APP) supporting video call, which can be simply referred to as "video call application". The mobile terminal can realize video call through these APPs. Among them, the scene of video call further includes the scene of actively initiating and entering video call and the scene of passively receiving and joining video call. Using the video call method provided by the embodiments of the present application, the mobile terminal can detect the face-swapping risk in real time and give a prompt during the video call process using the video call application.
[0100] In some embodiments, the method of the present application can be performed for a specified video call application. Specifically, the mobile terminal can obtain a pre-configured application white list (which can be simply referred to as "white list"). For the video call application in the white list, the method in the embodiments of the present application can be performed, that is, when the user uses the video call application in the white list to perform video call, the method in the embodiments of the present application can be performed to detect the face-swapping risk and give a prompt. Thus, the face-swapping risk control of video call is targeted and more flexible. In addition, without complex operation or version update, the face-swapping risk control of video call application can be more convenient by modifying the white list.
[0101] It should be understood that the traditional forgery detection process usually needs to rely on the cloud, and the processing timeliness is poor, the user privacy is not good, and there are also security problems. Or, it needs to rely on a device with high hardware configuration, which is high in cost and not suitable for the video call scene of mobile terminal, and cannot detect the risk in time when the user uses the mobile terminal for video call, so the security protection role is also limited.
[0102] Using the video call method provided by the embodiments of the present application, the mobile terminal can detect the face changing risk in time and in real time during the video call process and give a prompt, improving the timeliness of risk detection. Moreover, the face changing detection (face changing risk detection) can be realized on the mobile terminal, without the need for additional large memory and other system resources. In addition, the face changing detection is performed by the mobile terminal itself for video call, which can process the user's private data locally on the mobile terminal, improving the security and protecting the user's privacy.
[0103] Specifically, the video call method provided in the embodiments of the present application mainly performs face changing detection on the face image of the call opposite terminal in the video call process of the mobile terminal, to identify in real time whether the face image of the call opposite terminal is a fake synthesis, that is, whether the face image of the call opposite terminal has a face changing risk. Further, if there is a face changing risk, a risk prompt can be given, and if there is no face changing risk, no prompt or a no-risk prompt can be given.
[0104] Exemplarily, the mobile terminal can be a mobile terminal such as a mobile phone, a tablet, a PC, etc. that can provide a video call function, such as a mobile terminal installed with an APP supporting video call. The embodiments of the present application do not specially limit the specific form of the mobile terminal. Hereinafter, the mobile terminal is taken as a mobile phone as an example to illustrate the solution of the present application.
[0105] Referring to FIG. 1B, the mobile phone can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, a sensor 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0106] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the mobile phone. In other embodiments of the present application, the mobile phone can include more or fewer components than the illustration, or combine certain components, or split certain components, or different arrangement of components. The illustrated components can be implemented in hardware, software or a combination of software and hardware.
[0107] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors.
[0108] In some embodiments, the mobile phone can complete the video call method through the processor 110, so as to detect the face changing risk and give a prompt in the process of video call.
[0109] The wireless communication function of the mobile phone can be realized through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor, etc.
[0110] The mobile phone realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected with the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs, which execute program instructions to generate or change display information.
[0111] The display screen 194 is used to display images, videos, etc. In some embodiments, the mobile phone can display the interface of the video call, the prompt information of the face changing risk, etc. through the display screen 194.
[0112] The mobile phone can realize the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194 and the application processor, etc.
[0113] The mobile phone can realize the audio function through the audio module 170, the loudspeaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, music playing, recording, etc.
[0114] The key 190 can include a power-on key, a volume key, and the like. The key 190 can be a mechanical key. It can also be a touch key. The mobile phone can receive key input and generate key signal input related to user settings and function control of the mobile phone. The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompts and also for touch vibration feedback. The indicator 192 can be an indicator light and can be used to indicate a charging state, a power change, and also to indicate a message, a missed call, a notification, and the like. The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the mobile phone.
[0115] The software system of the mobile phone can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. Embodiments of the present application exemplarily illustrate the software structure of the mobile phone by taking a layered architecture as an example. It should be noted that embodiments of the present application are not limited to the software structure of the mobile phone.
[0116] FIG. 1C is a block diagram of a software structure of a mobile phone according to an embodiment of the present application.
[0117] It can be understood that the layered architecture can divide the software into several layers, each layer having a clear role and division of labor. The layers communicate with each other through software interfaces. As shown in FIG. 1C, the software system can include an application layer, an application framework layer (Framework layer), and a native service layer (Native layer).
[0118] The application layer can include a series of application packages. As shown in FIG. 1C, the application packages can include a third-party video call application, a perception module, and a detection module. The perception module is used for Activity behavior perception, for example, the perception module can perceive a video call event, and also can perceive a preset risk event, and the like. The detection module includes a face swapping detection module and a face swapping risk identification and handling module.
[0119] The face swapping detection module is used to identify whether the face displayed by the opposite end in the video call is forged by AI face swapping technology. The face swapping risk identification and handling module is used to confirm whether there is a risk based on the identification result of the face swapping detection module and to perform corresponding handling.
[0120] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the application programs of the application layer. The application framework layer includes some pre-defined functions.
[0121] The application framework layer can include a window manager (WindowManager). The window manager is used to manage window programs. The window manager can acquire a display screen size, determine whether there is a status bar, lock a screen, take a screenshot, and record a screen (for example, record a screen of a video call interface), and the like.
[0122] The application framework layer can also include other modules not shown in FIG. 1C, such as a view system and a notification manager, and the like.
[0123] The view system includes visual controls, such as a control for displaying text, a control for displaying a picture, and the like. The view system can be used to build an application program. A display interface can be composed of one or more views. For example, a display interface including a notification icon can include a view for displaying text and a view for displaying a picture.
[0124] The notification manager enables an application program to display notification information in a status bar, and can be used to convey a message of a notification type that can automatically disappear after a short stay without user interaction. For example, in an embodiment of the present application, a notification manager can be used to output prompt information to a user, such as a prompt that a video call counterpart has a face changing risk or a prompt to turn on a detection switch.
[0125] The local service layer is located between the application program framework layer and the kernel layer. The local service layer can include a graphics compositor (Surface Flinger). The graphics compositor is used to manage, compose, and render a graphical interface of a system, and ensures that an interface of an application program and an interface of the system can be correctly and efficiently displayed on a screen. For example, in an embodiment of the present application, a graphics compositor is used to provide a video frame in a video call process.
[0126] It should be noted that the software structure is not limited to the above hierarchical structure, and can also include a kernel layer and a driver layer (such as a display driver), and the like, which will not be described again.
[0127] For ease of understanding, the video call method in the present application will now be described in conjunction with the software structure of FIG. 1C, as follows:
[0128] 1. A perception module perceives a video incoming call event in a video call application.
[0129] 2. The perception module notifies a detection module of the video incoming call event perception result.
[0130] 3. In a video call process, a face changing detection module in the detection module captures video frame layer data from a graphics compositor through a screenshot interface.
[0131] 4. The graphics compositor returns the captured video frame to the face changing detection module.
[0132] 5. The face swapping detection module performs face spoofing detection inference on the video frame, and sends the confidence result of the inference to the face swapping risk identification and handling module.
[0133] The face swapping risk identification and handling module can identify face swapping risks and perform corresponding handling based on the confidence result of the face spoofing detection inference. It should be understood that the "handling" in the embodiments of the present application refers to processing the face swapping risk identification result.
[0134] The video call method provided by the embodiments of the present application can be implemented in a mobile phone with the above-mentioned hardware and software structure. The video call method provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0135] In some embodiments, the mobile phone can start the face swapping detection function after detecting the starting event of the face swapping detection. In the case of starting the face swapping detection function, the mobile phone can detect the face swapping risk and give a prompt during the video call. Wherein, the detection of the starting event indicates that the user allows the opening of the permissions and data required for face swapping detection. In this way, the mobile phone can execute the corresponding permissions and obtain data under the permission of the user.
[0136] Wherein, the starting event can be an operation event of skipping the risk protection setting in the out-of-box experience (OOBE) (hereinafter referred to as method one), an operation event of starting the risk protection switch in the settings application (hereinafter referred to as method two), an operation event of starting the face swapping detection switch in the settings application (hereinafter referred to as method three), an operation event of executing the starting detection on the prompt a of starting the risk protection (hereinafter referred to as method four), etc. The embodiments of the present application do not make specific limitations on this.
[0137] The following will specifically describe the above-mentioned several starting events and the process of triggering the starting of the face swapping detection function.
[0138] Method one, the starting event is an operation event of skipping the risk protection configuration in the OOBE.
[0139] OOBE refers to the process of configuring the mobile phone after the mobile phone is manufactured and started for the first time, such as configuring language, region, network, etc. OOBE can also be interpreted as boot guide, boot guide, etc.
[0140] In the OOBE, the mobile phone usually starts the risk protection by default. Various risk detection functions are provided in the risk protection, such as risk phone detection function, risk information detection function, face swapping detection function, etc. After starting the risk protection, it is equivalent to starting the various risk detection functions under the risk protection.
[0141] In the OOBE, the mobile phone detects an operation event of the user skipping the risk protection setting, such as a triggering operation on the confirmation skip control, indicating that the user supports starting the risk protection, and the mobile phone can start the risk protection, thereby starting the face changing detection function under the risk protection.
[0142] Referring to FIG. 1D, the configuration interface of the risk protection in the OOBE is the interface 101 shown in FIG. 1D, and the switch 1011 of the risk protection in the interface 101 is in the default start state. In this case, if the user does not turn off the switch 1011, but directly performs a triggering operation on the confirmation skip control "next step" 1012 in the interface 101, the mobile phone can skip the risk protection setting and start the risk protection by default in response to the triggering operation on the "next step" 1012.
[0143] In a second mode, the start event is an operation event of starting the risk protection switch in the setting application.
[0144] In the setting application of the mobile phone, there is a setting item 1 of the risk protection. In a specific implementation mode, the mobile phone further provides the setting item 1 in the security and / or privacy setting item of the setting application.
[0145] Referring to FIG. 2, the mobile phone can display an interface 201, which is an application interface of the setting application. "Security" 2011 of the interface 201 is a security setting item. In response to a triggering operation of the user on the "security" 2011, the mobile phone can display an interface 202. The interface 202 is a setting interface of the security setting item. "Risk protection" 2021 in the interface 202 is the setting item 1 of the risk protection. That is, the setting item 1 is provided in the security setting item,
[0146] In the setting interface of the security setting item shown in the interface 202, the "risk protection" 2021 is located in the area close to the top of the interface, and is larger than other setting items such as "SOS emergency help" 2022 and "emergency warning notification" 2023, so that the setting item 1 can be located in a more eye-catching position.
[0147] The mobile phone detects an operation event of starting the risk protection switch in the setting item 1, such as a triggering operation on the risk protection switch in the closed state, and can start the risk protection, thereby starting the face changing detection function under the risk protection.
[0148] Continuing to refer to FIG. 2, in response to a triggering operation of the user on the "risk protection" 2021, the mobile phone can display an interface 203, which is a setting interface of the risk protection. The "risk detection" switch 2031 in the interface 203 is a risk protection switch, and the "risk detection" switch 2031 in the interface 203 is in the closed state, indicating that the risk protection is not started.
[0149] In response to the triggering operation of the "risk detection" switch 2031 in the interface 203 by the user, the phone can start the risk protection. For example, after starting the risk protection, the phone can display an interface 204, which is also a setting interface of the risk protection. The interface 204 also includes the "risk detection" switch 2031, and the "risk detection" switch 2031 in the interface 204 is in the starting state, indicating that the risk protection has been started.
[0150] In a specific implementation, the risk protection is provided by a system manager service in the phone.
[0151] If the phone does not start the system manager service, and an operation event of starting the risk protection switch in the setting item 1 is detected, the phone can first prompt the user to start the system manager service (denoted as prompt b), and in response to the operation of starting the system manager service by the user in the prompt b, such as the triggering operation of the confirmation start control in the prompt b, the phone can start the system manager service and further start the risk protection.
[0152] Continuing to refer to FIG. 2, in the case where the system manager service is not started, in response to the triggering operation of the "risk detection" switch 2031 in the interface 203 by the user, the phone can display an interface 205, which adds a prompt 2051 in the setting interface of the risk protection. The prompt 2051 is the prompt a. For example, the prompt 2051 includes the text "When starting the risk detection, the "system manager service" needs to be started synchronously", prompting to start the system manager service. The "agree" 2052 in the prompt 2051 is a control for confirming to start the system manager service, and in response to the triggering operation of the "agree" 2052 by the user, the phone can start the system manager service and the risk protection, such as displaying the interface 204.
[0153] In response to the operation of canceling to start the system manager service by the user in the prompt b, such as the triggering operation of the confirmation cancel control (such as the "cancel" 2053 in the interface 205) in the prompt b, the phone can cancel to start the system manager service, and accordingly, the risk protection cannot be started.
[0154] If the phone has started the system manager service, and an operation event of starting the risk protection switch in the setting item 1 is detected, the risk protection can be started, such as the response process from the interface 203 to the interface 204 described above.
[0155] Generally, the system manager service is started by default in the OOBE. Alternatively, in the case where the system manager service is not started, in response to starting the system manager service, the phone can display the prompt b for starting the system manager service, and in response to the starting operation of the prompt b by the user, the phone can start the system manager service.
[0156] Referring to FIG. 3, in the case where the system manager service is not started, in response to starting the system manager application, the mobile phone can display an interface 301, and a prompt 3011 in the interface 301 is prompt b. The prompt 3011 prompts various capabilities provided by the system manager service, such as "the system manager service can provide notification management, battery management, privacy security, application startup management, permission management, cleaning and acceleration, traffic management, spam interception, virus detection, and risk protection functions". In addition, the prompt 3011 includes a control "agree" 3012 for starting the system manager service, and in response to a user triggering operation on the "agree" 3012, the mobile phone can start the system manager service.
[0157] In addition, the setting interface of the risk protection can further include setting items of a plurality of risk detection functions, and the setting items include switch states of the corresponding risk detection functions.
[0158] For example, the interface 203 includes a setting item "risk phone interception" 2032 of risk phone detection, a setting item "risk information interception" 2033 of risk information detection, and a setting item "AI face changing reminder" 2034 of face changing detection, and the switch states are all not started.
[0159] For example, in the interface 204, the switch states of the "risk phone interception" 2032, the "risk information interception" 2033, and the "AI face changing reminder" 2034 are all "started".
[0160] Further, in response to a user triggering operation on any setting item of a risk detection function in the setting interface of the risk protection, the mobile phone can display a setting interface of the risk detection function.
[0161] For example, in response to a user triggering operation on the "AI face changing reminder" 2034 in the interface 204 shown in FIG. 2 after starting the risk protection, the mobile phone can display an interface 401 shown in FIG. 4, and the interface 401 is a setting interface of face changing detection. In the interface 401, "AI face changing detection" 4011 is a face changing detection switch. Since the risk protection is started at this time, the switch state of the "AI face changing detection" 4011 is started.
[0162] In the setting interface of the face changing detection, a function introduction of the face changing detection can be further included, such as the text introduction "in a video call, the system can detect whether the other party is an AI face changing in real time to prevent face changing risks" in the interface 401, so that the user can clearly understand the function of the face changing detection.
[0163] In the setting interface of face swapping detection, a details entry, such as “Learn More” 4012 in interface 401, can also be included. In response to a triggering operation of the details entry by the user, the phone can display more detailed information of face swapping detection, such as the scenario of face swapping and the formula of alerting face swapping risk, so that the user can learn more about face swapping detection. Continuing to refer to FIG. 4, in response to a triggering operation of “Learn More” 4012 in interface 401 by the user, the phone can display interface 402, which includes the scenario of face swapping 4021, the formula of alerting face swapping risk 4022, and the like.
[0164] In addition, in the setting interface of risk protection, the protection records of multiple risk detection functions in a recent preset time period (such as 60 days or 1 month), such as the number of risks, can also be included. Referring to FIG. 5, the phone can display interface 501 (the same as interface 204 described above), and the protection records of face swapping detection, such as the number of risks, are recorded in region 5011 of interface 501.
[0165] In response to a triggering operation of the protection records of any risk detection function by the user, the phone can display the detailed detection records of the risk detection function. Referring to FIG. 5, in response to a triggering operation of region 5011 in interface 501 by the user, the phone can display interface 502, which is the record interface of face swapping detection and records the detailed detection records of face swapping detection. However, since the number of risks is 0 at this time, there are no detailed detection records in interface 502.
[0166] It should be noted that the names and layouts of the setting items in the setting application shown in FIGS. 2-5 are only exemplary. In actual implementation, they are not limited thereto. For example, the setting interface of the security setting item can also be as shown in interface 601 in FIG. 6, and “Risk Protection” 6011 in interface 601 is the first setting item of risk protection. Interface 601 is quite different from the layout of the setting items in interface 202 described above, such as the display position and size of “Risk Protection” 6011 in interface 601, which are different from “Risk Protection” 2021 in interface 202.
[0167] Method three: the opening event is an operation event of turning on the face swapping detection switch in the setting application.
[0168] The second setting item of face swapping detection is provided in the setting of the phone. The phone detects an operation event of turning on the face swapping detection switch in the second setting item, such as detecting a triggering operation of the face swapping detection switch in the off state, and the phone can turn on the face swapping detection.
[0169] Referring to FIG. 7, in the case where the risk protection is not started, the mobile phone can display interface 701 (the same as interface 203 described above), and the "AI face changing reminder" 7011 in interface 701 is setting item 2. In response to a triggering operation of the user on the "AI face changing reminder" 7011, the mobile phone can display the setting interface of the face changing detection shown in interface 702, and the "AI face changing detection" switch 7021 in interface 702 is the face changing detection switch. Since the risk protection is not started at this time, the switch state of the "AI face changing detection" switch 7021 in interface 702 is the off state. In response to a triggering operation of the user on the "AI face changing detection" switch 7021, the mobile phone can start the face changing detection, such as displaying interface 703 (the same as interface 401 described above), and the "AI face changing detection" switch 7021 in interface 703 is in the on state, indicating that the face changing detection has been started.
[0170] Based on the foregoing description, it can be known that the risk protection provides the face changing detection function. Based on this, if the mobile phone does not start the risk protection, after detecting the triggering operation of the user on setting item 2, the mobile phone can first prompt the user to start the risk protection (denoted as prompt c), and in response to the operation of starting the risk protection on prompt c, the mobile phone can start the risk protection and the face changing detection.
[0171] Continuing to refer to FIG. 7, in response to a triggering operation of the user on the "AI face changing reminder" 7011 in interface 701, the mobile phone can first display interface 704. The prompt 7041 in interface 704 is prompt c, which includes the prompt text "When starting the AI face changing detection, the "risk detection" switch will be started simultaneously", which is used to prompt the user to start the risk protection. The prompt 7041 also includes the control "start" 7042 for starting the risk protection. The operation of starting the risk protection can be a triggering operation of the user on the "start" 7042, and in response to the triggering operation of the user on the "start" 7042, the mobile phone can start the face changing detection, such as displaying interface 703.
[0172] If the mobile phone has started the risk protection, after detecting the operation event of starting the face changing detection switch in setting item 2, the mobile phone can start the face changing detection without displaying prompt c. It should be noted that in the case where the risk protection is started, the face changing detection is started by default, but in response to the operation event of the user separately closing the face changing detection, the mobile phone can separately close the face changing detection, so there can be a case where the risk protection is started but the face changing detection is not started. For specific implementation of separately closing the face changing detection, please refer to the description of mode 3 below, which will not be described in detail here.
[0173] As in the manner two, the system manager service in the mobile phone provides the risk protection. Based on this, if the mobile phone does not start the system manager service, and the operation event of starting the face changing detection switch in the setting item 2 is detected, the mobile phone can first prompt the user to start the system manager service (marked as prompt d), and in response to the operation of the user starting the system manager service in the prompt d, the mobile phone can start the system manager service and the face changing detection.
[0174] Continuing to refer to FIG. 7, in the case where the system manager service is not started, in response to the triggering operation of the user on the “AI face changing reminder” 7011 in the interface 701, the mobile phone can display the interface 705, and the prompt 7051 in the interface 705 is the prompt d. The prompt 7051 includes the text prompt “When starting the face changing detection, the “system manager service” needs to be started simultaneously”, prompting the user to start the system manager service. The prompt 7051 includes the control “Agree” 7052 for starting the system manager service, and the operation of starting the system manager service can be the triggering operation on the “Agree” 7052. In response to the triggering operation of the user on the “Agree” 7052, the mobile phone can start the system manager service and the face changing detection, as shown in the interface 703.
[0175] If the mobile phone starts the system manager service, and the operation event of starting the face changing detection switch in the setting item 2 is detected, the prompt d can not be displayed, and the face changing detection can be started, as shown in the response process of the interface 702 to the interface 703 described above.
[0176] In practice, if the system manager service is not started, it means that the risk protection is also not started. Therefore, in the case where the system manager service is not started, the response process of the mobile phone can be the interface 701, the interface 705, the interface 704 and the interface 703, so as to ensure that the face changing detection is started on the premise that the system manager service and the risk protection are started. In the case where the system manager service is started, the risk protection can not be started. In this case, the response process of the mobile phone can be the interface 701, the interface 704 and the interface 703.
[0177] Manner four, the starting event is the operation event of starting the detection on the prompt a for prompting to start the risk protection.
[0178] In the case where the risk protection is not started, the mobile phone can push the prompt a for prompting to start the face changing detection. Referring to FIG. 8, in the case where the risk protection is not started, the mobile phone can display the interface 801, and the prompt 8011 in the interface 801 is the prompt a. The prompt 8011 includes the prompt text “Newly added “AI face changing detection” to detect whether the AI face changing in the video call”, prompting that the face changing detection function is newly added, and the “Start risk detection” 8012 in the prompt 8011 prompts the user to start the risk protection, that is, the prompt 8011 can prompt the user to start the risk protection.
[0179] In response to the operation event of starting detection performed by the user on the prompt a, the phone can display a setting interface of the risk protection. In a specific implementation, the prompt a includes a control of starting the risk protection, such as the "start risk detection" 8012 in the interface 801, and the operation event of starting detection can be a triggering operation on the control of starting the risk protection.
[0180] Continuing to refer to FIG. 8, the control of starting the risk protection is "start risk protection" 8012 in the interface 801, and in response to the triggering operation of the user on the "start risk protection" 8012, the phone can display an interface 802 (the same as the interface 203 described above), which is a setting interface of the risk protection.
[0181] After the setting interface of the risk protection, the phone can start the risk protection in response to the operation event of the control of starting the risk protection, so as to start the face swapping detection. For details, refer to the description in the third manner described above, and details are not described herein again.
[0182] Of course, if the risk protection has been started, but the face swapping detection has not been started, in response to the operation event of starting detection performed by the user on the prompt a, the phone displays the setting interface of the risk protection, in which the control of starting the risk protection is in the started state, instead of the closed state in the interface 801, but the control corresponding to the setting item of the face swapping detection is still in the closed state. In this case, the phone can start the face swapping detection in response to the operation event of the control of starting the face swapping detection. For details, refer to the description in the third manner described above, and details are not described herein again.
[0183] After displaying the prompt a, in response to the operation event of starting detection, or in response to the triggering operation of the user on an area other than the prompt a, or in response to the display time of the prompt a reaching a time length 1 (such as 10 s), the phone can cancel the display of the prompt a.
[0184] Continuing to refer to FIG. 8, in response to the triggering operation of the user on the "start risk protection" 8012, the interface 802 displayed by the phone no longer includes the prompt 8011.
[0185] Further, the prompt a can be in the form of a card as shown in the prompt 8011 described above. After the display time of the prompt a reaches a time length 2, such as 5 s, the phone can first shrink the prompt a into a capsule 1 to avoid interfering with the user using the phone. After the display time of the capsule 1 reaches a time length 3, the phone can cancel the display of the capsule 1, so as to cancel the display of the prompt a.
[0186] Taking the time length 2 as 5s and the time length 3 as 10s as an example, after the display time length of the prompt 8011 in the interface 801 reaches 5s, the phone can shrink the prompt 8011 into a capsule 9011 in the interface 901 shown in FIG. 9, that is, the capsule 9011 is a capsule 1. After the display time length of the capsule 9011 reaches 10s, the phone can cancel the display of the capsule 9011, such as the display interface 902.
[0187] After shrinking the prompt a into the capsule 1, in response to the triggering operation of the user on the capsule 1, the phone can restore the display of the prompt a, such as in response to the triggering operation of the user on the capsule 9011 in the interface 901, the phone can restore the display of the interface 801 shown in FIG. 8.
[0188] It should be noted that in the foregoing introduction of the fourth mode, the phone displays the prompt a on the desktop as an example. In actuality, the phone can also display the prompt a on other interfaces. For example, when the phone detects that the push condition of the prompt a is met, the phone displays the prompt a in the current interface of the phone.
[0189] The application examples herein exemplarily introduce the push condition of the prompt a:
[0190] The first push condition is that the entering of the foregoing setting item 1 is detected in the case where the risk protection is not started.
[0191] When the phone detects the entering of the setting item 1, such as the entering of the foregoing interface 202, the phone can push the prompt a. In this way, the prompt a can be pushed after the risk protection is not started and the setting item 1 of the risk protection is viewed.
[0192] The second push condition is that a risk event is detected.
[0193] The risk event includes receiving a risk phone, receiving a risk information, receiving an overseas phone, downloading an unknown application and performing an operation of inputting a bank account, transferring, sharing a screen and the like.
[0194] When the phone detects the risk event, the phone pushes the prompt a. In this way, the user can be timely prompted to start the face changing detection in the case where the phone has a security risk.
[0195] The third push condition is that the face changing detection is not started and the interval time length 4, such as 7 days, 20 days, 60 days and the like, has passed since the last display of the prompt a.
[0196] After the display of the prompt a, if the phone does not detect the operation event of starting the detection, the phone will not start the face changing detection. If the face changing detection is not started all the time and a long time has passed since the last display of the prompt a, the phone can display the prompt a again to prompt the user again.
[0197] In actual implementation, the multiple push conditions can be combined.
[0198] For example, the mobile phone can display the prompt a for the first time after any one of the push condition one and the push condition two is reached, and then can display the prompt a again after the push condition three is reached.
[0199] For another example, the mobile phone can display the prompt a for the first time after any one of the push condition one and the push condition two is reached, and then can display the prompt a again after the push condition three is reached and any one of the push condition one and the push condition two is reached.
[0200] In addition, in the case that the risk protection is enabled in the OOBE, the mobile phone can push the prompt d, which is used to prompt the user to understand the face swapping detection. Based on the foregoing description of the first mode, it is known that the risk protection is enabled by default in the OOBE. If the risk protection is enabled in the OOBE, it is highly possible that the user does not understand the risk protection and the multiple risk detection functions. In this case, after the OOBE is completed, the mobile phone can prompt the user to understand the face swapping detection through the prompt d.
[0201] Referring to FIG. 10, in the case that the risk protection is enabled in the OOBE, the mobile phone can display an interface 1001, and a prompt 10011 of the interface 1001 is the prompt d. The prompt 10011 includes the prompt text “Add “AI face swapping detection” to detect whether AI face swapping in video call”, which prompts that the face swapping detection function is added, and the prompt 10011 further includes the text “Learn more” 10012, which prompts the user to learn more, that is, the prompt 10011 can prompt the user to understand the face swapping detection.
[0202] In response to the user performing the operation of learning more on the prompt d, such as the triggering operation of the control of learning more in the prompt d, the mobile phone can display a setting interface of the face swapping detection to enable the user to understand the face swapping detection. Continuing to refer to FIG. 10, taking the control of learning more 10012 in the interface 1001 as an example, in response to the user performing the triggering operation on the control of learning more 10012, the mobile phone can display an interface 1002 (which is the same as the interface 401 described above), and the interface 1002 is the setting interface of the face swapping detection.
[0203] The mobile phone can display a function introduction of the face swapping detection in the setting interface of the face swapping detection. In addition, the setting interface of the face swapping detection includes a details entry, which is used to trigger the mobile phone to display detailed information of the face swapping detection. Thus, the user can understand the face swapping detection. For details, refer to the description of the second mode described above, which will not be described herein again.
[0204] The disappearance, contraction, push condition, and the like of the prompt d can refer to the related descriptions of the prompt a described above, which will not be described herein again.
[0205] After starting the face swapping detection, the mobile phone can perform the face swapping detection and give a prompt during the video call.
[0206] Further, the mobile phone can actively or passively perform the face swapping detection. Active detection means that the detection can be started actively without triggering by the user. Passive detection means that the detection can be started passively based on the detection operation of the user. The following will be described respectively.
[0207] Scheme one, active detection.
[0208] In response to starting the video call, the mobile phone can start the face swapping detection actively.
[0209] In the case of detecting face swapping risk, the mobile phone displays a prompt e for prompting the existence of face swapping risk, that is, to realize risk feedback. In the case of detecting no face swapping risk, no prompt can be given. In this way, the mobile phone can give a prompt only in the case of face swapping risk, so as to reduce the interference of face swapping detection on the video call as much as possible.
[0210] Referring to FIG. 11, for the active detection mode: in the no-risk (i.e. no face swapping risk) scenario, the user has no awareness before, during and after the detection. In the risk (i.e. no face swapping risk) scenario, the user has no awareness before and during the detection, and the mobile phone displays the prompt e after the detection to realize risk feedback, and the user has awareness in this process.
[0211] After detecting the face swapping risk, the mobile phone can display the prompt e in the video call interface. Referring to FIG. 12, during the video call, the mobile phone can display the interface 1201. In the case of detecting the face swapping risk, the mobile phone can display the interface 1202, the prompt 12021 in the interface 1202 is the prompt e, and the prompt 12021 includes the prompt text “the other party is suspected of AI face swapping and identity fraud”, which is used to prompt the existence of face swapping risk.
[0212] The prompt e can include information indicating the existence of security risk, such as the prompt text “the other party is suspected of AI face swapping and identity fraud” in the prompt 12021.
[0213] Further, the prompt e can also include information indicating the risk degree, such as the text “AI face swapping synthesis probability 95%” in the prompt 12021, which is used to prompt the risk degree.
[0214] Further, the prompt e can also include information indicating the verification of the other party's identity, such as the text “please verify the authenticity of its identity through other ways” in the prompt 12021, which is used to prompt the verification of the identity.
[0215] Further, the prompt e can further include an entry for understanding face changing detection, such as the entry “Further understand AI face changing risk” 12022 in the prompt 12021, for triggering display of a setting interface of face changing detection. In response to a triggering operation of the entry for understanding face changing detection by the user, the phone can display the setting interface of face changing detection, such as the interface 401 in the foregoing.
[0216] Further, the prompt e can further include a re-detection control (which can also be referred to as “re-detection control”), such as “re-detect” 12023 in the prompt 12021, for triggering the phone to re-perform face changing detection. In response to a triggering operation of the re-detection control by the user, the phone can re-perform face changing detection. It should be noted that the phone triggers face changing detection due to the triggering operation of the re-detection control by the user, which belongs to a case of passive execution of face changing detection by the phone. Details about this part will be described in the part about passive execution of face changing detection by the phone below, and will not be described in detail here.
[0217] Further, the prompt e can further include a screen recording control (which can be referred to as “screen recording control”), such as “record and save video” 12024 in the prompt 12021, for triggering the phone to record the screen of the video call. In response to a triggering operation of the screen recording control by the user, the phone can record the screen of the video call until the end of the video call, and then stop recording and save the screen recording to the phone.
[0218] In actual implementation, after detecting the face changing risk, the phone can have ended the video call. For this case, the phone can display the prompt f in an interface after ending the video call, such as a chat list interface, for prompting the existence of the face changing risk.
[0219] Continuing to refer to FIG. 12, during the video call, the phone can display the interface 1201. After detecting the face changing risk and exiting the video call, the phone can display the interface 1203. The prompt 12031 in the interface 1203 is the prompt f, and the prompt 12031 includes the prompt text “opponent suspected AI face changing and false identity”, for prompting the existence of the face changing risk.
[0220] Similar to the prompt e, the prompt f can also include information indicating the existence of the security risk, information indicating the risk degree, information indicating verification of the identity of the opponent, an entry for understanding face changing detection, and the like. Details can be referred to the related description of the prompt e in the foregoing.
[0221] indicating the existence of the security risk, information indicating the risk degree, information indicating verification of the identity of the opponent, an entry for understanding face changing detection, and the like. Details can be referred to the related description of the prompt e in the foregoing.
[0222] At this point, it should be noted that the prompt f is displayed after exiting the video call, and at this time, it is not possible to re-detect and record the screen of the video call. Accordingly, the prompt f usually does not include the re-detection control and the screen recording control.
[0223] Further, the prompt f can include a control for canceling display, such as "I know" 12032 in prompt 12031, for triggering the phone to cancel display of the prompt f. In response to the user triggering operation on the control for canceling display, the phone can cancel display of the control 6.
[0224] In actual implementation, the phone can intercept a picture of the video call for face swapping detection. In some cases, the phone can not have intercepted a picture for face swapping detection before the video call ends. In this case, the phone cannot obtain a detection result. However, considering that the phone initiates detection, the phone does not prompt the user about the incomplete detection, thereby avoiding interference with the user.
[0225] In actual implementation, the phone can not detect a face and cannot obtain a detection result. Similarly, considering that the phone initiates detection, the phone does not prompt the user about the undetected face, thereby avoiding interference with the user.
[0226] Further, in response to starting a video call, the phone can initiate detection of face swapping risk when a detection condition of face swapping risk is met. In this way, face swapping risk can be detected in a targeted manner, without performing face swapping detection for each video call, thereby reducing power consumption of the phone.
[0227] The following lists several typical detection conditions:
[0228] Detection condition 1
[0229] The opposite party of the video call is a friend added within a time length of 5. Among them, for the same newly added friend, if a video call is made with the friend within a time length of 5 after the friend is added, the phone performs face swapping detection; if a video call is made with the friend after a time length of 5 after the friend is added, the phone does not perform face swapping detection.
[0230] Referring to FIG. 13, the phone can display interface 1301, which is a management interface of a new friend addition request in a chat application. The interface 1301 includes an addition request 13011 of David. In response to the user triggering operation on "accept" 13012 in the addition request 13011, the phone can add David as a friend. For example, after the addition, the phone returns to a chat list interface and displays interface 1302. The interface 1302 includes a conversation option 13021 of the newly added friend David.
[0231] Continuing to refer to FIG. 13, within a time length of 5 after David is added as a friend, the phone receives a video invitation from David and can display interface 1303, which includes an answer control 13031. In response to the user triggering operation on the answer control 13031, the phone can start a video call, such as displaying interface 1201 in the foregoing FIG. 12. At this time, the phone can start face swapping detection.
[0232] Conversely, within the time duration 5 after adding David as a friend, there is no call with David, but after that, a video invitation from David is received, interface 1303 can also be displayed, and after starting the video call, the phone can not perform face changing detection.
[0233] Generally, the face changing risk occurs in the video call with the recently added friend. In this way, the phone can perform face changing detection for the video call with the recently newly added friend, so as to more likely detect the video call with the face changing risk.
[0234] Detection condition 2
[0235] The first video call with the opposite party of the video call. That is to say, for the same friend, if the first video call with the friend, the phone performs face changing detection, and if the non-first video call with the friend, the phone does not perform face changing detection.
[0236] Generally, the face changing risk occurs in the first video call with a friend, and if there is no face changing risk in the first time, there is probably no face changing risk in the subsequent video call with the friend. In this way, the phone can detect whether the first video call with the friend has the face changing risk, so as to realize whether the video call with the friend has the face changing risk.
[0237] Detection condition 3
[0238] The opposite party of the video call is in the list 1, and the list 1 records the newly added number 1 of friends.
[0239] In practice, for the newly added friend, the phone can store the related information such as the name, the place of origin, etc. in a specific area for face changing detection. At the same time, some users may add a large number of friends in a short period of time due to the particularity of their profession, such as adding dozens of friends in a day. Correspondingly, the storage space required by the phone to store the related information is very large.
[0240] For this situation, the phone can store the related information of the number 1 of friends at most. After storing the related information of the number 1 of friends, if a newly added friend, the phone can delete the information stored in the specific area in the form of first-in first-out, and then store the related information of the newly added friend. Correspondingly, the list 1 is the list of friends whose related information is stored in the specific area. The phone can perform face changing detection on the stored friends. In this way, the phone always performs face changing detection on the newly added number 1 of friends, while avoiding the large occupation of storage space.
[0241] Detection condition 4
[0242] The detection condition of the face changing risk includes that the other party initiates the video call actively. Generally, the user who initiates the video call actively is the user who has obtained the trust of the mobile phone user and then guides the mobile phone user to perform some operations that may cause property loss or information leakage. Therefore, the mobile phone can perform face changing detection on the video call initiated by the other party actively, and can more likely detect the video call with the face changing risk.
[0243] In actual implementation, at least two of the above detection conditions can be combined. For example, the above detection conditions 1-4 can be combined, so that the mobile phone can perform face changing detection on the video call initiated by the new friend added in the list 1 within the time length 5.
[0244] Of course, in actual implementation, the detection condition of active detection is not limited to the above-mentioned detection conditions, for example, the detection condition can also be that the mobile phone has a risk event, which can be referred to the description of the second push condition in the foregoing.
[0245] Scheme two, passive detection.
[0246] In the process of the video call, the mobile phone can start face changing detection passively based on the detection operation of the user in the video call interface. In this way, the mobile phone can detect whether there is a face changing risk based on the detection demand of the user.
[0247] Among them, the possible cases of the detection operation include at least one of the following:
[0248] Case one, the detection operation is the triggering operation of the user on the re-detection control in the foregoing prompt e.
[0249] Taking the prompt e as the prompt 12021 in the interface 1202 shown in the foregoing FIG. 12, and the re-detection control as the “re-detection” 12023 in the prompt 12021 as an example, the detection operation can be the triggering operation of the user on the “re-detection” 12023.
[0250] Case two, in the case where the face changing detection has been started, the mobile phone can display a prompt g in the video call interface, for prompting the use of the face changing detection function. Correspondingly, the detection operation can be the triggering operation of the user on the prompt g.
[0251] In a specific implementation manner, the prompt g includes a detection control, and the detection operation can be the triggering operation of the user on the detection control.
[0252] Referring to FIG. 14, the mobile phone can display an interface 1401, and a switch state of a setting item "AI face changing reminder" 14011 of face changing detection in the interface 1401 is "turned on", indicating that the face changing detection is turned on. In this case, during the process of the video call, the mobile phone can display an interface 1402, which is a video call interface, and a prompt 14021 in the interface 1402 is prompt g. The prompt g includes prompt text "AI face changing detection can detect whether the face of the video call opposite party is impersonated" and "immediate detection" 14022, so as to prompt the user to use the face changing detection. The "immediate detection" 14022 is a detection control, and the detection operation can be a trigger operation of the user on the "immediate detection" 14022.
[0253] Further, in addition to the face changing detection being turned on, the push condition of the prompt g can further include one or more of the aforementioned push condition two, detection condition 1, detection condition 2, detection condition 3, and detection condition 4. The present application does not make specific limitations thereto.
[0254] The mobile phone can present different forms when pushing the prompt g for different times.
[0255] Among them, when the mobile phone displays the prompt g for the first to N (N>1, N is an integer) times, the mobile phone can first present the prompt g in the form of a card, such as the prompt 14021 in the interface 1402 described above. When the mobile phone displays the prompt g for the N+1 time and thereafter, the mobile phone can first present the prompt g in the form of a capsule, such as the capsule 14031 in the interface 1403 described above.
[0256] In this way, when the mobile phone displays the prompt g for the first few times, the mobile phone can present the details of the prompt to the user, so that the user can clearly understand the content of the prompt, and when the mobile phone displays the prompt g thereafter, the mobile phone can present a capsule to the user, thereby reducing the impact on the call.
[0257] Further, the card form and the capsule form can be switched to each other.
[0258] Among them, after the display time of the prompt g in the card form reaches the time length 7, such as 5s, the mobile phone can shrink the prompt g into the capsule form. Referring to FIG. 14, after the interface 1402 is displayed for 5s, the mobile phone can shrink the prompt 14021 in the interface 1402 into the capsule 14031 in the interface 1403, that is, the capsule 14031 is the shrunk prompt g.
[0259] Among them, in response to the click operation of the user on the prompt g in the capsule form, the mobile phone can expand the prompt g into the card form. Referring to FIG. 14, after the interface 1403 is displayed, in response to the click operation of the user on the capsule 14031, the mobile phone can expand the capsule 14031 into the prompt 14021 in the interface 1402.
[0260] In case three, when the face changing detection is not enabled, the mobile phone can display a prompt h in the video call interface, to prompt the user to enable and use the face changing detection function. Correspondingly, the detection operation can be a trigger operation of the user on the prompt h.
[0261] In a specific implementation, the prompt h includes an enable detection control, and the detection operation can be a trigger operation of the user on the enable detection control.
[0262] Referring to FIG. 15A, the mobile phone can display an interface 1501, and the switch state of the setting item “AI face changing reminder” 15011 of the face changing detection in the interface 1501 is “not enabled”, indicating that the face changing detection is not enabled. In this case, during the video call, the mobile phone can display an interface 1502, which is a video call interface, and the prompt 15021 in the interface 1502 is the prompt h. The prompt 15021 includes the prompt text “enable “AI face changing detection” to detect whether the video call opponent's face is impersonated” and “enable and detect” 15022, so as to prompt the user to enable and use the face changing detection. The “enable and detect” 15022 is an enable and detection control, and the detection operation can be a trigger operation of the user on the “enable and detect” 15022.
[0263] At this point, it should be noted that after the mobile phone enables the risk protection, the multiple risk detection functions under the risk protection will be enabled by default. Thereafter, the user can individually trigger to close one or more of the multiple risk detection functions, such as individually triggering to close the face changing detection. At this time, the situation shown in the interface 1501 described above will occur, i.e., the risk protection is enabled, but the face changing detection is not enabled. For details, please refer to the description of mode 3 below, which will not be described in detail here. In this way, the mobile phone can push the prompt h to specifically prompt the user to enable and use the face changing detection when the user individually closes the face changing detection.
[0264] It should be noted that if the face changing detection is not enabled due to the risk protection not being enabled, i.e., both the risk protection and the face changing detection are not enabled, the mobile phone will generally push the aforementioned prompt a, and will not push the prompt h. For details, please refer to the related description of mode four in the foregoing, which will not be described here.
[0265] Further, in addition to not enabling the face changing detection, the pushing condition of the prompt h can also include one or more of the aforementioned pushing condition two, detection condition 1, detection condition 2, detection condition 3, and detection condition 4. The present application does not make specific limitations in this regard.
[0266] Generally, the push condition of prompt h is stricter than the push condition of prompt g. In this way, the phone can push prompt h at a low frequency without starting face-swapping detection, reducing the impact on the user; and the phone can push prompt g at a high frequency with face-swapping detection started, improving the usage rate of face-swapping detection.
[0267] For example, a set of typical push conditions of prompt g and prompt h are as follows:
[0268] The push condition of prompt g includes the existence of a security risk (i.e., push condition two) and the opposite party of the video call being a friend added within a time length of 5 (i.e., detection condition 1).
[0269] The push condition of prompt h includes the existence of a security risk (i.e., push condition two), the opposite party of the video call being the first added friend, and the opposite party initiating a video call within a time length of 6 (less than the time length of 5) after adding the friend.
[0270] In addition, similar to prompt g, the phone can present prompt h in different forms (such as card form and capsule form) in different times of pushing prompt h. Further, the card form and the capsule form can be switched between each other. Referring back to FIG. 15B, after the phone displays interface 1502 for 5s, the phone can shrink prompt 15021 (i.e., prompt h in card form) in interface 1502 to capsule 15031 in interface 1503, i.e., capsule 15031 is the shrunk prompt h.
[0271] In response to the user's click operation on the capsule form of prompt h, the phone can expand prompt h to the card form. Referring back to FIG. 15B, after the phone displays interface 1503, in response to the user's click operation on capsule 15031, the phone can expand capsule 15031 to prompt 15021 in interface 1502.
[0272] Further, the detection condition triggering the foregoing active detection is generally looser than the push condition of prompt h and prompt g. In this way, the phone can trigger the execution of active detection with a lower perception degree (only detectable when there is a face-swapping risk) at a higher frequency, and trigger the execution of passive detection with a higher perception degree (full visual display) at a lower frequency.
[0273] In the passive detection scheme, in response to the detection operation, the phone can visually display the whole process, so that the user can clearly know the detection process.
[0274] Referring to FIG. 16, in the passive detection scheme: before detection, the phone can prompt to start detection. During detection, the phone can show the detection process. After detection, the phone can feedback the detection result. Among them, the visualization before detection and during detection can make the user clearly perceive that the phone is in detection, so that the user has a sense of certainty. The visualization after detection can make the user obtain the result of face swapping risk, so that the user has a sense of control.
[0275] The response processes before, during and after detection will be introduced respectively as follows:
[0276] First, before detection. Before detection can be understood as the stage before the phone obtains data for face swapping detection. For example, if the phone needs to obtain at least 10 frames of screenshots of video call for face swapping detection, then before obtaining at least 10 frames of screenshots, it belongs to the stage before detection.
[0277] Before detection, the phone can display prompt i for prompting to start detection.
[0278] Referring to FIG. 17, in response to the detection operation, the phone can display interface 1701, which is a video call interface. Capsule 17011 in interface 1701 is prompt i, which includes prompt text “start detection” for prompting to start detection.
[0279] Second, during detection. During detection can be understood as the stage from after the phone obtains data for face swapping detection to before obtaining the detection result.
[0280] During detection, the phone can display prompt j and detection animation, prompt j for prompting that detection is in progress.
[0281] Continuing to refer to FIG. 17, after displaying interface 1701, the phone can then display interface 1702, which is still a video call interface. Capsule 17021 in interface 1702 is prompt j, which includes prompt text “detection in progress” for prompting that detection is in progress. In addition, interface 1702 also includes detection animation of scanning from top to bottom (as shown by the arrow in interface 1702).
[0282] It should be noted that the form of detection animation can be various and is not limited to that shown in interface 1702. Further, in a detection process, multiple effects can be included, such as the detection animation of scanning from top to bottom in interface 1702, and the animation of scanning the face area of the opposite party (represented by the dot matrix on face 17031 in the figure) can also be displayed.
[0283] Third, after detection. After detection can be understood as the stage after the mobile phone detects the detection result. The detection result includes the existence of face swapping risk and the non-existence of face swapping risk.
[0284] After detection, the mobile phone can display prompt j, which is used to prompt that the detection has been completed.
[0285] Continuing to refer to FIG. 17, after the mobile phone displays interface 1702 and obtains the detection result, the mobile phone can then display interface 1704, which is still a video call interface. Capsule 17041 in interface 1704 is prompt j, which includes the prompt text “detection completed” and is used to prompt that the detection has been completed.
[0286] It should be noted that in the passive detection scheme, whether the face swapping risk exists or does not exist, the mobile phone can display prompt j, so that the user can obtain all the detection results.
[0287] In the above introduction of the visualization display before, during and after detection, the prompt in the form of a capsule is mainly taken as an example for illustration, such as prompt h, prompt i and prompt j, which are all capsules. In actual implementation, this is not limited thereto.
[0288] For example, prompt h-prompt j can also be in the form of a card. Referring to FIG. 18, prompt h can also be prompt 18011 in interface 1801, and prompt i can also be prompt 18021 in interface 1802.
[0289] Among them, the card-form prompt j can be divided into two cases: the existence of face swapping risk and the non-existence of face swapping risk.
[0290] Continuing to refer to FIG. 18, in the case of the existence of face swapping risk, prompt j can be prompt 18031 in interface 1803, which is used to prompt the existence of face swapping risk. The specific content of prompt j in the case of the existence of face swapping risk is basically the same as the content included by prompt e, and specific reference can be made to the related description of prompt e in the foregoing, which will not be described here again.
[0291] In actual implementation, in order to avoid malicious detection, if the number of times of repeatedly performing face swapping detection by the mobile phone for the video call does not reach number 1 (such as 3 times, 5 times, etc.), the mobile phone can provide a re-detection control in the card-form prompt j, such as “re-detection” 18032 in prompt 18031 in interface 1803. If the number of times of repeatedly performing face swapping detection by the mobile phone for the video call reaches number 1, the mobile phone can not provide a re-detection control in the card-form prompt j, such as the fact that prompt 18041 in interface 1804 does not include a re-detection control.
[0292] Continuing to refer to FIG. 18, in the case where there is no face swapping risk, the prompt j can also be a prompt 18051 in the interface 1805, and the prompt 18051 includes the prompt text "no AI face swapping risk detected" for prompting that there is no face swapping risk. Further, in the case where there is no security risk, the prompt j in the form of a card can also include a control for re-detection, a control for screen recording, and the like, which are not limited in the present application.
[0293] In a specific implementation, the phone can present the prompts h-j in the form of capsules by default. Subsequently, in response to a click operation of the user on the capsule-form prompt, the phone can switch to presenting in the form of a card.
[0294] In another specific implementation, the phone can present the prompts h and i in the form of capsules by default, and present the prompt j in the form of a card by default, thereby facilitating viewing of specific detection results.
[0295] If the detection result is that there is no face swapping risk, the phone can automatically shrink the prompt j into a capsule form after the display duration of the prompt j in the form of a card reaches a duration 8 (such as 5s), such as shrinking the card prompt 19011 in the interface 1901 shown in FIG. 19 into the capsule 19021 in the interface 1902. Subsequently, after the display duration of the prompt j in the form of a capsule reaches a duration 9 (such as 5s), the phone can automatically cancel display of the prompt j, such as automatically canceling display of the capsule 19021 in the interface 1902 to display the interface 1903. In this way, in the case where there is no face swapping risk, the phone can automatically first shrink and then cancel display of the prompt j, thereby avoiding affecting the call.
[0296] Of course, in the case of displaying the prompt j in the form of a card, in response to an upward swipe operation of the user on the prompt j or a click operation on an area other than the prompt j, the phone can also passively shrink the prompt j into a capsule form.
[0297] In actual implementation, when the detection result is obtained, the phone can have ended the video call. For this case, the phone can display a prompt k in an interface after ending the video call, such as a chat list interface, for prompting that the detection has been completed.
[0298] Similar to the prompt j in the foregoing, the prompt k can also be in the form of a capsule or a card, and corresponding to the two cases of existing face swapping risk and non-existing face swapping risk, the content of the prompt k in the form of a card is also different, and can further prompt the detection result, such as existing face swapping risk or non-existing face swapping risk.
[0299] Referring to Figure 20, the phone receives the test results only after the video call ends, at which point interface 2001 can be displayed. The prompt 20011 in interface 2001 is a capsule-shaped prompt k. Prompt 20011 includes the text "Detection Completed," indicating that the detection has been completed.
[0300] In cases where there is a risk of face swapping, refer to Figure 20. After the video call ends, if the phone detects a face swapping risk, interface 2002 will be displayed. The prompt 20021 in interface 2002 is a card-style prompt k, containing the text "The other party is suspected of using AI face swapping to impersonate someone," indicating that the detection has been completed and the result confirms a face swapping risk.
[0301] If there is no risk of face swapping, continue referring to Figure 20. After the video call ends, if the phone detects that there is no risk of face swapping, it can display interface 2003. The prompt 20031 in interface 2003 is a card-style prompt k, which includes the prompt text "No AI face swapping risk detected," indicating that the detection has been completed and the detection result is that there is no risk of face swapping.
[0302] Regarding the specific content of hint k, please refer to the previous explanation of hint j; it will not be repeated here.
[0303] It should be noted that the "k" prompt is only displayed after the video call ends. At this point, it is no longer possible to re-detect and record the video call screen. Consequently, the "k" prompt usually does not include controls for re-detection or screen recording.
[0304] In practice, the phone can capture screenshots of video calls for face-swapping detection. In some cases, the phone may not have captured the necessary image for detection before the video call ends. In this situation, the phone cannot obtain a detection result. Furthermore, considering a passive detection approach, user feedback is needed; the phone could indicate that the detection is incomplete. Similarly, the phone could display this notification in capsule or card format.
[0305] Referring to Figure 21, after the video call ends, before the phone has acquired the image for face-swapping detection, the phone can display the chat list interface shown in interface 2101. Interface 2101 includes prompt 21011, which contains the text "Detection incomplete," indicating that the detection is not yet complete. Alternatively, the phone can display the chat list interface shown in interface 2102, which includes prompt 21021, which contains the text "Incomplete," indicating that the detection is not yet complete.
[0306] In practice, the phone might not even detect the other person's face, in which case no detection result can be obtained. Furthermore, considering a passive detection approach, user feedback is necessary; the phone could indicate that the detection was incomplete. Similarly, the phone could display a capsule or card notification indicating incomplete detection, and the card notification could further specify that the incomplete detection was due to the failure to detect a face.
[0307] It should be noted that if no face is detected during a video call, a notification can be displayed on the video call interface. If no face is detected after the video call ends, a notification can be displayed on the interface after the video call ends, such as the chat list interface.
[0308] Taking a video call interface prompt as an example, as shown in Figure 22, during a video call, if the phone detects that the other party's face is not present, the phone can display the video call interface shown in interface 2202. Interface 2202 includes prompt 22021, which includes the text "Incomplete," indicating that the detection was not completed. Alternatively, the phone can display the video call interface shown in interface 2201, which includes prompt 22011. Prompt 22011 includes the text "Detection Incomplete" and "No Face Detected," indicating that the detection was incomplete because no face was detected.
[0309] Regarding Scheme 1 and Scheme 2 above, if a face-swapping risk is detected, the mobile phone can adopt at least one of the following response strategies to make users aware of the face-swapping risk.
[0310] Response strategy one: During video calls, the phone does not remove the warning about the risk of face swapping, such as warning e in the active detection and warning j in the passive detection, and the same applies below. In this way, the phone can continuously display the risk of face swapping throughout the video call.
[0311] In the proactive detection solution, the phone can continuously display the prompt "e" during a video call.
[0312] In the passive detection scheme, during a video call, the phone can switch between a capsule-shaped prompt j and a card-shaped prompt j in response to user actions. However, the phone will always keep prompt j displayed. For example, in response to the user's swipe up on prompt j or tapping an area outside of prompt j, the phone can passively shrink prompt j into a capsule shape, but will continue to display prompt j in capsule form. This is unlike the scenario shown in Figure 19 above, where there is no risk of face swapping, where the phone actively shrinks prompt j from card to capsule shape and eventually cancels its display.
[0313] Of course, in response to the user's action of canceling the display, the phone can passively cancel the display of the warning about the risk of face swapping. This cancellation action can be a swipe up action on the warning, or a trigger action on an area other than the warning.
[0314] The second response strategy involves triggering a face-swapping risk warning during a video call. In response, the phone will access the notification center and display Notification 1, indicating that a face-swapping risk has been detected. This way, even after accessing the notification center, the phone will continue to display the face-swapping risk warning.
[0315] Taking a passive detection scheme as an example, referring to Figure 23, after detecting a risk of face swapping, the phone can display interface 2301, where prompt 23011 is prompt j. In response to the user's swipe-down operation on interface 2301, such as the swipe in the direction indicated by the arrow in interface 2301, the phone can display interface 2302. Interface 2302 is a pull-down notification center, which includes various notifications received by the phone (not shown in the figure). Notification 23021 is notification 1, which includes the text "The other party is suspected of using AI face swapping to impersonate someone," thus notifying the user that a face swapping risk has been detected.
[0316] In addition, the content in notification 1 can be consistent with the content in prompt j. For example, notification 1 can also include controls for re-detection, screen recording, etc., so that users can easily trigger re-detection, screen recording, etc. in the pull-down notification center.
[0317] The third response strategy involves the phone displaying Notification 2 on the screen after a video call ends, such as in the chat list. Notification 2 can again indicate that a face-swapping risk has been detected and provide an entry point to the detection record to view the risk information. In response to the user's triggering of the detection record entry, the phone can display the face-swapping detection record interface. Thus, after detecting a face-swapping risk and ending the video call, the phone provides a quick access to the face-swapping detection record interface so that the user can view the risk information.
[0318] Taking the passive detection scheme as an example, as shown in Figure 24, after detecting a face-swapping risk, the phone can display interface 2401, with prompt 24011 being prompt j. After ending the video call, the phone can display interface 2402, which includes prompt 24021, notification 2. Prompt 24021 includes the text "The other party in the video call is suspected of being impersonated by someone using AI face-swapping. Please verify their identity through other means," thus re-notifying the phone of the face-swapping risk. Prompt 24021 also includes "Learn More" 24022, which is the entry point for the detection record. In response to the user's triggering of "Learn More" 24022, the phone can display interface 2403, which is the face-swapping detection record interface, including detection record 24031.
[0319] Additionally, if Notification 2 is displayed on the screen after a video call ends, the phone can display Notification 2 in the pull-down notification center in response to the notification being accessed. This way, even when the phone enters the pull-down notification center, it can still provide a warning about the risk of face-swapping and offer a quick access to the face-swapping detection log.
[0320] Referring again to Figure 24, in response to the user's swipe-down action on interface 2402, the phone can display interface 2404. Interface 2404 is the pull-down notification center, which includes notification 24041, which is notification 2. The "Learn More" option 24042 within notification 24041 is the entry point for the detection record. In response to the user's triggering action on "Learn More" 24042, the phone can also display interface 2403.
[0321] The fourth response strategy is to provide a second warning about the risk of face-swapping within a 9-hour window (e.g., 12 hours, 24 hours) after detecting a face-swapping risk and then detecting related risky behavior. This strengthens the warning after a risky behavior occurs.
[0322] Among the associated risky behaviors are accessing money transfer interfaces and receiving website links from friends in video calls that may involve deepfakes.
[0323] Taking the passive detection scheme as an example, as shown in Figure 25A, after detecting a risk of face swapping, the phone can display interface 2501, with prompt 25011 (prompt j) in interface 2501. If the phone then enters the transfer interface within the following 9 hours, interface 2502 will be displayed. Interface 2502 includes the prompt text "The system detected a risk of AI face swapping in your recent video call; continuing the payment may result in financial loss," thus reiterating the risk of face swapping.
[0324] In addition, after detecting a face-swapping risk or when the risk level reaches a preset threshold, such as a probability of 80%, the phone can send the risk information to a protection device. For example, the phone can provide a settings entry for a protection device (e.g., in the risk protection settings interface) and send the risk information to the designated protection device.
[0325] Referring to Figure 25B, the mobile phone can display interface 2511, which is the risk protection settings interface. Interface 2511 includes "Send a message to the guardian" 25111 and "Remote guardian reminder" 25112. "Send a message to the guardian" 25112 is a toggle option to send face-swapping risk information to the guardian device. "Remote guardian reminder" 25112 is the settings entry for the guardian device. After the "Send a message to the guardian" 25112 switch is turned on, if the mobile phone detects face-swapping risk or the risk level of face-swapping reaches a preset level value, it can send risk information to the guardian device set in "Remote guardian reminder" 25112.
[0326] Furthermore, after receiving risk information, the protection device can display a risk warning. Referring again to Figure 25B, after receiving risk information, the protection device can display interface 2512, which includes risk warning 25121, indicating that it has received risk information about face-swapping.
[0327] Similar to enabling the face-swapping detection function mentioned earlier, users can also disable it once they no longer need it. Specifically, the phone can disable the face-swapping detection function after detecting an event indicating that it has been disabled. When the face-swapping detection function is disabled, the phone will no longer detect the risk of face-swapping and provide any warnings during video calls.
[0328] The shutdown event can be an operation event that disables risk protection settings in OOBE, an operation event that disables the risk protection switch in the settings application, or an operation event that disables the face-swapping detection switch in the settings application, etc. This application embodiment does not specifically limit this.
[0329] The following sections will explain in detail the process of triggering the face-swapping detection function for the various events listed above.
[0330] Method 1: The shutdown event is the operation event for disabling risk protection settings in OOBE.
[0331] In OOBE, the phone usually has risk protection enabled by default. If the phone detects an operation event in OOBE where the user turns off the risk protection settings, such as turning off the switch 1011 that is enabled by default in interface 101 in Figure 1D above, it means that the user does not support enabling risk protection and the phone can turn off risk protection.
[0332] Method 2, the shutdown event is the operation event of turning off the risk protection switch in the settings application.
[0333] The phone detects a trigger operation on the risk protection switch that is currently enabled, and can then disable risk protection, thereby disabling the face-swapping detection function under risk protection.
[0334] Referring to Figure 26, with risk protection enabled, the phone displays interface 2601, which is the risk protection settings interface. The risk protection switch 26011 in interface 2601 is enabled. In response to the user's triggering of the risk protection switch 26011 in interface 2601, the phone displays interface 2602. In interface 2602, the risk protection switch 26011 is disabled, meaning risk protection is turned off. At this time, the "AI Face Swap Reminder" setting 26021 in interface 2602 is not enabled, indicating that the face swap detection function is also disabled.
[0335] Furthermore, disabling risk protection will disable various risk detection functions under risk protection. Based on this, in response to the user's triggering of the risk protection switch (which is currently enabled), the phone can first display prompt 'm', such as prompt 26031 in interface 2603, to indicate that disabling risk protection will render various risk detection functions unavailable. In response to the user's confirmation of prompt 'm', such as the triggering of "Still Off" 26032 in prompt 26031, the phone will then disable risk protection, thereby disabling face-swapping detection, as shown in interface 2602 after disabling.
[0336] Method 3, the off event is the operation event of turning off the face-swapping detection switch in the settings application.
[0337] The phone detects a trigger operation on the face-swapping detection switch, which is currently enabled, and can disable face-swapping detection.
[0338] Referring to Figure 27, when the face-swapping detection function is enabled, the phone displays interface 2701, which is the face-swapping detection settings interface. The face-swapping detection switch 27011 in interface 2701 is in the enabled state. In response to the user's triggering operation on switch 27011 in interface 2701, the phone displays interface 2702. In interface 2702, the face-swapping detection switch 27011 is in the disabled state, meaning face-swapping detection is turned off.
[0339] It should be noted that using method 3 only disables the face-swapping detection function; the risk protection function remains enabled. This can result in situations where risk protection is on but face-swapping detection is off. For example, in the risk protection settings interface, the risk protection switch may be on, but the face-swapping detection setting may be off.
[0340] Referring again to Figure 27, in response to the triggering operation of the return control 27021 in interface 2702, the phone can display interface 2703 (the same as interface 1501 mentioned above). Interface 2703 is the risk protection settings interface. In interface 2703, the risk protection switch 27031 is in the on state, indicating that risk protection has been enabled, and the switch state of the face-swapping detection setting item "AI face-swapping reminder" 27032 is off, indicating that face-swapping detection is not enabled.
[0341] Please refer to Figure 28. In a specific embodiment, a video call method is provided. Taking the application of this method to a mobile phone as an example, it includes at least the following steps S2801 to S2804:
[0342] Step S2801: During the video call, intercept the video stream and repeatedly acquire video frames.
[0343] It should be understood that the embodiment illustrated in Figure 28 is a face-swapping detection performed by active detection through steps S2801 to S2804 when the detection switch is turned on (i.e., user authorization is obtained).
[0344] The detection switch may include at least one of the following: a risk protection switch or a face-swapping detection switch.
[0345] It should be understood that in some scenarios, the face-swapping detection switch can be integrated into the risk protection switch, and the face-swapping detection function can be activated simply by turning on the risk protection switch. For example, in some scenarios, the risk protection switch can be decoupled from the face-swapping detection switch, so only turning on the face-swapping detection switch is needed to activate the face-swapping detection function. Furthermore, in some scenarios, both the risk protection switch and the face-swapping detection switch must be turned on simultaneously to activate the face-swapping detection function. In summary, in the embodiments of this application, as long as the detection switch can activate the face-swapping detection function, it is not limited to whether there is one or multiple detection switches, nor is it limited to the specific naming of the detection switches.
[0346] The timing for triggering the face-swapping detection switch can be found in the preceding description. Specifically, as mentioned earlier, after the phone detects the face-swapping detection activation event, it can enable the face-swapping detection function (i.e., turn on the detection switch). Activation events can include skipping risk protection settings in the Out-of-box experience (OOBE), enabling the risk protection switch in the Settings app, enabling the face-swapping detection switch in the Settings app, and performing the activation detection operation in response to prompt 1 indicating the need to enable risk protection. The process of triggering the face-swapping detection function for the aforementioned activation events can be found in the preceding description, and the relevant interfaces can be referred to Figures 2 to 8 and Figure 10 above, which will not be repeated here.
[0347] For example, the mobile phone can intercept video during a video call according to preset interception timing and / or interception frequency parameters to obtain video frames. For each video frame, steps S2802 to S2804 can be executed to obtain the face forgery detection result corresponding to each video frame.
[0348] The preset timing for data interception refers to the timing at which data interception is triggered. For example, data interception can start immediately after a video call is connected, or it can start 3 seconds after a video call is connected.
[0349] The throttling rate indicates the number of frames throttled per second. For example, 24 frames are throttled per second.
[0350] Furthermore, the preset throttling parameters can also include the total throttling duration. For example, a total throttling time of 3 seconds or 5 seconds.
[0351] It should be noted that all parameters related to face-swapping detection in the embodiments of this application can be configured, such as the aforementioned throttling parameters, the triggering parameters used for triggering detection (e.g., preset durations of 5 or 6), and the confidence threshold used below. Parameter configuration enables more flexible and accurate face-swapping detection.
[0352] Understandably, video calls that are passively established (i.e., the local end is called) are generally riskier. Since passively established video calls are based on video call events (i.e., video call events initiated by the other end of the call), mobile phones can identify the risks of video call events after receiving them. For video call events with risks, face-swapping detection can be initiated using either active or passive detection methods. For video call events without risks, face-swapping detection can be omitted.
[0353] In other embodiments, the mobile phone may also perform face-swapping detection for any video call, without being limited to the called party scenario, or limited to starting face-swapping detection only for video calls that pose a risk.
[0354] In some examples, the mobile phone can actively detect face-swapping for high-risk (which can be denoted as the first risk level) video call events (i.e., suspicious video calls with relatively high risk). For medium-risk (which can be denoted as the second risk level) video call events (i.e., suspicious video calls with moderate risk), a passive detection method is used for face-swapping detection. That is, interaction between the front end and the user is added during the detection process.
[0355] It should be noted that in other examples, the mobile phone can also use a passive detection method for face-swapping detection for high-risk (which can be denoted as the first risk level) video call events (i.e., suspicious video calls with relatively high risk). For medium-risk (which can be denoted as the second risk level) video call events (i.e., suspicious video calls with moderate risk), an active detection method can be used. Furthermore, the mobile phone can also uniformly use a default method (such as active or passive detection) for face-swapping detection for video call events of any risk level. There are no limitations on this.
[0356] Regarding the differences in interface presentation between active and passive detection methods, please refer to the interface descriptions in Figures 11 to 22 above. Figures 11 to 13 describe the interface under active detection, while Figures 14 to 22 describe the interface under passive detection. It should be noted that although Figures 23 to 25B are illustrated using passive detection as an example, they can also be applied to active detection. That is, when using active detection for face-swapping detection, if a risk is detected, the risk warning described in Figures 23 to 25B can also be used. Furthermore, the differences in data processing between active and passive detection will be described gradually below.
[0357] In some embodiments, the conditions for triggering face-swapping detection mentioned above may include meeting preset conditions and detecting a video call event within a preset time period after starting monitoring. Specifically, the mobile phone can start monitoring suspicious video calls (i.e., risky video call events) when the preset conditions are met. If a risky video call event is detected within a preset time period after starting monitoring, face-swapping risk detection is performed on the video call established based on that video call event. In other words, if the time difference between the time when the preset conditions are met and the start time of the video call is within the preset time period, it is determined that the conditions for triggering face-swapping detection are met. The start time of the video call may include the time when the video call event is received or the time when the video call is established.
[0358] Optionally, meeting the preset conditions may include at least one of the following: adding a friend event, detecting a preset risk event, or detecting a full-link trigger risk.
[0359] Referring to Figure 13 above, the following explanation uses the "add friend" event as an example. Referring to Figure 13, after receiving David's friend request and adding him as a friend, the system will start listening for David's initiated video call requests within a duration of 5. If David's video call request is received and a video call is established within 5 hours (i.e., a video call event is detected), it is determined that David's video call may pose a risk. That is, if the time difference between adding a friend and the start of the video call is within 5 hours, the conditions for triggering face-swapping detection are met. Therefore, face-swapping detection can be performed on this video call, i.e., face-swapping detection and handling can be carried out during the video call through steps S2801 to S2804.
[0360] In some embodiments, the preset risk events may include at least one of the following: receiving a risky phone call, receiving a risky text message, or downloading a risky application.
[0361] For example, after detecting at least one risk event, such as receiving a risky phone call, receiving a risky SMS message, or downloading a risky application, monitoring of suspicious video call events can be initiated within a duration of 10. Any video call event received within this duration, or even the first video call event, can be considered a risky video call event. In other embodiments, it is not limited to video call events; any video call established within the duration of 10 (e.g., a video call passively established through a video call event and / or an actively initiated video call) can also be considered risky. That is, if the time difference between the moment a risky event is detected on the phone and the start time of the video call is within a duration of 10, it is determined that the conditions for triggering face-swapping detection are met. Then, the phone can be triggered to perform face-swapping detection on the video call. It should be understood that a duration of 10 can be equal to or different from a duration of 5; it is only used to represent a preset duration.
[0362] Furthermore, risk events are not limited to the events listed above, and may also include receiving overseas calls / text messages, receiving unknown calls / text messages, downloading unknown applications and entering bank account numbers, etc., without limitation.
[0363] In some examples, risky phone calls, risky text messages, and risky apps have been identified or confirmed as posing a risk. Overseas phone calls / text messages, unknown phone calls / text messages, and unknown apps may pose a potential risk, but it has not yet been confirmed that a risk definitely exists.
[0364] End-to-end triggered risk: This refers to the determination of the existence of risk based on end-to-end risk detection. In other words, end-to-end triggered risk is an end-to-end risk detection result that indicates the presence of risk.
[0365] End-to-end risk detection refers to the comprehensive assessment of the existence of risk by combining multiple related risk behavior factors. Related risk behavior factors can refer to risk behavior factors that are related in their occurrence sequence. Therefore, detecting end-to-end triggered risk can be based on determining the presence of risk on the mobile terminal based on multiple related risk behavior factors.
[0366] For example, receiving an overseas call is a risk behavior factor, but it cannot be determined solely by this. However, if some risky behaviors are performed after receiving an overseas call, such as downloading unknown applications and sharing the screen, then it is very likely to be a risky call. Therefore, the end-to-end risk detection result can be determined to be risky, thus satisfying the end-to-end risk triggering condition. Consequently, the monitoring of suspicious video call events can be initiated.
[0367] Step S2802: Identify the face region of the other end of the call from the video frame.
[0368] It should be understood that in this embodiment of the application, the holder of the mobile phone is referred to as the user, and the other party in the video call with the user is referred to as the other party in the video call. As shown in Figure 29, the video call interface generally displays the faces of both parties in the call, that is, the video frame generally displays the faces of both parties (man and woman), as shown in Figure 29(a). Assuming that the woman is the image of the user on this end (such as the user using the mobile phone, which can also be called the mobile phone user), then the man in Figure 29(a) is the image of the other party in the call.
[0369] The main objective of the solution in this application embodiment is to identify whether the face of the other party in a video call is a fake face created using AI face-swapping technology to imitate the face of the user's real friend during the video call. Therefore, it is necessary to identify the face region of the other party in the video frame in step S2802, so that in step S2803 below, the other party can be identified as someone else impersonating the user's friend based on that face region.
[0370] In some examples, the phone can repeatedly capture video frames for face-swapping detection. For each current frame to be detected (i.e., the current video frame), the phone can perform face ROI region detection, obtaining at least two face regions: one on the phone itself and one on the other end of the call. The phone can then identify face regions that meet the constraints of the other end's face, which is the other end's face region (hereinafter referred to as the "other end face region"). It should be understood that the other end face region refers to the faces in the video call interface that need to be detected for face-swapping risk.
[0371] For example, the face constraints on the other end may include at least one of positional constraints or size constraints, which are described in detail below:
[0372] (I) Positional Constraints
[0373] For example, the positional constraint may include: the other end's face region is located within a preset positional range. This preset positional range refers to the positional range of the other end's face region / face image within a video frame.
[0374] Specifically, the mobile phone can compare the facial regions identified from the video frames with preset location ranges. If a facial region is located within the preset location range, the location constraint condition is met, and the facial region can be determined to be the opposite facial region.
[0375] (II) Size Constraints
[0376] Size constraints can also be called dimensional constraints. For example, a size constraint may include: the counterpart face region is a larger or smaller face region; that is, the size of the counterpart face region is greater than or less than the size of the local face region. In some embodiments, the size of the face region can be characterized by its resolution; therefore, a size constraint may include: the resolution of the counterpart face region is greater than or less than the resolution of the local face region.
[0377] It should be understood that after a video call is connected, the video call interface will be displayed in the default mode. The video call interface usually has at least two display windows. The display windows are used to display the images captured for both parties in the video call.
[0378] Taking two display windows as an example, as shown in Figure 29(a), one display window is used to display the image (i.e., the female) captured for the local user (i.e., the mobile phone user), and the other display window is used to display the image (i.e., the male) captured for the other end of the call (i.e., the user of the other end's device). Each display window has a default position. Therefore, by using the above position constraints, the face region located within the preset position range can be identified as the other end's face region.
[0379] Furthermore, in video call interfaces, the display window sizes for the local user and the remote user are typically different, resulting in different sizes of the face regions displayed within those windows. Therefore, the remote user's face region can also be identified based on these size constraints. For example, in the default display mode, the remote user's face region will be larger than the local user's face region. Therefore, the phone can identify the larger face region in the video frame as the remote user's face region. It should be noted that, assuming the remote user's face region is smaller than the local user's face region in the default display mode, the phone can also identify the smaller face region in the video frame as the remote user's face region.
[0380] Taking a video call between two people with two display windows as an example, this paper describes how to identify the face region of the other end from the perspective of data processing, using the following examples 1 and 2.
[0381] Example 1:
[0382] The phone can compare the sizes of two facial regions and identify the smaller one as the opposite facial region.
[0383] Example 2:
[0384] The display windows on different ends (this end and the other end) may have significant differences in display range. Therefore, during a video call, the display windows on different ends have different facial region size ranges (determined by the size of the display window). The facial region size range of this end is denoted as the first size range, and the facial region size range of the other end is denoted as the second size range. The mobile phone can compare the detected two facial regions with the second size range respectively, and identify the facial region that matches the second size range as the facial region of the other end.
[0385] It should be understood that during a video call, the user on the local end may switch the display windows of both parties, as shown in Figure 29(b). The user's image (woman) may be displayed in a larger window (i.e., a larger display window in the video call interface), while the image of the other party (man) may be displayed in a smaller window (i.e., a smaller display window in the video call interface). If the positional or size constraints of the other party's face are still applied according to the default display method, it is easy to mistakenly identify the user's face as the other party's face, thus affecting subsequent face-swapping detection of the other party's face region. Therefore, the mobile phone can adopt any one or more of the following two methods for detecting the other party's face region.
[0386] peer detection method 1:
[0387] Based on the peer face image in the target video frame that conforms to the default display method, identify the peer face region in the current frame.
[0388] The "destination face image" refers to the face image of the person on the other end of the call. The "target video frame" refers to the video frame displayed in the default video call interface.
[0389] It should be understood that after a video call is connected, in the first N (N≥1) video frames, such as the first frame after the call is connected, the video call interface will generally be presented according to the default display method. That is, the faces of the user and the recipient will be displayed in their respective positions in the video call interface according to the default display method. Furthermore, based on the phone's own processing characteristics, it can determine how many frames after the video call is connected will be presented according to the default display method. In other words, the phone can know which first few video frames after the call is connected conform to the default display method. Therefore, the phone can identify the recipient's face region from any of the first N target video frames (i.e., the video frames presented according to the default display method) after the video call is connected. Then, the phone can use this recipient's face region as a reference, denoted as the reference face region. When identifying the recipient's face region from subsequent captured video frames, the phone can match all the face regions identified in the subsequent captured video frames with this reference face region, and identify the face regions that match the reference face region as the recipient's face region in the video frame.
[0390] For example, if the first frame after a video call is connected conforms to the default display mode, the other party's face region can be identified from the first frame after the video call is connected, and used as the reference face region. Then, the face regions identified in the subsequent second, third, and so on video frames can be matched with the reference face region identified in the first frame, thereby identifying the other party's face region in the second, third, and so on video frames.
[0391] It should be noted that the target video frame is not limited to the first frame alone. If the first N video frames are all displayed in the default mode, then any frame can be selected as the target video frame, and the identified face region of the other end in the target video frame can be used as the reference face region. For example, if the first 3 frames are all displayed in the default mode, then the face region of the other end can be identified from any frame, such as the first, second, or third frame, as the reference face region.
[0392] Method 2 for detecting the other end:
[0393] The face constraint conditions of the peer corresponding to the current frame are determined based on the number of window switching.
[0394] The number of window switching times refers to the number of times the display windows of the two parties in the video call interface are switched.
[0395] Specifically, with user authorization or video call application authorization, the mobile phone can also listen to and obtain the number of window switching, and update the face constraint conditions of the other end corresponding to the current frame based on the number of window switching.
[0396] It should be understood that in a video call interface, when the display window switches, generally only the content displayed within the window is adjusted, without changing the size or position of the window. Taking two display windows as an example, suppose display window 1 is the main display window and display window 2 is the secondary display window (i.e., display window 1 is larger than display window 2), and by default, the other party's face is displayed in display window 1. Upon the first window switch, the display will be adjusted to show the other party's face in display window 2. Upon the second window switch, the display will revert to showing the other party's face in display window 1. Upon the third window switch, the display will again revert to showing the other party's face in display window 2.
[0397] In this way, the mobile phone can determine the target display window (i.e., the display window used to display the current frame) of the other party's face in the video call interface based on the number of window switching. Then, based on the position of the target display window, it can determine the constraints on the other party's face, such as the required positional range (i.e., positional constraints) and / or the required size (i.e., size constraints) of the other party's face region. Furthermore, based on the determined constraints on the other party's face corresponding to the current frame, the mobile phone can accurately identify the other party's face region in the current frame.
[0398] Step S2803: Perform face swapping recognition based on the face region to obtain the face forgery detection result corresponding to the video frame.
[0399] Specifically, the mobile phone can input the face region in each video frame into the forgery detection model for inference, and obtain the face forgery detection result corresponding to that video frame.
[0400] Step S2804: Based on the face forgery detection results corresponding to the video frames, determine the face-swapping risk identification results and corresponding handling for this video call.
[0401] When face-swapping detection is performed using an active detection method, as shown in Figure 11 above, the user is unaware of the process before and during detection (steps S2801 to S2803). After detection, if the face-swapping risk identification result obtained in S2804 indicates a face-swapping risk, the phone will display prompt 5, providing risk feedback. Only then is the user aware of this process. For details regarding prompt 5, please refer to the illustrative description in Figure 12 above.
[0402] Next, using Figure 30 as an example, we will illustrate the interface changes when using the active detection method. Referring to Figure 30, after receiving David's friend request and adding him as a friend, if a video call request from David is received within a preset time (e.g., 5 minutes), face-swapping detection is performed on the video call. If a face-swapping risk is detected, a prompt is displayed, including risk warning information (i.e., the specific content of the "Risk Reminder" on the interface) as well as controls for "Re-detect" and "Screen Recording and Keep Video". After receiving the user's trigger on the "Re-detect" control, the phone can re-execute the face-swapping detection process, i.e., re-intercept and re-identify the face-swapping risk. After receiving the user's trigger on the "Screen Recording and Keep Video" control, the phone can start the screen recording function to record the video call. If the number of re-detections by the user reaches a preset threshold, the "Re-detect" control will no longer be displayed; instead, a "OK" button will be shown.
[0403] When using a passive detection method for face-swapping detection, as shown in Figure 16 above, interaction with the user can occur before, during, and after detection, allowing the user to perceive the detection process. Specific interactions can be found in the descriptions of Figures 14 to 27 above.
[0404] Next, using Figure 31 as an example of adding a friend, we will provide an overall illustration of the interface changes when using the passive detection method. It should be noted that in the video call interface of Figure 31, the man is the other end of the call, and the woman is the user on the other end. This is equivalent to the user switching display windows, showing the other end's image in a large window and the user's own image in a small window. As shown in Figure 31, when a risk is detected in the video call, the phone can display a security prompt in the capsule. Once triggered by the user, the phone can display specific risk protection information, including a "Detect Now" control or a "Start and Detect" control. When the user triggers these controls, face-swapping detection is performed. The capsule displays a "Start Detection" prompt in a minimized manner on the video call interface, and displays "Detecting" during the detection process. When the capsule is triggered by the user, it can expand to display a larger prompt area for a more complete message, such as "AI Face-Swapping Detection in Progress." As shown in Figure 31, during detection, detection animations can be added to the detected face area of the other end (i.e., the man's face area in Figure 31). After the detection is complete, it can display "No AI face-swapping risk detected". Further, after the user triggers the function or within a certain time, the display area of the capsule can be reduced, and a message "Detection complete" can be displayed.
[0405] In some examples, the phone may determine that the video call is at risk of face swapping only if the face forgery detection results of a single frame indicate that face forgery is at risk.
[0406] In other examples, the mobile phone can also combine the face spoofing detection results of multiple video frames to determine the face-swapping risk identification result and corresponding actions for the current video call. For example, the face spoofing detection result includes a face spoofing confidence score. The mobile phone can fuse the face spoofing confidence scores corresponding to multiple video frames to obtain a final confidence score. Based on this final confidence score, it can determine whether the current video call carries a face-swapping risk, obtaining a more accurate face-swapping risk identification result. The mobile phone can then take appropriate actions based on the face-swapping risk identification result.
[0407] Optionally, the mobile phone can fuse the face forgery confidence scores corresponding to multiple video frames using either direct weighted fusion or heuristic fusion.
[0408] (a) Direct weighted fusion refers to directly weighting and summing the confidence scores of face forgery corresponding to multiple captured video frames.
[0409] (ii) Heuristic fusion refers to selecting a portion of the face forgery confidence scores from multiple video frames and then weighting and fusing the selected face forgery confidence scores. The selected face forgery confidence scores are higher than those not selected; therefore, heuristic fusion based on the selected face forgery confidence scores can more accurately calculate the confidence scores.
[0410] In some embodiments, the processing flow of the video call method is described in conjunction with another software system architecture provided in FIG32.
[0411] Referring to Figure 32, in the application layer, the detection module includes a first detection module and a second detection module. The first detection module coordinates the identification and handling of face-swapping risks. This first detection module can be the system management service mentioned earlier. The first detection module includes a face-swapping detection module, a face-swapping risk identification and handling module, and a full-link identification module. The face-swapping detection module includes a detection scheduling module and a traffic interception module. The second detection module provides specific capabilities for face forgery detection, such as a face ROI detection module for providing face ROI detection capabilities and a forgery detection model for providing face forgery detection capabilities, which are then scheduled and used by the first detection module.
[0412] As shown in Figure 32, the processing flow of the video call method includes the following steps:
[0413] Step 0: The sensing module detects the risk event at time point t.
[0414] It should be understood that Figure 32 only uses a preset risk event as an example for illustration, and is not limited to triggering monitoring of suspicious video calls only upon sensing a risk event. It can also trigger monitoring of suspicious video calls upon sensing / detecting a friend-adding event or sensing / detecting a risk triggered by the entire link. Among them, the risk triggered by the entire link can be sensed / detected by the entire link identification module in Figure 32.
[0415] Step 1: The sensing module detects the video call event within the time period [t, t+δ].
[0416] Here, δ refers to the preset duration, such as the duration of 5 in Figure 13 above, which is not limited.
[0417] Step 2: The sensing module sends the video call event sensing results to the first detection module.
[0418] Step 3: The cut-through module calls the screenshot interface of the Window Manager to instruct the capture of video layer data from the Surface Flinger for video cut-through.
[0419] Step 4: The graphics synthesizer returns the cutoff buffer to the cutoff module.
[0420] Step 5: The detection scheduling module sends the captured video frames to the second detection module.
[0421] Specifically, the face ROI detection module in the second detection module can detect face regions in video frames, and the forgery detection model can perform face forgery detection inference based on the face regions detected by the face ROI detection module.
[0422] For example, the face ROI detection module can return the detected face region location coordinates to the face swap detection module in the first detection module (not shown in Figure 32), so that the detection scheduling module in the face swap detection module can send the face region to the forgery detection model in the second detection module based on the face region location coordinates, so as to call the forgery detection model to perform face forgery detection inference.
[0423] In other embodiments, after the face ROI detection module detects the coordinates of the face region, it can communicate directly within the second detection module so that the forgery detection model can obtain the face region and perform face forgery detection inference, without being limited to being scheduled by the face swap detection module in the first detection model.
[0424] Step 6: The forgery detection model feeds back the confidence result of face forgery to the face swap risk identification and handling module to instruct on face swap risk identification and handling.
[0425] In some embodiments, to facilitate a better understanding of the processing flow of the video call method, the following description is provided in conjunction with the timing diagram in Figure 33:
[0426] (1) The first detection module recognizes that the detection switch is turned on.
[0427] It should be understood that the detection switch (e.g., risk protection switch / face-swapping detection switch, etc.) is turned on (i.e. the face-swapping detection function is enabled), and the user does not need to trigger the start of detection (i.e. the user does not need to click a button or control to indicate the use of the face-swapping detection function to start detection). In other words, there is no need to interact with the user. The face-swapping detection processing logic is automatically executed in the background without the user's awareness.
[0428] (2) After the first detection module determines that the preset conditions are met, it starts listening for suspicious video calls within a preset time period.
[0429] (3) The perception module performs video call page recognition.
[0430] Specifically, within a preset time period, the sensing module can identify video call events through page recognition. For example, video call pages have specific interface features, and the sensing module can identify whether a video call event has been received based on the interface features displayed in the video call application.
[0431] (4) The sensing module sends the sensing fence result (i.e. video call event) to the first detection module.
[0432] (5) The first detection module starts to intercept the stream and acquire video frames.
[0433] (6) The first detection module reports the video frames to be inspected to the second detection module.
[0434] (7) The second detection module performs face ROI detection based on video frames.
[0435] (8) The second detection module returns the coordinates of the face region to the first detection module.
[0436] (9) The first detection module sends the face region to the second detection module.
[0437] (10) The second detection module performs forgery detection model reasoning.
[0438] Specifically, the second detection module uses a forgery detection model to perform face forgery inference based on the face region.
[0439] (11) The second detection module returns the face forgery reasoning result of the current frame to the first detection module.
[0440] (12) After the conditions for stopping the flow are met, the first detection module ends the flow interception.
[0441] (13) The second detection module ends detection and reasoning.
[0442] (14) The first detection module performs face-swapping risk identification and handling (such as high-risk pop-ups) based on the face forgery reasoning results of video frames.
[0443] In some examples, in addition to face-swapping detection, the video call method in this application embodiment also includes a full-link identification step. Specifically, the full-link identification step can be seen in steps (15) to (17) in Figure 33. It should be noted that full-link identification is not a necessary step in the video call method in this application embodiment and can be omitted in some embodiments. The full-link identification will be illustrated below with reference to steps (15) to (17) in Figure 33.
[0444] (15) The first detection module reports the risk behavior sequence.
[0445] If a face-swapping risk is detected, the first detection module can report the video call event as a risky behavior event, and report other detected risky behaviors (risk events) to the decision-making module. For example, if a face-swapping risk is detected, and it is also detected that a mobile user opens a payment application and enters a transfer interface, then both should be reported.
[0446] (16) The decision module identifies based on full-link rules and models.
[0447] (17) The decision module feeds back the full-link identification results to the first detection module.
[0448] Specifically, the decision-making module can perform end-to-end risk identification based on end-to-end rules or a trained end-to-end risk identification model, and on the reported risk behavior sequence. After obtaining the end-to-end risk identification result, the decision-making module can feed back the end-to-end identification result to the first detection module. For example, based on the reported risk behavior sequence, if the decision-making module identifies that a user first had a video call with a counterparty with a risk of face swapping, and then opened a payment application to transfer money, it can determine that there is a transfer risk, and then return the identification result to prompt the user.
[0449] The above solution can detect the risk of face swapping during a video call and then track and identify the user's subsequent risky behavior in a comprehensive and end-to-end manner, thereby providing risk protection to users on a larger scale.
[0450] It should be understood that steps (5) to (11) belong to the single-frame (single video frame) processing steps of cyclic frame-by-frame detection. The cyclic frame-by-frame detection steps, combined with steps (12) to (14) in Figure 33, are collectively referred to as the face-swapping detection and processing steps. It should be noted that the face-swapping detection and processing steps are not limited to the description of steps (5) to (14) in Figure 33. The above is only illustrated with Figure 33 and should not be construed as limiting the face-swapping detection and processing.
[0451] Please refer to Figure 34 for a more detailed explanation of the face-swapping detection and processing steps, which include the following steps:
[0452] (a) The first detection module notifies the screenshot interface to start intercepting traffic.
[0453] (b) Screenshot interface for frame-by-frame screenshotting.
[0454] (c) The screenshot interface returns a screenshot of the video frame layer to the first detection module.
[0455] (d) The first detection module requests the face detection module to call the face detection module and inputs the original video frame.
[0456] (e) The face detection module loads and initializes the face detection model.
[0457] (f) The face detection module performs face ROI detection based on the face detection model.
[0458] (g) The face detection module returns the position of the face rectangle to the first detection module.
[0459] (h) The first detection module performs face screening and cutout based on the position of the face rectangle.
[0460] (i) The first detection module requests the forgery detection model from the forgery reasoning module and inputs the face cutout.
[0461] (j) Use a forgery detection model to perform face forgery detection inference.
[0462] (k) The forgery reasoning module returns the reasoning result (i.e., the confidence level of face forgery in the current frame) to the first detection module.
[0463] (l) The first detection module performs multi-frame confidence heuristic fusion.
[0464] (m) The first detection module identifies and handles risks based on the confidence level of heuristic fusion.
[0465] The heuristic fusion processing can specifically include: a first detection module filtering out face forgery confidence scores higher than a confidence threshold, obtaining face forgery confidence scores for multiple selected video frames. Further, the first detection module can perform a weighted summation of the face forgery confidence scores from the multiple selected videos to obtain the final confidence score used for risk identification.
[0466] For example, the mobile phone can initialize an array of all zeros, with an array length equal to a preset number of throttled frames. The mobile phone can detect frames one by one. After detecting a valid frame (i.e., a valid video frame), it can refresh the array to store the corresponding confidence value. That is, for each video frame whose face spoofing confidence is calculated, it is updated and stored in the array. The face spoofing confidence values stored in the array are then weighted and fused (e.g., heuristic fusion). Based on the confidence value calculated by the weighted fusion, the risk of this video call is identified to determine whether there is a risk of face swapping on the other party in this video call.
[0467] In some examples, the mobile phone can update and store the confidence scores of all video frames that meet the threshold for frame interception in an array. Then, it can perform weighted fusion (such as direct weighted fusion or heuristic fusion) based on multiple confidence scores in the array to obtain the final confidence score. Finally, risk identification can be performed based on the final confidence score.
[0468] In other examples, the phone can capture a video frame at a time, update the face spoofing confidence score of that frame to an array, and then perform weighted fusion (e.g., direct weighted fusion or heuristic fusion) on the currently stored face spoofing confidence scores in the array. Risk identification is then performed based on the current weighted fusion confidence score. The risk identification results can include high risk (referred to as the first risk level), moderate / medium risk (referred to as the second risk level), and low / no risk (referred to as the third risk level). If a low or high risk is identified, the screenshot interface is notified to stop further interception and detection. Furthermore, if a high risk is predicted for the current video call, appropriate action is required, such as a pop-up notification. This scheme eliminates the need to detect all frames to obtain face-spoofing risk identification results, saving system resources and improving detection efficiency.
[0469] If a general / medium risk is identified, the next frame is captured and sent to the second detection module for face swap detection. This process is repeated frame by frame until a low or high risk is detected, or the preset number of frames to be cut off is met, at which point the loop process ends.
[0470] In some embodiments, the specific process of risk identification based on confidence level may include any of the following methods:
[0471] Method 1: Compare the confidence level with multiple preset risk confidence level intervals to determine the target risk confidence level interval. The risk level corresponding to the target risk confidence level interval is the risk level corresponding to this video call.
[0472] Method 2: Compare the confidence level with the first confidence threshold, the second confidence threshold, and the third confidence threshold respectively. If the confidence level is greater than or equal to the first confidence threshold, it is judged as high risk. If the confidence level is less than the first confidence threshold but greater than or equal to the second confidence threshold, it is judged as moderate risk. If the confidence level is less than the third confidence threshold, it is judged as low risk or no risk.
[0473] Please refer to Figure 35, which illustrates the processing steps of a video call method, taking passive detection during a video call as an example.
[0474] It should be understood that the processing steps illustrated in Figure 33 (i.e., the processing before the end of detection, including before and during detection) do not involve interaction between the front end and the user. Therefore, active detection (i.e., actively performing face-swapping detection) is performed during the video call. As can be seen from Figures 33 and 35, one of the main differences between the two lies in whether there is interaction with the user before performing face-swapping detection.
[0475] As shown in Figure 35, before triggering face-swapping detection processing (i.e., face-swapping detection and handling in Figure 35) (referred to as "before detection"), user interaction is required. For example, if the detection switch is not turned on, the user needs to actively turn on the detection switch through front-end interaction to enable the face-swapping detection function, thereby triggering the execution of subsequent face-swapping detection processing logic, as shown in Figure 15A above. Alternatively, if all detection switches are turned on, after detecting a suspicious video call event, face-swapping detection processing may not start directly, but the user may be prompted to actively trigger the detection through front-end interaction, as shown in Figure 14 above.
[0476] It should be understood that in passive detection scenarios, interaction with the user is not limited to "before detection," but can also include interaction during detection and interaction after detection.
[0477] Please refer to Figure 36, which uses passive detection as an example to more clearly illustrate how to interact with the user on the front end before and after detection.
[0478] As shown in Figure 36, before detection, referring to steps (4.1) to (4.2), the first detection module can instruct the notification manager to prompt the user to actively enable the detection switch, whether the detection switch is not turned on or is already turned on. For example, if the detection switch is not turned on, the module prompts the user to actively turn on the detection switch; if the detection switch is already turned on, the module prompts the user to actively enable face-swapping detection (i.e., actively trigger face-swapping detection). The prompt format can be seen in the pop-up prompt to the user in step (4.1). It should be understood that the user prompted by the mobile phone is the mobile phone user.
[0479] In some examples, if the first detection module receives a callback result (i.e., event notification) from the perception module within a preset time after starting listening, it can determine whether the detection switch (such as the AI face-swapping detection switch) is on through Settings. If the detection switch is not on, or if it is on, the RemoteView remote control software can be used to construct a dynamic capsule card, which is then used to notify the user through the NotificationManager. It should be understood that the dynamic capsule card is a notification tool that can switch between capsule and card formats. The specific processing is as follows:
[0480] (a) Function not enabled: that is, the detection switch is not turned on.
[0481] The first detection module can determine whether the current video call was initiated by a newly added friend by using the newFriendInfo information. If so, it instructs the notification manager to prompt the user to enable the detection switch via a pop-up window (such as a capsule card), for example, to enable the face-swapping detection function.
[0482] (ii) Function enabled: That is, the detection switch is turned on.
[0483] The first detection module can also instruct the notification manager to pop up a prompt for the user to actively trigger the start of detection. For example, in the first passive detection, a capsule card can be used to prompt the user whether to perform face-swapping detection on the current video call.
[0484] Furthermore, after the user actively enables (e.g., actively turns on the detection switch or actively triggers the start of detection) by performing step (4.2), the notification manager can trigger the first detection module to perform interception detection, i.e., see step (4.3).
[0485] As shown in Figure 36, after the face-swapping detection (face forgery detection) is completed, the first detection module can instruct the notification manager to update the detection results to the user via a pop-up window, i.e., refer to step (14.1) of Figure 36 for the pop-up window update of the detection results. The specific updated detection results can be seen in Figure 18 or Figure 20 above, for example, displaying "risky" or "no risk". As mentioned earlier, updating the detection results can also be applied to active detection, and is not limited to active detection only.
[0486] It should be noted that although the interaction during the passive detection process is not shown in Figure 36, in practice, when using passive detection to detect face swapping, the first detection module can also instruct the notification manager to provide corresponding prompts through pop-up windows.
[0487] Furthermore, during the passive detection process (i.e., in passive detection), in addition to displaying prompts in the pop-up window, the first detection module can also instruct the addition of face recognition detection animations to the video call interface, as shown in Figure 17. It should be understood that adding face recognition detection animations requires, from the perspective of underlying data processing, identifying the other party's face region and then adding the animations to that region. For details on how to identify the other party's face region, please refer to the implementation scheme described above; it will not be repeated here.
[0488] Please refer to Figure 36. In step (18), you can cancel the perception subscription. It should be noted that the cancellation of the perception subscription is not limited to passive detection scenarios. The cancellation of the perception subscription is also applicable to active detection scenarios. Figure 36 is only used as an example.
[0489] The triggers for canceling the awareness subscription include:
[0490] Timing 1: If no video call event is detected or listened to within a preset time after starting suspicious video call monitoring, the subscription to the sensing fence will be canceled.
[0491] Timing 2: If a video call event is detected within a preset time period, and face swap detection is performed after the first video call occurs (i.e., after the first received video call event and the video call is established), then the perception fence subscription can be canceled after the face swap detection is completed.
[0492] This application also provides a mobile terminal, which may include a display screen, a memory, and one or more processors (such as a CPU, GPU, NPU, etc.). The display screen, memory, and processor are coupled. The memory is used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the mobile terminal can perform various functions or steps performed by the device in the above method embodiments.
[0493] This application also provides a chip system including at least one processor and at least one interface circuit. The processor and the interface circuit are interconnected via lines. For example, the interface circuit can be used to receive signals from other devices (e.g., the memory of a mobile terminal). As another example, the interface circuit can be used to send signals to other devices (e.g., the processor). Exemplarily, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the mobile terminal can perform the steps in the above embodiments. Of course, the chip system may also include other discrete devices, and this application does not specifically limit this.
[0494] This embodiment also provides a computer storage medium storing computer instructions. When the computer instructions are executed on a mobile terminal, the mobile terminal performs the aforementioned method steps to implement the image processing method described above.
[0495] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the image processing method described above.
[0496] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component, or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the image processing methods in the above-described method embodiments.
[0497] In this embodiment, the mobile terminal, computer storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0498] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0499] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0500] The unit described as a separate component may or may not be physically separate. The component shown as a unit can be one physical unit or multiple physical units, that is, it can be located in one place or distributed in multiple different places. Some or all of the units can be selected to achieve the purpose of the solution in this embodiment according to actual needs.
[0501] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0502] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0503] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this application without departing from the spirit and scope of the technical solutions of this application.
Claims
1. A video call method, characterized in that, Applied to mobile terminals, including: Establish a video call and display the video call interface; If the first condition is met, detect whether there is a risk of face swapping in the first face in the video call interface, and display the first prompt after detecting that there is a risk of face swapping; If the second condition is met, a second prompt will be displayed, which is used to alert the user to the risk of face swapping. In response to the first operation of the second prompt, the system detects whether there is a risk of face swapping in the second face in the video call interface, prompts the detection process, and displays a third prompt after obtaining the detection result. The detection result includes whether there is a risk of face swapping or whether there is no risk of face swapping. The first condition and the second condition are different.
2. The method according to claim 1, characterized in that, The first condition or the second condition includes at least one of the following: The time difference between the moment the first user is added as a friend and the start time of the video call is within a first duration; the first user is an account logged in on the other end device that establishes the video call with the mobile terminal; The first user is the first friend added; The video call is the first video call with the first user; The video call is a video call initiated by the first user. The first user is in the detection list of the mobile terminal, and the detection list records the first number of users recently added to the mobile terminal; The mobile terminal is experiencing a risk event; The time difference between the second moment when the risk event is detected on the mobile terminal and the start time of the video call is within a second duration; The mobile terminal is deemed to pose a risk based on multiple related risk behavior factors. as well as The time difference between the third moment and the start time of the video call is within a third duration; wherein, the third moment is the moment when the mobile terminal is determined to be at risk based on multiple related risk behavior factors.
3. The method according to claim 1 or 2, characterized in that, The display of the first prompt includes: If the video call has not ended when a risk of face swapping is detected, the first prompt is displayed on the video call interface. The first prompt includes a re-detection control and / or a screen recording control. If the video call has ended when a risk of face swapping is detected, the first prompt is displayed on the interface after the video call ends. The first prompt does not include the re-detection control or the screen recording control. The re-detection control is used to trigger the mobile terminal to re-execute face-swapping detection; the screen recording control is used to record the screen of the video call interface after being triggered.
4. The method according to any one of claims 1-3, characterized in that, The test results indicate that there is no risk of face swapping. The third prompt displayed after obtaining the test results includes: If the video call has not ended when the detection result is obtained, the third prompt is displayed on the video call interface. The third prompt includes a re-detection control and / or a screen recording control. If the video call has ended when the detection result is obtained, the third prompt is displayed on the interface after the video call ends. The third prompt does not include the re-detection control or the screen recording control.
5. The method according to claim 3 or 4, characterized in that, The video call interface displays a prompt including a re-detection control, including: If the number of times the video call is checked for face-swapping risk does not exceed the first time, a prompt including a re-check control will be displayed on the video call interface.
6. The method according to any one of claims 3-5, characterized in that, The method further includes: If the number of times the video call is checked for face-swapping risk exceeds the first number, a prompt without a re-detection control will be displayed on the video call interface.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: If the second condition is met and the second face is not detected, a message indicating that the detection was not completed will be displayed; or, If the second condition is met, and the video call ends before the video frame of the video call is acquired, a message indicating that the detection was not completed will be displayed on the interface for ending the video call. Specifically, if the first condition is met and the first face is not detected, or if the video call ends before the video frame of the video call is acquired, no prompt indicating that the detection is not completed will be displayed.
8. The method according to any one of claims 1-7, characterized in that, The notification detection process includes: Display at least one of the following messages: a prompt to start detection, a prompt that is detecting, and a detection animation.
9. The method according to any one of claims 1-8, characterized in that, The display of the second prompt includes: When the second prompt is displayed from the 1st to the Nth time, the second prompt is displayed in the form of a card, which includes a description of the face-swapping detection function and / or a control to confirm the detection. When the second prompt is displayed for the N+1th time, the second prompt is displayed in capsule form, wherein the capsule-form second prompt does not include a function introduction of the face-swapping detection function or a control for confirming the detection; Where N is a natural number greater than 1.
10. The method according to any one of claims 1-9, characterized in that, The second condition includes the first sub-condition and the second sub-condition; If the second condition is met, a second prompt will be displayed, including: If the first sub-condition is met, a second prompt indicating the risk of face swapping is displayed. The first sub-condition includes the face swapping detection switch on the mobile terminal being turned on. If the second sub-condition is met, a second prompt will be displayed to enable the face-swapping detection function and perform detection. The second sub-condition includes turning off the face-swapping detection switch on the mobile terminal.
11. The method according to any one of claims 1-10, characterized in that, The detection of whether the first face in the video call interface is at risk of face swapping includes: Capture video frames from the video call interface; The captured video frames are subjected to face region detection, and the face regions in the detected face regions that meet the face constraints of the other end are identified as the first face; Face-swapping risk detection is performed based on the first face in the video frame.
12. The method according to claim 11, characterized in that, The face constraints at the other end include preset position constraints and / or preset size constraints.
13. The method according to claim 11 or 12, characterized in that, Before identifying the face region in the detected face region that meets the opposite face constraint condition as the first face, the method further includes: The number of window switching times during the capture of the video frame is obtained; the number of window switching times refers to the number of times the display windows of the two parties in the video call interface are exchanged during the capture of the video frame. Based on the number of window switching, the face constraint conditions of the counterparty corresponding to the video frame are determined.
14. The method according to any one of claims 11-13, characterized in that, The captured video frames are multiple; the step of performing face-swapping risk detection based on the first face in the video frames includes: For each captured video frame, face-swapping recognition is performed based on the first face in the video frame to obtain the face forgery confidence level corresponding to the video frame; The confidence scores of face forgery corresponding to the multiple video frames are fused to obtain the target confidence score; The presence of face-swapping risk is determined based on the target confidence level.
15. The method according to claim 14, characterized in that, The step of fusing the face forgery confidence scores corresponding to the multiple video frames to obtain the target confidence score includes: From the confidence scores of face forgery corresponding to the multiple video frames, a subset of video frames are selected to have higher confidence scores for face forgery than those not selected. The selected confidence scores for face forgery are weighted and fused to obtain the target confidence score.
16. The method according to claim 14, characterized in that, The step of fusing the face forgery confidence scores corresponding to the multiple video frames to obtain the target confidence score includes: After capturing each video frame and calculating the face forgery confidence score of that video frame, the face forgery confidence score corresponding to that video frame is updated and stored in an array. The face forgery confidence scores currently stored in the array are then weighted and fused to obtain the target confidence score. The step of determining whether there is a risk of face swapping based on the target confidence level includes: The risk level of face-swapping risk is identified based on the target confidence level; If the risk level is medium risk, then continue to capture the next video frame; If the risk level is low or high, stop capturing video frames, and if the risk level is high, it is determined that there is a risk of face swapping.
17. A mobile terminal, characterized in that, include: A display screen, one or more processors, and one or more memories; the one or more processors are coupled to the display screen and the one or more memories; the one or more memories are used to store computer program code, the computer program code including computer instructions, which, when executed by the one or more processors, cause the mobile terminal to perform the method as described in any one of claims 1-16.
18. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instructions are executed on the mobile terminal, the mobile terminal causes the mobile terminal to perform the method as described in any one of claims 1-16.
19. A computer program product comprising computer instructions, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1-16.
Citation Information
Patent Citations
Method of using flash short message to prevent phone fraud
CN106657538A
Face recognition data processing method and device
CN111783617A
Image detection method and device, computer equipment and storage medium
CN112749686A
Method, device and equipment for detecting AI face changing video
CN117789311A
System and Method for Face Tracking
US20100021008A1