Risk analysis and identification method and system based on big data
By sending camera parameters to the terminal of the video call to adjust the command sequence and judging the video screen in real time, the problem of difficult AI fake videos in video calls is solved, and the prevention ability of telecom fraud and the authenticity identification ability of video calls is improved.
Patent Information
- Application Number
- CN202510241107.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The prior art is difficult to identify video images forged by AI in real time during video calls, making it difficult to prevent telecommunications fraud.
By sending a sequence of camera parameter adjustment commands to the other terminal in the video call, controlling its camera parameter adjustment, and obtaining video clips within the command execution cycle, making real shots and judgments on the video screen at the time of each command triggering to determine whether the video screen is real shot.
It has realized the effective identification of AI forged video images, improved the prevention ability of telecom fraud, enhanced the authenticity identification ability of video calls, and protected the safety of user property and social security.
Smart Images

Figure CN120017783A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of risk identification, and in particular to a risk analysis and identification method and system based on big data. Background Art
[0002] With the advent of the intelligent era, artificial intelligence technology (AI) is increasingly being used in various fields, but it is also being used by criminals. Especially in the field of telecommunications network fraud, fraudsters abuse AI technology to conduct extremely deceptive fraud activities by forging faces and voices.
[0003] In the process of telecommunications fraud, fraudsters often use AI technology to forge human faces and make video calls with the deceived, in order to gain trust and commit fraud. This fraud method is extremely deceptive, because it is often difficult for victims to tell with the naked eye whether the other party in the video call is a real person. Existing technologies can often only authenticate the authenticity of video images generated by AI technology, but it is difficult to perform real-time analysis during a video call. Summary of the invention
[0004] The purpose of this application is to provide a risk analysis and identification method and system based on big data, which can improve the above-mentioned problems.
[0005] The embodiment of the present application is implemented as follows: In the first aspect, the present application provides a risk analysis and identification method based on big data, which includes steps S1 to S3. Among them, S1, S2, etc. are only step identifiers, and the execution order of the method is not necessarily in the order from small to large numbers. For example, step S2 may be executed first and then step S1, and the present application does not limit it.
[0006] S1, in response to a risk identification operation, sending a camera parameter adjustment instruction sequence to a terminal of a party currently conducting a video call, for controlling the camera parameter adjustment of the terminal of the party; S2, in response to the other party terminal receiving the camera parameter adjustment instruction sequence, obtaining the other party's video clips within the execution period of the camera parameter adjustment instruction sequence, and performing real-shot judgment on the video pictures corresponding to each instruction triggering moment to determine whether the video pictures are real-shot pictures; S3, when it is determined that any of the video images in the other party's video clips is not a real-shot image, display risk warning information.
[0007] It can be understood that the present application provides a risk analysis and identification method based on big data, which controls the adjustment of camera parameters by sending a camera parameter adjustment command sequence to the other terminal of the video call. The other party's video clip is obtained during the command execution cycle, and the video screen at the time of each command triggering is judged as a real shot. The judgment is based on the picture change characteristics when the command is triggered, including wide-angle mode switching, telephoto mode switching, and white balance setting. If any picture is judged to be non-real shot, a risk warning message is displayed. This method can effectively identify AI-forged video pictures and improve the ability to prevent telecommunications fraud. By adjusting camera parameters in real time and analyzing picture changes, the authenticity identification ability of video calls is enhanced, providing a new technical means for combating telecommunications fraud, which helps to protect user property safety and social security.
[0008] In an optional embodiment of the present application, the camera parameter adjustment instruction sequence includes at least one of the following instructions.
[0009] Wide-angle mode switching command, used to instruct the camera to shorten its focal length to a short focal length preset value; Telephoto mode switching command, used to command the camera's focal length to increase to a telephoto preset value; The white balance setting command is used to instruct the camera to adjust the current white balance value to the preset white balance value.
[0010] The real-shot judgment of the video pictures corresponding to each instruction triggering moment includes obtaining a first picture corresponding to the wide-angle mode switching instruction triggering moment and a second picture that is a frame before the first picture, and judging that the first picture is a real-shot picture when all items of the following conditions 1-3 are met, otherwise, judging that the first picture is a non-real-shot picture.
[0011] Condition 1: There is non-pure color display content in other areas of the first picture except the second picture. After switching to wide-angle mode, the viewing angle of the picture captured by the real camera will increase significantly, so in addition to the content of the previous picture, more details of the surrounding environment will appear in the new picture. However, since it is difficult to simulate this change in viewing angle in a short period of time, the AI-generated picture may not be able to quickly generate the corresponding extended content. This judgment condition is based on the increase in the viewing angle of the picture in wide-angle mode, and determines whether it is a real shot by comparing the picture content before and after the switch.
[0012] Condition 2: If the short focal length preset value is less than the first focal length value, the edge of the first picture is distorted. In wide-angle mode, due to the shortened focal length, a certain degree of barrel distortion may occur at the edge of the picture. This is an inherent characteristic of optical lens imaging, which cannot be avoided by real cameras. However, AI-generated pictures may have difficulty in simulating such distortion, especially when switching focal lengths in a short period of time. Therefore, by detecting whether the expected distortion occurs at the edge of the picture, the authenticity of the picture can be judged.
[0013] Condition 3: Determine the first reference element in the second picture, the area ratio of the face area in the second picture to the first reference element area is a second ratio, the area ratio of the face area in the first picture to the first reference element area is a third ratio, and the second ratio is equal to the third ratio. Before and after the wide-angle mode is switched, the area ratio of the elements in the picture (such as faces, objects, etc.) to the background or other fixed reference elements should remain consistent. This is because the wide-angle mode only changes the viewing angle, but does not change the actual size relationship between objects. If the AI-generated picture fails to accurately maintain this proportional relationship when simulating the wide-angle effect, its authenticity can be judged by comparing the pictures before and after the switch.
[0014] In an optional embodiment of the present application, the real-shot judgment of the video pictures corresponding to each command triggering moment includes obtaining a third picture corresponding to the telephoto mode switching command triggering moment, a fourth picture that is one frame before the third picture, and a fifth picture that is n frames after the third picture, where n is a positive integer greater than or equal to 1. If all of the following conditions 4-6 are met, the third picture is judged to be a real-shot picture; otherwise, the third picture is judged to be a non-real-shot picture.
[0015] Condition 4, determining the second reference element in the third picture, the area ratio of the face area in the fourth picture to the second reference element area is a fourth ratio, the area ratio of the face area in the third picture to the second reference element area is a fifth ratio, and the fourth ratio is equal to the fifth ratio. It can be understood that the telephoto mode narrows the viewing angle by increasing the focal length of the camera. Although the viewing angle is narrowed, the area ratio of the elements in the picture (such as faces, objects, etc.) to the background or other fixed reference elements should remain consistent in real shooting. This is because the telephoto mode only magnifies the details in the picture without changing the actual size ratio relationship between objects. If the AI-generated picture fails to accurately maintain this proportional relationship when simulating the telephoto effect, its authenticity can be judged by comparing the pictures before and after the switch.
[0016] Condition 5: If the preset value of the telephoto focal length is greater than the second focal length value, the clarity of the third picture is less than the first clarity value. It can be understood that when the preset value of the telephoto focal length is greater than a certain threshold, the clarity of the actual picture will be significantly reduced. In the picture generated by AI simulated telephoto, this degree of blur may be difficult to generate in time. Therefore, this condition 5 can be used as a judgment item.
[0017] Condition 6: The content of the fifth picture is the same as that of the third picture, and the clarity of the fifth picture is greater than that of the third picture. It can be understood that when the preset value of the telephoto focal length is large, after the camera switches from wide-angle to telephoto mode, the initially captured picture may be slightly blurred due to the need to adjust the focal length to achieve clear focus. As the autofocus process is completed, the picture clarity will gradually improve. AI-generated pictures may have difficulty simulating such clarity changes because they find it difficult to simulate the camera's focusing process in real time.
[0018] In an optional embodiment of the present application, the real-shot judgment of the video pictures corresponding to each instruction triggering moment includes obtaining the sixth picture corresponding to the white balance setting instruction triggering moment and the seventh picture in the previous frame of the sixth picture, and judging that the sixth picture is a real-shot picture when the following condition 7 or condition 8 is met, otherwise, judging that the sixth picture is a non-real-shot picture.
[0019] Condition 7, if the white balance preset value is greater than the first color temperature value, the color value of each pixel in the sixth picture is extracted, the average value of the red color value of each pixel in the sixth picture is calculated as the first color value, the color value of each pixel in the seventh picture is extracted, and the average value of the red color value of each pixel in the seventh picture is calculated as the second color value, and the first color value is greater than the second color value. It can be understood that the white balance function of the camera is used to adjust the color balance to adapt to the shooting environment under different light sources. When the white balance value is increased, the picture will gradually tend to yellow, especially the red component will increase. This change can be naturally presented in a real camera because the camera sensor will adjust the color sensitivity accordingly. For the AI-generated picture, since its generation mechanism is different from that of a real camera, it may be difficult to accurately simulate the color change after the white balance adjustment, especially the increase in the red component. Therefore, by comparing the average red color before and after the white balance adjustment, the authenticity of the picture can be judged.
[0020] Condition 8: If the preset white balance value is less than the second color temperature value, extract the color value of each pixel in the sixth picture, calculate the average value of the blue color value of each pixel in the sixth picture as the third color value, extract the color value of each pixel in the seventh picture, calculate the average value of the blue color value of each pixel in the seventh picture as the fourth color value, and the third color value is greater than the fourth color value. It can be understood that, similar to judgment condition 7, when the white balance value of the camera is adjusted to a smaller value, the picture will gradually tend to blue, especially the blue component will increase. A real camera can naturally present this color change, while the picture generated by AI may be difficult to accurately simulate. By comparing the average value of the blue color before and after the white balance adjustment, the authenticity of the picture can be further verified.
[0021] In an optional embodiment of the present application, the risk analysis and identification method based on big data further includes the following steps between step S2 and step S3: S41, in response to the risk identification operation, obtaining the terminal model of the counterpart terminal, and if the terminal model supports the depth perception function, sending a depth perception function activation prompt to the counterpart terminal; S42, receiving the portrait area depth information and the structured light raw data of the portrait area fed back by the other party terminal, and calculating the calculated depth information of the portrait area according to the structured light raw data; S43, comparing the calculated depth information with the fed-back depth information of the portrait area, if the comparison is consistent, determining that the video picture is a real shot picture, if the comparison is inconsistent, determining that the video picture is a non-real shot picture.
[0022] It can be understood that this application has added the judgment and utilization of depth perception function to the above-mentioned AI raw image judgment technology. If the other terminal supports depth perception function, turn on this function and obtain the depth information and structured light raw data of the portrait area. By calculating the calculated depth information of the portrait area and comparing it with the feedback depth information, it is determined whether the video picture is a real shot. This method uses depth perception technology to further improve the accuracy of authenticity identification of video pictures, combines depth perception and structured light technology, provides a more reliable basis for authenticity identification of video calls, and effectively improves the prevention and combat capabilities of telecommunications fraud.
[0023] In an optional embodiment of the present application, if the other terminal has an infrared structured light projector and an infrared camera, it is determined that the terminal model supports the depth perception function; the structured light raw data includes the original sinusoidal wave pattern projected by the infrared structured light projector and the actual pattern captured by the infrared camera after being modulated by the human face.
[0024] In an optional embodiment of the present application, the calculating depth information of the portrait area according to the structured light raw data includes: calculating the depth difference of each pixel in the portrait area according to the following formula: ; in, Represents the first The depth difference of pixels relative to the infrared structured light projector and the infrared camera baseline, represents the wavelength of the light beam emitted by the infrared structured light projector, represents the highest spatial frequency that the infrared camera can clearly resolve, Represents the phase difference between the original sine wave pattern and the actual pattern.
[0025] It can be understood that one of the key steps in judging the authenticity of the video image through the depth perception function is to calculate the calculated depth information of the portrait area based on the original structured light data. This process mainly relies on the collaborative work of the infrared structured light projector and the infrared camera, as well as the application of the sine wave pattern. Specifically, the infrared structured light projector projects a sine wave pattern onto the face. This pattern is modulated on the surface of the object and then captured by the infrared camera. There is a phase difference between the actual pattern captured by the infrared camera and the original sine wave pattern, and this phase difference is directly related to the depth information of the surface of the object. In order to calculate the depth information, we can use the principle of triangulation. Based on the relative position of the infrared structured light projector and the infrared camera (i.e., the baseline distance), the wavelength of the projector's output light beam, and the highest spatial frequency that the infrared camera can clearly distinguish, we can calculate the depth difference of each pixel relative to the baseline based on the phase difference. This calculation process involves complex mathematical models and algorithms, but the core is to use the key parameter of phase difference to derive depth information. In this way, we can obtain detailed depth information of the portrait area and then compare it with the depth information fed back by the other terminal. If the two are consistent, it means that the video footage is likely to be real footage; if they are inconsistent, there is a risk of fraud.
[0026] In a second aspect, the present application discloses an electronic device, comprising: An instruction sequence generation module, configured to send a camera parameter adjustment instruction sequence to a terminal of a party currently conducting a video call in response to a risk identification operation, so as to control the camera parameter adjustment of the terminal of the party; A judgment module, configured to obtain, in response to the other terminal receiving the camera parameter adjustment instruction sequence, a video clip of the other party within the execution cycle of the camera parameter adjustment instruction sequence, and perform a real-shot judgment on the video picture corresponding to each instruction triggering moment to determine whether the video picture is a real-shot picture; The prompt module is used to display risk warning information when it is determined that any video image in the other party's video clip is not a real-shot image.
[0027] In a third aspect, the present application discloses a risk analysis and identification system based on big data, comprising a processor, an input device, an output device and a memory, wherein the processor, input device, output device and memory are interconnected, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute a method as described in any one of the first aspects.
[0028] Beneficial effects: The present application provides a risk analysis and identification method and system based on big data, which controls the adjustment of the camera parameters by sending a camera parameter adjustment command sequence to the terminal of the other party of the video call. The video clip of the other party is obtained during the command execution cycle, and the video screen at the time of each command triggering is judged as a real shot. The judgment is based on the characteristics of the picture changes when the commands such as wide-angle mode switching, telephoto mode switching and white balance setting are triggered. If any picture is judged to be not a real shot, a risk warning message is displayed. This method can effectively identify AI-forged video pictures and improve the ability to prevent telecommunications fraud. By adjusting the camera parameters in real time and analyzing the changes in the picture, the authenticity identification ability of the video call is enhanced, which provides a new technical means for combating telecommunications fraud and helps protect user property safety and social security.
[0029] This application adds the judgment and use of depth perception function to the above-mentioned AI raw image judgment technology. If the other terminal supports depth perception function, turn on this function and obtain the depth information and structured light raw data of the portrait area. By calculating the calculated depth information of the portrait area and comparing it with the feedback depth information, it is determined whether the video picture is a real shot. This method uses depth perception technology to further improve the accuracy of authenticity identification of video pictures, and combines depth perception and structured light technology to provide a more reliable basis for authenticity identification of video calls.
[0030] In order to make the above-mentioned objects, features and advantages of the present application more obvious and understandable, optional embodiments are specifically listed below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0032] Figure 1It is a flowchart of a risk analysis and identification method based on big data provided by this application; Figure 2 is a schematic diagram of a camera parameter adjustment instruction sequence provided by the present application; Figure 3 This is a schematic diagram of the risk warning scenario provided by this application; Figure 4 It is a flow chart of another risk analysis and identification method based on big data provided by this application. DETAILED DESCRIPTION
[0033] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0034] In the first aspect, the present application provides a risk analysis and identification method based on big data, which includes steps S1 to S3. Among them, S1, S2, etc. are only step identifiers, and the execution order of the method is not necessarily in the order from small to large numbers. For example, step S2 may be executed first and then step S1, and the present application does not limit it.
[0035] S1, in response to a risk identification operation, sending a camera parameter adjustment instruction sequence to a terminal of a party currently conducting a video call, so as to control the camera parameter adjustment of the terminal of the party.
[0036] The above-mentioned risk identification operation refers to an action actively triggered by the user, which is intended to start a risk analysis and identification method based on big data. For example, during a video call, when the user suspects that the other party's picture may be forged by AI, a specific operation (such as clicking a button) can be used to send a camera parameter adjustment command sequence to identify and prevent risks.
[0037] The above-mentioned camera parameter adjustment instruction sequence is a set of instructions for controlling the adjustment of the other party's terminal camera parameters during a video call. The instruction sequence includes multiple instructions, such as wide-angle mode switching instructions, telephoto mode switching instructions, and white balance setting instructions, etc., which aims to enhance the authenticity authentication capability of video calls by adjusting camera parameters in real time and analyzing picture changes. Figure 2 The camera parameter adjustment instruction sequence shown includes a wide-angle mode switching instruction set at time t1, a telephoto mode switching instruction set at time t2, and a white balance setting instruction set at time t3.
[0038] S2, in response to the other party's terminal receiving a camera parameter adjustment instruction sequence, obtaining the other party's video clips within the execution cycle of the camera parameter adjustment instruction sequence, and performing a real-shot judgment on the video pictures corresponding to each instruction triggering moment to determine whether the video pictures are real-shot pictures.
[0039] S3, when it is determined that any video frame in the other party's video clip is not a real shot frame, display risk warning information.
[0040] It can be understood that the present application provides a risk analysis and identification method based on big data, which controls the adjustment of camera parameters by sending a camera parameter adjustment command sequence to the other terminal of the video call. The other party's video clip is obtained during the command execution cycle, and the video screen at the time of each command triggering is judged as a real shot. The judgment is based on the picture change characteristics when the command is triggered, including wide-angle mode switching, telephoto mode switching, and white balance setting. If any picture is judged to be non-real shot, a risk warning message is displayed. This method can effectively identify AI-forged video pictures and improve the ability to prevent telecommunications fraud. By adjusting camera parameters in real time and analyzing picture changes, the authenticity identification ability of video calls is enhanced, providing a new technical means for combating telecommunications fraud, which helps to protect user property safety and social security.
[0041] In an optional embodiment of the present application, the camera parameter adjustment instruction sequence includes at least one of the following instructions.
[0042] The wide-angle mode switching instruction is used to command the camera's focal length to be shortened to a preset value of short focal length. The wide-angle mode switching instruction is sent to the camera control module of the other terminal device in the form of a data packet through a network communication protocol (such as TCP / IP). The instruction contains a clear instruction code that instructs the camera to switch from the current mode to the wide-angle mode. At the software level, after the operating system of the other terminal device receives the instruction, it passes it to the camera driver. The camera driver parses the instruction code and calls the corresponding hardware control function. At the hardware level, after the camera module receives the control signal, it drives the lens group to move through the internal motor or stepper motor to shorten the focal length, thereby achieving wide-angle shooting. This process is real-time and fast to ensure the smoothness of the video call. The wide-angle mode enables the camera to capture a wider field of view by changing the focal length of the camera. This change is based on the principle of optical lens imaging, and the focal length and viewing angle of the imaging are changed by adjusting the position of the lens group. In a video call, the authenticity of the other party's picture can be quickly verified by switching the wide-angle mode, because the picture generated by AI may not be able to adapt to this change in time when the focal length is quickly switched.
[0043] The telephoto mode switching instruction is used to command the camera's focal length to increase to the preset telephoto value. The telephoto mode switching instruction is similar to the wide-angle mode switching instruction, and is also sent to the camera control module of the other terminal device through the network communication protocol. The difference is that the instruction instructs the camera to switch to telephoto mode, that is, to increase the focal length to capture distant details. At the software level, the operating system receives the instruction and passes it to the camera driver, which parses the instruction and calls the hardware control function. At the hardware level, the camera module drives the lens group to move in the opposite direction through a motor or a stepper motor to increase the focal length. This process is also real-time and fast to ensure the continuity of the video call. The telephoto mode reduces the viewing angle by increasing the focal length of the camera, thereby improving the clarity and detail expression of distant objects. This principle is based on the imaging law of optical lenses, and the focal length and magnification of the imaging are changed by adjusting the position of the lens group. In a video call, switching to telephoto mode can further verify the authenticity of the other party's picture, because the AI-generated picture may not be able to present details consistent with the real scene when the focal length is increased.
[0044] The white balance setting instruction is used to command the camera's current white balance value to be adjusted to the white balance preset value. The white balance setting instruction is sent to the camera control module of the other terminal device through the network communication protocol. The instruction contains a specific white balance preset value, instructing the camera to adjust its white balance setting to match the value. At the software level, the operating system receives the instruction and passes it to the camera driver. The driver parses the instruction and calls the corresponding white balance adjustment function. At the hardware level, the camera module adjusts its internal color correction algorithm according to the instruction to compensate for the color temperature difference under different light sources. This process involves real-time color adjustment of the raw data captured by the image sensor. White balance is a key parameter for the accuracy of camera color reproduction. Different light sources have different color temperatures, and the white balance setting ensures that the camera can accurately restore colors under various lighting conditions. By adjusting the white balance value, the overall hue and color saturation of the image can be changed. In a video call, changing the white balance setting can observe the color changes of the picture, because the AI-generated picture may not fully imitate the color changes of the real scene when adjusting the color, thus providing a basis for judging the authenticity of the picture.
[0045] Perform real shot judgment on the video images corresponding to each command triggering moment, including obtaining the first image corresponding to the wide-angle mode switching command triggering moment and the second image in the previous frame of the first image, and judging the first image as a real shot image if all of the following conditions 1-3 are met, otherwise, judging the first image as a non-real shot image. Figure 2 As shown, the first picture P1 corresponding to the wide-angle mode switching instruction triggering moment and the second picture P2 in the previous frame of the first picture P1 are judged to be the real shot picture when all of the following conditions 1-3 are met.
[0046] Condition 1: There is non-pure color display content in other areas of the first picture except the second picture. It can be understood that after switching to wide-angle mode, the viewing angle of the picture captured by the real camera will increase significantly, so in addition to the content of the previous picture, more details of the surrounding environment will appear in the new picture. However, since it is difficult to simulate this change in viewing angle in a short period of time, the AI-generated picture may not be able to quickly generate the corresponding extended content. This judgment condition is based on the increase in the viewing angle of the picture in wide-angle mode, and determines whether it is a real shot by comparing the picture content before and after the switch.
[0047] Condition 2: If the short focal length preset value is less than the first focal length value, the edge of the first picture is distorted. It is understandable that in wide-angle mode, due to the shortened focal length, a certain degree of barrel distortion may occur at the edge of the picture. This is an inherent characteristic of optical lens imaging, which cannot be avoided by real cameras. The AI-generated picture may have difficulty in simulating this distortion, especially when switching focal lengths in a short period of time. Therefore, by detecting whether the expected distortion occurs at the edge of the picture, the authenticity of the picture can be judged.
[0048] Condition 3: Determine the first reference element in the second picture, the area ratio of the face area to the first reference element area in the second picture is the second ratio, the area ratio of the face area to the first reference element area in the first picture is the third ratio, and the second ratio is equal to the third ratio. It can be understood that before and after the wide-angle mode is switched, the area ratio of the elements in the picture (such as faces, objects, etc.) to the background or other fixed reference elements should remain consistent. This is because the wide-angle mode only changes the viewing angle, but does not change the actual size relationship between objects. If the AI-generated picture fails to accurately maintain this proportional relationship when simulating the wide-angle effect, its authenticity can be judged by comparing the pictures before and after the switch.
[0049] In an optional embodiment of the present application, the video images corresponding to each command triggering moment are judged to be real shot, including obtaining the third image corresponding to the telephoto mode switching command triggering moment, the fourth image in the previous frame of the third image, and the fifth image n frames after the third image, where n is a positive integer greater than or equal to 1. If all of the following conditions 4-6 are met, the third image is judged to be a real shot image, otherwise, the third image is judged to be a non-real shot image. Figure 2 As shown, the third picture P3 corresponding to the moment when the telephoto mode switching instruction is triggered, the fourth picture P4 in the previous frame of the third picture P3, and the fifth picture P5 in the next few frames after the third picture P3, are judged to be the real shot picture when all the following conditions 4-6 are met.
[0050] Condition 4, determine the second reference element in the third picture, the area ratio of the face area to the second reference element area in the fourth picture is the fourth ratio, the area ratio of the face area to the second reference element area in the third picture is the fifth ratio, and the fourth ratio is equal to the fifth ratio. It can be understood that the telephoto mode narrows the viewing angle by increasing the focal length of the camera. Although the viewing angle is narrowed, the area ratio of the elements in the picture (such as faces, objects, etc.) to the background or other fixed reference elements should remain consistent in real shooting. This is because the telephoto mode only magnifies the details in the picture without changing the actual size ratio relationship between objects. If the AI-generated picture fails to accurately maintain this proportional relationship when simulating the telephoto effect, its authenticity can be judged by comparing the pictures before and after the switch.
[0051] Condition 5: If the preset value of the telephoto focal length is greater than the second focal length value, the clarity of the third picture is less than the first clarity value. It can be understood that when the preset value of the telephoto focal length is greater than a certain threshold, the clarity of the actual picture will be significantly reduced. In the picture generated by AI simulated telephoto, this degree of blur may be difficult to generate in time. Therefore, this condition 5 can be used as a judgment item.
[0052] Condition 6: The content of the fifth picture is the same as that of the third picture, and the clarity of the fifth picture is greater than that of the third picture. It is understandable that when the preset value of the telephoto focal length is large, after the camera switches from wide-angle to telephoto mode, the initially captured picture may be slightly blurred because the focal length needs to be adjusted to achieve clear focus. As the autofocus process is completed, the picture clarity will gradually improve. AI-generated pictures may have difficulty simulating this change in clarity because they find it difficult to simulate the camera's focusing process in real time.
[0053] In an optional embodiment of the present application, the video images corresponding to each instruction triggering moment are judged to be real shot, including obtaining the sixth image corresponding to the white balance setting instruction triggering moment and the seventh image in the previous frame of the sixth image, and judging the sixth image to be a real shot image when the following condition 7 or condition 8 is met, otherwise, judging the sixth image to be a non-real shot image. Figure 2 As shown, the sixth picture P6 corresponding to the triggering moment of the white balance setting instruction and the seventh picture P7 in the previous frame of the sixth picture P6 are judged to be the real shot picture when the following condition 7 or condition 8 is met.
[0054] Condition 7, if the white balance preset value is greater than the first color temperature value, extract the color value of each pixel in the sixth picture, calculate the average value of the red color value of each pixel in the sixth picture as the first color value, extract the color value of each pixel in the seventh picture, calculate the average value of the red color value of each pixel in the seventh picture as the second color value, and the first color value is greater than the second color value. It can be understood that the white balance function of the camera is used to adjust the color balance to adapt to the shooting environment under different light sources. When the white balance value is increased, the picture will gradually tend to yellow, especially the red component will increase. This change can be naturally presented in a real camera because the camera sensor will adjust the color sensitivity accordingly. For the AI-generated picture, since its generation mechanism is different from that of a real camera, it may be difficult to accurately simulate the color change after the white balance adjustment, especially the increase in the red component. Therefore, by comparing the average red color before and after the white balance adjustment, the authenticity of the picture can be judged.
[0055] Condition 8: If the preset white balance value is less than the second color temperature value, extract the color value of each pixel in the sixth picture, calculate the average value of the blue color value of each pixel in the sixth picture as the third color value, extract the color value of each pixel in the seventh picture, calculate the average value of the blue color value of each pixel in the seventh picture as the fourth color value, and the third color value is greater than the fourth color value. It can be understood that, similar to judgment condition 7, when the white balance value of the camera is adjusted to a smaller value, the picture will gradually tend to blue, especially the blue component will increase. Real cameras can naturally present this color change, while AI-generated pictures may be difficult to accurately simulate. By comparing the average value of the blue color before and after the white balance adjustment, the authenticity of the picture can be further verified.
[0056] In an optional embodiment of the present application, in the risk analysis and identification method based on big data, the following steps S41 to S43 are also included between step S2 and step S3.
[0057] S41, in response to the risk identification operation, obtaining the terminal model of the other party's terminal, and if the terminal model supports the depth perception function, sending a depth perception function activation prompt to the other party's terminal.
[0058] In an optional embodiment of the present application, if the other terminal has an infrared structured light projector and an infrared camera, it is determined that the terminal model supports the depth perception function; the structured light raw data includes the original sinusoidal wave pattern projected by the infrared structured light projector and the actual pattern captured by the infrared camera after being modulated by the human face.
[0059] S42, receiving the portrait area depth information and the structured light raw data of the portrait area fed back by the other terminal, and calculating the calculated depth information of the portrait area according to the structured light raw data.
[0060] In an optional embodiment of the present application, calculating the depth information of the portrait area according to the structured light raw data includes: calculating the depth difference of each pixel in the portrait area according to the following formula: ; in, In the representative portrait area The depth difference of pixels relative to the infrared structured light projector and the infrared camera baseline, Represents the wavelength of the outgoing light beam of the infrared structured light projector, Represents the highest spatial frequency that an infrared camera can clearly resolve. Represents the phase difference between the original sine wave pattern and the actual pattern.
[0061] It can be understood that one of the key steps in judging the authenticity of the video image through the depth perception function is to calculate the calculated depth information of the portrait area based on the original structured light data. This process mainly relies on the collaborative work of the infrared structured light projector and the infrared camera, as well as the application of the sine wave pattern. Specifically, the infrared structured light projector projects a sine wave pattern onto the face. This pattern is modulated on the surface of the object and then captured by the infrared camera. There is a phase difference between the actual pattern captured by the infrared camera and the original sine wave pattern, and this phase difference is directly related to the depth information of the surface of the object. In order to calculate the depth information, we can use the principle of triangulation. Based on the relative position of the infrared structured light projector and the infrared camera (i.e., the baseline distance), the wavelength of the projector's output light beam, and the highest spatial frequency that the infrared camera can clearly distinguish, we can calculate the depth difference of each pixel relative to the baseline based on the phase difference. This calculation process involves complex mathematical models and algorithms, but the core is to use the key parameter of phase difference to derive depth information. In this way, we can obtain detailed depth information of the portrait area and then compare it with the depth information fed back by the other terminal. If the two are consistent, it means that the video footage is likely to be real footage; if they are inconsistent, there is a risk of fraud.
[0062] S43, comparing the calculated depth information with the fed-back portrait area depth information, if the comparison is consistent, determining that the video picture is a real shot picture, if the comparison is inconsistent, determining that the video picture is a non-real shot picture.
[0063] It can be understood that this application has added the judgment and utilization of depth perception function to the above-mentioned AI raw image judgment technology. If the other terminal supports depth perception function, turn on this function and obtain the depth information and structured light raw data of the portrait area. By calculating the calculated depth information of the portrait area and comparing it with the feedback depth information, it is determined whether the video picture is a real shot. This method uses depth perception technology to further improve the accuracy of authenticity identification of video pictures, combines depth perception and structured light technology, provides a more reliable basis for authenticity identification of video calls, and effectively improves the prevention and combat capabilities of telecommunications fraud.
[0064] In a second aspect, the present application discloses an electronic device, comprising: An instruction sequence generation module, configured to send a camera parameter adjustment instruction sequence to a terminal of a party currently conducting a video call in response to a risk identification operation, so as to control the camera parameter adjustment of the terminal of the party; A judgment module, configured to obtain, in response to the other terminal receiving the camera parameter adjustment instruction sequence, a video clip of the other party within the execution cycle of the camera parameter adjustment instruction sequence, and perform a real-shot judgment on the video picture corresponding to each instruction triggering moment to determine whether the video picture is a real-shot picture; The prompt module is used to display risk warning information when it is determined that any video image in the other party's video clip is not a real-shot image.
[0065] In a third aspect, the present application provides a risk analysis and identification system based on big data. The risk analysis and identification system based on big data includes one or more processors; one or more input devices, one or more output devices and a memory. The above-mentioned processor, input device, output device and memory are connected via a bus. The memory is used to store a computer program, which includes program instructions, and the processor is used to execute the program instructions stored in the memory. Among them, the processor is configured to call the program instructions to perform the operation of any method of the first aspect: It should be understood that in the embodiments of the present invention, the processor referred to may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0066] Input devices may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint direction information), a microphone, etc., and output devices may include a display (LCD, etc.), a speaker, etc.
[0067] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0068] In a specific implementation, the processor, input device, and output device described in the embodiments of the present invention may execute the implementation method described in any method of the first aspect, or may execute the implementation method of the terminal device described in the embodiments of the present invention, which will not be repeated here.
[0069] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the steps of any method of the first aspect are implemented.
[0070] The computer-readable storage medium may be an internal storage unit of the terminal device of any of the aforementioned embodiments, such as a hard disk or memory of the terminal device. The computer-readable storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of the terminal device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the terminal device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0071] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0072] In the several embodiments provided in the present application, it should be understood that the disclosed terminal device and method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.
[0073] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present invention.
[0074] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0075] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0076] The expressions "first", "second", "the first" or "the second" used in various embodiments of the present disclosure may modify various components regardless of order and / or importance, but these expressions do not limit the corresponding components. The above expressions are only configured for the purpose of distinguishing an element from other elements. For example, a first user device and a second user device represent different user devices, although both are user devices. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the present disclosure.
[0077] When one element (e.g., a first element) is referred to as being "(operably or communicatively) coupled" or "(operably or communicatively) coupled to" or "connected to" another element (e.g., a second element), it is understood that the one element is directly connected to the other element or the one element is indirectly connected to the other element via yet another element (e.g., a third element). Conversely, it is understood that when an element (e.g., a first element) is referred to as being "directly connected" or "directly coupled" to another element (the second element), no element (e.g., a third element) is interposed between the two.
[0078] It should be noted that, in this article, the terms "include", "comprises" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context of the specific embodiment.
[0079] The above description is only an optional embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other.
[0080] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0081] The above description is only an optional embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other.
[0082] The above description is only an optional embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A risk analysis and identification method based on big data, characterized in that: The following steps are involved: S1, in response to a risk identification operation, sending a camera parameter adjustment instruction sequence to a terminal of a party currently conducting a video call, for controlling the camera parameter adjustment of the terminal of the party; S2, in response to the other party terminal receiving the camera parameter adjustment instruction sequence, obtaining the other party's video clips within the execution period of the camera parameter adjustment instruction sequence, and performing real-shot judgment on the video pictures corresponding to each instruction triggering moment to determine whether the video pictures are real-shot pictures; S3, when it is determined that any of the video images in the other party's video clips is not a real-shot image, display risk warning information.
2. The risk analysis and identification method based on big data according to claim 1 is characterized in that: The camera parameter adjustment instruction sequence includes at least one of the following instructions: Wide-angle mode switching command, used to instruct the camera to shorten its focal length to a short focal length preset value; Telephoto mode switching command, used to command the camera's focal length to increase to a telephoto preset value; The white balance setting command is used to instruct the camera to adjust the current white balance value to the preset white balance value.
3. The risk analysis and identification method based on big data according to claim 2 is characterized in that: The real-shot judgment of the video pictures corresponding to each instruction triggering moment includes obtaining a first picture corresponding to the wide-angle mode switching instruction triggering moment and a second picture in a frame before the first picture, and judging that the first picture is a real-shot picture when all of the following conditions are met, otherwise, judging that the first picture is a non-real-shot picture: There is non-pure color display content in other picture areas of the first picture except the second picture; If the short focal length preset value is less than the first focal length value, the edge of the first picture is distorted; Determine a first reference element in the second picture, wherein an area ratio of a face region in the second picture to an area of the first reference element is a second ratio, an area ratio of a face region in the first picture to an area of the first reference element is a third ratio, and the second ratio is equal to the third ratio.
4. The risk analysis and identification method based on big data according to claim 2 is characterized in that: The real shot judgment of the video pictures corresponding to each command triggering moment includes obtaining a third picture corresponding to the telephoto mode switching command triggering moment, a fourth picture in a frame before the third picture, and a fifth picture n frames after the third picture, where n is a positive integer greater than or equal to 1, and judging that the third picture is a real shot picture when all of the following conditions are met, otherwise, judging that the third picture is a non-real shot picture: Determine a second reference element in the third picture, wherein an area ratio of a face region in the fourth picture to an area of the second reference element is a fourth ratio, an area ratio of a face region in the third picture to an area of the second reference element is a fifth ratio, and the fourth ratio is equal to the fifth ratio; If the preset long focal length value is greater than the second focal length value, the clarity of the third image is less than the first clarity value; The fifth picture has the same content as the third picture, and the definition of the fifth picture is greater than that of the third picture.
5. The risk analysis and identification method based on big data according to claim 4 is characterized in that: The real shot judgment of the video pictures corresponding to each instruction triggering moment includes obtaining a sixth picture corresponding to the white balance setting instruction triggering moment and a seventh picture which is a frame before the sixth picture, and judging that the sixth picture is a real shot picture when any of the following conditions is met, otherwise, judging that the sixth picture is a non-real shot picture: If the preset white balance value is greater than the first color temperature value, extracting the color value of each pixel in the sixth picture, calculating the average value of the red color value of each pixel in the sixth picture as the first color value, extracting the color value of each pixel in the seventh picture, calculating the average value of the red color value of each pixel in the seventh picture as the second color value, and the first color value is greater than the second color value; If the white balance preset value is less than the second color temperature value, the color value of each pixel in the sixth picture is extracted, and the average value of the blue color value of each pixel in the sixth picture is calculated as the third color value, the color value of each pixel in the seventh picture is extracted, and the average value of the blue color value of each pixel in the seventh picture is calculated as the fourth color value, and the third color value is greater than the fourth color value.
6. The risk analysis and identification method based on big data according to claim 1 is characterized in that: In the risk analysis and identification method based on big data, the following steps are also included between step S2 and step S3: S41, in response to the risk identification operation, obtaining the terminal model of the counterpart terminal, and if the terminal model supports the depth perception function, sending a depth perception function activation prompt to the counterpart terminal; S42, receiving the portrait area depth information and the structured light raw data of the portrait area fed back by the other party terminal, and calculating the calculated depth information of the portrait area according to the structured light raw data; S43, comparing the calculated depth information with the fed-back depth information of the portrait area, if the comparison is consistent, determining that the video picture is a real shot picture, if the comparison is inconsistent, determining that the video picture is a non-real shot picture.
7. The risk analysis and identification method based on big data according to claim 6 is characterized in that: If the other terminal has an infrared structured light projector and an infrared camera, it is determined that the terminal model supports a depth perception function; The structured light raw data includes the original sinusoidal wave pattern projected by the infrared structured light projector and the actual pattern captured by the infrared camera after being modulated by the human face.
8. The risk analysis and identification method based on big data according to claim 7 is characterized in that: The calculating depth information of the portrait area according to the structured light raw data includes: calculating the depth difference of each pixel in the portrait area according to the following formula: ; in, Represents the first The depth difference of pixels relative to the infrared structured light projector and the infrared camera baseline, represents the wavelength of the light beam emitted by the infrared structured light projector, represents the highest spatial frequency that the infrared camera can clearly resolve, Represents the phase difference between the original sine wave pattern and the actual pattern.
9. An electronic device, characterized in that: include: An instruction sequence generation module, configured to send a camera parameter adjustment instruction sequence to a terminal of a party currently conducting a video call in response to a risk identification operation, so as to control the camera parameter adjustment of the terminal of the party; A judgment module, configured to obtain, in response to the receiving of the camera parameter adjustment instruction sequence by the other terminal, the other party's video clips within the execution cycle of the camera parameter adjustment instruction sequence, and perform a real-shot judgment on the video pictures corresponding to each instruction triggering moment to judge whether the video pictures are real-shot pictures; The prompt module is used to display risk warning information when it is determined that any video image in the other party's video clip is not a real-shot image.
10. A risk analysis and identification system based on big data, characterized in that: The method comprises a processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Infrared camera weak light environment compensation method and device and electronic equipment
CN112861645A
Intelligent interaction method, device and system, electronic equipment and computer readable medium
CN114565449A
Method and system for identifying AI face change risk based on AIGC detection algorithm
CN118521936A
Video processing method and electronic device
WO2023124200A1