A risk analysis and identification method and system based on big data
By sending camera parameters adjustment command sequences to the terminal of the video call partner, the video picture characteristics are judged in real time, and the problem of AI forged picture recognition in video calls is solved, effectively preventing telecom fraud and improving the authenticity of video calls.
Patent Information
- Application Number
- CN202510241107.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The prior art is difficult to identify AI-fake video images in real time during video calls, making it difficult to prevent telecom fraud activities.
By sending camera parameter adjustment command sequences to the other terminal, controlling camera parameter changes, and judging screen characteristics in real time, including wide-angle mode switching, telephoto mode switching and white balance settings, combined with depth perception function, we can judge whether the video screen is real shot.
Effectively identify video images forged by AI, improve telecommunications fraud prevention capabilities, enhance the authenticity identification of video calls, and protect user property safety and social security.
Smart Images

Figure CN120017783B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of risk identification technology, and in particular to a risk analysis and identification method and system based on big data. Background Art
[0002] With the advent of the intelligent age, artificial intelligence (AI) is increasingly being used in various fields, but it is also being exploited by criminals. This is particularly true in the field of telecommunications fraud, where fraudsters abuse AI technology to forge faces and voices, carrying out highly deceptive scams.
[0003] In telecom fraud schemes, fraudsters often use AI technology to forge human faces and conduct video calls with victims, garnering their trust and carrying out the deception. This fraudulent method is extremely deceptive, as victims often find it difficult to visually verify whether the person on the video call is a real person. Existing technology can only verify the authenticity of AI-generated video footage, but is unable to perform real-time analysis during a video call. Summary of the Invention
[0004] The purpose of this application is to provide a risk analysis and identification method and system based on big data, which can improve the above problems.
[0005] The embodiment of the present application is implemented as follows:
[0006] First, this application provides a risk analysis and identification method based on big data, which includes steps S1 to S3. S1, S2, etc. are merely step identifiers, and the method execution order does not necessarily follow ascending numerical order. For example, step S2 may be executed before step S1, and this application does not impose any restrictions.
[0007] S1, in response to a risk identification operation, sending a camera parameter adjustment instruction sequence to a terminal of a party currently in a video call, for controlling the adjustment of camera parameters of the terminal of the party;
[0008] S2, in response to the counterpart terminal receiving the camera parameter adjustment instruction sequence, obtaining a video clip of the counterpart terminal within an execution period of the camera parameter adjustment instruction sequence, and performing a real-shot determination on a video screen corresponding to each instruction triggering moment to determine whether the video screen is a real-shot screen;
[0009] S3, when it is determined that any of the video images in the other party's video clips is not a real shot image, displaying risk warning information.
[0010] It can be understood that the present application provides a risk analysis and identification method based on big data, which controls the adjustment of camera parameters by sending a camera parameter adjustment instruction sequence to the terminal of the other party of the video call. The video clip of the other party is obtained during the instruction execution cycle, and the video screen at the time of each instruction triggering is judged as a real shot. The judgment is based on the characteristics of the picture change when the instructions such as wide-angle mode switching, telephoto mode switching and white balance setting are triggered. If any picture is judged to be not a real shot, a risk warning message is displayed. This method can effectively identify AI-forged video pictures, improve the prevention ability of telecommunications fraud, and enhance the authenticity identification ability of video calls by adjusting camera parameters in real time and analyzing picture changes. It provides a new technical means to combat telecommunications fraud and helps protect user property safety and social security.
[0011] In an optional embodiment of the present application, the camera parameter adjustment instruction sequence includes at least one of the following instructions.
[0012] Wide-angle mode switching command, used to instruct the camera's focal length to be shortened to the short focal length preset value;
[0013] Telephoto mode switching command, used to instruct the camera's focal length to increase to the telephoto preset value;
[0014] The white balance setting command is used to instruct the camera to adjust the current white balance value to the preset white balance value.
[0015] The real-shot judgment of the video pictures corresponding to each instruction triggering moment includes obtaining a first picture corresponding to the wide-angle mode switching instruction triggering moment and a second picture in the previous frame of the first picture, and judging that the first picture is a real-shot picture when all of the following conditions 1-3 are met; otherwise, judging that the first picture is a non-real-shot picture.
[0016] Condition 1: There is non-pure color display content in the other areas of the first picture except the second picture. After switching to wide-angle mode, the viewing angle of the picture captured by the real camera will be significantly increased. Therefore, in addition to the content of the previous picture, more details of the surrounding environment will appear in the new picture. However, since it is difficult for AI-generated pictures to simulate this change in perspective in a short period of time, it may not be able to quickly generate the corresponding extended content. This judgment condition is based on the increase in the viewing angle of the picture in wide-angle mode. By comparing the picture content before and after the switch, it is determined whether it is a real shot.
[0017] Condition 2: If the short focal length preset value is less than the first focal length value, the edges of the first image are distorted. In wide-angle mode, due to the shortened focal length, a certain degree of barrel distortion may occur at the edges of the image. This is an inherent characteristic of optical lens imaging and cannot be avoided by real cameras. However, AI-generated images may have difficulty simulating this distortion, especially when switching focal lengths in a short period of time. Therefore, by detecting whether the expected distortion occurs at the edges of the image, the authenticity of the image can be determined.
[0018] Condition 3: Determine the first reference element in the second picture, the area ratio of the face area in the second picture to the area of the first reference element is a second ratio, the area ratio of the face area in the first picture to the area of the first reference element is a third ratio, and the second ratio is equal to the third ratio. Before and after switching to wide-angle mode, the area ratio of elements in the picture (such as faces, objects, etc.) to the background or other fixed reference elements should remain consistent. This is because the wide-angle mode only changes the perspective but does not change the actual size relationship between objects. If the AI-generated picture fails to accurately maintain this proportional relationship when simulating the wide-angle effect, its authenticity can be judged by comparing the pictures before and after the switch.
[0019] In an optional embodiment of the present application, the real-shot judgment of the video pictures corresponding to each instruction triggering moment includes obtaining a third picture corresponding to the telephoto mode switching instruction triggering moment, a fourth picture that is one frame before the third picture, and a fifth picture that is n frames after the third picture, where n is a positive integer greater than or equal to 1. If all of the following conditions 4-6 are met, the third picture is judged to be a real-shot picture; otherwise, the third picture is judged to be a non-real-shot picture.
[0020] Condition 4: Determine the second reference element in the third image; the area ratio of the face region in the fourth image to the area of the second reference element is a fourth ratio; the area ratio of the face region in the third image to the area of the second reference element is a fifth ratio; and the fourth ratio is equal to the fifth ratio. It is understood that telephoto mode narrows the viewing angle by increasing the camera's focal length. Although the viewing angle is narrowed, the area ratio of elements in the image (such as faces, objects, etc.) to the background or other fixed reference elements should remain consistent in real-world photography. This is because telephoto mode only magnifies details in the image without changing the actual size ratio between objects. If the AI-generated image fails to accurately maintain this proportional relationship when simulating the telephoto effect, its authenticity can be determined by comparing the images before and after the switch.
[0021] Condition 5: If the preset telephoto value is greater than the second focal length value, the clarity of the third image is less than the first clarity value. It is understood that when the preset telephoto value is greater than a certain threshold, the clarity of the actual image will be significantly reduced. This level of blur may be difficult to generate in a timely manner in images generated by AI-simulated telephoto. Therefore, this condition 5 can be used as a judgment item.
[0022] Condition 6: The content of the fifth picture is the same as that of the third picture, and the clarity of the fifth picture is greater than that of the third picture. It can be understood that when the preset value of the telephoto focal length is large, after the camera switches from wide-angle to telephoto mode, the initially captured picture may be slightly blurred due to the need to adjust the focal length to achieve clear focus. As the autofocus process is completed, the picture clarity will gradually improve. AI-generated pictures may have difficulty in simulating such clarity changes because they find it difficult to simulate the camera's focusing process in real time.
[0023] In an optional embodiment of the present application, the real-shot judgment of the video pictures corresponding to each instruction triggering moment includes obtaining the sixth picture corresponding to the white balance setting instruction triggering moment and the seventh picture in the previous frame of the sixth picture, and judging that the sixth picture is a real-shot picture when the following condition 7 or condition 8 is met; otherwise, judging that the sixth picture is a non-real-shot picture.
[0024] Condition 7: If the preset white balance value is greater than the first color temperature value, extract the color values of each pixel in the sixth image, calculate the average red color value of each pixel in the sixth image as the first color value, extract the color values of each pixel in the seventh image, and calculate the average red color value of each pixel in the seventh image as the second color value, where the first color value is greater than the second color value. It is understood that the camera's white balance function is used to adjust color balance to suit shooting environments under different light sources. When the white balance value is increased, the image gradually shifts toward yellow, with an increase in red. This change is naturally apparent in a real camera because the camera sensor adjusts its color sensitivity accordingly. However, for AI-generated images, since their generation mechanism differs from that of real cameras, it may be difficult to accurately simulate the color change after white balance adjustment, particularly the increase in red. Therefore, by comparing the average red color value before and after white balance adjustment, the authenticity of the image can be determined.
[0025] Condition 8: If the white balance preset value is less than the second color temperature value, extract the color value of each pixel in the sixth picture, calculate the average value of the blue color value of each pixel in the sixth picture as the third color value, extract the color value of each pixel in the seventh picture, calculate the average value of the blue color value of each pixel in the seventh picture as the fourth color value, and the third color value is greater than the fourth color value. It can be understood that, similar to judgment condition 7, when the white balance value of the camera is adjusted to a smaller value, the picture will gradually shift towards blue, especially the blue component will increase. A real camera can naturally present this color change, but the picture generated by AI may be difficult to accurately simulate. By comparing the average value of the blue color before and after the white balance adjustment, the authenticity of the picture can be further verified.
[0026] In an optional embodiment of the present application, the risk analysis and identification method based on big data further includes the following steps between step S2 and step S3:
[0027] S41, in response to the risk identification operation, obtaining the terminal model of the counterpart terminal, and if the terminal model supports the depth perception function, sending a depth perception function activation prompt to the counterpart terminal;
[0028] S42, receiving the portrait area depth information and the structured light raw data of the portrait area fed back by the other terminal, and calculating the calculated depth information of the portrait area based on the structured light raw data;
[0029] S43: Compare the calculated depth information with the fed-back depth information of the portrait area. If the comparison is consistent, determine that the video picture is a real shot picture; if the comparison is inconsistent, determine that the video picture is not a real shot picture.
[0030] It can be understood that this application has added the judgment and utilization of depth perception function to the above-mentioned AI raw image judgment technology. If the other terminal supports depth perception function, turn on this function and obtain the depth information and structured light raw data of the portrait area. By calculating the calculated depth information of the portrait area and comparing it with the feedback depth information, it is determined whether the video picture is a real shot. This method uses depth perception technology to further improve the accuracy of authenticity identification of video pictures. It combines depth perception and structured light technology to provide a more reliable basis for authenticity identification of video calls, and effectively enhances the prevention and combat capabilities of telecommunications fraud.
[0031] In an optional embodiment of the present application, if the other terminal has an infrared structured light projector and an infrared camera, it is determined that the terminal model supports depth perception function; the structured light raw data includes the original sinusoidal wave pattern projected by the infrared structured light projector and the actual pattern modulated by the human face captured by the infrared camera.
[0032] In an optional embodiment of the present application, calculating the depth information of the portrait area based on the structured light raw data includes: calculating the depth difference of each pixel in the portrait area according to the following formula:
[0033] ;
[0034] in, Represents the first The depth difference of pixels relative to the infrared structured light projector and the infrared camera baseline, represents the wavelength of the light beam emitted by the infrared structured light projector, represents the highest spatial frequency that the infrared camera can clearly resolve, Represents the phase difference between the original sine wave pattern and the actual pattern.
[0035] As you can understand, one of the key steps in using depth sensing to determine the authenticity of video footage is calculating depth information for the portrait area based on the raw structured light data. This process relies primarily on the collaborative work of an infrared structured light projector and an infrared camera, as well as the application of a sinusoidal pattern. Specifically, the infrared structured light projector projects a sinusoidal pattern onto the face. This pattern modulates the surface and is then captured by the infrared camera. There is a phase difference between the actual pattern captured by the infrared camera and the original sinusoidal pattern, which is directly related to the depth information of the object's surface. To calculate depth information, we can use the principle of triangulation. Based on the relative position of the infrared structured light projector and infrared camera (i.e., baseline distance), the wavelength of the projector's output beam, and the highest spatial frequency that the infrared camera can clearly resolve, we can calculate the depth difference of each pixel relative to the baseline based on the phase difference. This calculation process involves complex mathematical models and algorithms, but its core lies in the use of the key parameter, phase difference, to derive depth information. This method allows us to obtain detailed depth information for the portrait area, which can then be compared with the depth information returned by the other terminal. If the two are consistent, it means that the video footage is likely to be real footage; if they are inconsistent, there is a risk of fraud.
[0036] In a second aspect, the present application discloses an electronic device, comprising:
[0037] An instruction sequence generation module, configured to, in response to a risk identification operation, send a camera parameter adjustment instruction sequence to a terminal of a party currently in a video call, for controlling the adjustment of camera parameters of the terminal of the party;
[0038] a judgment module, configured to, in response to the receiving of the camera parameter adjustment instruction sequence by the other terminal, obtain video clips of the other terminal within an execution period of the camera parameter adjustment instruction sequence, and perform a real-shot judgment on the video images corresponding to each instruction triggering moment to determine whether the video images are real-shot images;
[0039] The prompt module is used to display risk warning information when it is determined that any video image in the other party's video clip is not a real shot image.
[0040] In a third aspect, the present application discloses a risk analysis and identification system based on big data, comprising a processor, an input device, an output device and a memory, wherein the processor, input device, output device and memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method as described in any one of the first aspects.
[0041] Beneficial effects: The present application provides a risk analysis and identification method and system based on big data, which controls the adjustment of camera parameters by sending a camera parameter adjustment instruction sequence to the terminal of the other party of the video call. The video clip of the other party is obtained during the instruction execution cycle, and the video screen at the time of each instruction triggering is judged as a real shot. The judgment is based on the characteristics of the picture changes when the instructions are triggered, including wide-angle mode switching, telephoto mode switching and white balance setting. If any picture is judged to be not a real shot, a risk warning message is displayed. This method can effectively identify AI-forged video pictures, improve the prevention ability of telecommunications fraud, and enhance the authenticity identification ability of video calls by adjusting camera parameters in real time and analyzing picture changes, providing a new technical means to combat telecommunications fraud and helping to protect user property safety and social security.
[0042] This application adds the judgment and utilization of depth perception function to the above-mentioned AI raw image judgment technology. If the other terminal supports depth perception, the function is turned on and the depth information and structured light raw data of the portrait area are obtained. By calculating the calculated depth information of the portrait area and comparing it with the feedback depth information, it is determined whether the video picture is a real shot. This method uses depth perception technology to further improve the accuracy of video picture authenticity identification, and combines depth perception and structured light technology to provide a more reliable basis for authenticity identification of video calls.
[0043] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, optional embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0045] Figure 1 This is a flowchart of a risk analysis and identification method based on big data provided by this application;
[0046] Figure 2 is a schematic diagram of the camera parameter adjustment instruction sequence provided by this application;
[0047] Figure 3 This is a schematic diagram of the risk warning scenario provided by this application;
[0048] Figure 4 This is a flow chart of another risk analysis and identification method based on big data provided by this application. DETAILED DESCRIPTION
[0049] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0050] First, this application provides a risk analysis and identification method based on big data, which includes steps S1 to S3. S1, S2, etc. are merely step identifiers, and the method execution order does not necessarily follow ascending numerical order. For example, step S2 may be executed before step S1, and this application does not impose any restrictions.
[0051] S1, in response to a risk identification operation, sending a camera parameter adjustment instruction sequence to a terminal of a party currently in a video call, for controlling the camera parameter adjustment of the terminal of the party.
[0052] The aforementioned risk identification operation refers to a user-initiated action designed to initiate risk analysis and identification methods based on big data. For example, during a video call, if a user suspects the other party's video may be artificially manipulated by AI, they can perform a specific action (such as clicking a button) to send a sequence of camera parameter adjustment commands, thereby identifying and preventing risks.
[0053] The above-mentioned camera parameter adjustment command sequence is a set of commands used to control the camera parameter adjustment of the other terminal during a video call. This command sequence includes various commands, such as wide-angle mode switching command, telephoto mode switching command, and white balance setting command, etc. It aims to enhance the authenticity verification capability of video calls by adjusting camera parameters in real time and analyzing image changes. Figure 2 The camera parameter adjustment instruction sequence shown includes a wide-angle mode switching instruction set at time t1, a telephoto mode switching instruction set at time t2, and a white balance setting instruction set at time t3.
[0054] S2, in response to the other party's terminal receiving the camera parameter adjustment instruction sequence, obtain the other party's video clips within the execution cycle of the camera parameter adjustment instruction sequence, and perform real-shot judgment on the video pictures corresponding to each instruction triggering moment to determine whether the video pictures are real-shot pictures.
[0055] S3, when it is determined that any video frame in the other party's video clip is not a real shot frame, a risk warning message is displayed.
[0056] It can be understood that the present application provides a risk analysis and identification method based on big data, which controls the adjustment of camera parameters by sending a camera parameter adjustment instruction sequence to the terminal of the other party of the video call. The video clip of the other party is obtained during the instruction execution cycle, and the video screen at the time of each instruction triggering is judged as a real shot. The judgment is based on the characteristics of the picture change when the instructions such as wide-angle mode switching, telephoto mode switching and white balance setting are triggered. If any picture is judged to be not a real shot, a risk warning message is displayed. This method can effectively identify AI-forged video pictures, improve the prevention ability of telecommunications fraud, and enhance the authenticity identification ability of video calls by adjusting camera parameters in real time and analyzing picture changes. It provides a new technical means to combat telecommunications fraud and helps protect user property safety and social security.
[0057] In an optional embodiment of the present application, the camera parameter adjustment instruction sequence includes at least one of the following instructions.
[0058] The wide-angle mode switch command shortens the camera's focal length to a preset short focal length. This command is sent in a data packet to the camera control module of the other device via a network communication protocol (such as TCP / IP). This command contains a clear command code, instructing the camera to switch from its current mode to wide-angle mode. At the software level, the operating system of the other device receives the command and passes it to the camera driver. The camera driver interprets the command code and calls the corresponding hardware control function. At the hardware level, after receiving the control signal, the camera module uses an internal motor or stepper motor to move the lens assembly, shortening the focal length and achieving wide-angle shooting. This process is real-time and fast, ensuring smooth video calls. Wide-angle mode changes the camera's focal length, allowing the camera to capture a wider field of view. This change is based on the principle of optical lens imaging, where the focal length and viewing angle of the image are altered by adjusting the position of the lens assembly. During a video call, using the wide-angle mode switch can quickly verify the authenticity of the other party's image, as AI-generated images may not be able to adapt to the rapid focus change.
[0059] The telephoto mode switch command increases the camera's focal length to a preset telephoto value. Similar to the wide-angle mode switch command, the telephoto mode switch command is also sent to the camera control module of the other device via a network communication protocol. The difference is that this command instructs the camera to switch to telephoto mode, increasing the focal length to capture distant details. At the software level, the operating system receives the command and passes it to the camera driver, which interprets the command and calls hardware control functions. At the hardware level, the camera module uses a motor or stepper motor to drive the lens group in the opposite direction to increase the focal length. This process is also real-time and fast, ensuring the continuity of the video call. Telephoto mode increases the camera's focal length to narrow the angle of view, thereby improving the clarity and detail of distant objects. This principle is based on the imaging principles of optical lenses, which adjust the position of the lens group to change the focal length and magnification of the image. During a video call, switching to telephoto mode can further verify the authenticity of the other party's image, as AI-generated images may not display the same details as the real scene when the focal length is increased.
[0060] A white balance setting command instructs the camera to adjust its current white balance value to a preset white balance value. This command is sent to the camera control module of the other device via a network communication protocol. This command contains a specific preset white balance value, instructing the camera to adjust its white balance settings to match that value. At the software level, the operating system receives the command and passes it to the camera driver. The driver interprets the command and calls the corresponding white balance adjustment function. At the hardware level, the camera module adjusts its internal color correction algorithm based on the command to compensate for color temperature differences under different light sources. This process involves real-time color adjustment of the raw data captured by the image sensor. White balance is a key parameter for accurate color reproduction in a camera. Different light sources have different color temperatures, and the white balance setting ensures that the camera accurately reproduces colors under various lighting conditions. Adjusting the white balance value changes the overall hue and color saturation of the image. During a video call, changing the white balance setting allows you to observe color changes in the image. This provides a basis for judging the authenticity of the image, as AI-generated images may not fully mimic the color variations of real scenes when adjusted.
[0061] Perform real shot judgment on the video images corresponding to each command triggering moment, including obtaining the first image corresponding to the wide-angle mode switching command triggering moment and the second image in the previous frame of the first image. If all of the following conditions 1-3 are met, the first image is judged to be a real shot image; otherwise, the first image is judged to be a non-real shot image. Figure 2 As shown, the first picture P1 corresponding to the wide-angle mode switching instruction triggering moment and the second picture P2 in the previous frame of the first picture P1 are judged to be the real shot picture when all of the following conditions 1-3 are met.
[0062] Condition 1: Non-pure color display content exists in areas of the first image other than the second image. It is understandable that after switching to wide-angle mode, the viewing angle of the image captured by the real camera will increase significantly, so in addition to the content of the previous image, more details of the surrounding environment will appear in the new image. However, since AI-generated images cannot simulate this change in perspective in a short period of time, they may not be able to quickly generate the corresponding expanded content. This judgment condition is based on the increased viewing angle of the image in wide-angle mode, and determines whether it is a real shot by comparing the image content before and after the switch.
[0063] Condition 2: If the short focal length preset value is less than the first focal length value, distortion occurs at the edges of the first image. Understandably, in wide-angle mode, due to the shortened focal length, a certain degree of barrel distortion may occur at the edges of the image. This is an inherent characteristic of optical lens imaging and cannot be avoided by real cameras. AI-generated images may have difficulty simulating this distortion, especially when switching focal lengths within a short period of time. Therefore, by detecting whether the expected distortion occurs at the edges of the image, the authenticity of the image can be determined.
[0064] Condition 3: Determine the first reference element in the second image. The area ratio of the face region in the second image to the area of the first reference element is a second ratio. The area ratio of the face region in the first image to the area of the first reference element is a third ratio, and the second ratio is equal to the third ratio. It is understood that the area ratio of elements (such as faces, objects, etc.) in the image to the background or other fixed reference elements should remain consistent before and after switching to wide-angle mode. This is because wide-angle mode only changes the perspective, not the actual size relationship between objects. If the AI-generated image fails to accurately maintain this proportional relationship when simulating a wide-angle effect, its authenticity can be determined by comparing the images before and after the switch.
[0065] In an optional embodiment of the present application, a real shot judgment is performed on the video screen corresponding to each command triggering moment, including obtaining the third screen corresponding to the telephoto mode switching command triggering moment, the fourth screen in the previous frame of the third screen, and the fifth screen n frames after the third screen, where n is a positive integer greater than or equal to 1. If all of the following conditions 4-6 are met, the third screen is judged to be a real shot screen; otherwise, the third screen is judged to be a non-real shot screen. Figure 2 As shown, the third picture P3 corresponding to the moment when the telephoto mode switching instruction is triggered, the fourth picture P4 in the previous frame of the third picture P3, and the fifth picture P5 in the next few frames after the third picture P3 are judged to be the real shot picture when all of the following conditions 4-6 are met.
[0066] Condition 4: Determine the second reference element in the third frame. The ratio of the area of the face region to the area of the second reference element in the fourth frame is a fourth ratio. The ratio of the area of the face region to the area of the second reference element in the third frame is a fifth ratio, and the fourth ratio is equal to the fifth ratio. It can be understood that telephoto mode narrows the angle of view by increasing the camera's focal length. Although the angle of view is narrowed, the area ratio of elements in the frame (such as faces, objects, etc.) to the background or other fixed reference elements should remain consistent in real-life photography. This is because telephoto mode only magnifies details in the frame without changing the actual size ratios between objects. If the AI-generated image fails to accurately maintain this ratio when simulating the telephoto effect, its authenticity can be determined by comparing the images before and after the switch.
[0067] Condition 5: If the preset telephoto value is greater than the second focal length value, the clarity of the third image is less than the first clarity value. It is understood that when the preset telephoto value exceeds a certain threshold, the clarity of the actual image will be significantly reduced. This level of blur may be difficult to generate in a timely manner in images generated by AI-simulated telephoto. Therefore, this condition 5 can be used as a judgment item.
[0068] Condition 6: The fifth and third frames have the same content, and the fifth frame has greater clarity than the third. Understandably, when the preset telephoto focal length is large, after the camera switches from wide-angle to telephoto mode, the initial captured image may be slightly blurry due to the need to adjust the focal length to achieve clear focus. As the autofocus process completes, the image clarity gradually improves. AI-generated images may have difficulty simulating this clarity change because they struggle to simulate the camera's focusing process in real time.
[0069] In an optional embodiment of the present application, the video images corresponding to each instruction triggering moment are judged to be actually shot, including obtaining the sixth image corresponding to the white balance setting instruction triggering moment and the seventh image in the previous frame of the sixth image, and judging the sixth image to be an actually shot image if the following conditions 7 or 8 are met, otherwise, judging the sixth image to be a non-actual shot image. Figure 2 As shown, the sixth picture P6 corresponding to the white balance setting instruction triggering moment and the seventh picture P7 in the previous frame of the sixth picture P6 are judged to be the real shot picture when the following condition 7 or condition 8 is met.
[0070] Condition 7: If the preset white balance value is greater than the first color temperature, extract the color values of each pixel in the sixth frame, calculate the average red color value of each pixel in the sixth frame as the first color value, extract the color values of each pixel in the seventh frame, calculate the average red color value of each pixel in the seventh frame as the second color value, and the first color value is greater than the second color value. It is understood that the camera's white balance function is used to adjust color balance to suit shooting environments under different light sources. When the white balance value is increased, the image gradually shifts toward yellow, especially with an increase in red. This change is naturally apparent in a real camera because the camera sensor adjusts its color sensitivity accordingly. However, for AI-generated images, since their generation mechanism differs from that of real cameras, it may be difficult to accurately simulate the color change after white balance adjustment, especially the increase in red. Therefore, by comparing the average red color value before and after white balance adjustment, the authenticity of the image can be determined.
[0071] Condition 8: If the preset white balance value is less than the second color temperature value, extract the color value of each pixel in the sixth image, calculate the average blue color value of each pixel in the sixth image as the third color value, extract the color value of each pixel in the seventh image, calculate the average blue color value of each pixel in the seventh image as the fourth color value, and the third color value is greater than the fourth color value. It can be understood that similar to judgment condition 7, when the camera's white balance value is adjusted to a smaller value, the image will gradually shift towards blue, especially the blue component will increase. Real cameras can naturally present this color change, while AI-generated images may be difficult to accurately simulate. By comparing the average blue color value before and after the white balance adjustment, the authenticity of the image can be further verified.
[0072] In an optional embodiment of the present application, in the risk analysis and identification method based on big data, the following steps S41 to S43 are also included between step S2 and step S3.
[0073] S41, in response to the risk identification operation, obtaining the terminal model of the other terminal, and if the terminal model supports the depth perception function, sending a depth perception function activation prompt to the other terminal.
[0074] In an optional embodiment of the present application, if the other terminal has an infrared structured light projector and an infrared camera, it is determined that the terminal model supports depth perception function; the structured light raw data includes the original sinusoidal wave pattern projected by the infrared structured light projector and the actual pattern modulated by the human face captured by the infrared camera.
[0075] S42 , receiving depth information of the portrait area and raw structured light data of the portrait area fed back by the other terminal, and calculating calculated depth information of the portrait area based on the raw structured light data.
[0076] In an optional embodiment of the present application, calculating depth information of the portrait area based on the structured light raw data includes: calculating the depth difference of each pixel in the portrait area according to the following formula:
[0077] ;
[0078] in, In the representative portrait area The depth difference of pixels relative to the infrared structured light projector and the infrared camera baseline, Represents the wavelength of the infrared structured light projector's output beam, Represents the highest spatial frequency that the infrared camera can clearly distinguish. Represents the phase difference between the original sine wave pattern and the actual pattern.
[0079] As you can understand, one of the key steps in using depth sensing to determine the authenticity of video footage is calculating depth information for the portrait area based on the raw structured light data. This process relies primarily on the collaborative work of an infrared structured light projector and an infrared camera, as well as the application of a sinusoidal pattern. Specifically, the infrared structured light projector projects a sinusoidal pattern onto the face. This pattern modulates the surface and is then captured by the infrared camera. There is a phase difference between the actual pattern captured by the infrared camera and the original sinusoidal pattern, which is directly related to the depth information of the object's surface. To calculate depth information, we can use the principle of triangulation. Based on the relative position of the infrared structured light projector and infrared camera (i.e., baseline distance), the wavelength of the projector's output beam, and the highest spatial frequency that the infrared camera can clearly resolve, we can calculate the depth difference of each pixel relative to the baseline based on the phase difference. This calculation process involves complex mathematical models and algorithms, but its core lies in the use of the key parameter, phase difference, to derive depth information. This method allows us to obtain detailed depth information for the portrait area, which can then be compared with the depth information returned by the other terminal. If the two are consistent, it means that the video footage is likely to be real footage; if they are inconsistent, there is a risk of fraud.
[0080] S43: Compare the calculated depth information with the fed-back portrait area depth information. If the comparison is consistent, the video image is determined to be a real shot image. If the comparison is inconsistent, the video image is determined to be a non-real shot image.
[0081] It can be understood that this application has added the judgment and utilization of depth perception function to the above-mentioned AI raw image judgment technology. If the other terminal supports depth perception function, turn on this function and obtain the depth information and structured light raw data of the portrait area. By calculating the calculated depth information of the portrait area and comparing it with the feedback depth information, it is determined whether the video picture is a real shot. This method uses depth perception technology to further improve the accuracy of authenticity identification of video pictures. It combines depth perception and structured light technology to provide a more reliable basis for authenticity identification of video calls, and effectively enhances the prevention and combat capabilities of telecommunications fraud.
[0082] In a second aspect, the present application discloses an electronic device, comprising:
[0083] An instruction sequence generation module, configured to, in response to a risk identification operation, send a camera parameter adjustment instruction sequence to a terminal of a party currently in a video call, for controlling the adjustment of camera parameters of the terminal of the party;
[0084] a judgment module, configured to, in response to the receiving of the camera parameter adjustment instruction sequence by the other terminal, obtain video clips of the other terminal within an execution period of the camera parameter adjustment instruction sequence, and perform a real-shot judgment on the video images corresponding to each instruction triggering moment to determine whether the video images are real-shot images;
[0085] The prompt module is used to display risk warning information when it is determined that any video image in the other party's video clip is not a real shot image.
[0086] In a third aspect, the present application provides a big data-based risk analysis and identification system. The big data-based risk analysis and identification system includes one or more processors; one or more input devices; one or more output devices; and a memory. The processors, input devices, output devices, and memory are connected via a bus. The memory is used to store a computer program, which includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor is configured to invoke the program instructions to perform the operations of any of the methods of the first aspect:
[0087] It should be understood that in the embodiments of the present invention, the processor referred to may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0088] Input devices may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint direction information), a microphone, etc. Output devices may include a display (LCD, etc.), a speaker, etc.
[0089] The memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory may also include non-volatile random access memory. For example, the memory may also store information about the device type.
[0090] In specific implementations, the processor, input device, and output device described in the embodiments of the present invention can execute the implementation method described in any method of the first aspect, and can also execute the implementation method of the terminal device described in the embodiments of the present invention, which will not be repeated here.
[0091] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the steps of any method of the first aspect are implemented.
[0092] The computer-readable storage medium may be an internal storage unit of the terminal device of any of the aforementioned embodiments, such as a hard disk or memory of the terminal device. The computer-readable storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device. Furthermore, the computer-readable storage medium may include both an internal storage unit and an external storage device of the terminal device. The computer-readable storage medium is used to store the computer program and other programs and data required by the terminal device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0093] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed terminal devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or can be electrical, mechanical or other forms of connection.
[0095] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.
[0096] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0097] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0098] The terms "first," "second," "the first," or "the second" used in various embodiments of the present disclosure may modify various components regardless of order and / or importance, but these terms do not limit the corresponding components. The above terms are configured solely for the purpose of distinguishing an element from other elements. For example, a first user device and a second user device represent different user devices, even though both are user devices. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the present disclosure.
[0099] When one element (for example, a first element) is referred to as being “(operably or communicably) coupled” or “(operably or communicably) coupled to” or “connected to” another element (for example, a second element), it should be understood that the one element is directly connected to the other element or that the one element is indirectly connected to the other element via yet another element (for example, a third element). Conversely, it should be understood that when an element (for example, a first element) is referred to as being “directly connected” or “directly coupled” to another element (the second element), there is no element (for example, a third element) interposed therebetween.
[0100] It should be noted that, in this document, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context of the specific embodiment.
[0101] The above description is merely an optional embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also encompass other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.
[0102] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0103] The above description is merely an optional embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also encompass other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.
[0104] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A risk analysis and identification method based on big data, characterized in that: The following steps are involved: S1, in response to a risk identification operation, sending a camera parameter adjustment instruction sequence to a terminal of a party currently in a video call, for controlling the adjustment of camera parameters of the terminal of the party; S2, in response to the counterpart terminal receiving the camera parameter adjustment instruction sequence, obtaining a video clip of the counterpart terminal within an execution period of the camera parameter adjustment instruction sequence, and performing a real-shot determination on a video screen corresponding to each instruction triggering moment to determine whether the video screen is a real-shot screen; S3, when it is determined that any of the video images in the other party's video clips is not a real shot image, displaying risk warning information; The camera parameter adjustment instruction sequence includes at least one of the following instructions: Wide-angle mode switching command, used to instruct the camera's focal length to be shortened to the short focal length preset value; Telephoto mode switching command, used to instruct the camera's focal length to increase to the telephoto preset value; The white balance setting command is used to instruct the camera to adjust the current white balance value to the preset white balance value.
2. The risk analysis and identification method based on big data according to claim 1 is characterized in that: The performing of a real shot determination on the video images corresponding to each instruction triggering moment includes obtaining a first image corresponding to the wide-angle mode switching instruction triggering moment and a second image that is a frame before the first image, and determining that the first image is a real shot image if all of the following conditions are met; otherwise, determining that the first image is a non-real shot image: There is non-pure color display content in other picture areas of the first picture except the second picture; If the short focal length preset value is less than the first focal length value, the edge of the first image is distorted; Determine the first reference element in the second picture, the area ratio of the face area in the second picture to the first reference element area is a second ratio, the area ratio of the face area in the first picture to the first reference element area is a third ratio, and the second ratio is equal to the third ratio.
3. The risk analysis and identification method based on big data according to claim 1 is characterized in that: The performing real shot determination on the video pictures corresponding to each command triggering moment includes obtaining a third picture corresponding to the telephoto mode switching command triggering moment, a fourth picture that is one frame before the third picture, and a fifth picture that is n frames after the third picture, where n is a positive integer greater than or equal to 1, and determining that the third picture is a real shot picture if all of the following conditions are met; otherwise, determining that the third picture is a non-real shot picture: Determining a second reference element in the third picture, wherein an area ratio of a face region in the fourth picture to an area of the second reference element is a fourth ratio, an area ratio of the face region in the third picture to an area of the second reference element is a fifth ratio, and the fourth ratio is equal to the fifth ratio; If the long focal length preset value is greater than the second focal length value, the clarity of the third image is less than the first clarity value; The fifth picture has the same content as the third picture, and the definition of the fifth picture is greater than that of the third picture.
4. The risk analysis and identification method based on big data according to claim 3 is characterized in that: The performing of a real shot determination on the video pictures corresponding to each instruction triggering moment includes obtaining a sixth picture corresponding to the white balance setting instruction triggering moment and a seventh picture that is a frame before the sixth picture, and determining that the sixth picture is a real shot picture if any of the following conditions is met; otherwise, determining that the sixth picture is a non-real shot picture: If the preset white balance value is greater than the first color temperature value, extracting the color value of each pixel in the sixth picture, calculating the average red color value of each pixel in the sixth picture as the first color value, extracting the color value of each pixel in the seventh picture, calculating the average red color value of each pixel in the seventh picture as the second color value, and the first color value is greater than the second color value; If the white balance preset value is less than the second color temperature value, extract the color value of each pixel in the sixth picture, calculate the average of the blue color values of each pixel in the sixth picture as the third color value, extract the color value of each pixel in the seventh picture, calculate the average of the blue color values of each pixel in the seventh picture as the fourth color value, and the third color value is greater than the fourth color value.
5. The risk analysis and identification method based on big data according to claim 1 is characterized in that: The risk analysis and identification method based on big data further includes the following steps between step S2 and step S3: S41, in response to the risk identification operation, obtaining the terminal model of the counterpart terminal, and if the terminal model supports the depth perception function, sending a depth perception function activation prompt to the counterpart terminal; S42, receiving the portrait area depth information and the structured light raw data of the portrait area fed back by the other terminal, and calculating the calculated depth information of the portrait area based on the structured light raw data; S43: Compare the calculated depth information with the fed-back depth information of the portrait area. If the comparison is consistent, determine that the video picture is a real shot picture; if the comparison is inconsistent, determine that the video picture is not a real shot picture.
6. The risk analysis and identification method based on big data according to claim 5 is characterized in that: If the other terminal has an infrared structured light projector and an infrared camera, determining whether the terminal model supports depth perception; The structured light raw data includes the original sinusoidal wave pattern projected by the infrared structured light projector and the actual pattern captured by the infrared camera after being modulated by the human face.
7. The risk analysis and identification method based on big data according to claim 6 is characterized in that: The calculating depth information of the portrait area according to the structured light raw data includes: calculating the depth difference of each pixel in the portrait area according to the following formula: ; in, Represents the first The depth difference of pixels relative to the infrared structured light projector and the infrared camera baseline, represents the wavelength of the light beam emitted by the infrared structured light projector, represents the highest spatial frequency that the infrared camera can clearly resolve, Represents the phase difference between the original sine wave pattern and the actual pattern.
8. An electronic device, characterized in that: include: An instruction sequence generation module, configured to, in response to a risk identification operation, send a camera parameter adjustment instruction sequence to a terminal of a party currently in a video call, for controlling the adjustment of camera parameters of the terminal of the party; a judgment module, configured to, in response to the receiving of the camera parameter adjustment instruction sequence by the other terminal, obtain video clips of the other terminal within an execution period of the camera parameter adjustment instruction sequence, and perform a real-shot judgment on the video images corresponding to each instruction triggering moment to determine whether the video images are real-shot images; a prompt module, configured to display risk warning information when determining that any video image in the other party's video clip is not a real-shot image; The camera parameter adjustment instruction sequence includes at least one of the following instructions: Wide-angle mode switching command, used to instruct the camera's focal length to be shortened to the short focal length preset value; Telephoto mode switching command, used to instruct the camera's focal length to increase to the telephoto preset value; The white balance setting command is used to instruct the camera to adjust the current white balance value to the preset white balance value.
9. A risk analysis and identification system based on big data, characterized in that: The method comprises a processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Infrared camera weak light environment compensation method and device and electronic equipment
CN112861645A
Intelligent interaction method, device and system, electronic equipment and computer readable medium
CN114565449A