Image processing method, device, storage medium, and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-29
- Publication Date
- 2026-08-11
AI Technical Summary
这种通过运行环境是否异常识别伪造图像的方式,存在识别准确性较低的问题
Smart Images

Figure CN122550987A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of information processing technology, and in particular to an image processing method, device, storage medium, and program product. Background Technology
[0002] With the acceleration of digital transformation and the increasing popularity of online services, Electronic Know Your Customer (eKYC) has become a standard procedure in finance, government, and other fields. In the eKYC process, users can use their smartphone cameras to capture and upload images of their identification documents. The service provider's identity verification module then verifies these images. Once verification is successful, users can access relevant services online, improving service efficiency.
[0003] However, eKYC systems are facing increasingly serious security challenges, one of which is image injection attacks. This type of attack refers to attackers using technical means (such as using a modified operating system or exploiting application vulnerabilities) to bypass the normal camera capture process and provide a forged ID image to the authentication module in an attempt to pass verification.
[0004] To defend against image injection attacks, the identity verification module checks for anomalies in the phone's operating environment before performing identity verification. This includes checking if the phone is running in root mode or on an emulator. If anomalies are detected, the image is considered forged. However, this method of identifying forged images based on anomalies in the operating environment has a relatively low accuracy rate. Summary of the Invention
[0005] This specification provides an image processing method, apparatus, storage medium, and program product to improve the accuracy of identifying injected images.
[0006] This specification provides an image processing method applied to a terminal device with multiple cameras, comprising: displaying a target page, the target page including an image acquisition control; responding to a trigger operation on the image acquisition control, displaying a viewfinder, and displaying real-time image data of the main camera among the multiple cameras in the viewfinder; if a document to be acquired appears in the real-time image data, calling at least two cameras among the multiple cameras to acquire images of the document at different optical focal lengths, obtaining at least two document images; and sending the at least two document images to a server device, the server device determining that the at least two document images are injected images if the similarity between two document images exceeds a set first similarity threshold.
[0007] This specification also provides an image processing method applicable to a server device, comprising: receiving at least two document images sent by a terminal device; the at least two document images are obtained by capturing images of the document from cameras with different optical focal lengths, or are injected images used to replace real captured images during the image acquisition process; if the similarity between two document images exceeds a set first similarity threshold, then the at least two document images are determined to be injected images.
[0008] This specification also provides an image processing method for an application terminal device. The method includes: displaying a target page, the target page including an image acquisition control; responding to a trigger operation on the image acquisition control, displaying a viewfinder and displaying real-time image data from a main camera in the viewfinder; if a document to be acquired appears in the real-time image data, calling the main camera to acquire images of the document at least twice at different digital focal lengths to obtain at least two document images; and sending the at least two document images to a server device, wherein the server device determines the at least two document images as injected images if the similarity between two document images exceeds a set first similarity threshold.
[0009] This specification also provides an image processing method, which includes: displaying a target page, the target page including an image acquisition control; responding to a trigger operation on the image acquisition control, displaying a viewfinder, and displaying real-time image data of the main camera among multiple cameras in the viewfinder; if a document to be acquired appears in the real-time image data, calling at least two cameras among the multiple cameras to acquire images of the document at different optical focal lengths, obtaining at least two document images; if the similarity between two document images exceeds a set first similarity threshold, determining at least two document images as injected images.
[0010] This specification also provides an image processing method for an application terminal device. The method includes: displaying a target page, the target page including an image acquisition control; responding to a trigger operation on the image acquisition control, displaying a viewfinder, and displaying real-time image data from a main camera in the viewfinder; if a document to be acquired appears in the real-time image data, calling the main camera to acquire images of the document at least twice at different digital focal lengths to obtain at least two document images; if the similarity between two document images exceeds a set first similarity threshold, determining at least two document images as injected images.
[0011] This specification also provides an electronic device, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute one or more computer instructions to perform the steps in the method provided in this specification.
[0012] This specification also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps of the method provided in this specification.
[0013] This specification also provides a computer program product, including: a computer program / instructions, which, when executed by a processor, can implement the steps of the method provided in this specification.
[0014] In the embodiments of this specification, when a user photographs their ID card, at least two cameras (such as a main camera and a telephoto camera) on the terminal device are used to capture at least two ID card images at different optical focal lengths. Since at least two ID card images capture the same ID card using different optical focal lengths, and the at least two ID card images differ in field of view, perspective, and subject size, the similarity between the at least two ID card images is high, i.e., less than a first similarity threshold (e.g., 90% or 93%), but not close to the theoretical upper limit (e.g., 100%). To address the risk of injection, a sub-image is obtained by cropping the original injection image. This sub-image is part of the original injection image, and the similarity between the original injection image and the sub-image is close to the theoretical upper limit. Therefore, if the similarity between two ID card images exceeds the set first similarity threshold, then at least two ID card images are identified as injection images, improving the accuracy of image injection detection. Attached Figure Description
[0015] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the structure of an image processing system provided as an exemplary embodiment of this specification.
[0016] Figure 2 This is an internal schematic diagram of a terminal device and a server device provided in one embodiment of this specification.
[0017] Figure 3a This is a flowchart illustrating the implementation of an image processing method on a terminal device, which is an exemplary embodiment of this specification.
[0018] Figure 3b This is a flowchart illustrating the implementation of an image processing method on a server device, as provided in an exemplary embodiment of this specification.
[0019] Figure 4 This is a flowchart illustrating the implementation of an image processing method on a terminal device, which is an exemplary embodiment of this specification.
[0020] Figure 5This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this specification. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0022] It should be noted that, in the cases involving user information in this application's embodiments, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. The various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.
[0023] To address the aforementioned technical issues, in some embodiments of this specification, when a user photographs their ID card, at least two cameras (such as a main camera and a telephoto camera) on the terminal device are used to capture at least two ID card images at different optical focal lengths. Since at least two ID card images capture the same ID card using different optical focal lengths, and these images differ in field of view, perspective, and subject size, the similarity between the at least two ID card images is relatively high, i.e., less than a first similarity threshold (e.g., 90% or 93%), but not close to the theoretical upper limit (e.g., 100%). To mitigate injection risk, a sub-image is obtained by cropping the original injection image. This sub-image is part of the original injection image, and the similarity between the original injection image and the sub-image is close to the theoretical upper limit. Therefore, if the similarity between two ID card images exceeds the set first similarity threshold, then at least two ID card images are identified as injection images, improving the accuracy of image injection detection.
[0024] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0025] Figure 1 This is a schematic diagram of the structure of an image processing system provided in an exemplary embodiment of this specification, as shown below. Figure 1 The system shown includes: terminal device 101 and server device 102.
[0026] The implementation form of terminal device 101 is not limited. For example, terminal device 101 may include, but is not limited to: smartphones, tablets, laptops, identity verification all-in-one machines, smart visitor machines, and access control and attendance all-in-one machines. Terminal device 101 includes multiple cameras. Different cameras may include, but are not limited to: main cameras, ultra-wide-angle cameras, telephoto cameras, macro cameras, depth cameras, and time-of-flight (ToF) cameras. Terminal device 101 is mainly used to acquire images of documents using at least two of the multiple cameras at different optical focal lengths, obtaining at least two document images, and providing these at least two document images to server device 102 for server device 102 to identify whether the at least two document images are injected images. The documents may include, but are not limited to: driver's licenses, vehicle registration certificates, student IDs, access cards, public transport cards, social security cards, residence permits, work permits, and passes.
[0027] The server device 102 can be any computing device capable of providing network services, data processing, or authentication functions, including but not limited to: servers, server clusters, cloud platforms, virtual machines, container instances, edge computing nodes, or computing units in data centers.
[0028] In the embodiments described in this specification, a target application is installed on the terminal device. The target application may include, but is not limited to, financial payment applications, government service applications, and digital identity verification applications. The target application displays a target page where users can perform identity authentication. For ease of distinction and description, the following description will use the target page as the first page as an example. The first page may be a user interface (UI) page in the target application used to implement image acquisition functionality. The first page can assist users in accessing the camera on the terminal device, previewing live footage, and acquiring images of identification documents. The first page includes image acquisition controls, which are used to trigger the entry control that displays the viewfinder. For example, the first page may be a "document shooting page," a "QR code recognition page," or a "photo upload page," guiding users to upload photos of the front and back of their identification documents to complete the account opening process. The image acquisition controls may include, but are not limited to, text buttons, icon buttons, clickable image placeholder areas, interactive fields in forms, and function cards.
[0029] When the image acquisition control is triggered, the terminal device can respond to the trigger operation by displaying a viewfinder, which shows the real-time image data from the main camera among multiple cameras. The implementation of the viewfinder is not limited. For example, the viewfinder can be a partial UI area on a first page or a second page, embedded as a rectangle of fixed size within the first or second page. The second page differs from the first page. Alternatively, the viewfinder can be displayed full-screen on the terminal device's screen, with the screen occupied by the real-time image from the main camera.
[0030] In the embodiments of this specification, if a document to be captured appears in the real-time video data, at least two of the multiple cameras are used to capture images of the document at different optical focal lengths, resulting in at least two document images. The two document images are captured by at least two cameras at different optical focal lengths. Specifically, the at least two cameras can capture at least two document images simultaneously, or the time interval between capturing at least two document images can be less than a set time interval threshold.
[0031] Due to differences in optical focal length, at least two images can have inherent and predictable physical differences in field of view, perspective, and document size. In an "image injection" attack, regardless of which virtual camera is fed data, the content of the document image is identical, resulting in a high degree of similarity between the injected images. By comparing the differences between at least two images on the server-side device, it is possible to accurately determine whether the data collection originated from a genuine physical photograph or was a forged image injection.
[0032] Based on this, at least two document images can be sent to the server device. If the similarity between any two document images exceeds a set first similarity threshold, the server device determines that at least two document images are injected images.
[0033] Specifically, the server device can receive at least two document images sent by the terminal device; the at least two document images are obtained by cameras with different optical focal lengths capturing images of the document, or the at least two document images are injected images used to replace the real captured images during the image acquisition process. The server device can determine whether the at least two document images are real images captured in real time by a camera or an injection attack.
[0034] The server-side device calculates the similarity between any two ID document images from at least two available images. For real-life photography, due to differences in optical focal length, lens angle, distortion, and shooting distance, the size, shape, and perspective of the ID document vary between any two images. This difference is not random but determined by the laws of physical optics, conforming to the principles of optical imaging, such as pinhole camera models, lens distortion, and perspective projection. The similarity between any two ID document images in a real-life shooting scenario falls within a certain range, such as 0.7 to 0.9. For image injection, the injected at least two images are cropped from the original images. At the pixel level, the injected at least two images are completely identical, and their similarity approaches 1.0.
[0035] Therefore, if the similarity between two document images exceeds a set first similarity threshold, then at least two document images are determined to be injection images. For example, the first similarity threshold could be 0.85, 0.9, or 0.95. For instance, among the at least two document images, each pair of document images is considered a pair, and the calculated similarity of the pair is calculated. If the similarity between any pair of document images exceeds the set first similarity threshold, then at least two document images are determined to be injection images.
[0036] Optionally, if the similarity between any two document images is higher than a set second similarity threshold, and the similarity between any two document images is lower than or equal to a set first similarity threshold, then at least two document images are determined to be non-injected images. Non-injected images indicate that at least two document images were captured in a secure environment and can be used for subsequent identity verification. The second similarity threshold may include, but is not limited to, 0.68, 0.7, or 0.75. This means that in the at least two document images, every two document images form a pair, the similarity of each pair of document images is higher than the set second similarity threshold, and the similarity between any two document images is lower than or equal to the set first similarity threshold.
[0037] In the embodiments of this specification, when a user photographs their ID card, at least two cameras on the terminal device are used to capture at least two images of the ID card at different optical focal lengths. Because the same ID card is captured using different optical focal lengths, and the ID card images differ in field of view, perspective, and subject size, the similarity between the at least two ID card images is relatively high, i.e., less than a first similarity threshold (e.g., 90%), but not close to the theoretical upper limit (e.g., 100%). If there is a risk of injection, a sub-image is obtained by cropping the original injection image. The sub-image is part of the original injection image, and the similarity between the original injection image and the sub-image is close to the theoretical upper limit. Therefore, if the similarity between two ID card images exceeds the set first similarity threshold, the at least two ID card images are determined to be injection images, improving the accuracy of image injection detection.
[0038] In one optional embodiment, when at least two ID card images are captured, the content displayed in the viewfinder is not limited. For example, in one example, at least two cameras capturing at least two ID card images include the main camera. The viewfinder pauses the preview of the live video data and displays the static ID card image captured by the main camera, ensuring that the ID card image seen by the user is consistent with the operation feedback (what you see is what you get), avoiding confusion caused by sudden screen switching. The user adjusts the ID card position based on the main camera's view, while other cameras silently capture data in the background, without affecting the workflow. The main camera is used for interaction, and the ID card images captured by other cameras are used to compare the similarity between ID card images to identify injection attacks. In another example, the ID card image with higher clarity is selected from at least two ID card images and displayed in the viewfinder.
[0039] Optionally, the viewfinder may include two layers: the upper layer displays a semi-transparent guide frame, alignment lines, or status prompts to guide the user to place the document in the appropriate position, and the lower layer displays real-time image data. For example, status prompts may include, but are not limited to, "Please place your ID card in the frame" or "Keep it stable."
[0040] In one optional embodiment, the terminal device uses at least two of the multiple cameras to capture images of the document at different optical focal lengths, and the implementation method for obtaining at least two images of the document is not limited. An exemplary description follows.
[0041] In one example, metadata for each of multiple cameras is obtained, including at least functional information and optical focal length; based on the functional information of each of the multiple cameras, candidate cameras with imaging capabilities are selected from the multiple cameras, including the main camera; from the other cameras in the candidate cameras excluding the main camera, the candidate camera whose optical focal length differs from that of the main camera by a value greater than a set difference threshold is selected as the target camera; the main camera and the target camera are invoked to capture images of the document to obtain at least two images of the document.
[0042] The camera's metadata describes its basic capabilities and current status. This metadata may include, but is not limited to: optical focal length, functional information, whether it supports autofocus, whether it has zoom capability, image sensor size, and current availability. For example, the main camera's optical focal length could be 26mm, the ultra-wide-angle camera's focal length could be 13mm, and the telephoto camera's focal length could be 52mm. Functional information may include, but is not limited to: main camera, ultra-wide-angle, telephoto, depth of field, and macro capabilities. For example, a camera with successful functionality may include, but is not limited to, main camera, ultra-wide-angle, and telephoto. Current availability may include, but is not limited to: whether it is occupied, whether it is enabled, and whether it is damaged.
[0043] The threshold value for the difference is not limited. For example, the threshold value may include, but is not limited to, 8mm, 10mm, and 15mm. When the difference in optical focal length exceeds the set threshold value, there will be a significant difference in the viewing angle and imaging scale of the document image.
[0044] The number of target cameras can be one or more, and there is no limit to the number.
[0045] In another example, the metadata of each of the multiple cameras is obtained, including at least functional information and optical focal length; based on the functional information of each of the multiple cameras, candidate cameras with imaging capabilities are selected from the multiple cameras; according to the order of the optical focal length of the multiple cameras from smallest to largest, the first camera with a shorter optical focal length is selected from the multiple cameras, and the last camera with a longer optical focal length is selected; the two selected cameras are used to capture images of the document to obtain two images of the document.
[0046] In an optional embodiment, if the document to be captured appears in the real-time video data, before calling at least two of the multiple cameras to capture images of the document at different optical focal lengths and obtaining at least two document images, the method further includes at least one of the following operations when the document to be captured appears in the real-time video data: In one example, if the document to be captured coincides with the document guide frame included in the viewfinder, at least two cameras from multiple cameras are used to capture images of the document at different optical focal lengths, resulting in at least two document images. For example, a semi-transparent document guide frame, such as a rectangular area, is superimposed on the viewfinder to fit the document's size ratio. The system detects whether there are approximately rectangular objects with the same aspect ratio as the document in the real-time image data. When the detected document outline coincides with the position, scale, and angle of the document guide frame, it is determined that the document is in a suitable position, and the document image can be captured.
[0047] In another example, if the document to be captured includes specific visual elements, at least two of the multiple cameras are used to capture images of the document at different optical focal lengths, resulting in at least two document images. Specific visual elements may include, but are not limited to, icons, geometric images, text blocks, and blocks with specific color combinations unique to the document. For example, an image detection model can be used to perform feature recognition on the document appearing in real-time video data to identify key visual elements on the document (e.g., specific patterns, specific geometric shapes, text areas, etc.). If at least two of the above types of visual elements are detected in real-time video data across consecutive frames, and the confidence level of each type of visual element is higher than a preset threshold (e.g., 0.85), at least two of the multiple cameras are used to capture images of the document at different optical focal lengths, resulting in at least two document images.
[0048] In another example, if a trigger operation is detected on the shooting control corresponding to the viewfinder, at least two of the multiple cameras are used to capture images of the document at different optical focal lengths, resulting in at least two images of the document. For example, if the viewfinder includes a "Shoot" button, and the user clicks the "Shoot" button, the terminal device uses at least two of the multiple cameras to capture images of the document at different optical focal lengths and checks whether the captured document image includes the document. If not, it prompts "No document detected, please re-capture"; if detected, it sends at least two document images to the server device.
[0049] In one optional embodiment, the method of sending at least two ID card images to the server device is not limited. In one example, at least two ID card images are encrypted, and the encrypted at least two ID card images are sent to the server device. In another example, at least two metadata corresponding to each of the at least two ID card images are obtained, and the at least two metadata and at least two ID card images are encrypted to obtain an encrypted data packet. The encrypted data packet is sent to the server device so that the server device can perform pre-verification of the at least two ID card images based on the at least two metadata. Pre-verification refers to the verification process performed on at least two ID card images before identifying whether the at least two ID card images are the result of an injection attack. Pre-verification may include, but is not limited to: verifying whether the at least two ID card images come from a real physical camera, verifying whether the at least two ID card images conform to the physical laws of multi-view optical imaging, and verifying the temporal consistency of the at least two ID card images.
[0050] Optionally, the implementation of the terminal device receiving at least two document images sent by the terminal device is not limited.
[0051] In one example, an implementation of receiving at least two document images sent by a terminal device includes: receiving at least two document images sent by the terminal device.
[0052] In another example, the implementation of receiving at least two document images sent by the terminal device includes: receiving an encrypted data packet and decrypting the encrypted data packet to obtain at least two document images and their respective corresponding metadata.
[0053] Optionally, a method for determining at least two document images as injection images if the similarity between any two document images exceeds a set first similarity threshold includes: performing at least one pre-verification on the at least two document images based on the metadata corresponding to each of the at least two document images; if the verification results are both positive, then determining at least two document images as injection images if the similarity between any two document images exceeds the set first similarity threshold.
[0054] Further, optionally, if any verification result is negative, then at least two document images are identified as injected images.
[0055] Optionally, the metadata corresponding to any document image shall include at least: camera identification information, optical focal length, physical photosensitive area size, acquisition timestamp, and terminal device model, camera configuration information, and image processing unit. The physical photosensitive area size describes the physical size of the area in the image sensor used to receive light and convert it into electrical signals, and is generally expressed in millimeters (mm) as its diagonal length, width, and height.
[0056] Based on this, an implementation method for verifying whether at least two ID card images are from real-time acquisition by a real camera, based on the metadata corresponding to at least two ID card images respectively, includes: verifying whether at least two ID card images are from at least two cameras based on camera identification information and optical focal length; verifying whether the imaging parameters of at least two ID card images satisfy the multi-view optical imaging law based on the physical photosensitive area size and optical focal length; verifying whether the acquisition time interval of at least two ID card images is less than a set time interval threshold based on the acquisition timestamp; and verifying whether the metadata corresponding to at least two ID card images is real data based on the terminal device model, camera configuration information, and image processing unit.
[0057] The implementation method for verifying whether at least two ID card images originate from at least two cameras based on camera identification information and optical focal length is not limited. For example, if the camera identification information in the metadata corresponding to two ID card images is the same, it indicates that the two ID card images may originate from the same camera, and the two ID card images are considered to pose an injection risk. As another example, if two ID card images are a telephoto image and a wide-angle image respectively, theoretically the optical focal lengths in the metadata of the two ID card images are different, but the optical focal lengths are the same in the data parsed from the actual data packet, then the two ID card images are considered to pose an injection risk.
[0058] The implementation method for verifying whether the imaging parameters of at least two document images satisfy the multi-view optical imaging law based on the physical photosensitive area size and optical focal length is not limited. For example, the physical photosensitive area size, optical focal length, resolution, etc. are obtained from the metadata of any document image. Based on the physical photosensitive area size, optical focal length, and resolution of any two document images, the theoretical field of view (FOV) and theoretical scaling ratio of any two document images are calculated. The centers of any two document images are aligned, and one document image is scaled to the scale of the other image according to the theoretical scaling ratio. Feature extraction is performed on the two document images of the same scale to obtain common visual feature points (such as the four corners of the document, the edges of text blocks, etc.). If any two document images are from a real camera, the common visual feature points of any two document images have no rotation, nonlinear distortion, or perspective abrupt change, satisfying the multi-view optical imaging law. If the following situations occur, it is determined that any two document images do not satisfy the multi-view optical imaging law: inconsistent perspective relationship (e.g., abrupt change in the tilt angle of the document), edge position offset exceeding the set error range, etc.
[0059] The implementation method of verifying whether the time interval between the acquisition of at least two document images is less than a set time interval threshold based on the acquisition timestamp is not limited. The acquisition time interval threshold for at least two images includes, but is not limited to, 1ms, 0.5ms, or 2ms. If the acquisition time interval is greater than or equal to the set time interval threshold, at least two document images may be forged images.
[0060] The implementation method for verifying whether the metadata corresponding to at least two ID document images is genuine, based on the terminal device's model, camera configuration information, and image processing unit, is not limited. For example, the metadata corresponding to any ID document image may include the device model, camera configuration, and image signal processing (ISP) version. If the terminal device submits abnormal metadata, it is determined that the metadata corresponding to at least two ID document images is fake data. For example, abnormal metadata includes, but is not limited to: the presence of cross-manufacturer or cross-brand parameters in the metadata, such as the device model being from a first manufacturer but the camera parameters being from a second manufacturer; and camera parameter specifications exceeding the preset parameter specifications of the terminal device, such as a first-model terminal device having camera parameters corresponding to a second-model terminal device.
[0061] In an optional embodiment, the server device is further configured to calculate the similarity between any two document images. For example, for any two document images, feature extraction is performed on the two document images to obtain the feature points and their corresponding descriptor vectors included in each of the two document images; based on the feature points and descriptor vectors included in each of the two document images, feature point matching is performed on the two document images to obtain the successfully matched target feature points; based on the number of target feature points and the total number of feature points included in the two document images, the similarity between the two document images is calculated.
[0062] Feature extraction methods can include, but are not limited to, Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), and Oriented Fast and Rotated BRIEF (ORB). Feature extraction includes keypoint detection and descriptor generation. Keypoint detection is used to locate key feature points in the document image that are salient, stable, and repeatable, such as corner points, edge intersections, and texture-rich feature points. Descriptor generation is used to calculate a numerical vector (i.e., a descriptor vector) for each detected feature point, which describes the local appearance information of the feature point's neighborhood, making the feature point have certain invariances, such as robustness to rotation, scale, and illumination changes.
[0063] The implementation method of matching feature points between any two document images based on their respective feature points and descriptor vectors to obtain successfully matched target feature points is not limited. For example, the distance between the descriptor vectors corresponding to feature points in any two document images can be calculated, such as Euclidean distance or Hamming distance. For example, for each feature point in document image A corresponding to descriptor vector X1, based on the distance between descriptor vector X1 and each descriptor vector in document image B, a descriptor vector X2 that is similar to descriptor vector X1 is found in document image B, and the feature point corresponding to descriptor vector X2 is successfully matched with the feature point corresponding to descriptor vector X1.
[0064] The method for calculating the similarity between any two ID card images based on the number of target feature points and the total number of feature points included in any two ID card images is not limited. For example, the ratio of the number of target feature points to the total number of feature points included in any two ID card images can be used as the similarity between any two ID card images. Another example is using the product of the ratio of the number of target feature points to the total number of feature points included in any two ID card images and the matching coverage rate as the similarity between any two ID card images. The matching coverage rate is an indicator used to measure the proportion of spatially or content-wise effectively matched regions between two images. It considers not only how many pairs of feature points are matched, but also whether these matched points cover important areas of the image. For example, the ID card image can be divided into several grids (e.g., 4×4, a total of 16 grids); the number of grids containing at least one target feature point can be counted; the coverage rate = number of grids with matched points / total number of grids; for example, 12 / 16 = 0.75, that is, the coverage rate is 75%.
[0065] In one alternative embodiment, the terminal device includes an image processing client, such as Figure 2 As shown in the illustration, the image processing client includes: an acquisition control module, a multi-camera detection and invocation module, an image acquisition and preprocessing module, a data packaging and encryption module, a secure transmission module, and wide-angle, main, and telephoto cameras. The image processing client can be implemented as a Software Development Kit (SDK) for easy integration and invocation by target applications.
[0066] The acquisition control module is used to schedule the entire acquisition process, display the viewfinder to the user, and trigger acquisition commands. For example, it can automatically take a picture when the document to be acquired is detected in the real-time image data, or trigger an acquisition command in response to the user's shooting command.
[0067] The multi-camera detection and retrieval module is used to detect whether the terminal device has multiple cameras with different focal lengths (such as a rear camera) during initialization. When acquiring images, it requests and obtains document images from different cameras according to a preset strategy (e.g., simultaneously retrieval of the main camera and the wide-angle camera).
[0068] Image acquisition and preprocessing module: Acquires raw image frames from camera hardware and performs necessary preprocessing, such as size normalization and compression, to optimize transmission efficiency.
[0069] The data packaging and encryption module is used to package multiple ID images (e.g., image A, image B) from different cameras, along with information from the terminal device, collection timestamps, and other metadata, into a single data structure and encrypt it to prevent middleware attacks.
[0070] The secure transmission module is used to upload encrypted data packets to the server device using secure protocols such as HyperText Transfer Protocol Secure (HTTPS).
[0071] For example, such as Figure 2 As shown, taking a server-side device running a server-side SDK as an example, the server-side SDK includes: a data receiving and decryption module, a multi-image parsing module, a multi-focal length image analysis module, a decision-making module, and an authentication result generation module.
[0072] The data receiving and decryption module is used to receive data packets uploaded by the client and decrypt them using a predefined key.
[0073] The multi-image parsing module is used to parse at least two document images (image A, image B, etc.) from the decrypted data packet.
[0074] The multifocal length image analysis module is used for: For any two document images, feature extraction is performed on the two document images to obtain the feature points and their corresponding descriptor vectors for each of the two document images.
[0075] Based on the feature points and their descriptor vectors included in any two document images, feature point matching is performed on any two document images to obtain the target feature points that are successfully matched.
[0076] The similarity between any two document images is calculated based on the number of target feature points and the total number of feature points included in any two document images.
[0077] The decision module is used for: If the similarity between two document images exceeds a set first similarity threshold, then at least two document images are identified as injected images.
[0078] If the similarity between any two document images is higher than the set second similarity threshold, and the similarity between any two document images is lower than or equal to the set first similarity threshold, then at least two document images are determined to be non-injected images.
[0079] The authentication result generation module integrates all analysis steps to generate the final authentication result for this request, and returns the result to the terminal device. The terminal device then displays the authentication result on the target application's application page. Examples include: authentication successful, authentication failed - information mismatch, and authentication failed - injection risk detected.
[0080] In the above embodiments, the terminal device uses at least two cameras with different optical focal lengths to capture images of the document, obtaining at least two document images. In an optional embodiment, a method is also provided whereby the terminal device uses a single camera with different digital focal lengths to capture images of the document at least twice, obtaining at least two document images. Specifically, the terminal device displays a target page, which includes an image capture control; in response to a trigger operation on the image capture control, a viewfinder is displayed, and real-time image data from the main camera is shown in the viewfinder; if the document to be captured appears in the real-time image data, the main camera is invoked to capture images of the document at least twice with different digital focal lengths, obtaining at least two document images; the at least two document images are sent to a server device, so that the server device can determine that the at least two document images are injected images if the similarity between any two document images exceeds a set first similarity threshold.
[0081] For a description of the target page, please refer to the aforementioned embodiments, which will not be repeated here.
[0082] In the process of calling the main camera to capture images of the document at least twice with different digital focal lengths to obtain at least two document images, the terminal device quickly captures at least two document images, and the time interval between capturing at least two document images is less than a set time interval threshold, for example, the time interval threshold is in the millisecond range.
[0083] In this method, at least two ID card images can be captured in the same capture session. Within the same capture session, camera parameters such as exposure time, resolution, and white balance, as well as optical characteristics such as focal length, aperture, and lens distortion, are consistent. However, within the same capture session, the camera output is still affected by various random noises, leading to slight differences between images captured repeatedly in the same scene. For example, random noise can include, but is not limited to, thermal noise, readout noise, and photon shot noise. Therefore, the similarity between at least two ID card images captured by a real camera is within a certain range, for example, 0.7 to 0.9. Image injection, on the other hand, injects a fixed-resolution original ID card image. The sub-ID card image obtained by cropping the original ID card image is a part of the original ID card image. The similarity between the original ID card image and the sub-ID card image is high, approaching 1.0.
[0084] Therefore, if the similarity between two document images exceeds a set first similarity threshold, the server device determines that at least two document images are injected images. If the similarity between any two document images is higher than a set second similarity threshold, and the similarity between any two document images is lower than or equal to the set first similarity threshold, then at least two document images are determined to be non-injected images.
[0085] In one optional embodiment, a method of calling the main camera to capture images of an ID card at least twice at different digital focal lengths to obtain at least two ID card images includes: controlling the main camera to capture images of the ID card at a first digital focal length to obtain a first ID card image; changing the focal length used by the main camera from the first digital focal length to a second digital focal length while keeping the relative position of the ID card and the camera unchanged; and controlling the main camera to capture images of the ID card at the second digital focal length to obtain a second ID card image.
[0086] The time interval between at least two image acquisitions must be less than a set time interval threshold, such as 1 ms (milliseconds). The at least two image acquisitions can belong to the same acquisition session. The first digital focal length and the second digital focal length are not limited. For example, the first digital focal length corresponds to a first digital zoom ratio, such as 1.0x, and the second digital focal length corresponds to a second digital zoom ratio, such as 1.5x.
[0087] Optionally, displaying real-time image data from the main camera in the viewfinder includes: displaying real-time image data captured by the main camera at the first digital focal length; after the main camera captures images of the document at least twice at different digital focal lengths to obtain at least two document images, it also includes: displaying the document image captured by the main camera at the first digital focal length in the viewfinder. Displaying the static document image captured by the main camera ensures that the document image seen by the user is consistent with the operation feedback (what you see is what you get), avoiding confusion caused by sudden screen switching. Meanwhile, image capture at the second digital focal length is performed silently in the background without affecting the preview stream.
[0088] In addition to the image processing system described above, in the embodiments of this specification, such as Figure 3a As shown, an image processing method is also provided, applicable to a terminal device with multiple cameras, such as... Figure 3a As shown, the method includes: S301. Display the target page, which includes image acquisition controls.
[0089] S302. Respond to the trigger operation of the image acquisition control, display the viewfinder, and display the real-time image data of the main camera among multiple cameras in the viewfinder; S303. If the document to be captured appears in the real-time video data, at least two of the multiple cameras are called to capture images of the document with different optical focal lengths, so as to obtain at least two images of the document. S304. Send at least two document images to the server device. The server device is used to determine at least two document images as injection images if the similarity between two document images exceeds a set first similarity threshold.
[0090] In an optional embodiment, when at least two cameras, including a main camera, are used to capture images of the document at at least two of the multiple cameras with different optical focal lengths to obtain at least two document images, the method further includes: displaying the document image captured by the main camera in the viewfinder.
[0091] In one optional embodiment, at least two of the multiple cameras are invoked to capture images of the document at different optical focal lengths to obtain at least two document images. This includes: acquiring metadata for each of the multiple cameras, the metadata including at least functional information and optical focal length; selecting candidate cameras with imaging capabilities from the multiple cameras based on the functional information of each camera, the candidate cameras including the main camera; selecting, from the other cameras besides the main camera, a candidate camera whose optical focal length differs from that of the main camera by a value greater than a set difference threshold as the target camera; and invoking the main camera and the target camera to capture images of the document to obtain at least two document images.
[0092] In one optional embodiment, before activating at least two of the multiple cameras to capture images of the document at different optical focal lengths and obtaining at least two document images, the method further includes at least one of the following operations: if the document to be captured coincides with the document guide frame included in the viewfinder, then activating at least two of the multiple cameras to capture images of the document at different optical focal lengths and obtaining at least two document images; if it is detected that the document to be captured includes a set visual element, then activating at least two of the multiple cameras to capture images of the document at different optical focal lengths and obtaining at least two document images; if a trigger operation is detected for the shooting control corresponding to the viewfinder, then activating at least two of the multiple cameras to capture images of the document at different optical focal lengths and obtaining at least two document images.
[0093] In one optional embodiment, sending at least two document images to a server device includes: encrypting at least two document images and at least two corresponding metadata to obtain an encrypted data packet, and sending the encrypted data packet to the server device, wherein the server device is used to perform pre-verification on at least two document images based on at least two metadata.
[0094] In addition to the image processing system described above, in the embodiments of this specification, such as Figure 3b As shown, an image processing method is also provided, suitable for server-side devices, such as... Figure 3b As shown, the method includes: T301. Receive at least two document images sent by the terminal device; the at least two document images are obtained by cameras with different optical focal lengths capturing images of the document, or are injected images used to replace the real captured images during the image acquisition process.
[0095] T302. If the similarity between two document images exceeds the set first similarity threshold, then at least two document images are determined to be injected images.
[0096] In an optional embodiment, the method of this application further includes: if the similarity between any two document images is higher than a set second similarity threshold, and the similarity between any two document images is lower than or equal to a set first similarity threshold, then at least two document images are determined to be non-injected images.
[0097] In an optional embodiment, the method provided in this specification further includes: extracting features from any two document images to obtain feature points and their corresponding descriptor vectors for each of the two document images; performing feature point matching on the two document images based on the feature points and their descriptor vectors for each of the two document images to obtain successfully matched target feature points; and calculating the similarity between the two document images based on the number of target feature points and the total number of feature points included in the two document images.
[0098] In one optional embodiment, receiving at least two document images sent by a terminal device includes: receiving an encrypted data packet, decrypting the encrypted data packet to obtain at least two document images and their corresponding metadata; if the similarity between any two document images exceeds a set first similarity threshold, then determining at least two document images as injection images includes: performing at least one pre-verification on the at least two document images based on their respective metadata; if all verification results are positive, then determining at least two document images as injection images when the similarity between any two document images exceeds the set first similarity threshold.
[0099] Optionally, the metadata corresponding to any ID image includes at least: camera identification information, optical focal length, physical photosensitive area size, acquisition timestamp, and terminal device model, camera configuration information, and image processing unit; based on the metadata corresponding to at least two ID images respectively, at least one pre-verification is performed on the at least two ID images, including: verifying whether the at least two ID images come from at least two cameras based on camera identification information and optical focal length; verifying whether the imaging parameters of the at least two ID images meet the multi-view optical imaging law based on physical photosensitive area size and optical focal length; verifying whether the acquisition time interval of the at least two ID images is less than a set time interval threshold based on the acquisition timestamp; and verifying whether the metadata corresponding to the at least two ID images is real data based on the terminal device model, camera configuration information, and image processing unit.
[0100] In addition to the image processing system described above, this specification also provides an image processing method applied to a terminal device, such as... Figure 4 As shown, the method includes: 401. Display the target page, which includes image capture controls.
[0101] 402. Respond to the trigger operation of the image acquisition control, display the viewfinder, and show the real-time image data of the main camera in the viewfinder.
[0102] 403. If the document to be captured appears in the real-time video data, call the main camera to capture the document at least twice with different digital focal lengths to obtain at least two images of the document.
[0103] 404. Send at least two document images to the server device. The server device is used to determine that at least two document images are injected images if the similarity between two document images exceeds a set first similarity threshold.
[0104] In one optional embodiment, the main camera is invoked to capture images of the document at least twice using different digital focal lengths to obtain at least two document images. This includes: controlling the main camera to capture images of the document using a first digital focal length to obtain a first document image; changing the focal length used by the main camera from the first digital focal length to a second digital focal length while keeping the relative position of the document and the camera unchanged; and controlling the main camera to capture images of the document using the second digital focal length to obtain a second document image.
[0105] Optionally, displaying real-time image data from the main camera in the viewfinder includes: displaying real-time image data captured by the main camera at a first digital focal length in the viewfinder; after calling the main camera to capture images of the document at least twice at different digital focal lengths to obtain at least two document images, it further includes: displaying the document image captured by the main camera at the first digital focal length in the viewfinder.
[0106] In addition to the image processing system described above, this specification also provides an image processing method applied to a terminal device, the method comprising: Display the target page, which includes image capture controls; In response to the trigger operation of the image acquisition control, the viewfinder is displayed, and the real-time image data of the main camera among multiple cameras is displayed in the viewfinder; If the document to be captured appears in the real-time video data, at least two of the multiple cameras are used to capture images of the document with different optical focal lengths, so as to obtain at least two images of the document. If the similarity between two document images exceeds a set first similarity threshold, at least two document images are identified as injection images.
[0107] In addition to the image processing system described above, this specification also provides an image processing method applied to a terminal device, the method comprising: Display the target page, which includes image capture controls; In response to a trigger operation on the image acquisition control, the viewfinder is displayed, and the real-time image data from the main camera is shown in the viewfinder; If the document to be captured appears in the real-time video data, the main camera is used to capture images of the document at least twice at different digital focal lengths to obtain at least two images of the document. If the similarity between two document images exceeds a set first similarity threshold, at least two document images are identified as injection images.
[0108] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 101 to 104 can be device A; or the execution subject of steps 101 and 102 can be device A, and the execution subject of step 103 can be device B; and so on.
[0109] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0110] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving data sent by B, or it can be understood as A indirectly receiving data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending data directly to A, or it can be understood as B indirectly sending data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.
[0111] This specification provides a schematic diagram of the structure of an image processing apparatus according to an exemplary embodiment. The apparatus includes: The display module is used to display the target page, which includes image capture controls; The display module is used to respond to the trigger operation of the image acquisition control, display the viewfinder, and display the real-time image data of the main camera among multiple cameras in the viewfinder; The calling module is used to call at least two of the multiple cameras to capture images of the document if the document to be captured appears in the real-time video data, so as to obtain at least two images of the document. The sending module is used to send at least two document images to the server device. The server device is used to determine that at least two document images are injection images if the similarity between two document images exceeds a set first similarity threshold.
[0112] The detailed implementation methods and beneficial effects of each step in the above-described apparatus have been described in detail in the foregoing embodiments, and will not be elaborated upon here.
[0113] This specification provides a schematic diagram of another image processing apparatus according to exemplary embodiments, the apparatus comprising: The receiving module is used to receive at least two document images sent by the terminal device; the at least two document images are obtained by cameras with different optical focal lengths capturing images of the document, or are injected images used to replace the real captured images during the image acquisition process. The determination module is used to determine at least two document images as injection images if the similarity between two document images exceeds a set first similarity threshold.
[0114] The detailed implementation methods and beneficial effects of each step in the above-described apparatus have been described in detail in the foregoing embodiments, and will not be elaborated upon here.
[0115] In this specification embodiment, an image processing apparatus is also provided, the apparatus comprising: The display module is used to display the target page, which includes image capture controls; The display module is used to respond to the trigger operation of the image acquisition control, display the viewfinder, and display the real-time image data of the main camera in the viewfinder; The calling module is used to call the main camera to capture images of the document at least twice at different digital focal lengths if the document to be captured appears in the real-time video data, so as to obtain at least two images of the document. The sending module is used to send at least two document images to the server device. The server device is used to determine that at least two document images are injection images if the similarity between two document images exceeds a set first similarity threshold.
[0116] In this specification embodiment, an image processing apparatus is also provided, the apparatus comprising: The display module is used to display the target page, which includes image capture controls; The display module is used to respond to the trigger operation of the image acquisition control, display the viewfinder, and display the real-time image data of the main camera among multiple cameras in the viewfinder; The calling module is used to call at least two of the multiple cameras to capture images of the document if the document to be captured appears in the real-time video data, so as to obtain at least two images of the document. The determination module is used to determine at least two document images as injection images if the similarity between two document images exceeds a set first similarity threshold.
[0117] In this specification embodiment, an image processing apparatus is also provided, the apparatus comprising: The display module is used to display the target page, which includes image capture controls; The display module is used to respond to the trigger operation of the image acquisition control, display the viewfinder, and display the real-time image data of the main camera in the viewfinder; The calling module is used to call the main camera to capture images of the document at least twice at different digital focal lengths if the document to be captured appears in the real-time video data, so as to obtain at least two images of the document. The determination module is used to determine at least two document images as injection images if the similarity between two document images exceeds a set first similarity threshold.
[0118] The detailed implementation methods and beneficial effects of each step in the above-described apparatus have been described in detail in the foregoing embodiments, and will not be elaborated upon here.
[0119] Figure 5 This specification illustrates a schematic diagram of an electronic device provided in an exemplary embodiment, which is applicable to the image processing method provided in the foregoing embodiments. Figure 5As shown, the electronic device 700 mainly consists of a communication interface 702, a user interface 704, a processor 706, and a memory 708. These components are interconnected and communicate with each other through a system bus, network, or other connection mechanism 410. The communication interface 702 enables the device 700 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 702 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 702 can also be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi (Wireless Fidelity), Bluetooth, Global Positioning System (GPS), or wide-area wireless interface such as WiMAX (Wireless Maximum) or LTE (Long Term Evolution). Of course, the communication interface 702 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 702 may also include multiple physical communication interfaces, such as a Wi-Fi interface, a Bluetooth interface, and a wide-area wireless interface.
[0120] User interface 704 includes receiving user input and providing output to the user. Therefore, user interface 704 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT (Cathode Ray Tube), LCD (Liquid Crystal Display), LED (Light Emitting Diode), display using DLP (Digital Light Processing) technology, printer, and other known or future similar devices. User interface 704 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other known or future similar devices. In some embodiments, user interface 704 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, electronic device 700 may support remote access from other devices via communication interface 702 or another physical interface (not shown). User interface 704 can be configured to receive user input, the position and movement of which can be indicated by an indicator or cursor described herein. User interface 704 can also be configured as a display device for rendering or displaying text fragments.
[0121] Processor 706 may include one or more general-purpose processors and / or special-purpose processors. Memory 708 may include one or more volatile and / or non-volatile memory components and may be integrated wholly or partially with processor 706. Memory 708 may include removable and non-removable components.
[0122] The processor 706 is capable of executing program instructions 718 (e.g., compiled or uncompiled program logic and / or machine code) stored in memory 708 to perform the various functions described herein.
[0123] Memory 708 may contain non-transitory computer-readable media, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Memory 708 stores program instructions that, when executed by device 700, enable device 700 to perform any of the methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Processor 706 executing program instructions 718 may cause processor 706 to use data 712.
[0124] For example, program instructions 718 may include an operating system 722 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 700 and one or more applications 720 (e.g., a browser, social application, or game application). Similarly, data 712 may include operating system data 716 and application data 714. Operating system data 716 is primarily accessible to the operating system 722, while application data 714 is primarily accessible to one or more applications 720. Application data 714 may reside in a file system visible or hidden from the user of device 700.
[0125] Application 720 can communicate with operating system 722 through one or more application programming interfaces (APIs). These APIs help application 720 read and / or write application data 714, transmit or receive information via communication interface 702, receive or display information on user interface 704, etc.
[0126] In some terminology, application 720 may be simply referred to as "app". Furthermore, application 720 can be downloaded to device 700 through one or more online app stores or app markets. However, applications can also be installed on device 700 in other ways, such as through a web browser or a physical interface on electronic device 700 (e.g., a USB port).
[0127] Accordingly, embodiments of this specification also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Accordingly, embodiments of this specification also provide a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above-described method embodiments. It should be understood that each step or combination of steps in the above-described method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above-described method embodiments.
[0128] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, product, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, product, or apparatus that includes that element.
[0129] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.
[0130] The terminology used in the embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” used in the embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. “Multiple” generally includes at least two, but does not exclude the inclusion of at least one. “A plurality” generally includes at least two, but does not exclude the inclusion of at least one.
[0131] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0132] The above are merely embodiments of this specification and are not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. An image processing method, characterized in that, Applied to a terminal device having multiple cameras, the method includes: Display the target page, which includes image acquisition controls; In response to a trigger operation on the image acquisition control, a viewfinder is displayed, and real-time image data of the main camera among the multiple cameras is displayed in the viewfinder; If the document to be collected appears in the real-time video data, at least two of the multiple cameras are used to collect images of the document at different optical focal lengths to obtain at least two document images. The at least two document images are sent to a server device, which determines that the at least two document images are injected images if the similarity between two document images exceeds a set first similarity threshold.
2. The method of claim 1, wherein, In the case where the at least two cameras include the main camera, after calling at least two of the multiple cameras to capture images of the document at different optical focal lengths and obtaining at least two document images, the method further includes: The viewfinder displays the image of the document captured by the main camera.
3. The method of claim 1, wherein, At least two of the multiple cameras are used to capture images of the document at different optical focal lengths, resulting in at least two document images, including: Obtain metadata for each of the plurality of cameras, wherein the metadata includes at least functional information and optical focal length; Based on the functional information of each of the multiple cameras, a candidate camera with imaging function is selected from the multiple cameras, and the candidate camera includes the main camera; From the candidate cameras other than the main camera, select the candidate camera whose optical focal length difference from the optical focal length of the main camera is greater than a set difference threshold as the target camera; The main camera and the target camera are used to capture images of the document to obtain at least two images of the document.
4. The method according to any one of claims 1 to 3, characterized in that, Before acquiring at least two images of the document by calling at least two of the multiple cameras with different optical focal lengths, the process includes at least one of the following operations: If the document to be captured coincides with the document guide frame included in the viewfinder, then at least two of the multiple cameras are used to capture images of the document at different optical focal lengths to obtain at least two document images. If the document to be collected is detected to include a set visual element, then at least two of the multiple cameras are called to collect images of the document at different optical focal lengths to obtain at least two document images. If a trigger operation is detected on the shooting control corresponding to the viewfinder, at least two of the multiple cameras are invoked to capture images of the document at different optical focal lengths, resulting in at least two images of the document.
5. The method according to any one of claims 1 to 3, characterized in that, Sending the at least two document images to the server device includes: The at least two document images and their corresponding at least two metadata are encrypted to obtain an encrypted data packet, and the encrypted data packet is sent to a server device. The server device is used to perform pre-verification of the at least two document images based on the at least two metadata.
6. An image processing method characterized by, Applicable to server-side devices, including: The receiving terminal device sends at least two document images; the at least two document images are obtained by cameras with different optical focal lengths capturing images of the document, or are injected images used to replace the real captured images during the image acquisition process; If the similarity between two document images exceeds a set first similarity threshold, then the at least two document images are determined to be injected images.
7. The method of claim 6, wherein, Also includes: If the similarity between any two document images is higher than the set second similarity threshold, and the similarity between any two document images is lower than or equal to the set first similarity threshold, then the at least two document images are determined to be non-injected images.
8. The method according to claim 6 or 7, characterized in that, Also includes: For any two document images, feature extraction is performed on the two document images to obtain the feature points and their corresponding descriptor vectors included in each of the two document images; Based on the feature points and their descriptor vectors included in any two document images, feature point matching is performed on the two document images to obtain the target feature points that are successfully matched. The similarity between any two document images is calculated based on the number of target feature points and the total number of feature points included in any two document images.
9. The method according to claim 6 or 7, characterized in that, Receiving terminal device sends at least two document images, including: Receive encrypted data packets, decrypt the encrypted data packets to obtain at least two document images and their corresponding metadata; if the similarity between any two document images exceeds a set first similarity threshold, then determine that the at least two document images are injected images, including: Based on the metadata corresponding to the at least two document images, at least one pre-verification is performed on the at least two document images; if the verification results are both yes, then if the similarity between any two document images exceeds a set first similarity threshold, the at least two document images are determined to be injected images.
10. The method of claim 9, wherein, The metadata corresponding to any document image includes at least: camera identification information, optical focal length, physical photosensitive area size, acquisition timestamp, as well as the model of the terminal device, camera configuration information, and image processing unit; Based on the metadata corresponding to the at least two document images, at least one pre-verification is performed on the at least two document images, including: Based on the camera identification information and the optical focal length, verify whether the at least two document images come from at least two cameras; Based on the physical photosensitive area size and the optical focal length, verify whether the imaging parameters of the at least two document images meet the multi-view optical imaging rules; Based on the acquisition timestamp, verify whether the acquisition time interval of the at least two document images is less than a set time interval threshold; Based on the model of the terminal device, camera configuration information, and image processing unit, verify whether the metadata corresponding to the at least two document images is genuine data.
11. An image processing method, characterized by, The method, which involves using an application terminal device, includes: Display the target page, which includes image acquisition controls; In response to a trigger operation on the image acquisition control, a viewfinder is displayed, and real-time image data from the main camera is shown in the viewfinder; If the document to be captured appears in the real-time video data, the main camera is invoked to capture images of the document at least twice with different digital focal lengths, so as to obtain at least two images of the document. The at least two document images are sent to a server device, which determines that the at least two document images are injected images if the similarity between two document images exceeds a set first similarity threshold.
12. The method of claim 11, wherein, The main camera is used to capture images of the document at least twice at different digital focal lengths, resulting in at least two images of the document, including: The main camera is controlled to capture an image of the document at a first digital focal length to obtain a first image of the document; While keeping the relative positions of the document and the camera unchanged, the focal length used by the main camera is changed from the first digital focal length to the second digital focal length. The main camera is controlled to capture an image of the document at a second digital focal length, thereby obtaining a second image of the document.
13. The method of claim 12, wherein, Displaying real-time image data from the main camera within the viewfinder includes: displaying real-time image data captured by the main camera at a first digital focal length within the viewfinder; After using the main camera to capture images of the document at least twice at different digital focal lengths to obtain at least two document images, the method further includes: displaying the document image captured by the main camera at the first digital focal length in the viewfinder.
14. An electronic device, comprising: include: A memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions for performing the steps of the method according to any one of claims 1-5, 6-10, and 11-13.
15. A computer readable storage medium storing a computer program, characterized in that, When executed by a processor, the computer program is capable of implementing the steps of the method described in any one of claims 1-5, 6-10, and 11-13.
16. A computer program product, characterized in that, include: A computer program / instruction, which, when executed by a processor, is capable of implementing the steps of the method according to any one of claims 1-5, 6-10, and 11-13.