Image recognition method and related apparatus
By automatically identifying card and document types using image recognition methods on electronic devices, the problem of users having to manually select document types is solved, improving the efficiency and accuracy of card and document scanning and optimizing the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, users need to manually select the card type, which makes the card scanning process cumbersome and reduces the user experience.
By using image recognition methods, electronic devices can automatically identify card types, highlight target objects, and automatically select control states based on image content, reducing unnecessary computation and improving recognition accuracy.
It improves the efficiency and accuracy of card and document scanning, optimizes the user experience, and reduces the tedious process of users selecting document types.
Smart Images

Figure CN120259606B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to image recognition methods and related devices. Background Technology
[0002] Some applications in electronic devices offer card scanning capabilities. These applications may include photo-taking apps and document editing apps. Card scanning can be used to reconstruct various types of cards, ensuring that the printed document closely resembles the original in size and shape.
[0003] However, in some applications, users are required to actively select the document type, and because there are many document types to choose from, the search process is cumbersome, which reduces the user experience. Summary of the Invention
[0004] The image recognition method and related apparatus provided in this application can intelligently identify the type of card and document during the card and document scanning process, and actively match the corresponding card and document type without requiring the user to actively select the document type. This improves the efficiency of card and document scanning, optimizes the card and document scanning process, and thus enhances the user experience.
[0005] In a first aspect, the image recognition method provided in the embodiments of this application includes:
[0006] The system displays a first interface, which includes a first control; receives a first operation that triggers the first control; responds to the first operation and displays a second interface, which includes a first area for displaying an image captured by the electronic device's camera; if the first area displays the first image, a third interface is displayed, where the first area of the third interface displays the first image, and the third interface also includes a second area where a second control is selected and a third control is unselected, with the second control being automatically selected by the electronic device based on the image content of the first image; if the first area displays the second image, a fourth interface is displayed, where the first area of the fourth interface displays the second image, the third control in the second area of the fourth interface is selected and the second control in the second area is unselected, with the third control being automatically selected by the electronic device based on the image content of the second image. This improves the efficiency of card scanning, optimizes the card scanning process, and thus enhances the user experience.
[0007] In one possible implementation, the electronic device stores image feature information corresponding to each image type. After the second image is displayed in the first area of the fourth interface, the method further includes: outputting one or more sets of image types and corresponding confidence scores for the second image based on its content; recognizing the text content of the second image; and selecting a third control in the second area of the fourth interface, specifically when the matching degree between the image feature information of the first image type and the text content of the second image is greater than or equal to a first threshold. The first image type corresponds to a first confidence score, which is the maximum value among one or more sets of confidence scores, and the third control matches the first image type. This results in a relatively high calculated consistency, eliminating the need to continue comparing card types with lower confidence scores, thus reducing unnecessary computation.
[0008] In one possible implementation, the second image includes multiple target objects, among which a first target object is included. The method further includes highlighting the first target object in a first region of the fourth interface, wherein the center of the first region is located within the region containing the first target object. This increases the probability of selecting the card or document the user wants to scan, thereby improving the accuracy of recognition and enhancing the user experience.
[0009] In one possible implementation, the second image includes multiple target objects, among which a first target object is included. The method further includes highlighting the first target object in a first region of the fourth interface, wherein the distance between the center of the first region and the first target object is less than the distance between the center of the first region and other target objects, and the difference in distance is greater than or equal to a second threshold. In this way, since users are more likely to place the card to be recognized close to the center of the first region, the user can more likely select the card they want to scan, increasing the probability of correct recognition.
[0010] In one possible implementation, the second image includes multiple target objects, including a first target object and a second target object. The distance between the center of the first region and the first target object is a first distance, and the distance between the center of the first region and the second target object is a second distance. The method further includes highlighting the first target object in the first region of the fourth interface. The area of the region containing the first target object is larger than the area containing the second target object. The difference between the second distance and the first distance is less than a second threshold, and the second distance is less than the distance between the center of the first region and other target objects. In this way, since the card to be recognized is more likely to occupy a larger area in the first region, the card that the user wants to scan can be selected with a higher probability, increasing the probability of correct recognition.
[0011] In one possible implementation, the fourth interface further includes a fourth control, and the method further includes: receiving a second operation that triggers the fourth control; responding to the second operation, displaying a second interface; and, if the first area contains the third image, displaying a fifth interface, wherein the first area of the fifth interface displays the third image, the third control in the second area of the fifth interface is selected, the second control in the second area is unselected, and the third control is automatically selected by the electronic device based on the image content of the third image, and the similarity between the third image and the second image is greater than or equal to a third threshold. In this way, even after the user accidentally clicks the back button, the application can still correctly identify the card type.
[0012] In one possible implementation, after the first area of the fifth interface displays the third image, it further includes: based on the image content of the third image, outputting one or more sets of image types and corresponding confidence scores for the third image; the third control in the second area of the fifth interface is selected, including: when the second confidence score of the first image type of the third image is greater than the confidence scores corresponding to other image types of the third image, and the difference between the second confidence score and other confidence scores is greater than or equal to a fourth threshold, the third control in the second area of the fifth interface is selected. Thus, by judging the confidence threshold, the accuracy of the application's recognition can be determined. If the difference between the second confidence score and other confidence scores is greater than or equal to the fourth threshold, it indicates that the confidence difference is large, and the application is likely to be accurate in recognition.
[0013] In one possible implementation, the method further includes: receiving a third operation that triggers the fourth control; displaying a second interface in response to the third operation; and, if the first area contains the fourth image, displaying a sixth interface, wherein the first area of the sixth interface displays the fourth image, the fifth control in the second area of the sixth interface is selected, the third control in the second area is unselected, and the fifth control is automatically selected by the electronic device based on the image content of the fourth image, and the similarity between the fourth image and the second image is greater than or equal to a third threshold. In this way, even after a user accidentally clicks the back button, the application can identify the card as another type, thereby determining the type the user wants to identify.
[0014] In one possible implementation, after the fourth image is displayed in the first area of the sixth interface, the method further includes: based on the image content of the fourth image, outputting one or more sets of image types and corresponding confidence scores for the fourth image; increasing the value of the third confidence score if the second confidence score of the first image type of the fourth image is greater than the third confidence score corresponding to the second image type of the fourth image, the difference between the second and third confidence scores is less than a fourth threshold, and the third confidence score is greater than the confidence scores of other image types; and selecting the fifth control in the second area of the sixth interface, including selecting the fifth control in the second area of the sixth interface when the third confidence score is greater than or equal to the second confidence score, and matching the fifth control with the second image type. In this way, the application can achieve the ability to learn autonomously based on user feedback, thereby improving the user experience.
[0015] In one possible implementation, the method further includes: receiving a fourth operation that triggers the fourth control; responding to the fourth operation, displaying a second interface; and, if the first area contains the fifth image, displaying a seventh interface, where the first area of the seventh interface displays the fifth image, the sixth control in the second area of the seventh interface is selected, the third control in the second area is unselected, and the sixth control is automatically selected by the electronic device based on the image content of the fifth image, and the similarity between the fifth image and the second image is less than a third threshold. In this way, the application can determine whether the user has changed their card based on the similarity between the fifth image and the second image, enabling more accurate image recognition based on the user's selection, thereby improving the user experience.
[0016] In one possible implementation, when the first area is the second image, before displaying the fourth interface, the method further includes: displaying an eighth interface, which includes prompt information and a seventh control. The prompt information is used to indicate whether the image content of the second image has been correctly recognized; receiving a fifth operation that triggers the seventh control; and displaying the second interface in response to the fifth operation. In this way, the user can determine whether the currently recognized card type is correct based on the application's prompts, and the application can more accurately recognize cards based on the user's actions.
[0017] In one possible implementation, the second image includes multiple target objects, with the first object among the multiple target objects highlighted. The method further includes: receiving a sixth operation that triggers the second object among the multiple target objects; responding to the sixth operation, highlighting the second object, and setting an eighth control in the second area of the fourth interface to a selected state. The eighth control is automatically selected by the electronic device based on the image content of the second object. In this way, when the user clicks on other objects, the application can display the object to be recognized according to the user's selection, which can improve the image recognition accuracy.
[0018] Secondly, embodiments of this application provide an image recognition apparatus, which may be an electronic device, a chip or chip system within an electronic device. The apparatus may include a processing unit and a display unit. The processing unit is used to implement any processing-related method executed by the electronic device in the first aspect or any possible implementation of the first aspect. The display unit is used to implement any display-related method executed by the electronic device in the first aspect or any possible implementation of the first aspect. When the apparatus is an electronic device, the processing unit may be a processor. The apparatus may further include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to cause the electronic device to implement the methods described in the first aspect or any possible implementation of the first aspect. When the apparatus is a chip or chip system within an electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to cause the electronic device to implement the methods described in the first aspect or any possible implementation of the first aspect. The storage unit may be a storage unit within the chip (e.g., a register, cache, etc.), or a storage unit located outside the chip within the electronic device (e.g., a read-only memory, random access memory, etc.).
[0019] For example, the display unit is used to display a first interface, a second interface, a third interface, and a fourth interface. The processing unit is used to receive a first operation that triggers the first control.
[0020] In one possible implementation, the processing unit is configured to output one or more sets of image types of the second image and the confidence scores corresponding to the image types based on the image content of the second image; it is also configured to identify the text content of the second image; specifically, it is configured to select the third control in the second area of the fourth interface when the matching degree between the image feature information of the first image type of the second image and the text content of the second image is greater than or equal to a first threshold.
[0021] In one possible implementation, a display unit is used to highlight a first target object in a first area of the fourth interface, wherein the center of the first area is located in the area where the first target object is located.
[0022] In one possible implementation, the display unit is used to highlight the first target object in the first area of the fourth interface, wherein the distance between the center of the first area and the first target object is less than the distance between the center of the first area and other target objects, and the difference in distance is greater than or equal to a second threshold.
[0023] In one possible implementation, the display unit is used to highlight the first target object in the first area of the fourth interface, wherein the area where the first target object is located is larger than the area where the second target object is located, the difference between the second distance and the first distance is less than a second threshold, and the second distance is less than the distance between the center of the first area and other target objects.
[0024] In one possible implementation, a processing unit is used to receive a second operation that triggers the fourth control. A display unit is used to display a second interface and also to display a fifth interface.
[0025] In one possible implementation, the processing unit is configured to output one or more sets of image types and corresponding confidence scores of the third image based on the image content of the third image; and to select the third control in the second area of the fifth interface when the second confidence score of the first image type of the third image is greater than the confidence scores of other image types of the third image, and the difference between the second confidence score and other confidence scores is greater than or equal to a fourth threshold.
[0026] In one possible implementation, a processing unit is used to receive a third operation that triggers the fourth control. A display unit is used to display a second interface and also to display a sixth interface.
[0027] In one possible implementation, the processing unit is configured to output one or more sets of image types and corresponding confidence levels of the fourth image based on the image content of the fourth image; it is also configured to increase the value of the third confidence level; specifically, it is also configured to select the fifth control in the second area of the sixth interface when the third confidence level is greater than or equal to the second confidence level.
[0028] In one possible implementation, a processing unit is used to receive a fourth operation that triggers the fourth control. A display unit is used to display the second interface and also to display a seventh interface.
[0029] In one possible implementation, a processing unit is used to receive a fifth operation that triggers the seventh control. A display unit is used to display an eighth interface and also to display a second interface.
[0030] In one possible implementation, a processing unit is used to receive a sixth operation that triggers a second object among multiple target objects. A display unit is used to highlight the second object.
[0031] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory for storing code instructions, and the processor for running the code instructions to perform the methods described in the first aspect or any possible implementation of the first aspect.
[0032] Fourthly, this application provides a chip or chip system including at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a circuit. The at least one processor is used to run computer programs or instructions to perform the methods described in the first aspect or any possible implementation thereof. The communication interface in the chip can be an input / output interface, pins, or circuits, etc.
[0033] In one possible implementation, the chip or chip system described above in this application further includes at least one memory storing instructions. The memory can be an internal storage unit of the chip, such as a register or cache, or it can be a storage unit of the chip itself (e.g., read-only memory, random access memory, etc.).
[0034] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program or instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0035] In a sixth aspect, embodiments of this application provide a computer program product including a computer program, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0036] It should be understood that the second to sixth aspects of this application correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0038] Figure 2 A schematic diagram of the software structure of an electronic device provided in an embodiment of this application;
[0039] Figure 3 A schematic diagram of the interface of a camera application provided in an embodiment of this application;
[0040] Figure 4 This is a schematic diagram of an interface for selecting card type provided in an embodiment of this application;
[0041] Figure 5 A schematic diagram of an interface for providing more certificates in an embodiment of this application;
[0042] Figure 6A schematic diagram of a card scanning interface provided for an embodiment of this application;
[0043] Figure 7 A schematic diagram of a card recognition process provided in an embodiment of this application;
[0044] Figure 8 A schematic diagram of a card recognition pop-up window provided in an embodiment of this application;
[0045] Figure 9 A schematic diagram of a document application interface provided in an embodiment of this application;
[0046] Figure 10 A flowchart of an image recognition method provided in an embodiment of this application;
[0047] Figure 11 A flowchart illustrating a card / certificate object selection method provided in this application embodiment;
[0048] Figure 12 This is a schematic diagram of multiple card object areas in an image provided in an embodiment of this application;
[0049] Figure 13 A flowchart of confidence adjustment provided for embodiments of this application;
[0050] Figure 14 A schematic diagram illustrating an image recognition method provided in an embodiment of this application;
[0051] Figure 15 This is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation
[0052] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:
[0053] 1. Manhattan distance: Manhattan distance can be understood as the distance between two points in the north-south direction plus the distance between two points in the east-west direction. For example, on a plane, Manhattan distance can be the distance between two points in the x-axis direction plus the distance between two points in the y-axis direction.
[0054] For example, suppose the coordinates of the first point are (x1, y1) and the coordinates of the second point are (x2, y2). Then the distance between the two points in the x-axis direction is |x1-x2| and the distance between the two points in the y-axis direction is |y1-y2|. Therefore, the Manhattan distance d can satisfy the following formula: d=|x1-x2|+|y1-y2|.
[0055] 2. Terminology
[0056] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with substantially the same function and purpose. For example, "first chip" and "second chip" are used only to distinguish different chips and do not limit their order of execution. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.
[0057] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0058] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, and c can be single or multiple.
[0059] 3. Electronic equipment
[0060] The electronic devices in this application embodiment can also be any form of terminal device. For example, electronic devices may include: mobile phones, tablet computers, handheld computers, laptops, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, in-vehicle devices, wearable devices, electronic devices in 5G networks, or future evolved public land mobile communication networks (PLANs). The embodiments of this application do not limit the scope of electronic devices in a mobile network (PLMN).
[0061] By way of example and not limitation, in this embodiment, the electronic device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.
[0062] Furthermore, in this application embodiment, the electronic device can also be an electronic device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.
[0063] The electronic equipment in the embodiments of this application may also be referred to as: user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.
[0064] In this embodiment, the electronic device or various network devices include a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on top of the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also called main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.
[0065] For example, Figure 1 A schematic diagram of the electronic device is shown.
[0066] The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0067] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may include hardware, software, or a combination of software and hardware.
[0068] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0069] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the aforementioned memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0070] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a limitation on the structure of the electronic device. In other embodiments of this application, the electronic device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0071] Internal memory 121 can be used to store executable program code, including instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc. The data storage area may store data created during the use of the electronic device, etc. Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of the electronic device by running instructions stored in internal memory 121 and / or instructions stored in memory disposed within the processor.
[0072] Camera 193 is used to capture still images or videos. In some embodiments, an electronic device may include one or N cameras 193, where N is a positive integer greater than 1. For example, in embodiments of this application, camera-type applications or document editing applications can implement scanning functions based on images captured by camera 193.
[0073] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. In some embodiments, the electronic device may include one or N displays screens 194, where N is a positive integer greater than 1. The electronic device implements display functions through a GPU, display screen 194, and application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. For example, in an embodiment of this application, display screen 194 can be used to display images captured by camera 193.
[0074] The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information. Electronic devices can implement shooting functions through ISPs, cameras 193, video codecs, GPUs, displays 194, and application processors.
[0075] Figure 2 This is a software structure block diagram of an electronic device according to an embodiment of this application.
[0076] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: the application layer, the application framework layer, the Android runtime, the algorithm engine layer and system libraries, the hardware adaptation layer (HAL), and the kernel layer.
[0077] The application layer, also known as the application layer, can include a series of application packages. For example... Figure 2 As shown, the application package can include applications such as camera, gallery, notes, and documents. Applications can include system applications and third-party applications.
[0078] The application framework layer, also known as the framework layer, provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The framework layer can include some predefined functions.
[0079] like Figure 2 As shown, the Framework layer can include the Camera framework, resource manager, ActivityManager, and notification manager, etc.
[0080] The Camera framework can be used to manage cameras, acquire camera equipment and information, and implement related functions such as previewing, taking photos, and recording videos. In this embodiment, the Camera framework can provide system support for the scanning function, manage the camera's lifecycle and data acquisition, etc.
[0081] The file explorer can provide applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0082] ActivityManager can be used to manage the lifecycle of an application and determine its running status, such as whether it is started, paused, stopped, or destroyed. When an application goes into the background or is closed, ActivityManager can release the corresponding memory resources, ensuring the stability of the electronic device system.
[0083] The notification manager allows applications to display notification information in the status bar. It can be used to convey informational messages, or it can disappear automatically after a short time without user interaction.
[0084] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the control and management of the Android system.
[0085] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0086] The application layer and framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection. For example, in the embodiments of this application, the virtual machine can be used to perform functions such as multi-object detection, text detection and recognition, image classification, and image similarity calculation.
[0087] The algorithm engine layer may include multi-object detection algorithms, text detection and recognition algorithms, information comparison algorithms, image classification algorithms, image similarity algorithms, etc. It is understood that these algorithms include, but are not limited to, machine learning models, deep learning models, and corresponding traditional algorithms. In this embodiment, the algorithm engine layer can also be understood as a card / certificate recognition model.
[0088] In this embodiment of the application, the multi-target detection algorithm can detect whether a card or certificate object exists, and can also acquire one or more card or certificate objects in the image, and identify the location information of the card or certificate object, the type of the card or certificate object, and the confidence level of the type of the card or certificate object, etc.
[0089] Text detection and recognition algorithms can be used to detect text information in card and document objects.
[0090] Information comparison algorithms can compare the textual information of a card or certificate with prior knowledge.
[0091] Image classification algorithms can be used to classify card and document objects. Card and document objects can be categorized into types such as documents, cards and documents, and presentations (PowerPoint presentations, PPTs). Documents can include printed documents, books, business cards, etc.; cards and documents can include ID cards, bank cards, household registration books, etc.; and PPTs can include presentations shown on a conference screen or computer monitor.
[0092] It should be noted that the embodiments of this application obtain card information with user permission. After completing the card scanning function, the relevant information of the card will be cleared or saved based on the user's instructions. In other words, the user information (including but not limited to card information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0093] Image similarity algorithms can be used to calculate the similarity between images. For example, image similarity algorithms can include feature point matching algorithms, etc.
[0094] It is understood that the image recognition method used by the card scanning function provided by the application in this embodiment can be provided by the electronic device system. For ease of description, the following description will focus on the application as the execution subject.
[0095] In one possible implementation, the application can call the interface provided by the algorithm engine layer of the electronic device system to use multi-target detection algorithms, text detection and recognition algorithms, information comparison algorithms, image classification algorithms, image similarity algorithms, etc., thereby implementing the image recognition method of the embodiments of this application.
[0096] In another possible implementation, the application can integrate a software development kit (SDK), which may include relevant functions provided by the algorithm engine layer of the electronic device system, thereby implementing the image recognition method of the embodiments of this application.
[0097] The system library, also known as the Native layer, can include multiple functional modules. Examples include the Camera Service module and the OpenCV library.
[0098] In this embodiment of the application, the Camera Service module can be used to process requests from the Camera framework and implement camera shooting and other related functions.
[0099] The Hardware Abstraction Layer (HAL) is a layer of abstraction situated between the kernel layer and the Android runtime. The HAL can be a wrapper around a hardware driver, providing a unified interface for calls from upper-layer applications. The HAL may include modules such as a camera module, a sensor module, and an audio module. For example, in this embodiment, the camera module can be used to control the camera driver to capture images, thereby enabling functions such as scanning cards and ID cards.
[0100] The kernel layer is the layer between hardware and software. The kernel layer may include camera drivers, display drivers, audio drivers, and power management modules, etc. For example, in this embodiment, the camera driver can instruct the camera device to acquire images.
[0101] It should be noted that the embodiments of this application are only illustrated using the Android system. In other operating systems (such as Windows system, iOS system, etc.), as long as the functions implemented by each functional module are similar to those in the embodiments of this application, the solution of this application can also be implemented.
[0102] Some applications in electronic devices can provide card scanning functionality. These applications may include photo-taking apps and document editing apps. Card scanning can be used to trim and correct various types of cards, and to reproduce them 1:1, ensuring that the printed document is nearly identical in size and shape to the original.
[0103] In some implementations, card scanning may also be referred to as document scanning, card recognition, document identification, or card collection. For ease of description, this application will use card scanning as an example for illustration.
[0104] For example, Figures 3 to 6 The diagram illustrates the interface of a photo-taking application that provides card / ID scanning functionality. It can be understood that the interface for this function may include... Figures 3 to 6 Any one or more interfaces in the application, the display of which can be set by each application, and this application embodiment does not limit it.
[0105] like Figure 3 As shown in Figure a, interface 301 is the camera's shooting interface. Interface 301 can include various shooting modes such as large aperture mode, portrait mode, photo mode, and video mode. Interface 301 can also include more controls 302, etc.
[0106] When the user triggers more controls (302), in response to the user's triggering action, such as... Figure 3 As shown in b, the electronic device can display more interfaces 303 of the camera. More interfaces 303 may include functions such as HDR, slow motion, short film, time-lapse photography, live photos, and card scanning 304.
[0107] Optionally, interfaces 301 and 303 are exemplary interfaces. It is understood that different applications may have different interface displays, and interfaces 301 and 303 may each include more or less content. The specific content included in interfaces 301 and 303 is not limited in this embodiment.
[0108] When a user triggers a card scan 304 error, in response to the user's trigger action, such as... Figure 4 As shown, the electronic device can display a card / document type selection interface 401. The card / document type selection interface 401 may include a document type selection area 402, a document type area 403, and more document controls 404, etc.
[0109] Optionally, the card type selection interface 401 is an exemplary interface. It is understood that different applications may have different interface displays, and the card type selection interface 401 may include more or less content. The specific content included in the card type selection interface 401 is not limited in this embodiment.
[0110] Document type selection area 402 can include multiple document type options such as photo capture, serial number recognition, document scanning, exam paper / assignment scanning, and photo translation. Users can select the document type they want to scan.
[0111] For example, if a user selects the "scan document" option in the document type selection area 402, the document type area 403 can display the document type corresponding to the "scan document" option. The document type area 403 can include document types such as household registration book, driver's license, vehicle registration certificate, bank card, and ID card. The document type area 403 can also include more document controls 404.
[0112] Users can select the type of document they want to scan in the document type area 403.
[0113] In one possible scenario, if the card / document type area 403 does not contain the desired document type, the user can trigger the additional document control 404. In response to the user's trigger action, such as... Figure 5 As shown, the electronic device can display a more documents interface 501. The more documents interface 501 may include one or more document options 502, document sizes corresponding to the document options 503, search controls 504, return controls 505, etc.
[0114] The document option 502 can include social security cards, citizen cards, military officer ID cards, bankbooks, social security cards, passes, student ID cards, graduation certificates, etc. It's understandable that different document options 502 can correspond to different document sizes 503. This allows the application to crop the scanned document based on the size selected by the user, ensuring the scanned document's size closely matches the actual document size. The application can also stretch the border lines of the scanned document to make its size and shape closely approximate the original.
[0115] The 504 search control can be used to quickly search for the type of document a user is looking for.
[0116] The return control 505 can be used to exit the more documents interface 501. When the user triggers the return control 505, in response to the user's triggering operation, the application can return to the card type selection interface 401. The user can then reselect the card type.
[0117] Optionally, the "More Documents" interface 501 is an exemplary interface. It is understood that different applications may display different interfaces, and the "More Documents" interface 501 may include more or less content. The specific content included in the "More Documents" interface 501 is not limited in this embodiment.
[0118] In another possible scenario, when a user selects a document type in the document type area 403, in response to the user's triggering action, such as... Figure 6 As shown in Figure a, the electronic device can display a card scanning interface 601. The card scanning interface 601 may include prompts 602, a viewfinder 603, a zoom area 604, an image preview 605, and shooting controls 606, etc.
[0119] The prompt 602 may include "Please place the document completely within the viewfinder". The specific content of the prompt 602 is not limited in this embodiment.
[0120] The viewfinder 603 is the area for scanning cards and certificates. Optionally, in some implementations, the card scanning interface 601 may not include the viewfinder 603, and the application may use the border of the electronic device as the viewfinder. The specific display style of the viewfinder 603 is not limited in this embodiment.
[0121] Zoom area 604 can be used to adjust the shooting focal length. Image preview 605 can be used to preview pictures or videos in the album.
[0122] The shooting control 606 can be used to photograph cards and documents. When the user triggers the shooting control 606, in response to the user's triggering action, such as... Figure 6As shown in b, the electronic device can display a card preview interface 607. The card preview interface 607 may include a card preview area 608, a confirmation control 609, a return control 610, etc.
[0123] When the user triggers the confirmation control 609, in response to the user's triggering operation, the application can crop and correct the content in the card preview area 608, and can also execute the printing process.
[0124] When the user triggers the return control 610, the application can return to the card scanning interface 601 in response to the user's triggering operation. The user can then re-trigger the card scanning operation or perform other operations.
[0125] Optionally, the card scanning interface 601 and the card preview interface 607 are exemplary interfaces. It is understood that different applications may have different interface displays, and the card scanning interface 601 and the card preview interface 607 may each include more or less content. The specific content included in the card scanning interface 601 and the card preview interface 607 is not limited in this embodiment.
[0126] As can be seen from the above process, when the application scans cards and documents, users need to actively select the document type. Since there are many document types to select, the search process is cumbersome, which reduces the user experience.
[0127] In view of this, the image recognition method provided in this application embodiment can intelligently identify the type of card and document during the card and document scanning process, and actively match the corresponding card and document type, without requiring the user to actively select the document type, thereby improving the efficiency of card and document scanning, optimizing the card and document scanning process, and thus improving the user experience.
[0128] Figure 7 A schematic diagram of the card recognition interface provided in an embodiment of this application is shown.
[0129] like Figure 7 As shown in Figure a, the electronic device can display a card scanning interface 701. The card scanning interface 701 may include prompts 702, a viewfinder 703, a zoom area 704, an image preview 705, and shooting controls 706, etc.
[0130] The prompt 702 may include "Please place the document completely within the viewfinder". The specific content of the prompt 702 is not limited in this embodiment.
[0131] The viewfinder 703 is the area for scanning cards and certificates. Optionally, in some implementations, the card scanning interface 701 may not include the viewfinder 703, and the application can use the frame of the electronic device as the viewfinder. The specific display style of the viewfinder 703 is not limited in this embodiment.
[0132] The zoom area 704 can be used to adjust the shooting focal length. The image preview 705 can be used to save pictures or videos from the album. The shooting control 706 can be used to shoot cards or ID cards.
[0133] When the viewfinder 703 detects a card or certificate, the application can automatically identify the type of card or certificate and automatically switch to the corresponding card or certificate type interface 707, such as... Figure 7 As shown in b. The card type interface 707 may include a card type area 708, a shooting control 709, a return control 710, etc.
[0134] In the card type area 708, the application can highlight the recognized card type. For example, such as... Figure 7 As shown in b, taking the ID card type as an example, highlighting can include adding a border to the icon corresponding to the ID card, bolding the text "ID card", changing the font size of the text "ID card", and / or changing the color of the text "ID card", etc. The specific highlighting method is not limited in this application embodiment.
[0135] Optionally, after recognizing the card type, the application can also display a prompt message on the interface, which can be used to identify the currently recognized card type. For example, the application can display a toast message on the card type interface 707, which may include "Card type recognized as ID card". The prompt message may also contain other content, and the specific content of the prompt message is not limited in this embodiment.
[0136] It is understood that the card type interface 707 may only highlight the identified card type, or only display prompt information, such as a toast notification, or both highlight the identified card type and display prompt information. This application embodiment does not limit this.
[0137] In addition, the above-mentioned prompts can also be displayed on the card scanning interface 701, which is not limited in this embodiment.
[0138] When the application recognizes an incorrect card type, the user can trigger the return control 710. In response to the user's trigger, the application can return to the card scanning interface 701. The user can then re-trigger the card scanning operation or perform other actions.
[0139] When the application recognizes the correct card type, the user can trigger the shooting control 709. In response to the user's trigger action, such as... Figure 7 As shown in Figure c, the electronic device can display a card preview interface 711. The card preview interface 711 may include a card preview area 712, a confirmation control 713, a return control 714, etc.
[0140] When the user triggers the confirmation control 713, in response to the user's trigger operation, the application can crop and correct the content in the card preview area 712 and perform a 1:1 restoration, so that the size and shape of the printed document are close to the original. In addition, the application can also perform enhancement processing such as removing shadows and highlights from the card, which is not limited in this embodiment.
[0141] When the user triggers the return control 714, the application can return to the card scanning interface 701 in response to the user's triggering operation. The user can then re-trigger the card scanning operation or perform other operations.
[0142] Optional, Figure 7 The card scanning interface 701, card type interface 707, and card preview interface 711 shown are exemplary interfaces. It is understood that different applications may display different interfaces. Figure 7 The interface shown can include more or less content, respectively. Specifically... Figure 7 The content included in the interface shown is not limited in this embodiment of the application.
[0143] Optionally, in the card scanning interface 701 described above, when the application recognizes the card type, it may not automatically switch to the card type interface 707, but instead display a pop-up prompt. For example... Figure 8 As shown, interface 801 is a pop-up prompt interface, and interface 801 can display pop-up 802. Pop-up 802 may include: a prompt message 803 to prompt the user whether to switch the ID card mode, a "Do not remind again" button 804, a prompt message 805 to indicate that the user will not be reminded again, an "OK" button 806, and a "Cancel" button 807, etc.
[0144] The prompt message 803 may include "The document type is identified as A. Do you want to switch to A mode?", and the prompt message 805 may include "Do not remind again". The specific content of the prompt message 803 and the prompt message 805 is not limited in this embodiment.
[0145] Understandably, if the user triggers the "Don't remind me again" button 804, in response to this action, the application can automatically switch to the card type interface 707 the next time it recognizes a card type, instead of displaying the pop-up 802. If the user does not trigger the "Don't remind me again" button 804, the application will not automatically switch to the card type interface 707 the next time it recognizes a card type, but will instead display a pop-up prompt.
[0146] When the user presses the OK button 806, in response to the user's action, the application can cancel the display of the pop-up window 802 and jump to the card type interface 707 corresponding to document A. When the user presses the Cancel button 807, in response to the user's action, the application can cancel the display of the pop-up window 802 and return to the card scanning interface 701.
[0147] Optionally, in the card scanning interface 701 described above, when the application recognizes the card type, it can jump to the card preview interface 711 instead of displaying the card type interface 707. Understandably, this approach can be adopted when the application's accuracy in recognizing card types is high. After recognizing the card type, the application can perform processes such as cropping, correction, image enhancement, and content recognition on the card. This reduces user interaction and speeds up the card recognition process.
[0148] It is understandable that, in addition to the above Figure 3 In addition to camera apps providing card scanning functionality, document editing apps can also offer card scanning functionality.
[0149] For example, Figure 9 The diagram shows a schematic of the card scanning function in a document editing application.
[0150] like Figure 9 As shown in Figure a, interface 901 is a document editing interface, which may include document title, document editing time, document content, style controls, list controls, add control 902, recording control, etc.
[0151] When the user triggers the addition of control 902, in response to the user's trigger action, such as Figure 9 As shown in b, the electronic device can display a document adding interface 903. The document adding interface 903 may include table controls, hyperlink controls, document import controls, document scanning controls, table extraction controls, card scanning controls 904, etc.
[0152] It is understood that the icon and name of each control can be different, and the icon and name of each control can be set by the application. This application embodiment does not limit this.
[0153] Optionally, the document editing interface 901 and the document adding interface 903 are exemplary interfaces. It is understood that different applications may have different interface displays, and the document editing interface 901 and the document adding interface 903 may each include more or less content. The specific content included in the document editing interface 901 and the document adding interface 903 is not limited in this embodiment.
[0154] When the user triggers the card scanning control 904, in response to the user's triggering action, as described above... Figure 7 As shown, the electronic device can display a card scanning interface 701. The specific card scanning interface 701 can be found above. Figure 7 The relevant descriptions in the corresponding embodiments will not be repeated here.
[0155] Figure 10 A flowchart of an image recognition method according to an embodiment of this application is shown.
[0156] S1001, Viewfinder Preview Stream.
[0157] In the aforementioned card scanning interface 701, the application can start a camera preview stream to acquire the image in the viewfinder 703.
[0158] S1002, capture images by extracting frames at preset intervals.
[0159] The application can extract images from the camera preview stream at preset intervals. The specific preset interval can be customized by the application, and this embodiment does not limit it.
[0160] S1003. The application performs multi-target detection. Does it detect at least one object?
[0161] The application can perform multi-object detection on the acquired image and determine whether at least one object has been detected.
[0162] In this embodiment, cards, certificates, and documents awaiting scanning can be understood as targets or objects. For ease of description, the following description will use cards as an example.
[0163] If the application does not detect at least one card object, it can continue to acquire images from the camera preview stream for detection.
[0164] If the application detects at least one card object, it can execute step S1004 to select a card object.
[0165] S1004. The application automatically selects a card or certificate object, or the user selects it manually.
[0166] In one possible implementation, when the application detects more than one card object, it can select one of the multiple card objects as the object to be identified. The specific process of selecting the card object can be referred to below. Figure 11 The relevant descriptions of the corresponding embodiments will not be repeated here.
[0167] In another possible implementation, when the application detects more than one card object, the user can manually select one of the multiple card objects as the object to be identified.
[0168] After selecting the object to be identified, the application can perform step S1005.
[0169] S1005. Capture the image of the selected object area.
[0170] The application can extract the image corresponding to the object to be identified. On one hand, the application can execute step S1006, inputting the object to be identified into the image classification model. On the other hand, the application can execute step S1009, inputting the object to be identified into the text recognition model.
[0171] S1006, Image classification model identifies object categories.
[0172] Image classification models can categorize objects to be identified, thus identifying the object category. Object categories can be understood as the card / certificate types mentioned above, such as household registration booklets, driver's licenses, vehicle registration certificates, bank cards, and ID cards.
[0173] It is understandable that different objects to be identified have different object characteristics. For example, the object characteristics of an ID card may include a human image on the right side of the card.
[0174] In possible implementations, the image classification model can perform image recognition based on object features. The image classification model can perform image recognition in any possible way, and the embodiments of this application are not limited thereto.
[0175] After the image classification model classifies the object to be identified, step S1007 can be executed.
[0176] S1007. Output object categories are sorted by confidence level.
[0177] Image classification models can output the object category of the object to be identified and the corresponding confidence score. The confidence score can also be called the reliability score.
[0178] Optionally, the application can also sort the output content according to the confidence level, such as sorting in ascending order, descending order, or other sorting methods. This application embodiment does not limit the sorting.
[0179] Understandably, for the same object to be identified, an image classification model can output one or more sets of object categories and corresponding confidence scores. For example, the output of an image classification model may include: object category 1, confidence score 1; object category 2, confidence score 2; object category 3, confidence score 3, etc.
[0180] The above-mentioned object category 1 can also be called card 1, and confidence level 1 can also be called the confidence level of card 1; object category 2 can also be called card 2, and confidence level 2 can also be called the confidence level of card 2; object category 3 can also be called card 3, and confidence level 3 can also be called the confidence level of card 3.
[0181] Taking an ID card as an example, the image classification model's output can include: ID card (0.9); bank card (0.4); business card (0.2), etc. That is, the probability that the object is an ID card is 0.9, the probability that it is a bank card is 0.4, and the probability that it is a business card is 0.2, making the ID card the most likely candidate. Understandably, since the object is an ID card, the image classification model will output the ID card category with the highest confidence, which aligns with image recognition results.
[0182] After outputting the object category of the object to be identified and the confidence level corresponding to the object category, the application can execute step S1008.
[0183] S1008. Is it a predefined card type?
[0184] Understandably, due to the large number of document types, some applications cannot exhaustively list all documents. Therefore, applications can predefine some document types for certain card objects. For example, if a card recognition model can support the recognition of many document types, the application can select one or more document types as predefined card types from among the many document types that the card recognition model can recognize.
[0185] If the card type of the object to be identified is a predefined card type, then step S1010 can be executed.
[0186] If the card type of the object to be identified is not a predefined card type, it means that there is a high probability that the card type is incorrectly identified, and the application can re-execute step S1002.
[0187] S1009, The text recognition model identifies the text content of the object.
[0188] Text recognition models can identify the text content within an object to be recognized. Taking an ID card as an example, a text recognition model can identify relevant information on the ID card, such as name, number, issuing authority, and validity period.
[0189] After the text content in the object to be identified is detected, step S1010 can be executed.
[0190] S1010. Compare the actual object recognition text information with the prior object key information and check whether the consistency is greater than the preset value A.
[0191] Applications can record key information corresponding to different card types. For example, key information for an ID card may include name, number, issuing authority, and validity period, while key information for a bank card may include the bank's Chinese name, English name, and number. This key information corresponding to these card types can be referred to as prior knowledge.
[0192] After the text information is recognized by the text recognition model and the card / certificate type is identified by the image classification model, the application can compare the text information with the prior knowledge corresponding to the card / certificate type and determine the consistency of the information. Consistency can also be understood as similarity or matching degree.
[0193] Understandably, among the card types identified by the image classification model, card types with higher confidence can be prioritized for comparison. This results in a relatively high level of consistency, eliminating the need to continue comparing card types with lower confidence, thus reducing unnecessary computation.
[0194] The preset value A can be determined by the application based on experience or laboratory testing, and this application embodiment does not limit it.
[0195] If the consistency between the text information and the prior knowledge corresponding to the card type is greater than the preset value A, it indicates that the matching degree between the text information and the prior knowledge corresponding to the card type is very high. In other words, the accuracy of the document recognized by the application is high, and step S1011 can be executed.
[0196] If the consistency between the text information and the prior knowledge corresponding to the card type is less than or equal to the preset value A, it indicates that the matching degree between the text information and the prior knowledge corresponding to the card type is low. In other words, the document recognized by the application may not be accurate, and step S1002 can be executed to re-identify the object.
[0197] Optionally, the case where the value is equal to the preset value A can also be judged as having high consistency, but this application embodiment does not limit this.
[0198] S1011, Output card type.
[0199] The interface corresponding to the output card type can be found above. Figure 7 The card type corresponding to b is interface 707, which will not be described further.
[0200] S1012, Perform the scan operation.
[0201] S1013. Enter the card type as the subsequent input.
[0202] Once the user confirms that the card type output by the application is correct, a scanning operation can be performed, using the card type as subsequent input to scan and print the card object.
[0203] Optionally, step S1004 above may not be performed. The application may not need to automatically select a card or certificate object, or the user may need to select it manually. Instead, it may process multiple documents or cards or certificates at the same time. This application embodiment does not limit this.
[0204] Optionally, step S1008 can be omitted. The application can skip determining whether it is a predefined card type and instead, after obtaining the card type and its corresponding confidence level in step S1007, execute step S1010 to perform a consistency comparison between the text information and prior knowledge. This reduces the code's judgment process, thereby saving computing power.
[0205] Figure 11 A flowchart of the card object selection method is shown.
[0206] S1101, Open camera preview stream.
[0207] After the application opens the camera preview stream, it can detect multiple card objects included in the image in the camera preview stream and execute step S1102.
[0208] For example, such as Figure 12 As shown in Figure a, in image 1201 extracted from the camera preview stream, the application can detect region 1202 corresponding to object 1, region 1203 corresponding to object 2, and region 1204 corresponding to object 3. The center point of image 1201 is 1205. The center point of image 1201 can also be referred to as the camera preview stream center point.
[0209] Or, such as Figure 12 As shown in b, in the image 1201 extracted from the camera preview stream, the application can detect the region 1206 corresponding to object 4 and the region 1207 corresponding to object 5.
[0210] S1102. Is the center point of the camera preview stream within a certain object area?
[0211] The application can determine whether the center point of the camera preview stream is within a certain object area.
[0212] If the center point of the camera preview stream is located within a certain object area, for example Figure 12 If the center point 1205 of the camera preview stream is located in the region 1203 corresponding to object 2, then step S1103 can be executed.
[0213] If the center point of the camera preview stream is not within a certain object area, for example Figure 12If, in step b, the center point 1205 of the camera preview stream is not in the region 1206 corresponding to object 4, nor in the region 1207 corresponding to object 5, then step S1104 can be executed.
[0214] S1103, The card object at the center point of the preview stream is selected by default.
[0215] Select the card object located at the center of the preview stream as the object to be identified, and highlight the object to be identified in the interface. Highlighting can be understood as highlighting the border of the object to be identified, or providing text prompts about the object. The specific method of highlighting the object is not limited in this embodiment, as long as it effectively informs the user of the selected object to be identified.
[0216] After the application has selected the card object at the center of the preview stream by default, step S1108 can be executed.
[0217] S1104. Calculate the Manhattan distance from the center point of the preview stream to the center point of each object, and sort them.
[0218] For example, with Figure 12 Taking b as an example, the application can calculate the Manhattan distance from the center point of the preview stream 1205 to the center point of each object area. Understandably, using Manhattan distance can improve the stability of the calculated data. Assume the Manhattan distance from the center point of the preview stream 1205 to the center point of area 1206 is d1, and the Manhattan distance from the center point of the preview stream 1205 to the center point of area 1207 is d2. Of course, if there are more card objects, more Manhattan distances can be calculated, such as d3, d4, d5, etc.
[0219] After calculating the Manhattan distance between the center point 1205 of the preview stream and the center point of each object area, step S1105 can be executed.
[0220] Optionally, other methods can be used to calculate the distance between the preview stream center point 1205 and the center point of each object area. For example, the straight-line distance between the preview stream center point 1205 and the center point of each object area can be calculated. The specific method for calculating the distance between the preview stream center point 1205 and each object is not limited in this embodiment.
[0221] S1105. Whether the Manhattan distance difference between each object is less than or equal to the preset value B.
[0222] Understandably, this step can be viewed as a fault-tolerance mechanism. For example, using... Figure 12For example, if the Manhattan distances corresponding to objects 4 and 5 differ from the preset value B, it means that the distance between object 4 and the center point 1205 of the preview stream is similar to the distance between object 5 and the center point 1205 of the preview stream. This can be considered as the Manhattan distances corresponding to objects 4 and 5 being the same, and the Manhattan distance will no longer be used as the standard for selecting objects.
[0223] The preset value B can be set to 10% or other values. The specific value of the preset value B can be determined by the application based on experience or laboratory testing, and this application embodiment does not limit it.
[0224] Assume the Manhattan distance between the center point 1205 of the preview stream and the center point of object 4 is d4, and the Manhattan distance between the center point 1205 of the preview stream and the center point of object 5 is d5.
[0225] In one possible implementation, the application can calculate the relationship between the value of |d4-d5| and the preset value B.
[0226] In another possible implementation, the application can calculate the relationship between the value of |d4-d5| / d4 and the preset value B, or calculate the relationship between the value of |d4-d5| / d5 and the preset value B.
[0227] It is understood that the application can calculate the relationship between the Manhattan distance of each object and the preset value B in any way, as long as it can identify the proximity of the Manhattan distances between objects. For example, the application can also calculate the relationship between multiples of |d4-d5| and the preset value B, or calculate the relationship between multiples of |d4-d5| / d4 and the preset value B, or calculate the relationship between multiples of |d4-d5| / d5 and the preset value B, etc., and the embodiments of this application are not limited thereto.
[0228] If the Manhattan distance difference between all objects is less than or equal to the preset value B, step S1106 can be executed. If the Manhattan distance difference between all objects is greater than the preset value B, step S1107 can be executed.
[0229] Optionally, step S1107 can also be performed if the value is equal to the preset value B, but this embodiment of the application is not limited to this.
[0230] S1106: By default, select the object with the largest area among multiple objects.
[0231] by Figure 12 Taking b as an example, if the Manhattan distance difference between object 4 and object 5 is less than or equal to the preset value B, the application can select the object with the larger area in image 1201, either object 4 or object 5. For example, if region 1206 of object 4 is larger than region 1207 of object 5, the application can select object 4 as the object to be identified.
[0232] It is understandable that when the areas of object 4 and object 5 are similar, the second largest or larger object can also be selected, and this application embodiment does not limit this.
[0233] After the application selects the object to be identified, step S1108 can be executed.
[0234] S1107. By default, the object with the smallest Manhattan distance among all objects is selected.
[0235] If the Manhattan distances of all objects differ by more than a preset value B, the application can select the object corresponding to the minimum or smaller Manhattan distance as the object to be identified.
[0236] After the application selects the object to be identified, step S1108 can be executed.
[0237] S1108, Receive the operation of clicking on other areas of the preview stream.
[0238] It should be noted that after the application selects the object to be identified, the object can be specially marked to indicate the selected object to the user. This special marking may include highlighting the border of the object and / or displaying related prompts on the interface. The specific method of special marking is not limited in this embodiment, as long as the user can understand the object selected by the application.
[0239] In some scenarios, the object selected by the application for recognition may not be the document the user wants to identify. In this case, the user can click on other areas in image 1201 to reselect.
[0240] After the user clicks on other areas in image 1201, the application can execute step S1109 in response to the user's click operation.
[0241] S1109. The clicked area is the object area.
[0242] The application can determine whether the area clicked by the user is an object area. For example, an object area can include the area where the card or document object is located, while a non-object area can include blank areas in image 1201 or areas where no card or document object is displayed.
[0243] If the area clicked by the user is a non-object area, it means that there is no object that can be identified. In this case, the application will still use the object selected in step S1103, step S1106 or step S1107 above and execute step S1110.
[0244] If the area clicked by the user is an object area, it means that the user has reselected the object to be identified, and then step S1111 can be executed.
[0245] S1110, The object selection box does not change.
[0246] S1111: Select the object chosen by the user.
[0247] S1112, Receive the operation of clicking the preview stream area.
[0248] It is understandable that the user can click on the object area or non-object area multiple times in image 1201, that is, the user can repeatedly execute the above step S1108, and the application can execute steps S1109-S1111 multiple times.
[0249] It is understandable that, in the above Figure 7 In the card type interface 707 corresponding to b, the user can trigger the return control 710. Alternatively, as described above... Figure 8 In the pop-up window 801, the user can trigger the cancel control 807. In response to the user's trigger action, the application can return to... Figure 7 The corresponding card scanning interface 701 is for card 'a'. When the application rescans the card, it can adjust the confidence level based on the user's actions, thereby improving the application's ability to recognize cards.
[0250] Figure 13 A flowchart of confidence level adjustment is shown.
[0251] S1301, Start scanning to acquire image 1.
[0252] When the application scans on the card scanning interface 701, it can acquire image 1 and execute step S1302.
[0253] S1302, classification algorithm, output classification results, obtain confidence ranking.
[0254] The application can calculate and rank the confidence scores of the objects to be identified in Image 1. The specific implementation can be found above. Figure 10 The relevant descriptions of steps S1007, etc., will not be repeated here.
[0255] The application can obtain confidence information for image 1, such as the confidence of card 1, card 2, card 3, and card 3 corresponding to image 1.
[0256] S1303, the interface jumps to the interface corresponding to card 1.
[0257] After recognizing the card type of the card object, the application can control the jump from the card scanning interface 701 to the card type interface 707, or to the pop-up interface 801.
[0258] S1304. Did the user click "back"?
[0259] Understandably, if the user does not click the back button, it means that the application's identification is accurate and no confidence level adjustment is needed.
[0260] If the user clicks "back," one possible scenario is that the user no longer wants to continue the card scanning process and can click "back." Understandably, in this scenario, the application's judgment of the type of the object to be identified is likely accurate.
[0261] In another possible scenario, if the user believes that the application has identified the wrong card type, they can also click "back".
[0262] If it is a card / certificate type interface 707, the user can trigger the return control 710. If it is a pop-up window interface 801, the user can trigger the cancel control 807. In response to the user's trigger operation, step S1305 can be executed.
[0263] S1305, Start scanning to acquire image 2.
[0264] The application can rescan at the card scanning interface 701 to obtain image 2 and execute step S1302.
[0265] S1306. Algorithm compares the similarity between image 1 and image 2.
[0266] Since the application cannot determine the reason why the user clicked the back button, it can calculate the similarity between image 1 and image 2 and execute step S1307. In this way, the application can determine whether the user has changed their card by comparing the similarity of the two images.
[0267] It is understood that the application can save the previously acquired image information, which may include the image itself and / or image characteristic information, etc. The specific image information is not limited in this application embodiment, as long as it allows for comparison of image similarity.
[0268] S1307. Is the similarity greater than the similarity threshold?
[0269] If the similarity between Image 1 and Image 2 exceeds the similarity threshold, it indicates a high degree of similarity between them, suggesting the user has likely not changed their card. In this case, the application may have previously misidentified the card. Therefore, the application can proceed to step S1308 to calculate the confidence level and adjust it accordingly.
[0270] If the similarity between Image 1 and Image 2 is less than or equal to the similarity threshold, it indicates that the similarity between Image 1 and Image 2 is low, and the user has most likely changed their card. In this case, the application likely had no problem recognizing the card last time. Thus, the application can execute step S1302 to rescan the card and calculate the confidence level.
[0271] Optionally, if the similarity is equal to the similarity threshold, it can also be judged that image 1 and image 2 have a high similarity. This application embodiment does not limit this.
[0272] It is understood that the similarity threshold can be determined by the application based on experience or laboratory testing, and this application does not limit it.
[0273] S1308, classification algorithm, outputs classification results, and obtains confidence ranking.
[0274] The application can calculate and rank the confidence scores of the objects to be identified in image 2. The specific implementation can be found above. Figure 10 The relevant descriptions of steps S1007, etc., will not be repeated here.
[0275] The application can obtain the confidence information of image 2, such as: card 1, confidence of card 1, card 2, confidence of card 2, card 3, confidence of card 3, etc. It can be understood that since image 2 has been re-acquired, subsequent card 1, card 2, and card 3 refer to image 2.
[0276] S1309. Is the confidence level of Card 1 greater than that of Card 2, and is the difference between their confidence levels greater than the confidence threshold?
[0277] In this step, determining the confidence threshold can also be understood as a fault-tolerance mechanism. If the confidence levels of card 1 and card 2 are similar, it means the two cards are not significantly different and cannot be clearly distinguished. The application cannot accurately determine whether the current document is card 1 or card 2. Therefore, the application needs to adjust the confidence level of card 2.
[0278] For example, suppose the confidence information for image 2 may include: Card 1, 0.85; Card 2, 0.75; Card 3, 0.22. In this case, since the confidence scores of Card 1 and Card 2 are small, it indicates that the two cards are not significantly different and are not easily distinguishable. The application cannot accurately determine whether the current document is Card 1 or Card 2. Therefore, the application needs to adjust the confidence score of Card 2.
[0279] Therefore, if the confidence level of card 1 is less than or equal to the confidence level of card 2, or the difference between their confidence levels is less than or equal to the confidence threshold, the application can execute step S1311.
[0280] For example, suppose the confidence information for image 2 may include: Card 1, 0.85; Card 2, 0.34; Card 3, 0.22. In this case, since the confidence level of Card 1 is greater than that of Card 2, and the difference between the two confidence levels is significant, it indicates that the two cards are quite different and easily distinguishable. Therefore, the application can accurately determine whether the current document is Card 1 or Card 2.
[0281] Therefore, if the confidence level of card 1 is greater than that of card 2, and the difference between their confidence levels is greater than the confidence threshold, the application does not need to adjust the confidence level of card 2 and can proceed to step S1310.
[0282] It is understood that the confidence threshold can be determined by the application based on experience or laboratory testing, and this application does not limit it.
[0283] S1310, The interface jumps to the interface corresponding to Card 1.
[0284] It is understandable that step S1304 can still be executed on the interface corresponding to card 1. If the user clicks back, steps S1304-S1313 can be executed again, which will not be described again.
[0285] S1311, Update the confidence level of card 2: original confidence level * (1 + default value 3).
[0286] It is understandable that the preset value 3 can be interpreted as... Figure 13 The value of a% and the preset value of 3 can be determined by the application based on experience or laboratory testing, and this application embodiment does not limit it.
[0287] Based on the confidence level calculation, the confidence level of Card 2 is improved, increasing the probability of it being recognized. When the confidence level of Card 2 is higher than that of Card 1, the application can recognize the document as Card 2.
[0288] It is understandable that the confidence level of card 2 can be updated in a way that is not limited to the above implementation. Other calculation methods can also be used to update the confidence level, as long as the confidence level gradually increases during the continuous update process. In this way, the confidence level of card 2 may be higher than that of card 1, allowing the application to re-identify cards and correct incorrectly identified cards.
[0289] After updating the confidence level of card 2, step S1312 can be performed.
[0290] S1312. Is the confidence level of the updated card 2 greater than that of card 1?
[0291] The application can determine whether the confidence level of the updated card 2 is greater than that of card 1.
[0292] If the confidence level of the updated card 2 is less than or equal to the confidence level of card 1, the application will still recognize the document as card 1 and execute step S1310.
[0293] If the confidence level of the updated card 2 is greater than that of card 1, the application can recognize the document as card 2 and execute step S1313.
[0294] Optionally, if the confidence level of the updated card 2 is equal to the confidence level of card 1, the application can also recognize the document as card 2. This application embodiment does not limit this.
[0295] S1313, The interface jumps to the interface corresponding to card 2.
[0296] It is understandable that step S1304 can still be executed on the interface corresponding to card 2. If the user clicks back, steps S1304-S1313 can be executed again, which will not be described again.
[0297] In this embodiment of the application, by performing the above... Figure 13 According to the confidence adjustment process in the corresponding embodiment, the application can correct misidentified cards to correctly identified cards. In other words, the application can achieve self-learning based on user feedback, thereby improving the user experience.
[0298] The methods of this application will be described in detail below through specific embodiments. The following embodiments can be combined with each other or implemented independently, and the same or similar concepts or processes may not be described again in some embodiments.
[0299] Figure 14 An image recognition method according to an embodiment of this application is illustrated. Applied to an electronic device, the method includes:
[0300] S1401. Display the first interface, which includes the first control.
[0301] In this embodiment of the application, the first interface can be understood as an interface capable of providing card scanning functionality. For example, the first interface may include the above-mentioned... Figure 3 The additional interfaces shown in b, 303, may also include the above. Figure 9 The document addition interface shown in b is 903.
[0302] The first control can be understood as the control for card scanning. For example, the first control may include the above-mentioned... Figure 3 The card scan 304 shown in b may also include the above. Figure 9 The card scanning control 904 shown in b is shown.
[0303] S1402, Received the first operation that triggered the first control.
[0304] In this embodiment, the first operation can be understood as clicking the first control, or it may include other operations that trigger the first control. This embodiment does not limit the scope of the operation.
[0305] S1403. In response to the first operation, a second interface is displayed. The second interface includes a first area for displaying images captured by the camera of the electronic device.
[0306] In this embodiment, the second interface can be understood as a card / certificate scanning interface. For example, the second interface may include the above-mentioned... Figure 4 The card type selection interface 401 shown may also include the above. Figure 7 The card scanning interface 701 shown in figure a.
[0307] The first region may include the above. Figure 6 The viewfinder 603 shown may also include the above-mentioned... Figure 7 Viewfinder 703 is shown in a.
[0308] S1404. When the first area is the first image, a third interface is displayed. The first area of the third interface displays the first image. The third interface also includes a second area. The second control in the second area is selected, and the third control in the second area is unselected. The second control is automatically selected by the electronic device based on the image content of the first image.
[0309] In this embodiment, the third interface can be understood as the interface that displays the first image. For example, the third interface may include the above-mentioned... Figure 7 The card scanning interface 701 shown in figure a.
[0310] The second area can be understood as the area displaying the document type. For example, the second area may include the above-mentioned... Figure 7 The card type area shown in b is 708.
[0311] The second control can be understood as a display control for the document type corresponding to the first image. For example, the second control may include the above-mentioned... Figure 7 The highlighted option (b) shows the identified card type. For specific highlighting methods, please refer to the above. Figure 7 The relevant descriptions in the embodiments corresponding to b are not repeated here.
[0312] The third control can be understood as a display control for the document type corresponding to the second image. It's understandable that, since the current interface displays the interface corresponding to the first image, the third control is not highlighted.
[0313] S1405. When the first area is the second image, a fourth interface is displayed. The first area of the fourth interface displays the second image. The third control in the second area of the fourth interface is selected. The second control in the second area is not selected. The third control is automatically selected by the electronic device based on the image content of the second image.
[0314] In this embodiment, the fourth interface can be understood as the interface that displays the second image. For example, the fourth interface may include the one described above. Figure 7 The card scanning interface 701 shown in figure a. It is understandable that, since the current interface displays the interface corresponding to the second image, the third control is highlighted, while the second control is not highlighted.
[0315] The image recognition method provided in this application embodiment can intelligently identify the type of card and document during the card and document scanning process, and actively match the corresponding card and document type without requiring the user to actively select the document type. This improves the efficiency of card and document scanning, optimizes the card and document scanning process, and thus enhances the user experience.
[0316] Optional, in Figure 14 Based on the corresponding embodiments, the electronic device stores image feature information corresponding to each image type. After the second image is displayed in the first area of the fourth interface, it may further include: outputting one or more sets of image types of the second image and the confidence scores corresponding to the image types based on the image content of the second image; recognizing the text content of the second image; the third control in the second area of the fourth interface is in a selected state, which may include: when the matching degree between the image feature information of the first image type of the second image and the text content of the second image is greater than or equal to a first threshold, the third control in the second area of the fourth interface is in a selected state, wherein the first image type corresponds to a first confidence score, the first confidence score is the maximum value among one or more sets of confidence scores, and the third control matches the first image type.
[0317] In this embodiment of the application, the image type can be understood as described above. Figure 10 In the corresponding embodiment, the image feature information corresponding to the object category and image type in step S1006 can be understood as described above. Figure 10 The object features in step S1006 of the corresponding embodiment will not be described again.
[0318] The output of one or more sets of second images, along with the corresponding confidence scores, can be referenced above. Figure 10 In step S1007 of the corresponding embodiment, the description of one or more sets of object categories and the confidence levels corresponding to the object categories output by the image classification model will not be repeated.
[0319] The above can be used as a reference for recognizing the text content of the second image.Figure 10 In step S1009 of the corresponding embodiment, the relevant description of the text content being recognized by the text recognition model will not be repeated here.
[0320] Matching the image feature information of the first image type with the text content of the second image can be done as described above. Figure 10 The relevant description in step S1010 of the corresponding embodiment, the first threshold can be understood as the preset value A in step S1010, and will not be repeated here.
[0321] Understandably, among the card types identified by the image classification model, card types with higher confidence are prioritized for comparison. This results in a relatively high consistency, eliminating the need to continue comparing card types with lower confidence, thus reducing unnecessary computation.
[0322] Optional, in Figure 14 Based on the corresponding embodiment, the second image includes multiple target objects, among which a first target object is included. The method may further include: highlighting the first target object in a first area of the fourth interface, wherein the center of the first area is located in the area where the first target object is located.
[0323] In this embodiment of the application, multiple target objects can be understood as described above. Figure 11 In the corresponding embodiments, the multiple card objects, the first target object can be understood as described above. Figure 11 The object to be identified in the corresponding embodiment.
[0324] The area where the first target object is located can be understood as described above. Figure 12 In the embodiment where a corresponds to object 2, the region 1203 is shown.
[0325] When the center of the first region is located within the region containing the first target object, the first target object is selected as the object to be identified. This determination process can be referenced above. Figure 11 The relevant descriptions in steps S1102 and S1103 of the corresponding embodiments will not be repeated here.
[0326] Since users are more likely to place the card or document to be scanned in the center of the first area, the application can default to selecting the object in the center of the first area as the object to be scanned. This increases the probability of selecting the card or document the user wants to scan, thus improving the accuracy of recognition and enhancing the user experience.
[0327] Optional, in Figure 14Based on the corresponding embodiment, the second image includes multiple target objects, among which a first target object is included. The method may further include: highlighting the first target object in a first area of the fourth interface, wherein the distance between the center of the first area and the first target object is less than the distance between the center of the first area and other target objects, and the distance difference is greater than or equal to a second threshold.
[0328] In this embodiment, the distance between the center of the first region and the first target object can be understood as described above. Figure 11 The Manhattan distance in the corresponding embodiments can of course be understood as a straight-line distance or other distances, and this application does not limit it.
[0329] The method for determining target objects based on distance can be referred to above. Figure 11 The relevant descriptions in steps S1104 and S1105 of the corresponding embodiments will not be repeated here.
[0330] The second threshold can be understood as described above. Figure 11 The preset value B in step S1105 of the corresponding embodiment will not be described again.
[0331] When there is no target object in the center of the first area, the application can select the nearest object as the object to be recognized. This way, since users are more likely to place the card or document to be recognized close to the center of the first area, the application can more likely select the card or document the user wants to scan, increasing the probability of correct recognition and thus improving the user experience.
[0332] Optional, in Figure 14 Based on the corresponding embodiment, the second image includes multiple target objects, including a first target object and a second target object. The distance between the center of the first region and the first target object is a first distance, and the distance between the center of the first region and the second target object is a second distance. The method may further include: highlighting the first target object in the first region of the fourth interface, wherein the area of the region where the first target object is located is larger than the area where the second target object is located, the difference between the second distance and the first distance is less than a second threshold, and the second distance is less than the distance between the center of the first region and other target objects.
[0333] In this embodiment of the application, the first target object can be understood as described above. Figure 12 In the embodiment, b corresponds to object 4. The second target object can be understood as the above. Figure 12 The object 5 in the embodiment corresponds to b.
[0334] The method for determining target objects based on area can be referred to the above. Figure 11 The relevant descriptions in step S1106 of the corresponding embodiment will not be repeated here.
[0335] When there is no target object in the center of the first area, and two objects are close to the center of the first area and at approximately the same distance, the application can select the object with the largest or most large area as the object to be recognized. In this way, since the card to be recognized will likely occupy a larger area in the first area, the application can more likely select the card the user wants to scan, increasing the probability of correct recognition and thus improving the user experience.
[0336] Optional, in Figure 14 Based on the corresponding embodiment, the fourth interface further includes a fourth control, and the method may further include: receiving a second operation that triggers the fourth control; responding to the second operation, displaying a second interface; and displaying a fifth interface when the first area is a third image, wherein the first area of the fifth interface displays the third image, the third control in the second area of the fifth interface is selected, the second control in the second area is unselected, and the third control is automatically selected by the electronic device based on the image content of the third image, and the similarity between the third image and the second image is greater than or equal to a third threshold.
[0337] In this embodiment, the fourth control can be understood as a control that cancels the application of the currently selected document type. For example, the fourth control may include the above-mentioned... Figure 7 The return control 710 in the card type interface 707 of b.
[0338] The second operation can be understood as clicking the fourth control, or it may include other operations that trigger the fourth control. This application embodiment does not limit the scope of the operation.
[0339] The fifth interface can be understood as the interface that displays the third image. For example, the fifth interface may include the above-mentioned... Figure 7 The card scanning interface 701 shown in figure a.
[0340] The third threshold can be understood as the above. Figure 13 The similarity threshold in step S1307 of the corresponding embodiment will not be described again.
[0341] Understandably, since the third image is highly similar to the second image, indicating that the user has not changed their card, the application may still recognize the third image as a third control. This allows the application to correctly identify the card type even after the user accidentally clicks the back button.
[0342] Optional, in Figure 14Based on the corresponding embodiment, after the first area of the fifth interface displays the third image, it may further include: outputting one or more sets of image types of the third image and the confidence scores corresponding to the image types based on the image content of the third image; the third control in the second area of the fifth interface being selected may include: the third control in the second area of the fifth interface being selected when the second confidence score of the first image type of the third image is greater than the confidence scores corresponding to other image types of the third image, and the difference between the second confidence score and other confidence scores is greater than or equal to a fourth threshold.
[0343] In this embodiment of the application, the fourth threshold can be understood as described above. Figure 13 The confidence threshold in step S1309 of the corresponding embodiment will not be described again.
[0344] If the difference between the second confidence level and other confidence levels is greater than or equal to the fourth threshold, it indicates that the confidence difference is large, the two cards are significantly different, and the distinction is relatively clear. The application can accurately determine the type of the current document, so there is no need to adjust the confidence level.
[0345] The accuracy of the application's recognition can be determined by judging the confidence threshold. If the difference between the second confidence level and other confidence levels is greater than or equal to the fourth threshold, it means that the confidence difference is large and the application is accurate in recognition with a high probability.
[0346] Optional, in Figure 14 Based on the corresponding embodiments, the method may further include: receiving a third operation that triggers the fourth control; responding to the third operation, displaying a second interface; and displaying a sixth interface when the first area is the fourth image, wherein the first area of the sixth interface displays the fourth image, the fifth control in the second area of the sixth interface is selected, the third control in the second area is unselected, and the fifth control is automatically selected by the electronic device based on the image content of the fourth image, and the similarity between the fourth image and the second image is greater than or equal to a third threshold.
[0347] In this embodiment, the third operation can be understood as clicking the fourth control, or it may include other operations that trigger the fourth control. This embodiment does not limit the scope of the operation.
[0348] The sixth interface can be understood as the interface that displays the fourth image. For example, the sixth interface may include the one described above. Figure 7 The card scanning interface 701 shown in figure a.
[0349] The fifth control can be understood as a display control for the document type corresponding to the fourth image. It's understandable that the card type can be updated after re-identification; therefore, the image type of the fifth control object is different from the image type of the third control object. However, the image type corresponding to the fifth control and the image type of the second control object may be the same or different.
[0350] Since the fourth image is highly similar to the second image, it indicates that the user has not changed the card. Therefore, the application may be making a recognition error. Thus, after the user accidentally clicks the back button, the application can identify the card as another type, thereby determining the type the user wants to identify.
[0351] Optional, in Figure 14 Based on the corresponding embodiment, after the first area of the sixth interface displays the fourth image, it may further include: outputting one or more sets of image types and corresponding confidence levels of the fourth image based on the image content of the fourth image; increasing the value of the third confidence level when the second confidence level of the first image type of the fourth image is greater than the third confidence level corresponding to the second image type of the fourth image, the difference between the second confidence level and the third confidence level is less than the fourth threshold, and the third confidence level is greater than the confidence level of other image types; the fifth control in the second area of the sixth interface is in a selected state, which may include: when the third confidence level is greater than or equal to the second confidence level, the fifth control in the second area of the sixth interface is in a selected state, and the fifth control matches the second image type.
[0352] In this embodiment of the application, the confidence adjustment process can be referred to the above. Figure 13 The relevant descriptions in steps S1309-S1313 of the corresponding embodiment will not be repeated.
[0353] To increase the third confidence level, refer to the above. Figure 13 The relevant descriptions in step S1311 of the corresponding embodiment will not be repeated here.
[0354] By adjusting the confidence level of card 2, it's possible to make its confidence level higher than that of card 1. This allows the application to learn autonomously based on user feedback, thereby improving the user experience.
[0355] Optional, in Figure 14Based on the corresponding embodiments, the method may further include: receiving a fourth operation that triggers the fourth control; responding to the fourth operation, displaying a second interface; and displaying a seventh interface when the first area is the fifth image, wherein the first area of the seventh interface displays the fifth image, the sixth control in the second area of the seventh interface is selected, the third control in the second area is unselected, and the sixth control is automatically selected by the electronic device based on the image content of the fifth image, and the similarity between the fifth image and the second image is less than a third threshold.
[0356] In this embodiment, the fourth operation can be understood as clicking the fourth control, or it may include other operations that trigger the fourth control. This embodiment does not limit the scope of the operation.
[0357] The seventh interface can be understood as the interface that displays the fifth image. For example, the seventh interface may include the above-mentioned... Figure 7 The card scanning interface 701 shown in figure a.
[0358] The sixth control can be understood as a display control for the document type corresponding to the fifth image. It's understandable that, since the third and second images have low similarity, indicating a user has changed their card, the application can re-identify the new card type. Therefore, the image type of the sixth control object is different from the image type of the third control object. However, the image type corresponding to the sixth control and the image type of the second control object may be the same or different.
[0359] The application can determine whether a user has changed their card based on the similarity between the fifth image and the second image. This allows for more accurate image recognition based on the user's selection, thereby improving the user experience.
[0360] Optional, in Figure 14 Based on the corresponding embodiment, when the first area is the second image, before displaying the fourth interface, it may further include: displaying an eighth interface, the eighth interface including prompt information and a seventh control, the prompt information being used to indicate whether the image content of the second image is correctly recognized; receiving a fifth operation that triggers the seventh control; and displaying the second interface in response to the fifth operation.
[0361] In this embodiment, the eighth interface can be understood as an interface that displays pop-up prompts. For example, the eighth interface may include the above-mentioned... Figure 8 The pop-up notification interface 801 is shown. The notification message 803 can be understood as the notification message itself, and the seventh control can be understood as the cancel button 807.
[0362] The fifth operation can be understood as clicking the seventh control, or it can include other operations that trigger the seventh control. This application does not limit the specific operation.
[0363] When the application recognizes a card type, it may not automatically switch to the card type interface, but instead display a pop-up prompt. Users can then use this prompt to determine if the currently recognized card type is correct. This allows the application to more accurately identify cards based on user actions.
[0364] Optional, in Figure 14 Based on the corresponding embodiment, the second image includes multiple target objects, and the first object among the multiple target objects is highlighted. The method may further include: receiving a sixth operation that triggers the second object among the multiple target objects; in response to the sixth operation, highlighting the second object, and the eighth control in the second area of the fourth interface is in a selected state, wherein the eighth control is automatically selected by the electronic device based on the image content of the second object.
[0365] In this embodiment, the sixth operation can be understood as clicking the second object, or it may include other operations that trigger the second object. This embodiment does not limit the scope of the operation.
[0366] The eighth control can be understood as a display control for the document type corresponding to the second object. It's understandable that, since the second and first objects may be of the same or different types, the image type of the eighth control object may be different from or the same as the image type of the third control object.
[0367] When a user clicks on another object, the application can display the object to be recognized according to the user's selection. This can improve the accuracy of image recognition.
[0368] The foregoing primarily describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the aforementioned functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the method steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0369] This application embodiment can divide the apparatus for implementing the method into functional modules based on the above method examples. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0370] like Figure 15 The diagram shows a chip structure according to an embodiment of this application. The chip 1500 includes one or more processors 1501, communication lines 1502, communication interfaces 1503, and memory 1504.
[0371] In some implementations, memory 1504 stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof.
[0372] The methods described in the embodiments of this application can be applied to, or implemented by, processor 1501. Processor 1501 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the hardware of processor 1501 or by instructions in software form. Processor 1501 may be a general-purpose processor (e.g., a microprocessor or conventional processor), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. Processor 1501 can implement or execute the various processing-related methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0373] The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in mature storage media in the art, such as random access memory, read-only memory, programmable read-only memory, or electrically erasable programmable read-only memory (EEPROM). This storage medium is located in memory 1504, and processor 1501 reads information from memory 1504 and, in conjunction with its hardware, completes the steps of the above method.
[0374] The processor 1501, memory 1504 and communication interface 1503 can communicate with each other via communication line 1502.
[0375] In the above embodiments, the instructions stored in the memory for execution by the processor can be implemented in the form of a computer program product. This computer program product can be pre-written into the memory, or it can be downloaded and installed into the memory as software.
[0376] This application also provides a computer program product comprising one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from a website site, computer, server, or data center to another website site, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. For example, available media may include magnetic media (e.g., floppy disk, hard disk, or magnetic tape), optical media (e.g., digital versatile disc (DVD)), or semiconductor media (e.g., solid-state disk (SSD)).
[0377] This application also provides a computer-readable storage medium. The methods described in the above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. The computer-readable medium may include computer storage media and communication media, and may also include any medium capable of transferring a computer program from one place to another. The storage medium can be any target medium accessible by a computer.
[0378] As one possible design, computer-readable media may include compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; computer-readable media may also include disk storage or other disk storage devices. Furthermore, any connecting cable may also be appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. As used herein, disks and optical discs include optical discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers.
[0379] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
Claims
1. An image recognition method, characterized in that, Applied to electronic devices, the method includes: Display a first interface, the first interface including a first control; Receive the first operation that triggers the first control; In response to the first operation, a second interface is displayed, the second interface including a first area, the first area being used to display images captured by the camera of the electronic device; When the first area is the first image, a third interface is displayed. The first area of the third interface displays the first image. The third interface also includes a second area. The second control in the second area is selected, and the third control in the second area is unselected. The second control is automatically selected by the electronic device based on the image content of the first image. When the first area is the second image, a fourth interface is displayed. The first area of the fourth interface displays the second image. The third control in the second area of the fourth interface is selected, and the second control in the second area is unselected. The third control is automatically selected by the electronic device based on the image content of the second image. The electronic device stores image feature information corresponding to various image types. After the first area of the fourth interface displays the second image, it further includes: Based on the image content of the second image, output one or more sets of image types of the second image and the confidence scores corresponding to the image types; Identify the text content of the second image; The third control in the second area of the fourth interface is in a selected state, including: When the matching degree between the image feature information of the first image type of the second image and the text content of the second image is greater than or equal to the first threshold, the third control in the second area of the fourth interface is selected, wherein the first image type corresponds to the first confidence level, the first confidence level is the maximum value among one or more sets of confidence levels, and the third control matches the first image type.
2. The method according to claim 1, characterized in that, The second image includes multiple target objects, among which a first target object is included. The method further includes: The first target object is highlighted in the first area of the fourth interface, wherein the center of the first area is located in the area where the first target object is located.
3. The method according to claim 1, characterized in that, The second image includes multiple target objects, among which a first target object is included. The method further includes: The first target object is highlighted in the first area of the fourth interface, wherein the distance between the center of the first area and the first target object is less than the distance between the center of the first area and other target objects, and the difference in distance is greater than or equal to a second threshold.
4. The method according to claim 1, characterized in that, The second image includes multiple target objects, among which are a first target object and a second target object. The distance between the center of the first region and the first target object is a first distance, and the distance between the center of the first region and the second target object is a second distance. The method further includes: The first target object is highlighted in the first area of the fourth interface, wherein the area of the first target object is larger than the area of the second target object, the difference between the second distance and the first distance is less than a second threshold, and the second distance is less than the distance between the center of the first area and other target objects.
5. The method according to any one of claims 1-4, characterized in that, The fourth interface also includes a fourth control, and the method further includes: Received a second operation that triggers the fourth control; In response to the second operation, the second interface is displayed; When the first area is the third image, a fifth interface is displayed. The first area of the fifth interface displays the third image. The third control in the second area of the fifth interface is selected, and the second control in the second area is unselected. The third control is automatically selected by the electronic device based on the image content of the third image. The similarity between the third image and the second image is greater than or equal to a third threshold.
6. The method according to claim 5, characterized in that, After the third image is displayed in the first area of the fifth interface, the following is also included: Based on the image content of the third image, output one or more sets of image types of the third image and the confidence scores corresponding to the image types; The third control in the second area of the fifth interface is in a selected state, including: If the second confidence level of the first image type of the third image is greater than the confidence level of other image types of the third image, and the difference between the second confidence level and other confidence levels is greater than or equal to the fourth threshold, the third control in the second area of the fifth interface is selected.
7. The method according to any one of claims 1-4, 6, characterized in that, The method further includes: Received the third operation that triggered the fourth control; In response to the third operation, the second interface is displayed; When the first area is the fourth image, a sixth interface is displayed. The first area of the sixth interface displays the fourth image. The fifth control in the second area of the sixth interface is selected, and the third control in the second area is unselected. The fifth control is automatically selected by the electronic device based on the image content of the fourth image. The similarity between the fourth image and the second image is greater than or equal to a third threshold.
8. The method according to claim 7, characterized in that, After the first area of the sixth interface displays the fourth image, it also includes: Based on the image content of the fourth image, output one or more sets of image types of the fourth image and the confidence scores corresponding to the image types; If the second confidence level of the first image type of the fourth image is greater than the third confidence level corresponding to the second image type of the fourth image, the difference between the second confidence level and the third confidence level is less than the fourth threshold, and the third confidence level is greater than the confidence level of other image types, then the value of the third confidence level is increased. The fifth control in the second area of the sixth interface is in a selected state, including: When the third confidence level is greater than or equal to the second confidence level, the fifth control in the second area of the sixth interface is selected, and the fifth control matches the second image type.
9. The method according to any one of claims 1-4, 6, and 8, characterized in that, The method further includes: Received the fourth operation that triggered the fourth control; In response to the fourth operation, the second interface is displayed; When the first area is the fifth image, a seventh interface is displayed. The first area of the seventh interface displays the fifth image, the sixth control in the second area of the seventh interface is selected, the third control in the second area is unselected, and the sixth control is automatically selected by the electronic device based on the image content of the fifth image. The similarity between the fifth image and the second image is less than a third threshold.
10. The method according to any one of claims 1-4, 6, and 8, characterized in that, If the first area is the second image, before displaying the fourth interface, the method further includes: The eighth interface is displayed, which includes prompt information and a seventh control. The prompt information is used to indicate whether the image content of the second image has been correctly recognized. The fifth operation that triggers the seventh control is received; In response to the fifth operation, the second interface is displayed.
11. The method according to any one of claims 1-4, 6, and 8, characterized in that, The second image includes multiple target objects, and a first object among the multiple target objects is highlighted. The method further includes: The sixth operation, which triggers the second object among the plurality of target objects, is received. In response to the sixth operation, the second object is highlighted, and the eighth control in the second area of the fourth interface is selected. The eighth control is automatically selected by the electronic device based on the image content of the second object.
12. An electronic device, characterized in that, include: Memory and processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the electronic device to perform the method as described in any one of claims 1-11.
13. A chip system, characterized in that, It includes at least one processor and a communication interface, the communication interface and the at least one processor being interconnected via a line, the at least one processor being configured to run a computer program or instructions to perform the method as described in any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, cause a computer to perform the method as described in any one of claims 1-11.
15. A computer program product, characterized in that, Includes a computer program that, when run, causes an electronic device to perform the method as described in any one of claims 1-11.
Citation Information
Patent Citations
Card image recognition method based on deep learning
CN110909809A
Certificate image classification method and device, computer equipment and readable storage medium
CN111046879A