Image recognition method and related device
By automatically identifying the card type in the card certificate scanning application, the problem of users needing to manually select the card type is solved, improving scanning efficiency and user experience.
Patent Information
- Application Number
- CN202311820511.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-12-26
AI Technical Summary
In the prior art, card and certificate scanning applications require users to actively select the document type, resulting in a cumbersome selection process and reducing user experience.
By automatically identifying the card type and matching the corresponding document type during the card card scanning process, the user selection steps are reduced and scanning efficiency is improved.
Improve the efficiency and user experience of card certificate scanning, and optimize the card certificate scanning process, so that users do not need to manually select the document type.
Smart Images

Figure CN120259606A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to an image recognition method and related devices. Background Art
[0002] Some applications in electronic devices can provide a card and certificate scanning function. For example, the applications can include photo-taking applications, document editing applications, etc. The card and certificate scanning function can be used to restore various types of cards and certificates, so that the size and shape of the printed document can be close to those of the original.
[0003] However, in some applications, the user needs to actively select the type of certificate, and since there are many types of certificates to be selected, the search process is rather cumbersome, reducing the user experience. Summary of the Invention
[0004] The image recognition method and related devices provided in the embodiments of this application can, during the process of card and certificate scanning, intelligently identify the type of card and certificate and actively match it to the corresponding type of card and certificate, without the user having to actively select the type of certificate, improving the efficiency of card and certificate scanning, optimizing the process of card and certificate scanning, and thus enhancing the user experience.
[0005] In a first aspect, the image recognition method provided in the embodiments of this application includes:
[0006] Display a first interface, where the first interface includes a first control; receive a first operation that triggers the first control; in response to the first operation, display a second interface, where the second interface includes a first area for displaying an image captured by the camera of the electronic device; when the first area is a first image, display a third interface, where the first area of the third interface displays the first image, the third interface further includes a second area, the second control in the second area is in a selected state, the third control in the second area is in an unselected state, and the second control is automatically selected by the electronic device based on the image content of the first image; when the first area is a second image, display a fourth interface, where the first area of the fourth interface displays the second image, the third control in the second area of the fourth interface is in a selected state, the second control in the second area is in an unselected state, and the third control is automatically selected by the electronic device based on the image content of the second image. In this way, the efficiency of card and certificate scanning is improved, the process of card and certificate scanning is optimized, and thus the user experience is enhanced.
[0007] In a possible implementation, the electronic device stores the image feature information corresponding to each image type. After the second image is displayed in the first area of the fourth interface, the following steps are further included: based on the image content of the second image, output one or more groups of the image type of the second image and the confidence level corresponding to the image type; identify the text content of the second image; when the third control in the second area of the fourth interface is in a selected state, it includes: when the matching degree between the image feature information of the first image type of the second image and the text content of the second image is greater than or equal to the first threshold, the third control in the second area of the fourth interface is in a selected state, where the first image type corresponds to the first confidence level, and the first confidence level is the maximum value among one or more groups of confidence levels, and the third control matches the first image type. In this way, the calculated consistency is relatively high, so there is no need to continue comparing with the card types with lower confidence levels, thereby reducing unnecessary computing amounts.
[0008] In a possible implementation, the second image includes multiple target objects, and the multiple target objects include a first target object. The method further includes: highlighting the first target object in the first area of the fourth interface, where the center of the first area is located in the area where the first target object is located. In this way, the card that the user wants to scan can be selected with a high probability, increasing the probability of correct recognition, thereby enhancing the user experience.
[0009] In a possible implementation, the second image includes multiple target objects, and the multiple target objects include a first target object. The method further includes: highlighting the first target object in the first area of the fourth interface, where the distance between the center of the first area and the first target object is less than the distance between the center of the first area and other target objects, and the difference in distance is greater than or equal to the second threshold. In this way, since the user will place the card to be recognized closer to the center position of the first area with a high probability, the card that the user wants to scan can be selected with a high probability, increasing the probability of correct recognition.
[0010] In a possible implementation, the second image includes multiple target objects, and the multiple target objects include a first target object and a second target object. The distance between the center of the first area and the first target object is the first distance, and the distance between the center of the first area and the second target object is the second distance. The method further includes: highlighting the first target object in the first area of the fourth interface, where the area of the region where the first target object is located is larger than the area of the region where the second target object is located, the difference between the second distance and the first distance is less than the second threshold, and the second distance is less than the distance between the center of the first area and other target objects. In this way, since the card to be recognized will occupy a relatively large area in the first area with a high probability, the card that the user wants to scan can be selected with a high probability, increasing the probability of correct recognition.
[0011] In a possible implementation, the fourth interface further includes a fourth control, and the method further includes: receiving a second operation that triggers the fourth control; in response to the second operation, displaying a second interface; when the first area is a third image, displaying a fifth interface, where the first area of the fifth interface displays the third image, the third control in the second area of the fifth interface is in a selected state, the second control in the second area is in an unselected state, and the third control is automatically selected by the electronic device based on the image content of the third image, and the similarity between the third image and the second image is greater than or equal to a third threshold. In this way, after the user accidentally clicks the back button, the application can still correctly identify the type of the card or certificate.
[0012] In a possible implementation, after the first area of the fifth interface displays the third image, it further includes: outputting one or more groups of the image type of the third image and the confidence level corresponding to the image type based on the image content of the third image; the third control in the second area of the fifth interface being in a selected state includes: when the second confidence level of the first image type of the third image is greater than the confidence levels corresponding to other image types of the third image, and the difference between the second confidence level and the other confidence levels is greater than or equal to a fourth threshold, the third control in the second area of the fifth interface is in a selected state. In this way, by judging the confidence level threshold, the accuracy of the application's recognition can be determined. If the difference between the second confidence level and the other confidence levels is greater than or equal to the fourth threshold, it indicates that the confidence level difference is relatively large, and the application is likely to be accurate in recognition.
[0013] In a possible implementation, the method further includes: receiving a third operation that triggers the fourth control; in response to the third operation, displaying a second interface; when the first area is a fourth image, displaying a sixth interface, where the first area of the sixth interface displays the fourth image, the fifth control in the second area of the sixth interface is in a selected state, the third control in the second area is in an unselected state, and the fifth control is automatically selected by the electronic device based on the image content of the fourth image, and the similarity between the fourth image and the second image is greater than or equal to a third threshold. In this way, after the user accidentally clicks the back button, the application can identify the card or certificate as another type, so as to determine the type that the user wants to identify.
[0014] In a possible implementation, after the fourth image is displayed in the first area of the sixth interface, it further includes: outputting one or more groups of the image type of the fourth image and the confidence corresponding to the image type based on the image content of the fourth image; when the second confidence of the first image type of the fourth image is greater than the third confidence corresponding to the second image type of the fourth image, the difference between the second confidence and the third confidence is less than the fourth threshold, and the third confidence is greater than the confidence of other image types, increasing the value of the third confidence; when the fifth control in the second area of the sixth interface is in a selected state, it includes: when the third confidence is greater than or equal to the second confidence, the fifth control in the second area of the sixth interface is in a selected state, and the fifth control matches the second image type. In this way, the application can implement the ability of autonomous learning based on user feedback, thereby improving the user experience.
[0015] In a possible implementation, the method further includes: receiving a fourth operation that triggers the fourth control; in response to the fourth operation, displaying the second interface; when the first area is the fifth image, displaying the seventh interface, where the first area of the seventh interface displays the fifth image, the sixth control in the second area of the seventh interface is in a selected state, the third control in the second area is not in a selected state, and the sixth control is automatically selected by the electronic device based on the image content of the fifth image, and the similarity between the fifth image and the second image is less than the third threshold. In this way, the application can determine whether the user has replaced the card or certificate based on the similarity between the fifth image and the second image, and can perform image recognition more accurately based on the user's selection, thereby improving the user experience.
[0016] In a possible implementation, before displaying the fourth interface when the first area is the second image, it further includes: displaying the eighth interface, which includes a prompt message and a seventh control, and the prompt message is used to prompt whether the image content of the second image is recognized correctly; receiving a fifth operation that triggers the seventh control; in response to the fifth operation, displaying the second interface. In this way, the user can determine whether the currently recognized card or certificate type is correct based on the prompt of the application, and the application can perform card or certificate recognition more accurately based on the user's operation.
[0017] In a possible implementation, the second image includes multiple target objects, and the first object among the multiple target objects is highlighted. The method further includes: receiving a sixth operation that triggers the second object among the multiple target objects; in response to the sixth operation, highlighting the second object, and the eighth control in the second area of the fourth interface is in a selected state, and the eighth control is automatically selected by the electronic device based on the image content of the second object. In this way, when the user clicks on other objects, the application can display the object to be recognized according to the user's selection, which can improve the accuracy of image recognition.
[0018] Second aspect, an embodiment of the present application provides an image recognition device, which may be an electronic device, or a chip or a chip system within the electronic device. The device may include a processing unit and a display unit. The processing unit is used to implement any processing-related method executed by the electronic device in the first aspect or any possible implementation manner of the first aspect. The display unit is used to implement any display-related method executed by the electronic device in the first aspect or any possible implementation manner of the first aspect. When the device is an electronic device, the processing unit may be a processor. The device may further include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to enable the electronic device to implement the method described in the first aspect or any possible implementation manner of the first aspect. When the device is a chip or a chip system within the electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to enable the electronic device to implement the method described in the first aspect or any possible implementation manner of the first aspect. The storage unit may be a storage unit within the chip (such as a register, a cache, etc.), or a storage unit outside the chip within the electronic device (such as a read-only memory, a random access memory, etc.).
[0019] Exemplarily, the display unit is used to display a first interface, and is further used to display a second interface, and is further used to display a third interface, and is further used to display a fourth interface. The processing unit is used to receive a first operation that triggers a first control.
[0020] In a possible implementation manner, the processing unit is used to output one or more groups of image types of the second image and the confidence levels corresponding to the image types based on the image content of the second image; and is further used to recognize the text content of the second image; specifically, when the matching degree between the image feature information of the first image type of the second image and the text content of the second image is greater than or equal to a first threshold, the third control in the second area of the fourth interface is in a selected state.
[0021] In a possible implementation manner, the display unit is used to highlight a first target object in the first area of the fourth interface, where the center of the first area is located in the area where the first target object is located.
[0022] In a possible implementation manner, the display unit is used to highlight a first target object in the first area of the fourth interface, where the distance between the center of the first area and the first target object is less than the distance between the center of the first area and other target objects, and the difference in distance is greater than or equal to a second threshold.
[0023] In a possible implementation, a display unit is configured to highlight a first target object in a first area of a fourth interface, where the area of the region where the first target object is located is larger than the region where a second target object is located, the difference between a second distance and a first distance is less than a second threshold, and the second distance is less than the distance between the center of the first area and other target objects.
[0024] In a possible implementation, a processing unit is configured to receive a second operation that triggers a fourth control. The display unit is configured to display a second interface and is further configured to display a fifth interface.
[0025] In a possible implementation, a processing unit is configured to output one or more groups of image types of a third image and confidences corresponding to the image types based on the image content of the third image; and is further configured to, when a second confidence of a first image type of the third image is greater than the confidences corresponding to other image types of the third image and the difference between the second confidence and the other confidences is greater than or equal to a fourth threshold, set a third control in a second area of a fifth interface to a selected state.
[0026] In a possible implementation, a processing unit is configured to receive a third operation that triggers a fourth control. The display unit is configured to display a second interface and is further configured to display a sixth interface.
[0027] In a possible implementation, a processing unit is configured to output one or more groups of image types of a fourth image and confidences corresponding to the image types based on the image content of the fourth image; and is further configured to increase the value of a third confidence; specifically, when the third confidence is greater than or equal to a second confidence, set a fifth control in a second area of a sixth interface to a selected state.
[0028] In a possible implementation, a processing unit is configured to receive a fourth operation that triggers a fourth control. The display unit is configured to display a second interface and is further configured to display a seventh interface.
[0029] In a possible implementation, a processing unit is configured to receive a fifth operation that triggers a seventh control. The display unit is configured to display an eighth interface and is further configured to display a second interface.
[0030] In a possible implementation, a processing unit is configured to receive a sixth operation that triggers a second object among multiple target objects. The display unit is configured to highlight the second object.
[0031] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory. The memory is configured to store code instructions, and the processor is configured to run the code instructions to execute the method described in the first aspect or any possible implementation of the first aspect.
[0032] Fourthly, this application provides a chip or a chip system, which includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a circuit. The at least one processor is configured to run a computer program or instruction to execute the method described in the first aspect or any possible implementation manner of the first aspect. Among them, the communication interface in the chip can be an input / output interface, a pin, a circuit, etc.
[0033] In a possible implementation, the chip or chip system described above in this application further includes at least one memory, and instructions are stored in the at least one memory. The memory can be a storage unit inside the chip, such as a register, a cache, etc., or it can be a storage unit of the chip (such as a read-only memory, a random access memory, etc.).
[0034] Fifthly, an embodiment of this application provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction runs on a computer, the computer is enabled to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0035] Sixthly, an embodiment of this application provides a computer program product including a computer program. When the computer program runs on a computer, the computer is enabled to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0036] It should be understood that the second to sixth aspects of this application correspond to the technical solutions of the first aspect of this application. The beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar and will not be elaborated herein. Description of the Drawings
[0037] Figure 1 It is a schematic structural diagram of an electronic device provided by an embodiment of this application;
[0038] Figure 2 It is a schematic software structure diagram of an electronic device provided by an embodiment of this application;
[0039] Figure 3 It is a schematic interface diagram of a camera application provided by an embodiment of this application;
[0040] Figure 4 It is a schematic interface diagram of a card type selection provided by an embodiment of this application;
[0041] Figure 5 It is a schematic interface diagram of more certificates provided by an embodiment of this application;
[0042] Figure 6A schematic diagram of the interface for card and certificate scanning provided by an embodiment of the present application;
[0043] Figure 7 A schematic diagram of the interface for the card and certificate recognition process provided by an embodiment of the present application;
[0044] Figure 8 A schematic diagram of the interface for the card and certificate recognition pop-up window provided by an embodiment of the present application;
[0045] Figure 9 A schematic diagram of the interface for document application provided by an embodiment of the present application;
[0046] Figure 10 A flowchart of an image recognition method provided by an embodiment of the present application;
[0047] Figure 11 A flowchart of a method for selecting card and certificate objects provided by an embodiment of the present application;
[0048] Figure 12 A schematic diagram of multiple card and certificate object regions in an image provided by an embodiment of the present application;
[0049] Figure 13 A flowchart of confidence adjustment provided by an embodiment of the present application;
[0050] Figure 14 A schematic diagram of an image recognition method provided by an embodiment of the present application;
[0051] Figure 15 A schematic diagram of the structure of a chip provided by an embodiment of the present application. Detailed implementation manners
[0052] To facilitate a clear description of the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:
[0053] 1. Manhattan distance: The Manhattan distance can be understood as the distance between two points in the north-south direction plus the distance between two points in the east-west direction. For example, in a plane, the Manhattan distance can be the distance between two points in the x-axis direction plus the distance between two points in the y-axis direction.
[0054] Exemplarily, assuming that the coordinates of the first point are (x1, y1) and the coordinates of the second point are (x2, y2), then the distance between the two points in the x-axis direction is |x1 - x2|, and the distance between the two points in the y-axis direction is |y1 - y2|. Therefore, the Manhattan distance d can satisfy the following formula: d = |x1 - x2| + |y1 - y2|.
[0055] 2. Terms
[0056] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. For example, the first chip and the second chip are only used to distinguish different chips, and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and the terms such as "first" and "second" do not necessarily mean different.
[0057] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0058] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (s) or plural item (s). For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple.
[0059] 3. Electronic device
[0060] The electronic device in the embodiments of the present application can also be any form of terminal device. For example, the electronic device can include: mobile phone, tablet computer, palm computer, laptop computer, mobile internet device (MID), wearable device, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, cellular phone, cordless phone, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA), handheld device with wireless communication function, computing device or other processing device connected to a wireless modem, in-vehicle device, wearable device, electronic device in a 5G network or electronic device in a future evolved public land mobile network (PLMN), etc. The embodiments of the present application are not limited thereto.
[0061] By way of example and not limitation, in the embodiments of the present application, the electronic device can also be a wearable device. A wearable device can also be referred to as a wearable intelligent device, which is a general term for devices developed by applying wearable technology to the intelligent design of daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is either directly worn on the body or integrated into the user's clothing or accessories. A wearable device is not only a hardware device, but also realizes powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable intelligent devices include those with complete functions and large sizes that can realize complete or partial functions without relying on a smart phone, such as smart watches or smart glasses, etc., and those that only focus on a certain type of application function and need to cooperate with other devices such as smart phones, such as various smart bracelets and smart jewelry for physical sign monitoring.
[0062] In addition, in the embodiments of the present application, the electronic device may also be an electronic device in an Internet of Things (IoT) system. The IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, so as to achieve an intelligent network of human-machine interconnection and object-object interconnection.
[0063] The electronic device in the embodiments of the present application may also be referred to as: user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile terminal, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user device, etc.
[0064] In the embodiments of the present application, the electronic device or each network device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and a memory (also called main memory). The operating system can be any one or more computer operating systems that implement service processing through processes. For example, Linux operating system, Unix operating system, Android operating system, iOS operating system, or Windows operating system, etc. The application layer includes applications such as a browser, an address book, a word processing software, and an instant messaging software.
[0065] Exemplarily, Figure 1 The schematic structural diagram of the electronic device is shown.
[0066] The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0067] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented by hardware, software, or a combination of software and hardware.
[0068] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. The controller may generate operation control signals according to the instruction operation code and the timing signal to complete the control of fetching instructions and executing instructions.
[0069] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the above-mentioned memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0070] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are only illustrative and do not constitute a structural limitation on the electronic device. In other embodiments of the present application, the electronic device may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0071] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc. The data storage area can store data created during the use of the electronic device. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 121, and / or the instructions stored in the memory provided in the processor.
[0072] The camera 193 is used to capture still images or videos. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1. For example, in the embodiments of the present application, camera applications or document editing applications can implement scanning functions based on the images captured by the camera 193.
[0073] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device may include one or N display screens 194, where N is a positive integer greater than 1. The electronic device realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. For example, in the embodiments of the present application, the display screen 194 can be used to display the images captured by the camera 193.
[0074] The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information. The electronic device can implement the shooting function through the ISP, camera 193, video codec, GPU, display screen 194, application processor, etc.
[0075] Figure 2 It is a software structure block diagram of the electronic device according to an embodiment of the present application.
[0076] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom are the application layer, application framework layer, Android runtime, algorithm engine layer and system library, hardware adaptation layer (HAL), and kernel layer.
[0077] The application layer can also be called the app layer, and the app layer can include a series of application packages. Such as Figure 2 shown, the application packages can include applications such as camera, gallery, notes, documents, etc. The applications can include system applications and third-party applications.
[0078] The application framework layer can also be called the Framework layer. The Framework layer can provide application programming interfaces (APIs) and programming frameworks for the applications in the app layer. The Framework layer can include some predefined functions.
[0079] Such as Figure 2 shown, the Framework layer can include the camera Camera framework, resource manager, activity manager ActivityManager, notification manager, etc.
[0080] The Camera framework can be used to manage the camera, obtain camera devices and information, and implement related functions such as preview, taking pictures, and recording videos. In the embodiments of the present application, the Camera framework can provide system support for the scanning function, manage the life cycle of the camera and data acquisition, etc.
[0081] The resource manager can provide various resources for the applications, such as localized strings, icons, pictures, layout files, video files, etc.
[0082] The ActivityManager can be used to manage the lifecycle of applications, determine the running status of applications, such as the startup, suspension, stop, and destruction of applications. When the application enters the background or is closed, the ActivityManager can release the corresponding memory resources to ensure the stability of the electronic device system.
[0083] The notification manager enables the application to display notification information in the status bar, which can be used to convey notification-type messages and can also disappear automatically after a short stay without user interaction.
[0084] The Android runtime includes core libraries and a virtual machine. The Android runtime is responsible for the control and management of the Android system.
[0085] The core libraries consist of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.
[0086] The application layer and the Framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the Framework layer as binary files. The virtual machine is used to perform functions such as the management of object lifecycles, stack management, thread management, security and exception management, and garbage collection. For example, in the embodiments of the present application, the virtual machine can be used to perform functions such as multi-object detection, text detection and recognition, image classification, and calculation of image similarity.
[0087] The algorithm engine layer can include multi-object detection algorithms, text detection and recognition algorithms, information comparison algorithms, image classification algorithms, image similarity algorithms, etc. It can be understood that each of these algorithms includes, but is not limited to, machine learning models, deep learning models, and traditional algorithms corresponding to each algorithm. In the embodiments of the present application, the algorithm engine layer can also be understood as a card and certificate recognition model.
[0088] In the embodiments of the present application, the multi-object detection algorithm can detect whether there are card and certificate objects, and can also obtain one or more card and certificate objects in the image, and identify the position information of the card and certificate objects, the category of the card and certificate objects, and the confidence level of the category of the card and certificate objects, etc.
[0089] The text detection and recognition algorithm can be used to detect the text information in the card and certificate objects.
[0090] The information comparison algorithm can compare the text information of the card and certificate objects with prior knowledge.
[0091] Image classification algorithms can be used to classify card and certificate objects. Card and certificate objects can be classified into types such as documents, cards and certificates, presentation (PowerPoint, PPT), etc. Among them, documents can include paper manuscripts, books, business cards, etc.; cards and certificates can include ID cards, bank cards, household registers, etc.; PPT can include meeting projection PPT, computer monitor PPT, etc.
[0092] It should be noted that in the embodiments of this application, with the user's permission, information of card and certificate objects is obtained. After the card and certificate scanning function is executed, relevant information of card and certificate objects will be cleared or saved based on the user's instructions. That is to say, the user information (including but not limited to information of card and certificate objects) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0093] Image similarity algorithms can be used to calculate the similarity of images. For example, image similarity algorithms can include feature point matching algorithms, etc.
[0094] It can be understood that in the embodiments of this application, the image recognition method used by the provided card and certificate scanning function of the application can be provided by the electronic device system. For the convenience of description, subsequent descriptions will all be based on the application as the execution entity.
[0095] In one possible implementation, the application can call the interfaces provided by the algorithm engine layer of the electronic device system to use multi-object detection algorithms, text detection and recognition algorithms, information comparison algorithms, image classification algorithms, image similarity algorithms, etc., so as to implement the image recognition method of the embodiments of this application.
[0096] In another possible implementation, the application can integrate a software development kit (SDK), and the SDK can include relevant functions provided by the algorithm engine layer of the electronic device system, so as to implement the image recognition method of the embodiments of this application.
[0097] The system library can also be called the Native layer, and the Native layer can include multiple functional modules. For example: Camera Service module of the camera service, OpenCV library, etc.
[0098] In the embodiments of this application, the Camera Service module can be used to process requests of the Camera framework and implement related functions such as camera shooting.
[0099] The hardware abstraction layer is an abstracted layer structure between the kernel layer and the Android runtime. The hardware abstraction layer can be a encapsulation of hardware drivers, providing a unified interface for upper-layer applications to call. The hardware abstraction layer can include a camera module, a sensor module, an audio module, etc. For example, in the embodiments of the present application, the camera module can be used to control the camera driver to collect images, so as to implement functions such as scanning cards and certificates.
[0100] The kernel layer is the layer between hardware and software. The kernel layer can include a camera driver, a display driver, an audio driver, a power management module, etc. For example, in the embodiments of the present application, the camera driver can instruct the camera device to collect images.
[0101] It should be noted that the embodiments of the present application only take the Android system as an example for illustration. In other operating systems (such as the Windows system, the IOS system, etc.), as long as the functions implemented by each functional module are similar to those of the embodiments of the present application, the solution of the present application can also be implemented.
[0102] Some applications in the electronic device can provide a card and certificate scanning function. For example, the application can include a photo-taking application, a document editing application, etc. The card and certificate scanning function can be used to trim and correct various types of cards and certificates, and perform a 1:1 restoration, so that the size and shape of the printed document can be close to those of the original.
[0103] In some implementations, card and certificate scanning can also be referred to as document scanning, card and certificate recognition, document recognition, or card and certificate collection, etc. For the sake of convenience of description, the embodiments of the present application are all described by taking card and certificate scanning as an example.
[0104] Exemplarily, Figures 3 to 6 shows a schematic diagram of a related interface where a photo-taking application provides a card and certificate scanning function. It can be understood that the related interface of the card and certificate scanning function can include Figures 3 to 6 any one or more of the interfaces. The display of the specific interface can be set by each application, and the embodiments of the present application do not make any limitations.
[0105] As Figure 3 shown in a of, the interface 301 is the photo-taking interface of the camera. The interface 301 can include various photo-taking modes such as the large aperture mode, the portrait mode, the photo-taking mode, and the video recording mode. The interface 301 can also include more controls 302, etc.
[0106] When the user triggers the more controls 302, in response to the user's trigger operation, as Figure 3 shown in b of, the electronic device can display more interfaces 303 of the camera. The more interfaces 303 can include functions such as HDR, slow motion, micro movie, time-lapse photography, live photos, and card and certificate scanning 304.
[0107] Optionally, interface 301 and interface 303 are exemplary interfaces. It can be understood that different applications may have different interface displays, and interface 301 and interface 303 may respectively include more or less content. The specific content included in interface 301 and interface 303 is not limited in the embodiments of the present application.
[0108] When the user triggers the card and certificate scanning 304, in response to the user's trigger operation, such as Figure 4 shown, the electronic device can display a card and certificate type selection interface 401. The card and certificate type selection interface 401 may include a file type selection area 402, a certificate type area 403, a more certificate control 404, etc.
[0109] Optionally, the card and certificate type selection interface 401 is an exemplary interface. It can be understood that different applications may have different interface displays, and the card and certificate type selection interface 401 may include more or less content. The specific content included in the card and certificate type selection interface 401 is not limited in the embodiments of the present application.
[0110] The file type selection area 402 may include multiple file type options such as shooting, stream recognition, scanning certificates, test paper operations, photo translation, etc. The user can select the file type they want to scan.
[0111] Exemplarily, taking the user's selection of the scan certificate option in the file type selection area 402 as an example, the certificate type area 403 can display the certificate types corresponding to the scan certificate option. The certificate type area 403 may include certificate types such as household register, driver's license, vehicle license, bank card, ID card, etc., and the certificate type area 403 may also include a more certificate control 404.
[0112] The user can select the certificate type they want to scan in the certificate type area 403.
[0113] In a possible scenario, if the certificate type area 403 does not have the certificate type the user wants, the user can trigger the more certificate control 404. In response to the user's trigger operation, such as Figure 5 shown, the electronic device can display a more certificate interface 501. The more certificate interface 501 may include one or more certificate options 502, the certificate sizes 503 corresponding to the certificate options, a search control 504, a return control 505, etc.
[0114] Among them, the document option 502 may include a social security card, a citizen card, an officer's certificate, a bankbook, a social security card, a permit, a student ID card, a graduation certificate, etc. It can be understood that different document options 502 may have corresponding different document sizes 503. In this way, the application can crop the scanned document based on the document size corresponding to the document selected by the user, so that the size of the scanned document is restored as close as possible to 1:1 of the actual document. The application can also stretch the border lines of the scanned document so that the size and shape of the document can be close to those of the original.
[0115] The search control 504 can be used to quickly search for the document type that the user wants to find.
[0116] The return control 505 can be used to exit the more documents interface 501. When the user triggers the return control 505, in response to the user's trigger operation, the application can return to the card and certificate type selection interface 401. The user can re-select the card and certificate type.
[0117] Optionally, the more documents interface 501 is an exemplary interface. It can be understood that different applications may have different interface displays, and the more documents interface 501 may include more or less content. The specific content included in the more documents interface 501 is not limited in the embodiments of the present application.
[0118] In another possible scenario, when the user selects a certain document type in the document type area 403, in response to the user's trigger operation, as Figure 6 shown in a of, the electronic device can display the card and certificate scanning interface 601. The card and certificate scanning interface 601 may include a prompt 602, a viewfinder 603, a zoom area 604, an image preview 605, a shooting control 606, etc.
[0119] The prompt 602 may include "Please place the document completely within the viewfinder". The specific content included in the prompt 602 is not limited in the embodiments of the present application.
[0120] The viewfinder 603 is the area for scanning the card and certificate. Optionally, in some implementations, the card and certificate scanning interface 601 may also not include the viewfinder 603, and the application can use the border of the electronic device as the viewfinder. The specific display style of the viewfinder 603 is not limited in the embodiments of the present application.
[0121] The zoom area 604 can be used to adjust the shooting focal rate. The image preview 605 can be used to preview pictures or videos in the album, etc.
[0122] The shooting control 606 can be used to shoot the card and certificate. When the user triggers the shooting control 606, in response to the user's trigger operation, as Figure 6As shown in Figure b, the electronic device can display a card preview interface 607. The card preview interface 607 may include a card preview area 608, a confirmation control 609, a return control 610, etc.
[0123] When the user triggers the confirmation control 609, in response to the user's trigger operation, the application can trim and correct the content in the card preview area 608, and can also execute the printing process.
[0124] When the user triggers the return control 610, in response to the user's trigger operation, the application can return to the card scanning interface 601. The user can re-trigger the card scanning operation or other operations.
[0125] Optionally, the card scanning interface 601 and the card preview interface 607 are exemplary interfaces. It can be understood that different applications may have different interface displays, and the card scanning interface 601 and the card preview interface 607 may respectively include more or less content. The specific content included in the card scanning interface 601 and the card preview interface 607 is not limited in the embodiments of the present application.
[0126] From the above process, it can be seen that during the card scanning process by the application, the user needs to actively select the document type, and since there are many document types to be selected, the search process is rather cumbersome, reducing the user experience.
[0127] In view of this, in the image recognition method provided by the embodiments of the present application, during the card scanning process, the application can intelligently recognize the type of the card, and actively match the corresponding card type, without the user having to actively select the document type, improving the efficiency of card scanning, optimizing the card scanning process, and thus enhancing the user experience.
[0128] Figure 7 Shows a schematic diagram of the interface for card recognition provided by the embodiments of the present application.
[0129] As Figure 7 As shown in Figure a, the electronic device can display a card scanning interface 701. The card scanning interface 701 may include a prompt 702, a viewfinder 703, a zoom area 704, an image preview 705, a shooting control 706, etc.
[0130] The prompt 702 may include "Please place the document completely within the viewfinder", and the specific content included in the prompt 702 is not limited in the embodiments of the present application.
[0131] The viewfinder 703 is the area for scanning the card. Optionally, in some implementations, the card scanning interface 701 may not include the viewfinder 703, and the application can use the border of the electronic device as the viewfinder. The specific display style of the viewfinder 703 is not limited in the embodiments of the present application.
[0132] The zoom area 704 can be used to adjust the shooting focal rate. The image preview 705 can be used to reserve pictures or videos in the album, etc. The shooting control 706 can be used to shoot cards and certificates.
[0133] When the viewfinder 703 detects a card or certificate, the application can automatically identify the type of the card or certificate and automatically switch to the corresponding card or certificate type interface 707, as Figure 7 shown in b of. The card or certificate type interface 707 may include a card or certificate type area 708, a shooting control 709, a return control 710, etc.
[0134] In the card or certificate type area 708, the application can prominently display the identified card or certificate type. Exemplarily, as Figure 7 shown in b of, taking the card or certificate type as an ID card as an example, the prominent display may include adding a border to the icon corresponding to the ID card, bolding the text of "ID card", modifying the font size of the text of "ID card", and / or modifying the color of the text of "ID card", etc. The specific way of prominent display is not limited in the embodiments of the present application.
[0135] Optionally, after the application identifies the card or certificate type, it can also give a prompt message to the user in the interface, and the prompt message can be used to identify the currently identified card or certificate type. Exemplarily, the application can pop up a toast prompt message in the card or certificate type interface 707, and the prompt message can include "The identified card or certificate type is an ID card", and the prompt message can also be other content. The specific content of the prompt message is not limited in the embodiments of the present application.
[0136] It can be understood that in the card or certificate type interface 707, only the identified card or certificate type can be prominently displayed, or only the prompt message can be displayed, such as a toast prompt, or both the identified card or certificate type can be prominently displayed and the prompt message can be displayed. The embodiments of the present application do not make a limitation.
[0137] In addition, the above-mentioned prompt message can also be displayed on the card or certificate scanning interface 701, and the embodiments of the present application do not make a limitation.
[0138] When the card or certificate type identified by the application is incorrect, the user can trigger the return control 710. In response to the user's trigger operation, the application can return to the card or certificate scanning interface 701. The user can re-trigger the card or certificate scanning operation or other operations.
[0139] When the card or certificate type identified by the application is correct, the user can trigger the shooting control 709. In response to the user's trigger operation, as Figure 7 shown in c of, the electronic device can display a card or certificate preview interface 711. The card or certificate preview interface 711 may include a card or certificate preview area 712, a confirmation control 713, a return control 714, etc.
[0140] When the user triggers the confirmation control 713, in response to the user's trigger operation, the application can perform edge trimming and correction on the content in the card preview area 712, and perform a 1:1 restoration, so that the size and shape of the printed document can be close to those of the original. In addition, the application can also perform enhancement processing such as shadow removal and highlight removal on the card, which is not limited in the embodiments of the present application.
[0141] When the user triggers the return control 714, in response to the user's trigger operation, the application can return to the card scanning interface 701. The user can re-trigger the card scanning operation or other operations.
[0142] Optionally, Figure 7 The shown card scanning interface 701, card type interface 707, and card preview interface 711 are exemplary interfaces. It can be understood that different applications can have different interface displays, Figure 7 and the shown interfaces can respectively include more or less content. Specifically Figure 7 the content included in the shown interfaces is not limited in the embodiments of the present application.
[0143] Optionally, in the above-mentioned card scanning interface 701, when the application recognizes the card type, the application may not automatically switch to the card type interface 707, but instead display a pop-up prompt. As Figure 8 shown, the interface 801 is a pop-up prompt interface. The interface 801 can display a pop-up window 802, and the pop-up window 802 can include: a prompt message 803 for prompting the user whether to switch the document mode, a "Don't remind again" button 804, a prompt message 805 for prompting not to remind again, an "OK" button 806, and a "Cancel" button 807, etc.
[0144] Among them, the prompt message 803 can include "The recognized document type is A. Do you want to switch to A mode?", and the prompt message 805 can include "Don't remind again". The specific content of the prompt message 803 and the prompt message 805 is not limited in the embodiments of the present application.
[0145] It can be understood that if the user triggers the operation of the "Don't remind again" button 804, in response to this operation of the user, when the application recognizes the card type next time, the application can automatically switch to the card type interface 707 without displaying the pop-up window 802. If the user does not trigger the operation of the "Don't remind again" button 804, when the application recognizes the card type next time, the application does not automatically switch to the card type interface 707, but instead displays a pop-up prompt.
[0146] When the user triggers the operation of the OK button 806, in response to this operation of the user, the application can cancel the display of the pop-up window 802 and jump to the card type interface 707 corresponding to Certificate A. When the user triggers the operation of the Cancel button 807, in response to this operation of the user, the application can cancel the display of the pop-up window 802 and return to the card scanning interface 701.
[0147] Optionally, in the above-mentioned card scanning interface 701, when the application recognizes the card type, the application can jump to the card preview interface 711 instead of displaying the card type interface 707. It can be understood that when the accuracy of the application in recognizing the card type is relatively high, this solution can be adopted. After recognizing the card type, the application can perform processes such as cutting, rectifying, image enhancement, and content recognition on the card. This can reduce the interaction with the user and speed up the card recognition process.
[0148] It can be understood that in addition to the Figure 3 camera application mentioned above that can provide the card scanning function, document editing applications can also provide the card scanning function.
[0149] Exemplarily, Figure 9 shows a schematic diagram of the interface of the card scanning function of a document editing application.
[0150] As Figure 9 shown in a of, the interface 901 is a document editing interface, and the interface 901 may include a document title, document editing time, document content, style controls, list controls, an add control 902, a recording control, etc.
[0151] When the user triggers the add control 902, in response to the triggering operation of the user, as Figure 9 shown in b of, the electronic device can display a document addition interface 903. The document addition interface 903 may include table controls, hyperlink controls, imported document controls, document scanning controls, table extraction controls, a card scanning control 904, etc.
[0152] It can be understood that the icons and control names corresponding to each control can be different, and the icons and control names corresponding to each control can be set by the application, which is not limited in the embodiments of the present application.
[0153] Optionally, the document editing interface 901 and the document addition interface 903 are exemplary interfaces. It can be understood that different applications may have different interface displays, and the document editing interface 901 and the document addition interface 903 may respectively include more or less content. The specific content included in the document editing interface 901 and the document addition interface 903 is not limited in the embodiments of the present application.
[0154] When the user triggers the card and certificate scanning control 904, in response to the user's triggering operation, as described above Figure 7 shown, the electronic device can display the card and certificate scanning interface 701. For the specific card and certificate scanning interface 701, reference can be made to the relevant description in the corresponding embodiment above Figure 7 and will not be elaborated here.
[0155] Figure 10 The flowchart of the image recognition method according to the embodiment of the present application is shown.
[0156] S1001. Viewfinder preview stream.
[0157] In the above-mentioned card and certificate scanning interface 701, the application can start the camera preview stream to obtain the image in the viewfinder 703.
[0158] S1002. Extract frames at each preset interval to obtain images.
[0159] The application can extract the images in the camera preview stream at each preset interval. The specific preset interval can be custom-set by the application and is not limited in the embodiment of the present application.
[0160] S1003. The application performs multi-object detection to determine whether at least one object is detected.
[0161] The application can perform multi-object detection on the obtained images to determine whether at least one object is detected.
[0162] In the embodiment of the present application, the cards, documents and other certificates waiting to be scanned can be understood as targets or objects. For the convenience of description, the card object will be used as an example for subsequent description.
[0163] If the application does not detect at least one card object, it can continue to obtain the images in the camera preview stream for detection.
[0164] If the application detects at least one card object, it can execute step S1004 to select the card object.
[0165] S1004. The application automatically selects a certain card object, or the user manually selects.
[0166] In a possible implementation, when the application detects more than one card object, it can select one of the multiple card objects as the object to be recognized. The specific process of selecting the card object can be referred to the relevant description in the corresponding embodiment below Figure 11 and will not be elaborated here.
[0167] In another possible implementation, when the application detects more than one card object, the user can manually select one of the multiple card objects as the object to be recognized.
[0168] After selecting the object to be recognized, the application can execute step S1005.
[0169] S1005: Capture an image of the selected object area.
[0170] The application can capture the image corresponding to the object to be recognized. On the one hand, the application can execute step S1006 and input the object to be recognized into the image classification model. On the other hand, the application can execute step S1009 and input the object to be recognized into the text recognition model.
[0171] S1006: The image classification model recognizes the object category.
[0172] The image classification model can classify the object to be recognized and recognize the object category of the object to be recognized. The object category can be understood as the above-mentioned card types. For example, the object category includes household register, driver's license, vehicle license, bank card, ID card, etc.
[0173] It can be understood that different objects to be recognized have different object features. For example, the object features of an ID card may include a portrait on the right side of the card object, etc.
[0174] In a possible implementation, the image classification model can perform image recognition on the object to be recognized based on the object features. The image classification model can use any possible implementation method for image recognition, which is not limited in the embodiments of the present application.
[0175] After the image classification model classifies the object to be recognized, it can execute step S1007.
[0176] S1007: Output the object categories sorted by confidence.
[0177] The image classification model can output the object category of the object to be recognized and the confidence corresponding to the object category. The confidence can also be referred to as the credibility.
[0178] Optionally, the application can also sort the output content according to the size of the confidence. For example, it can be sorted in ascending order, descending order or other sorting methods, which are not limited in the embodiments of the present application.
[0179] It can be understood that for the same object to be recognized, the image classification model can output one set or multiple sets of object categories and the confidences corresponding to the object categories. Exemplarily, the content output by the image classification model may include: object category 1, confidence 1; object category 2, confidence 2; object category 3, confidence 3, etc.
[0180] The above object category 1 can also be referred to as card 1, and the confidence level 1 can also be referred to as the confidence level of card 1; object category 2 can also be referred to as card 2, and the confidence level 2 can also be referred to as the confidence level of card 2; object category 3 can also be referred to as card 3, and the confidence level 3 can also be referred to as the confidence level of card 3.
[0181] Taking the object to be recognized as an ID card as an example, the content output by the image classification model can include: ID card, 0.9; bank card, 0.4; business card, 0.2, etc. That is to say, the probability that the object to be recognized is an ID card is 0.9, the probability that the object to be recognized is a bank card is 0.4, the probability that the object to be recognized is a business card is 0.2, and the probability that the object to be recognized is an ID card is the highest. It can be understood that since the object to be recognized is an ID card, among the object categories output by the image classification model, the confidence level of the ID card is the highest, which is consistent with the result of image recognition.
[0182] After outputting the object category of the object to be recognized and the confidence level corresponding to the object category, the application can execute step S1008.
[0183] S1008, Whether it is a predefined card type.
[0184] It can be understood that since there are many types of certificates, some applications cannot enumerate all certificates. Therefore, the application can predefine some card types of card objects. For example, the card recognition model can support recognizing a relatively large number of certificate types. Among the relatively large number of certificate types that the card recognition model can recognize, the application can select one or more certificate types as the predefined card types.
[0185] If the card type of the object to be recognized is a predefined card type, step S1010 can be executed.
[0186] If the card type of the object to be recognized is not a predefined card type, it means that there is a high probability that the recognized card type is incorrect. Then the application can re-execute step S1002.
[0187] S1009, The text recognition model recognizes the text content of the object.
[0188] The text recognition model can recognize the text content in the object to be recognized. Taking the ID card as an example, the text recognition model can recognize the relevant information on the ID card, such as name, number, issuing authority, expiration date, etc.
[0189] After recognizing the text content in the object to be recognized, step S1010 can be executed.
[0190] S1010, Compare the actual object recognition text information with the prior object key information, and whether the consistency is greater than the preset value A.
[0191] The application can record the key information corresponding to different types of cards and certificates. For example, the key information corresponding to an ID card may include name, number, issuing authority, expiration date, etc., and the key information corresponding to a bank card may include the Chinese name of the bank, the English name of the bank, number, etc. The key information corresponding to these types of cards and certificates can be referred to as prior knowledge.
[0192] After the text information recognized by the text recognition model and the type of card and certificate recognized by the image classification model, the application can compare the text information with the prior knowledge corresponding to the type of card and certificate, and judge the consistency of the information. Consistency can also be understood as similarity or matching degree.
[0193] It can be understood that among the types of cards and certificates recognized by the image classification model, the type of card and certificate with a higher confidence level can be preferentially used for comparison. In this way, the calculated consistency is relatively high, and there is no need to continue using the type of card and certificate with a lower confidence level for comparison, thereby reducing unnecessary calculation amounts.
[0194] The preset value A can be obtained by the application through means such as experience or laboratory tests, and the embodiments of the present application do not make limitations.
[0195] If the consistency between the text information and the prior knowledge corresponding to the type of card and certificate is greater than the preset value A, it indicates that the matching degree between the text information and the prior knowledge corresponding to the type of card and certificate is very high. That is to say, the correct rate of the certificate recognized by the application is relatively high, and step S1011 can be executed.
[0196] If the consistency between the text information and the prior knowledge corresponding to the type of card and certificate is less than or equal to the preset value A, it indicates that the matching degree between the text information and the prior knowledge corresponding to the type of card and certificate is relatively low. That is to say, the certificate recognized by the application may not be very accurate, and step S1002 can be executed to re-perform object recognition.
[0197] Optionally, the case of being equal to the preset value A can also be determined as having a relatively high consistency, and the embodiments of the present application do not make limitations.
[0198] S1011. Output the type of card and certificate.
[0199] The interface corresponding to the output type of card and certificate can refer to the interface 707 of the type of card and certificate corresponding to b in the above Figure 7 and will not be elaborated here.
[0200] S1012. Perform a scanning operation.
[0201] S1013. Use the type of card and certificate as subsequent input.
[0202] When the user determines that the type of card and certificate output by the application is correct, a scanning operation can be performed, and the type of card and certificate can be used as subsequent input to scan and print the card and certificate object.
[0203] Optionally, the above step S1004 may not be executed. The application may not automatically select a certain card object, or the user may select it manually. Instead, multiple documents or cards may be processed simultaneously. The embodiments of the present application do not make any limitations in this regard.
[0204] Optionally, the above step S1008 may not be executed. The application may not determine whether it is a predefined card type. Instead, after executing step S1007 and obtaining the card type and the confidence level corresponding to the card type, step S1010 may be executed to perform a consistency comparison between the text information and the prior knowledge. This can reduce the judgment process of the code and thus save computing power.
[0205] Figure 11 The flowchart of the card object selection method is shown.
[0206] S1101. Open the camera preview stream.
[0207] After the application opens the camera preview stream, it can detect multiple card objects included in the image in the camera preview stream and execute step S1102.
[0208] Exemplarily, as shown in Figure 12 a of FIG. 1201, in the image 1201 extracted from the camera preview stream, the application can detect the region 1202 corresponding to object 1, the region 1203 corresponding to object 2, and the region 1204 corresponding to object 3. The center point of the image 1201 is 1205. Among them, the center point of the image 1201 can also be referred to as the camera preview stream center point.
[0209] Or, as shown in Figure 12 b of FIG. 1201, in the image 1201 extracted from the camera preview stream, the application can detect the region 1206 corresponding to object 4 and the region 1207 corresponding to object 5.
[0210] S1102. Whether the camera preview stream center point is within a certain object region.
[0211] The application can determine whether the camera preview stream center point is within a certain object region.
[0212] If the camera preview stream center point is within a certain object region, for example, Figure 12 in a of FIG. 1201, the camera preview stream center point 1205 is within the region 1203 corresponding to object 2, then step S1103 can be executed.
[0213] If the camera preview stream center point is not within a certain object region, for example, Figure 12In b, if the center point 1205 of the camera preview stream is not within the area 1206 corresponding to object 4 nor within the area 1207 corresponding to object 5, step S1104 can be executed.
[0214] S1103. By default, select the card object where the center point of the preview stream is located.
[0215] Select the card object where the center point of the preview stream is located as the object to be recognized, and highlight the object to be recognized on the interface. Here, the highlighting can be understood as highlighting the border of the object to be recognized or providing a text prompt for the object to be recognized. The specific way of highlighting the object to be recognized is not limited in the embodiments of the present application, as long as it can prompt the user of the selected object to be recognized.
[0216] After the application defaults to selecting the card object where the center point of the preview stream is located, step S1108 can be executed.
[0217] S1104. Calculate the Manhattan distance between the center point of the preview stream and the center points of each object, and sort them.
[0218] Exemplarily, taking Figure 12 b as an example, the application can calculate the Manhattan distance between the center point 1205 of the preview stream and the center points of each object area. It can be understood that using the Manhattan distance can make the calculated data more stable. Assume that the Manhattan distance between the center point 1205 of the preview stream and the center point of area 1206 is d1, and the Manhattan distance between the center point 1205 of the preview stream and the center point of area 1207 is d2. Of course, if there are more card objects, more Manhattan distances can be calculated, such as d3, d4, d5...
[0219] After calculating the Manhattan distance between the center point 1205 of the preview stream and the center points of each object area, step S1105 can be executed.
[0220] Optionally, other methods can also be used to calculate the distance between the center point 1205 of the preview stream and the center points of each object area. For example, the straight-line distance between the center point 1205 of the preview stream and the center points of each object area can be calculated. The specific method of calculating the distance between the center point 1205 of the preview stream and each object is not limited in the embodiments of the present application.
[0221] S1105. Whether the difference between the Manhattan distances of each object is less than or equal to a preset value B.
[0222] It can be understood that this step can be regarded as a kind of fault tolerance mechanism. Exemplarily, taking Figure 12Taking b as an example, if the Manhattan distance difference between object 4 and object 5 is less than or equal to the preset value B, it indicates that the distance between object 4 and the center point 1205 of the preview stream is about the same as the distance between object 5 and the center point 1205 of the preview stream. It can be considered that the Manhattan distances corresponding to object 4 and object 5 are the same, and then the Manhattan distance will no longer be used as the standard for selecting objects.
[0223] Among them, the preset value B can be set to 10%, or it can be set to other values. The specific value of the preset value B can be obtained by the application according to experience or laboratory tests, etc. The embodiments of the present application do not make any limitations.
[0224] Suppose the Manhattan distance from the center point 1205 of the preview stream to the center point of object 4 is d4, and the Manhattan distance from the center point 1205 of the preview stream to the center point of object 5 is d5.
[0225] In one possible implementation, the application can calculate the relationship between the value of |d4 - d5| and the preset value B.
[0226] In another possible implementation, the application can calculate the relationship between the value of |d4 - d5| / d4 and the preset value B, or calculate the relationship between the value of |d4 - d5| / d5 and the preset value B.
[0227] It can be understood that the application can adopt any method to calculate the relationship between the Manhattan distance of each object and the preset value B, as long as it can identify the distance relationship of the Manhattan distances between objects. For example, the application can also calculate the relationship between the multiple of |d4 - d5| and the preset value B, or calculate the relationship between the multiple of |d4 - d5| / d4 and the preset value B, or calculate the relationship between the multiple of |d4 - d5| / d5 and the preset value B, etc. The embodiments of the present application do not make any limitations.
[0228] If the Manhattan distance difference of each object is less than or equal to the preset value B, step S1106 can be executed. If the Manhattan distance difference of each object is greater than the preset value B, step S1107 can be executed.
[0229] Optionally, the case equal to the preset value B can also execute step S1107. The embodiments of the present application do not make any limitations.
[0230] S1106. By default, select the object with the largest area among multiple objects.
[0231] Taking Figure 12 Taking b as an example, if the Manhattan distance difference between object 4 and object 5 is less than or equal to the preset value B, the application can select the object with the larger area occupied by object 4 or object 5 in the image 1201. For example, if the area 1206 of object 4 is larger than the area 1207 of object 5, the application can select object 4 as the object to be recognized.
[0232] It can be understood that when the areas of object 4 and object 5 are close, an object with the second largest or larger area can also be selected, and the embodiments of the present application do not make any limitations in this regard.
[0233] After the application selects the object to be recognized, step S1108 can be executed.
[0234] S1107. By default, select the object corresponding to the smallest Manhattan distance among all objects.
[0235] If the difference in Manhattan distances of all objects is greater than a preset value B, the application can select the object corresponding to the minimum or smaller Manhattan distance as the object to be recognized.
[0236] After the application selects the object to be recognized, step S1108 can be executed.
[0237] S1108. Receive an operation of clicking on other areas of the preview stream.
[0238] It should be noted that after the application selects the object to be recognized, the object to be recognized can be specially marked to prompt the user of the object selected by the application. Among them, the special marking can include highlighting the border of the object to be recognized, and / or displaying prompt information related to the object to be recognized in the interface, etc. The specific way of special marking is not limited in the embodiments of the present application, as long as the user can understand the object selected by the application.
[0239] In a possible scenario, the object to be recognized selected by the application may not be the document that the user wants to recognize. In this case, the user can click on other areas in Image 1201 to make a new selection.
[0240] After the user clicks on other areas in Image 1201, in response to the user's click operation, the application can execute step S1109.
[0241] S1109. The clicked area is an object area.
[0242] The application can determine whether the area clicked by the user is an object area. For example, the object area can include the area where the card object is located, and the non-object area can include the blank area in Image 1201 or the area where no card object is displayed.
[0243] If the area clicked by the user is a non-object area, it means that there is currently no object that can be used as the object to be recognized. Then, the application still uses the object to be recognized selected in step S1103, step S1106, or step S1107 above, and executes step S1110.
[0244] If the area clicked by the user is an object area, it means that the user has reselected the object to be recognized, and then step S1111 can be executed.
[0245] S1110. The object selection box remains unchanged.
[0246] S1111. Select the object selected by the user.
[0247] S1112. Receive the operation of clicking on the preview stream area.
[0248] It can be understood that the user can click on the object area or the non-object area in the image 1201 multiple times. That is to say, the user can repeat the above step S1108, and the application can execute steps S1109 - S1111 multiple times.
[0249] It can be understood that in the card type interface 707 corresponding to b in the above Figure 7 , the user can trigger the return control 710. Or, in the pop-up window interface 801 in the above Figure 8 , the user can trigger the cancel control 807. In response to the user's trigger operation, the application can return to Figure 7 the card scanning interface 701 corresponding to a. When the application re-scans the card, the application can adjust the confidence based on the user operation, thereby improving the application's ability to identify the card.
[0250] Figure 13 Shows the flowchart of confidence adjustment.
[0251] S1301. Start scanning to obtain Image 1.
[0252] When the application scans in the card scanning interface 701, it can obtain Image 1 and execute step S1302.
[0253] S1302. Classification algorithm, output the classification result, and obtain the confidence ranking.
[0254] The application can calculate and rank the confidence of the object to be recognized in Image 1. The specific implementation can refer to the relevant descriptions of step S1007 etc. in the above Figure 10 and will not be elaborated here.
[0255] The application can obtain the confidence information of Image 1, for example, including: Card 1 corresponding to Image 1, the confidence of Card 1, Card 2, the confidence of Card 2, Card 3, the confidence of Card 3, etc.
[0256] S1303. The interface skips frames to the interface corresponding to Card 1.
[0257] After recognizing the card type of the card object, the application can control to jump from the card scanning interface 701 to the card type interface 707, or jump to the pop-up window interface 801.
[0258] S1304. Whether the user clicks the back button.
[0259] It can be understood that if the user does not click the back button, it means that the application's recognition is accurate and no confidence adjustment is required.
[0260] If the user clicks the back button, in one possible scenario, the user does not want to continue the process of scanning the card and certificate and can click the back button. It can be understood that in this scenario, the application may accurately judge the type of the object to be recognized.
[0261] In another possible scenario, the user believes that the card and certificate type recognized by the application is incorrect and can also click the back button.
[0262] If it is the card and certificate type interface 707, the user can trigger the back control 710. If it is the pop-up window interface 801, the user can trigger the cancel control 807. In response to the user's trigger operation, step S1305 can be executed.
[0263] S1305. Start scanning to obtain Image 2.
[0264] The application can rescan on the card and certificate scanning interface 701 to obtain Image 2 and execute step S1302.
[0265] S1306. The algorithm compares the similarity between Image 1 and Image 2.
[0266] Since the application cannot determine the reason why the user clicks the back button, the application can calculate the similarity between Image 1 and Image 2 and execute step S1307. In this way, the application can judge whether the user has the operation of replacing the card and certificate by comparing the similarity of the two images.
[0267] It can be understood that the application can save the image information obtained last time, and the image information can include the image itself and / or the characteristic information of the image, etc. The specific image information is not limited in the embodiments of the present application as long as the similarity of the images can be compared.
[0268] S1307. Whether the similarity is greater than the similarity threshold.
[0269] If the similarity between Image 1 and Image 2 is greater than the similarity threshold, it means that the similarity between Image 1 and Image 2 is relatively high and the user probably has not replaced the card and certificate. In this case, it is possible that the application recognized the card and certificate incorrectly last time. In this way, the application can execute step S1308, calculate the confidence level, and perform confidence adjustment.
[0270] If Image 1 and Image 2 are less than or equal to the similarity threshold, it indicates that the similarity between Image 1 and Image 2 is low, and the user has probably replaced the card or certificate. In this case, it is likely that there is no problem with applying the previously recognized card or certificate. Thus, the application can execute step S1302 to rescan the card or certificate and calculate the confidence level.
[0271] Optionally, the case of being equal to the similarity threshold can also be determined as having a high similarity between Image 1 and Image 2, which is not limited in the embodiments of the present application.
[0272] It can be understood that the similarity threshold can be obtained by the application through means such as experience or laboratory tests, which is not limited in the embodiments of the present application.
[0273] S1308. Classification algorithm, output the classification result, and obtain the confidence level ranking.
[0274] The application can calculate and rank the confidence levels of the objects to be recognized in Image 2. The specific implementation can refer to the relevant descriptions of step S1007 and the like above Figure 10 and will not be elaborated here.
[0275] The application can obtain the confidence level information of Image 2, for example, including for Image 2: Card 1, the confidence level of Card 1, Card 2, the confidence level of Card 2, Card 3, the confidence level of Card 3, etc. It can be understood that since Image 2 is newly obtained, the subsequent Card 1, Card 2, and Card 3 refer to those in Image 2.
[0276] S1309. Whether the confidence level of Card 1 is greater than that of Card 2, and whether the difference between their confidence levels is greater than the confidence level threshold.
[0277] In this step, determining the confidence level threshold can also be understood as a kind of fault tolerance mechanism. If the difference between the confidence levels of Card 1 and Card 2 is small, it indicates that the difference between the two cards is small and the distinction is not obvious. The application cannot accurately determine whether the current certificate is Card 1 or Card 2. Then the application needs to adjust the confidence level of Card 2.
[0278] Exemplarily, assume that the confidence level information of Image 2 can include: Card 1, 0.85; Card 2, 0.75; Card 3, 0.22. In this case, since the difference between the confidence levels of Card 1 and Card 2 is small, it indicates that the difference between the two cards is small and the distinction is not obvious. The application cannot accurately determine whether the current certificate is Card 1 or Card 2. Then the application needs to adjust the confidence level of Card 2.
[0279] Therefore, if the confidence level of Card 1 is less than or equal to that of Card 2, or the difference between their confidence levels is less than or equal to the confidence level threshold, the application can execute step S1311.
[0280] Exemplarily, assume that the confidence information of Image 2 may include: Card 1, 0.85; Card 2, 0.34; Card 3, 0.22. In this case, since the confidence of Card 1 is greater than that of Card 2, and the difference in their confidences is relatively large. This indicates that the difference between the two cards is relatively large and the distinction is obvious, and the application can more accurately determine whether the current certificate is Card 1 or Card 2.
[0281] Therefore, if the confidence of Card 1 is greater than that of Card 2, and the difference in their confidences is greater than the confidence threshold, the application does not need to adjust the confidence of Card 2 and can execute step S1310.
[0282] It can be understood that the confidence threshold can be obtained by the application through means such as experience or laboratory tests, and the embodiments of the present application do not make any limitations.
[0283] S1310. The interface jumps to the interface corresponding to Card 1.
[0284] It can be understood that in the interface corresponding to Card 1, step S1304 can also be executed. If the user clicks to return, steps S1304 - S1313 can be executed again, which will not be elaborated here.
[0285] S1311. Update the confidence of Card 2: original confidence * (1 + preset value 3).
[0286] It can be understood that the preset value 3 can be understood as Figure 13 a% in [reference], and the preset value 3 can be obtained by the application through means such as experience or laboratory tests, and the embodiments of the present application do not make any limitations.
[0287] Through confidence calculation, the confidence of Card 2 is increased, and the probability of being recognized is increased. When the confidence of Card 2 is higher than that of Card 1, the application can recognize the certificate as Card 2.
[0288] It can be understood that the update method of the confidence of Card 2 is not limited to the above implementation method, and the confidence can also be updated based on other calculation methods, as long as the confidence can be gradually increased during the continuous update process. In this way, the confidence of Card 2 has the opportunity to be higher than that of Card 1, enabling the application to re-recognize the card and adjust the mis-recognized card to the correctly recognized card.
[0289] After updating the confidence of Card 2, step S1312 can be executed.
[0290] S1312. Whether the updated confidence of Card 2 is greater than that of Card 1.
[0291] The application can determine whether the updated confidence of Card 2 is greater than that of Card 1.
[0292] If the confidence level of the updated card certificate 2 is less than or equal to the confidence level of card certificate 1, the application will still identify the certificate as card certificate 1 and execute step S1310.
[0293] If the confidence level of the updated card certificate 2 is greater than the confidence level of card certificate 1, the application may identify the certificate as card certificate 2 and execute step S1313.
[0294] Optionally, in the case where the confidence level of the updated card certificate 2 is equal to the confidence level of card certificate 1, the application may also identify the certificate as card certificate 2, which is not limited in the embodiments of the present application.
[0295] S1313. The interface jumps to the interface corresponding to card certificate 2.
[0296] It can be understood that in the interface corresponding to card certificate 2, step S1304 can also be executed. If the user clicks to return, steps S1304 - S1313 can be executed again, which will not be elaborated here.
[0297] In the embodiments of the present application, by executing the above Figure 13 confidence level adjustment process of the corresponding embodiment, the application can adjust the mis-identified card certificate to the correctly identified card certificate. That is to say, the application can achieve the ability of autonomous learning based on user feedback, thereby improving the user experience.
[0298] The method of the embodiments of the present application will be described in detail below through specific embodiments. The following embodiments can be combined with each other or implemented independently. For the same or similar concepts or processes, they may not be elaborated in some embodiments.
[0299] Figure 14 An image recognition method of the embodiments of the present application is shown. Applied to an electronic device, the method includes:
[0300] S1401. Display a first interface, and the first interface includes a first control.
[0301] In the embodiments of the present application, the first interface can be understood as an interface capable of providing a card certificate scanning function. For example, the first interface may include the more interface 303 shown in b above, and may also include the document adding interface 903 shown in b above. Figure 3 of Figure 9 b above, and may also include the document adding interface 903 shown in b above.
[0302] The first control can be understood as a control for card certificate scanning. For example, the first control may include the card certificate scanning 304 shown in b above, and may also include the card certificate scanning control 904 shown in b above. Figure 3 of Figure 9 b above, and may also include the card certificate scanning control 904 shown in b above.
[0303] S1402. Receive a first operation that triggers the first control.
[0304] In the embodiments of the present application, the first operation may be understood as an operation of clicking on the first control, or may also include other operations that trigger the first control. The embodiments of the present application do not make any limitations in this regard.
[0305] S1403. In response to the first operation, display a second interface. The second interface includes a first area for displaying an image captured by the camera of the electronic device.
[0306] In the embodiments of the present application, the second interface may be understood as a card and certificate scanning interface. For example, the second interface may include the card and certificate type selection interface 401 shown above, or may also include the card and certificate scanning interface 701 shown as a in the above. Figure 4 shown above Figure 7 and the card and certificate scanning interface 701 shown as a in the above.
[0307] The first area may include the viewfinder 603 shown above, or may also include the viewfinder 703 shown as a in the above. Figure 6 shown above Figure 7 and the viewfinder 703 shown as a in the above.
[0308] S1404. When the first area shows a first image, display a third interface. The first area of the third interface shows the first image. The third interface further includes a second area. The second control in the second area is in a selected state, and the third control in the second area is in an unselected state. Moreover, the second control is automatically selected by the electronic device based on the image content of the first image.
[0309] In the embodiments of the present application, the third interface may be understood as an interface for displaying the first image. For example, the third interface may include the card and certificate scanning interface 701 shown as a in the above. Figure 7 shown as a in the above
[0310] The second area may be understood as an area for displaying the document type. For example, the second area may include the card and certificate type area 708 shown as b in the above. Figure 7 shown as b in the above
[0311] The second control may be understood as a display control for the document type corresponding to the first image. For example, the second control may include highlighting the recognized card and certificate type shown as b in the above. For the specific highlighting method, reference may be made to the relevant description in the embodiment corresponding to b in the above, and details will not be elaborated here. Figure 7 shown as b in the above Figure 7 shown as b in the above
[0312] The third control may be understood as a display control for the document type corresponding to the second image. It can be understood that since the current interface shows the interface corresponding to the first image, the third control is not highlighted.
[0313] S1405. When the second image is in the first area, display the fourth interface. The second image is displayed in the first area of the fourth interface. The third control in the second area of the fourth interface is in a selected state, the second control in the second area is in an unselected state, and the third control is automatically selected by the electronic device based on the image content of the second image.
[0314] In the embodiments of the present application, the fourth interface can be understood as an interface for displaying the second image. For example, the fourth interface may include the card scanning interface 701 shown in a of the above Figure 7 . It can be understood that since the current interface is the interface corresponding to the second image, the third control is highlighted and the second control is not highlighted.
[0315] In the image recognition method provided by the embodiments of the present application, during the process of card scanning, the application can intelligently identify the type of the card and actively match the corresponding card type, without the user having to actively select the document type, improving the efficiency of card scanning, optimizing the process of card scanning, and thus enhancing the user experience.
[0316] Optionally, on the basis of the Figure 14 corresponding embodiments, the electronic device stores the image feature information corresponding to each image type. After the second image is displayed in the first area of the fourth interface, it may further include: outputting one or more groups of the image type of the second image and the confidence level corresponding to the image type; identifying the text content of the second image; the third control in the second area of the fourth interface being in a selected state may include: when the matching degree between the image feature information of the first image type of the second image and the text content of the second image is greater than or equal to the first threshold, the third control in the second area of the fourth interface is in a selected state, where the first image type corresponds to the first confidence level, the first confidence level is the maximum value among one or more groups of confidence levels, and the third control matches the first image type.
[0317] In the embodiments of the present application, the image type can be understood as the object category in step S1006 of the above Figure 10 corresponding embodiments, and the image feature information corresponding to the image type can be understood as the object feature in step S1006 of the above Figure 10 corresponding embodiments, which will not be elaborated here.
[0318] Outputting one or more groups of the image type of the second image and the confidence level corresponding to the image type can refer to the relevant descriptions of the 1 group or more groups of object categories and the confidence levels corresponding to the object categories output by the image classification model in step S1007 of the above Figure 10 corresponding embodiments, which will not be elaborated here.
[0319] Identifying the text content of the second image can refer to the aboveFigure 10 In step S1009 of the corresponding embodiment, the relevant description of the text recognition model for recognizing text content will not be elaborated.
[0320] The matching of the image feature information of the first image type with the text content of the second image can refer to the relevant description in step S1010 of the above Figure 10 corresponding embodiment. The first threshold can be understood as the preset value A in step S1010, and will not be elaborated.
[0321] It can be understood that among the card types recognized by the image classification model, the card types with relatively high confidence are preferentially used for comparison. In this way, the calculated consistency is relatively high, and there is no need to continue using the card types with relatively low confidence for comparison, thus reducing unnecessary computational workload.
[0322] Optionally, on the basis of the Figure 14 corresponding embodiment, the second image includes multiple target objects, and the multiple target objects include a first target object. The method may further include: highlighting the first target object in the first area of the fourth interface, where the center of the first area is located in the area where the first target object is located.
[0323] In the embodiments of the present application, the multiple target objects can be understood as the multiple card objects in the above Figure 11 corresponding embodiment, and the first target object can be understood as the object to be recognized in the above Figure 11 corresponding embodiment.
[0324] The area where the first target object is located can be understood as the area 1203 corresponding to object 2 in the a corresponding embodiment of the above Figure 12 When the center of the first area is located in the area where the first target object is located, the first target object is selected as the object to be recognized. This judgment process can refer to the relevant description in steps S1102 and S1103 of the above
[0325] corresponding embodiment and will not be elaborated. Figure 11 Since the user will place the card to be recognized at the center of the first area with a relatively high probability, the application can default to select the object where the center of the first area is located as the object to be recognized. In this way, the card that the user wants to scan can be selected with a relatively high probability, increasing the probability of correct recognition and thus improving the user experience.
[0326] Optionally, in
[0327] Optionally, in Figure 14Based on the corresponding embodiment, the second image includes multiple target objects, and the multiple target objects include a first target object. The method may further include: highlighting the first target object in a first area of the fourth interface, where the distance between the center of the first area and the first target object is less than the distance between the center of the first area and other target objects, and the difference in distance is greater than or equal to a second threshold.
[0328] In the embodiments of the present application, the distance between the center of the first area and the first target object can be understood as the Manhattan distance in the corresponding Figure 11 embodiment above. Of course, it can also be understood as the straight-line distance, or other distances, which are not limited in the embodiments of the present application.
[0329] The manner of determining the target object based on the distance can refer to the relevant descriptions in steps S1104 and S1105 of the corresponding Figure 11 embodiment above, and will not be elaborated here.
[0330] The second threshold can be understood as the preset value B in step S1105 of the corresponding Figure 11 embodiment above, and will not be elaborated here.
[0331] When there is no target object at the center position of the first area, the application can select the object with the closest distance as the object to be recognized. In this way, since the user will probably place the card or certificate to be recognized closer to the center position of the first area, the card or certificate that the user wants to scan can be selected with a high probability, increasing the probability of correct recognition and thus improving the user experience.
[0332] Optionally, based on the corresponding Figure 14 embodiment, the second image includes multiple target objects, and the multiple target objects include a first target object and a second target object. The distance between the center of the first area and the first target object is a first distance, and the distance between the center of the first area and the second target object is a second distance. The method may further include: highlighting the first target object in a first area of the fourth interface, where the area of the region where the first target object is located is larger than the area of the region where the second target object is located, the difference between the second distance and the first distance is less than the second threshold, and the second distance is less than the distance between the center of the first area and other target objects.
[0333] In the embodiments of the present application, the first target object can be understood as object 4 in the corresponding Figure 12 embodiment of b above. The second target object can be understood as object 5 in the corresponding Figure 12 embodiment of b above.
[0334] The manner of determining the target object based on the area can refer to the relevant descriptions in step S1106 of the corresponding Figure 11 embodiment above, and will not be elaborated here.
[0335] When there is no target object at the center position of the first area, and there are two objects that are relatively close to the center position of the first area and have similar distances, the application can select the object with the largest or relatively large area as the object to be recognized. In this way, since the card or certificate to be recognized will probably occupy a relatively large area in the first area, it is possible to select with a relatively high probability the card or certificate that the user wants to scan, increase the probability of correct recognition, and thus improve the user experience.
[0336] Optionally, based on the Figure 14 corresponding embodiment, the fourth interface further includes a fourth control. The method may further include: receiving a second operation that triggers the fourth control; in response to the second operation, displaying a second interface; when the first area is the third image, displaying a fifth interface, where the first area of the fifth interface displays the third image, the third control in the second area of the fifth interface is in a selected state, the second control in the second area is in an unselected state, and the third control is automatically selected by the electronic device based on the image content of the third image, and the similarity between the third image and the second image is greater than or equal to a third threshold.
[0337] In the embodiments of the present application, the fourth control can be understood as a control for canceling the currently selected certificate type of the application. For example, the fourth control may include the return control 710 of the certificate type interface 707 in b of the above Figure 7 .
[0338] The second operation can be understood as an operation of clicking the fourth control, and may also include other operations that trigger the fourth control, which are not limited in the embodiments of the present application.
[0339] The fifth interface can be understood as an interface for displaying the third image. For example, the fifth interface may include the certificate scanning interface 701 shown in a of the above Figure 7 .
[0340] The third threshold can be understood as the similarity threshold in step S1307 of the corresponding embodiment above, and will not be elaborated here. Figure 13 It can be understood that since the similarity between the third image and the second image is relatively high, it indicates that the user has not replaced the card or certificate, so the application may still recognize the third image as the third control. This can enable the application to still correctly recognize the type of the card or certificate after the user clicks the return button due to a misoperation.
[0341] Optionally, in
[0342] Optionally, in Figure 14Based on the corresponding embodiment, after the third image is displayed in the first area of the fifth interface, it may further include: outputting one or more groups of image types of the third image and the confidence levels corresponding to the image types based on the image content of the third image; when the third control in the second area of the fifth interface is in a selected state, it may include: when the second confidence level of the first image type of the third image is greater than the confidence levels corresponding to other image types of the third image, and the difference between the second confidence level and other confidence levels is greater than or equal to the fourth threshold, the third control in the second area of the fifth interface is in a selected state.
[0343] In the embodiments of the present application, the fourth threshold can be understood as the confidence threshold in step S1309 of the corresponding embodiment above, which will not be elaborated here. Figure 13 When the difference between the second confidence level and other confidence levels is greater than or equal to the fourth threshold, it indicates that the difference in confidence levels is large, the difference between the two certificates is large, the distinction is obvious, and the application can more accurately determine the type of the current certificate, so there is no need to adjust the confidence level.
[0344] By judging the confidence threshold, the accuracy of application recognition can be determined. If the difference between the second confidence level and other confidence levels is greater than or equal to the fourth threshold, it indicates that the difference in confidence levels is large, and the application is likely to be accurate in recognition.
[0345]
[0346] Figure 14 Optionally, based on the corresponding embodiment, the method may further include: receiving a third operation that triggers the fourth control; in response to the third operation, displaying the second interface; when the first area is the fourth image, displaying the sixth interface, where the fourth image is displayed in the first area of the sixth interface, the fifth control in the second area of the sixth interface is in a selected state, the third control in the second area is not in a selected state, and the fifth control is automatically selected by the electronic device based on the image content of the fourth image, and the similarity between the fourth image and the second image is greater than or equal to the third threshold.
[0347]
[0348] In the embodiments of the present application, the third operation can be understood as an operation of clicking the fourth control, or may include other operations that trigger the fourth control, which are not limited in the embodiments of the present application.
[0348] The sixth interface can be understood as an interface for displaying the fourth image. For example, the sixth interface may include the card certificate scanning interface 701 shown in a of the above. Figure 7
[0349] The fifth control can be understood as a display control for the document type corresponding to the fourth image. It can be understood that after re-identification, the type of the card can be updated. Therefore, the image type of the fifth control object is different from that of the third control object, but the image type corresponding to the fifth control may be the same as or different from that of the second control object.
[0350] Since the similarity between the fourth image and the second image is relatively high, it indicates that the user has not replaced the card. Then there may be a situation of mis-identification in the application. Therefore, after the user clicks the return button due to misoperation, the application can identify the card as another type, so as to determine the type that the user wants to identify.
[0351] Optionally, on the basis of the Figure 14 corresponding embodiment, after the fourth image is displayed in the first area of the sixth interface, it may further include: outputting one or more groups of image types of the fourth image and the confidence levels corresponding to the image types based on the image content of the fourth image; when the second confidence level of the first image type of the fourth image is greater than the third confidence level corresponding to the second image type of the fourth image, the difference between the second confidence level and the third confidence level is less than the fourth threshold, and the third confidence level is greater than the confidence levels of other image types, increasing the value of the third confidence level; when the fifth control in the second area of the sixth interface is in a selected state, it may include: when the third confidence level is greater than or equal to the second confidence level, the fifth control in the second area of the sixth interface is in a selected state, and the fifth control matches the second image type.
[0352] In the embodiments of the present application, the process of confidence adjustment may refer to the relevant descriptions in step S1309-step S1313 of the above-mentioned Figure 13 corresponding embodiment, which will not be elaborated here.
[0353] Increasing the value of the third confidence level may refer to the relevant description in step S1311 of the above-mentioned Figure 13 corresponding embodiment, which will not be elaborated here.
[0354] By adjusting the confidence level of card 2, the confidence level of card 2 has the opportunity to be higher than that of card 1. In this way, the application can achieve the ability of autonomous learning based on user feedback, thereby improving the user experience.
[0355] Optionally, in Figure 14Based on the corresponding embodiment, the method may further include: receiving a fourth operation that triggers the fourth control; in response to the fourth operation, displaying a second interface; when the first region is the fifth image, displaying a seventh interface, where the fifth image is displayed in the first region of the seventh interface, the sixth control in the second region of the seventh interface is in a selected state, the third control in the second region is in an unselected state, and the sixth control is automatically selected by the electronic device based on the image content of the fifth image, and the similarity between the fifth image and the second image is less than a third threshold.
[0356] In the embodiments of the present application, the fourth operation can be understood as an operation of clicking the fourth control, or can also include other operations that trigger the fourth control, which are not limited in the embodiments of the present application.
[0357] The seventh interface can be understood as an interface for displaying the fifth image. For example, the seventh interface may include the Figure 7 card scanning interface 701 shown in a of the above.
[0358] The sixth control can be understood as a display control for the document type corresponding to the fifth image. It can be understood that since the similarity between the third image and the second image is relatively low, it indicates that the user has replaced the card or document. The application can re-identify the type of the card or document after replacement. Therefore, the image type of the sixth control object is different from that of the third control object, but the image type corresponding to the sixth control may be the same as or different from that of the second control object.
[0359] The application can determine whether the user has replaced the card or document based on the similarity between the fifth image and the second image. In this way, image recognition can be more accurately based on the user's selection, thereby improving the user experience.
[0360] Optionally, based on the corresponding embodiment, before displaying the fourth interface when the first region is the second image, it may further include: displaying an eighth interface, where the eighth interface includes a prompt message and a seventh control, and the prompt message is used to prompt whether the image content of the second image is recognized correctly; receiving a fifth operation that triggers the seventh control; in response to the fifth operation, displaying the second interface. Figure 14 In the embodiments of the present application, the eighth interface can be understood as an interface for displaying a pop-up prompt. For example, the eighth interface may include the
[0361] pop-up prompt interface 801 shown above. Among them, the prompt message can be understood as the prompt message 803, and the seventh control can be understood as the cancel button 807. Figure 8
[0362] The fifth operation can be understood as an operation of clicking the seventh control, or can also include other operations that trigger the seventh control, which are not limited in the embodiments of the present application.
[0363] When the application recognizes the type of card or certificate, instead of automatically switching to the card or certificate type interface, the application can also pop up a prompt. The user can determine whether the currently recognized card or certificate type is correct based on the prompt of the application. In this way, the application can more accurately identify the card or certificate based on the user's operation.
[0364] Optionally, based on Figure 14 the corresponding embodiment, the second image includes multiple target objects, and the first object among the multiple target objects is highlighted. The method may further include: receiving a sixth operation for triggering a second object among the multiple target objects; in response to the sixth operation, highlighting the second object, and an eighth control in a second area of the fourth interface is in a selected state, and the eighth control is automatically selected by the electronic device based on the image content of the second object.
[0365] In the embodiments of the present application, the sixth operation can be understood as an operation of clicking on the second object, or may also include other operations for triggering the second object, which is not limited in the embodiments of the present application.
[0366] The eighth control can be understood as a display control corresponding to the certificate type of the second object. It can be understood that since the second object and the first object may be objects of the same type or different types, therefore, the image type of the eighth control object and the image type of the third control object may be different or the same.
[0367] When the user clicks on other objects, the application can display the object to be recognized according to the user's selection. This can improve the recognition accuracy of the image.
[0368] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of the method. To implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the method steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0369] The embodiments of the present application can divide the device for implementing the method into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0370] Such as Figure 15 Shown is a schematic structural diagram of a chip provided by an embodiment of the present application. The chip 1500 includes one or more than two (including two) processors 1501, a communication line 1502, a communication interface 1503, and a memory 1504.
[0371] In some embodiments, the memory 1504 stores the following elements: executable modules or data structures, or subsets thereof, or extended sets thereof.
[0372] The methods described in the above embodiments of the present application can be applied to the processor 1501 or implemented by the processor 1501. The processor 1501 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 1501. The above-mentioned processor 1501 may be a general-purpose processor (e.g., a microprocessor or a conventional processor), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. The processor 1501 can implement or execute the various processing-related methods, steps, and logic block diagrams disclosed in the embodiments of the present application.
[0373] The steps of the method disclosed in the embodiments of the present application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. Among them, the software module can be located in a mature storage medium in the art such as a random access memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable read-only memory (EEPROM). This storage medium is located in the memory 1504, and the processor 1501 reads the information in the memory 1504 and combines its hardware to complete the steps of the above method.
[0374] Communication can be carried out between the processor 1501, the memory 1504, and the communication interface 1503 through the communication line 1502.
[0375] In the above embodiments, the instructions stored in the memory for the processor to execute can be implemented in the form of a computer program product. Among them, the computer program product can be pre-written in the memory in advance, or downloaded and installed in the memory in software form.
[0376] The embodiments of the present application also provide a computer program product including one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are fully or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a server or a data center including one or more available media integrated. For example, the available medium can include magnetic media (such as floppy disks, hard disks, or magnetic tapes), optical media (such as digital versatile discs (DVDs)), or semiconductor media (such as solid state disks (SSDs)).
[0377] The embodiments of the present application also provide a computer-readable storage medium. The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. A computer-readable medium may include a computer storage medium and a communication medium, and may also include any medium that can transfer a computer program from one place to another. A storage medium may be any target medium accessible by a computer.
[0378] As a possible design, a computer-readable medium may include a compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; a computer-readable medium may include a magnetic disk storage or other magnetic disk storage device. Moreover, any connecting line may also be appropriately referred to as a computer-readable medium. For example, if software is transmitted using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave) from a website, server, or other remote source, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, magnetic disks and optical discs include optical discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where magnetic disks typically reproduce data magnetically, while optical discs use lasers to optically reproduce data.
[0379] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processing unit of the computer or other programmable data processing device generate means for implementing the functions specified in Figure 1 one or more of the flows or multiple flows and / or blocks Figure 1 one or more of the blocks or multiple blocks.
Claims
1. An image recognition method, characterized in that, Applied to an electronic device, the method includes: Displaying a first interface, the first interface including a first control; Receiving a first operation that triggers the first control; In response to the first operation, displaying a second interface, the second interface including a first area for displaying an image captured by a camera of the electronic device; When the first area is a first image, displaying a third interface, the first area of the third interface displaying the first image, the third interface further including a second area, a second control in the second area being in a selected state, a third control in the second area being in an unselected state, and the second control being automatically selected by the electronic device based on the image content of the first image; When the first area is a second image, displaying a fourth interface, the first area of the fourth interface displaying the second image, the third control in the second area of the fourth interface being in a selected state, the second control in the second area being in an unselected state, and the third control being automatically selected by the electronic device based on the image content of the second image.
2. The method according to claim 1, characterized in that After the electronic device stores image feature information corresponding to each image type and the first area of the fourth interface displays the second image, it further includes: Based on the image content of the second image, outputting one or more groups of the image type of the second image and the confidence level corresponding to the image type; Identifying the text content of the second image; The third control in the second area of the fourth interface being in a selected state includes: When the matching degree between the image feature information of the first image type of the second image and the text content of the second image is greater than or equal to a first threshold, the third control in the second area of the fourth interface is in a selected state, where the first image type corresponds to a first confidence level, the first confidence level being the maximum value among one or more groups of confidence levels, and the third control matches the first image type.
3. The method according to claim 1 or 2, characterized in that The second image includes multiple target objects, and the multiple target objects include a first target object. The method further includes: Highlighting the first target object in the first area of the fourth interface, where the center of the first area is within the area where the first target object is located.
4. The method according to claim 1 or 2, characterized in that, The second image includes multiple target objects, and the multiple target objects include a first target object. The method further includes: Highlighting the first target object in the first area of the fourth interface, where the distance between the center of the first area and the first target object is less than the distance between the center of the first area and other target objects, and the difference in distance is greater than or equal to a second threshold.
5. The method according to claim 1 or 2, characterized in that, The second image includes multiple target objects, and the multiple target objects include a first target object and a second target object. The distance between the center of the first area and the first target object is a first distance, and the distance between the center of the first area and the second target object is a second distance. The method further includes: Highlight the first target object in the first area of the fourth interface, where the area of the region where the first target object is located is larger than the region where the second target object is located, the difference between the second distance and the first distance is less than a second threshold, and the second distance is less than the distance between the center of the first area and other target objects.
6. The method according to any one of claims 1-5, characterized in that, The fourth interface further includes a fourth control, and the method further includes: Receiving a second operation that triggers the fourth control; In response to the second operation, display the second interface; When the first area is a third image, display a fifth interface, where the third image is displayed in the first area of the fifth interface, the third control in the second area of the fifth interface is in a selected state, the second control in the second area is in an unselected state, and the third control is automatically selected by the electronic device based on the image content of the third image, and the similarity between the third image and the second image is greater than or equal to a third threshold.
7. The method according to claim 6, characterized in that, After the third image is displayed in the first area of the fifth interface, it further includes: Outputting one or more groups of the image types of the third image and the confidence levels corresponding to the image types based on the image content of the third image; The third control in the second area of the fifth interface being in a selected state includes: When the second confidence level of the first image type of the third image is greater than the confidence levels corresponding to other image types of the third image, and the difference between the second confidence level and other confidence levels is greater than or equal to a fourth threshold, the third control in the second area of the fifth interface is in a selected state.
8. The method according to any one of claims 1 to 7, characterized in that The method further includes: Receiving a third operation that triggers the fourth control; In response to the third operation, display the second interface; When the first area is a fourth image, display a sixth interface, where the fourth image is displayed in the first area of the sixth interface, the fifth control in the second area of the sixth interface is in a selected state, the third control in the second area is in an unselected state, and the fifth control is automatically selected by the electronic device based on the image content of the fourth image, and the similarity between the fourth image and the second image is greater than or equal to a third threshold.
9. The method according to claim 8, wherein After the fourth image is displayed in the first area of the sixth interface, it further includes: Outputting one or more groups of the image types of the fourth image and the confidence levels corresponding to the image types based on the image content of the fourth image; When the second confidence level of the first image type of the fourth image is greater than the third confidence level of the second image type of the fourth image, the difference between the second confidence level and the third confidence level is less than a fourth threshold, and the third confidence level is greater than the confidence levels of other image types, increase the value of the third confidence level; The fifth control in the second area of the sixth interface being in a selected state includes: When the third confidence level is greater than or equal to the second confidence level, the fifth control in the second area of the sixth interface is in a selected state, and the fifth control matches the second image type.
10. The method according to any one of claims 1-9, characterized in that, The method further includes: receiving a fourth operation that triggers a fourth control; in response to the fourth operation, displaying the second interface; when the first area is a fifth image, displaying a seventh interface, where the first area of the seventh interface displays the fifth image, the sixth control in the second area of the seventh interface is in a selected state, the third control in the second area is not in a selected state, and the sixth control is automatically selected by the electronic device based on the image content of the fifth image, and the similarity between the fifth image and the second image is less than a third threshold.
11. The method according to any one of claims 1-10, characterized in that, Before displaying the fourth interface when the first area is the second image, it further includes: displaying an eighth interface, where the eighth interface includes a prompt message and a seventh control, and the prompt message is used to prompt whether the image content of the second image is recognized correctly; receiving a fifth operation that triggers the seventh control; in response to the fifth operation, displaying the second interface.
12. The method according to any one of claims 1-11, characterized in that, The second image includes multiple target objects, and a first object among the multiple target objects is highlighted. The method further includes: receiving a sixth operation that triggers a second object among the multiple target objects; in response to the sixth operation, highlighting the second object, and the eighth control in the second area of the fourth interface is in a selected state, and the eighth control is automatically selected by the electronic device based on the image content of the second object.
13. An electronic device, characterized in that, including: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the electronic device executes the method according to any one of claims 1-12.
14. A chip system, characterized in that, including at least one processor and a communication interface, the communication interface and the at least one processor are interconnected by a line, and the at least one processor is used to run a computer program or instruction to execute the method according to any one of claims 1-12.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, and when the instructions are executed, the computer is caused to execute the method according to any one of claims 1-12.
16. A computer program product, characterized in that, including a computer program, and when the computer program is run, the electronic device is caused to execute the method according to any one of claims 1-12.
Citation Information
Patent Citations
Card image recognition method based on deep learning
CN110909809A
Certificate image classification method and device, computer equipment and readable storage medium
CN111046879A
Application interaction method and device, electronic equipment and storage medium
CN112306601A
Card recognition method and device, electronic equipment and storage medium
CN113111882A
Card text recognition method and device and storage medium
CN115050037A