Material Identification Method, Device, Electronic Device, Medium, and Financial Interaction Device
Through video stream processing and automatic identification technology, the stay-free material identification and processing of financial outlet personnel is achieved, which solves the problems of increased business processing time and reduced service efficiency, and improves identification efficiency and service efficiency.
Patent Information
- Application Number
- CN202210613455.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-31
AI Technical Summary
When handling business, financial outlets need to manually identify and process a variety of materials, resulting in an increase in business processing time and a decrease in service efficiency.
By extracting the target material image from the video stream, performing document classification and text recognition, automatically identifying the material type and selecting the corresponding text recognition model, realizing stop-free material recognition and processing.
It greatly improves the efficiency of material identification, reduces business processing time, reduces labor costs, and improves service efficiency and customer experience.
Smart Images

Figure CN114913539B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to artificial intelligence, NLP, computer vision, deep learning, and intelligent finance technologies. Specifically, it relates to a method, apparatus, electronic device, medium, program product, and financial interaction device for material recognition. Background Art
[0002] During the process of handling various services for customers, financial network staff need to use financial machines with visual perception capabilities, such as high-speed document scanners, to scan various certificates, materials, or bills for identification and archiving.
[0003] For different business types, the types of materials that customers need to provide are also different. Moreover, for the same business, the materials that usually need to be scanned, identified, and archived are not unique. Therefore, in the face of a variety of materials, how to improve the processing efficiency of network staff in using these financial machines to identify materials directly affects the duration of service handling and the service efficiency and effect of the network. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, electronic device, medium, program product, and financial interaction device for material recognition.
[0005] According to one aspect of the present disclosure, there is provided a method for material recognition, including:
[0006] Extracting a target material image from a currently captured video stream;
[0007] Performing document classification on the target material image to determine the target type of the target material;
[0008] Selecting a target text recognition model corresponding to the target type to perform text recognition on the target material image to obtain the structured target data of the target material.
[0009] According to another aspect of the present disclosure, there is provided a material recognition apparatus, including:
[0010] An image extraction module for extracting a target material image from a currently captured video stream;
[0011] A document classification module for performing document classification on the target material image to determine the target type of the target material;
[0012] A text recognition module for selecting a target text recognition model corresponding to the target type to perform text recognition on the target material image to obtain the structured target data of the target material
[0013] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0014] A photographing device for photographing a video stream of a material to be recognized;
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the material recognition method according to any embodiment of the present disclosure.
[0018] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the material recognition method according to any embodiment of the present disclosure.
[0019] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the material recognition method according to any embodiment of the present disclosure.
[0020] According to another aspect of the present disclosure, there is provided a financial interaction device including:
[0021] A photographing device for photographing a video stream of a material to be recognized;
[0022] A video stream processing device for processing the video stream and extracting a target material image from the video stream;
[0023] A material recognition device for classifying the target material image into a document category, determining a target type of the target material, and selecting a target character recognition model corresponding to the target type to perform character recognition on the target material image to obtain structured target data of the target material.
[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0025] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0026] Figure 1 is a schematic diagram of a material recognition method according to an embodiment of the present disclosure;
[0027] Figure 2 is a schematic diagram of a material recognition method according to an embodiment of the present disclosure;
[0028] Figure 3 It is a schematic diagram of a material recognition method according to an embodiment of the present disclosure;
[0029] Figure 4 It is a schematic diagram of a material recognition method according to an embodiment of the present disclosure;
[0030] Figure 5 It is a scenario diagram that can implement a material recognition method according to an embodiment of the present disclosure;
[0031] Figure 6 It is another scenario diagram that can implement a material recognition method according to an embodiment of the present disclosure;
[0032] Figure 7 It is a schematic structural diagram of a material recognition device according to an embodiment of the present disclosure;
[0033] Figure 8 It is a schematic structural diagram of a financial interaction device according to an embodiment of the present disclosure;
[0034] Figure 9 It is a block diagram of an electronic device for implementing the material recognition method according to an embodiment of the present disclosure. Detailed implementation manners
[0035] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0036] Figure 1 It is a flow schematic diagram of a material recognition method according to an embodiment of the present disclosure. This embodiment is applicable to the situation of scanning and recognizing multiple materials, such as the situation where financial network personnel use financial devices such as high-speed document scanners to scan and recognize materials. It relates to the field of computer technology, especially artificial intelligence, NLP, computer vision, deep learning, and intelligent finance technologies. This method can be executed by a material recognition device, which is implemented in software and / or hardware, preferably configured in an electronic device, such as a financial device like a high-speed document scanner, or any other electronic device with a photographing device. As Figure 1 shown, this method specifically includes the following:
[0037] S101. Extract a target material image from the currently captured video stream.
[0038] Taking a document scanner in financial devices as an example, a businessperson places the material to be recognized in the scanning area of the document scanner, and a photographing device (such as a camera) on the document scanner photographs the scanning area to obtain a video stream. Then, using image processing technology, the target material image can be extracted from the video stream. It should be noted that financial devices are just an example. In any scenario, any electronic device equipped with a photographing device is applicable in the embodiments of the present disclosure.
[0039] S102. Perform document classification on the target material image to determine the target type of the target material.
[0040] S103. Select a target character recognition model corresponding to the target type to perform character recognition on the target material image to obtain the structured target data of the target material.
[0041] Among them, performing character recognition on the target material image can be any character recognition technology in the prior art. In one implementation, machine learning methods can be used to pre-train character recognition models for different types of materials in advance. Special training of a corresponding character recognition model for each type of material can improve the character recognition accuracy of each material. For example, a character recognition model for identity cards can be trained for identity cards, and a special character recognition model for business licenses can be trained for business licenses.
[0042] Therefore, in the embodiments of the present disclosure, first perform document classification on the target material to determine the target type of the target material, and then select a target character recognition model corresponding to the target type to perform character recognition on the target material image to obtain more accurate structured target data of the target material. Among them, identifying the type of the target material in the target material image can be achieved by using image processing technology. For example, using a pre-trained image classification model for identification, or extracting keywords from the target material and comparing them with the preset materials to be recognized to determine its type, etc. The embodiments of the present disclosure do not make any limitations on the specific implementation manner of document classification.
[0043] It should be noted that in the prior art, for example, in the scenario of handling business at financial outlets, the types of materials that customers need to provide are usually different for different business types. In the process of handling them, the outlet staff need to place the materials one by one, then press the scan or shoot button, and manually select the recognition model. If the number of materials to be identified is large, this process will directly affect the recognition efficiency of the materials, which in turn affects the length of business handling, and also restricts the service efficiency and effect of the outlets. In the embodiment of the present disclosure, after the device is turned on, it is only necessary to place the material to be identified in the field of view of the shooting device without any other manual operation, and the video stream of the material can be automatically obtained, and the target material image can be extracted from it, and then its target type is determined through document classification, and the corresponding target text recognition model is selected for text recognition. Especially in the case of a large number of materials to be identified, it is possible to feed and identify without stopping, which greatly improves the recognition efficiency of the material.
[0044] The technical solution of the disclosed embodiment extracts the material image through video stream processing, and then classifies the material image into documents, automatically identifies the material type, and can automatically select the text recognition model corresponding to the type to perform text recognition on the material, avoiding the business personnel from manually selecting the recognition model of the material. The business personnel can scan and recognize the material without stopping, which can effectively reduce the operation time of material shooting in different business processes and improve the recognition efficiency. At the same time, there is no need for business personnel to pay attention to the material type. All types of materials can be operated together without stopping, which saves time and effort and greatly reduces labor costs.
[0045] Figure 2 2 is a flow chart of a material identification method according to an embodiment of the present disclosure, and this embodiment is further optimized on the basis of the above embodiment. Figure 2 As shown, the method specifically includes the following:
[0046] S201. Extract frames from a currently captured video stream to obtain a current frame image.
[0047] S202: Perform target detection on the current frame image to obtain the current target material on the current frame image.
[0048] Among them, any image recognition algorithm in the prior art can be used to implement object detection. For example, first determine which materials to be recognized, use the historical data of these materials to be recognized as training samples, and use machine learning methods to train an image recognition model, so that the image recognition model can perform object detection on any material to be recognized that appears in the image. In addition, since hands, irrelevant papers, accidentally dropped coins, stationery and other items often appear in the scanning area during the document intake operation, the model can also learn the characteristics of these items to better perform object detection and accurately process and judge when the material is blocked.
[0049] During object detection, for these materials to be recognized and other items that have been trained, the model can identify them from the current frame image and mark the corresponding detection boxes. Then identify the current target material from these detection boxes. For example, when a businessperson delivers materials in the scanning area, a hand and the current target material appear in the current frame image, and the model can identify them and determine which detection box is the current target material.
[0050] It should be noted that when the number of materials to be recognized is more than two, the current target material can be any one of the materials to be recognized. In terms of the recognition order, compared with any other material among the materials to be recognized, the recognition order of the target material includes prior recognition or subsequent recognition. That is, the businessperson can deliver the materials in any order without being aware of the recognition order. Thus, the businessperson can more efficiently achieve a non-stop document intake, making the entire interaction process of material recognition smoother and more convenient.
[0051] For example, when there are two or more materials to be recognized, one way is that the businessperson places these multiple materials in the scanning area one by one in any order for recognition. Another way is that the businessperson stacks these multiple materials in any order and places them together in the scanning area. After each material is recognized, one material is taken away from the top layer until all materials are recognized. That is, the video stream can be a continuous video stream of a single material. In this scenario, a single frame of the video stream includes a single material. In addition, the video stream can also be a video stream of multiple materials. In this scenario, a single frame of the video stream can include a single material, or multiple stacked materials, or some single frames include a single material and some single frames include multiple stacked materials. The present disclosure does not make any limitation in this regard. It should be noted that in the case where a single frame includes multiple materials, only the material located at the top layer of the stacked materials needs to be recognized each time.
[0052] S203. Determine whether the current target material is blocked. If it is blocked, return to execute S201; otherwise, continue to execute S204.
[0053] If the current target material is occluded, accurate text recognition cannot be performed on it. Therefore, as long as there is occlusion, S201 needs to be re-executed to extract the next frame image of the current frame for recognition.
[0054] In one implementation, the material recognition method of the embodiments of the present disclosure may further include: if a current target material without occlusion cannot be obtained within a preset time period, a preset first warning prompt sound is played. That is, in order to avoid always being unable to extract a suitable material image and affecting the material recognition efficiency, a waiting duration, that is, the preset time period, can be set. If a current target material without occlusion cannot be obtained within this preset time period, an alarm is given to the operator through a prompt sound (such as "there is occlusion") so that the operator can adjust the occluder in time to avoid long waiting.
[0055] Regarding how to determine whether an image is occluded, specifically, any existing algorithm in image processing technology can be used to implement it. In one implementation, the process of determining whether the current target material is occluded may include: if, through target detection, a target object different from the current target material is detected in the current frame image, the overlapping area of the detection frames of the current target material and the target object is determined; if the overlapping area of the detection frames is higher than a preset threshold, it is determined that the current target material is occluded.
[0056] That is to say, if foreign objects such as hands, stationery, or coins appear in the scanning area and are recognized through target detection, then whether there is occlusion is determined according to the overlapping area between the detection frames of these foreign objects and the detection frame of the target material. A threshold can be preset in advance. If the overlapping area does not exceed the preset threshold, it means that it does not affect the text recognition of the target material, and it is determined that there is no occlusion. Otherwise, it means that it will affect the text recognition of the target material, and it is determined that there is occlusion.
[0057] In the single-piece feeding method, since only one piece of material is placed in the scanning area each time, the foreign objects that may appear in the current frame image are other objects except at least one material to be recognized. In the stacked feeding method, the operator stacks a plurality of current materials to be recognized together and places them in the scanning area. Therefore, the current material image detected in the current frame image is the material located at the top layer, and only the material located at the top layer is not occluded. The embodiments of the present disclosure do not make any limitations on the feeding method and can be applicable to any feeding method. In this way, the operator does not need to be restricted by the feeding method, thereby further improving the speed of business processing.
[0058] S204. Determine whether the exposure parameter of the current frame image meets the preset standard. If not, return to execute S201; otherwise, continue to execute S205.
[0059] S205. Use the current frame image as the target material image.
[0060] To further improve the clarity of the target material image, after determining that there is no occlusion, it is also necessary to further determine whether its exposure parameter meets the preset standard. If it does not meet the standard, return to execute S201, continue to extract the next frame image of the current frame for processing. If it meets the standard, the current frame image can be used as the target material image.
[0061] In one implementation, if the exposure parameter of the current frame image does not meet the preset standard, a preset second warning prompt sound can be played so that the operator can assist in adjusting the light source around the shooting device, enabling a clear image with an exposure parameter that meets the standard to be obtained in a timely manner. Of course, in the embodiments of the present disclosure, the device can also adjust the light source around the shooting device through self-adjustment, which can be implemented according to the existing technology and will not be elaborated here.
[0062] S206. Classify the target material image into documents to determine the target type of the target material.
[0063] S207. Select a target character recognition model corresponding to the target type to perform character recognition on the target material image, and obtain the structured target data of the target material.
[0064] The technical solution of the embodiments of the present disclosure extracts the material image through video stream processing, then classifies the material image into documents, automatically identifies the material type, and thus can automatically select a character recognition model corresponding to this type to perform character recognition on the material, avoiding the business personnel from manually selecting the recognition model of the material. The business personnel can perform document scanning and recognition without pause, and can effectively reduce the operation time of material shooting in different business processes, improving the recognition efficiency. Moreover, during the recognition process, it can automatically judge whether the material is occluded and whether the exposure parameter meets the requirements, and interact with the operator through a prompt sound when there is occlusion or the exposure condition does not meet the requirements, allowing the operator to correct it in a timely manner. During the process of placing the material, the operator does not need to pay attention to the order of the materials to be recognized, occlusions or light sources, etc., and only needs to adjust according to the prompt sound. This not only reduces the labor cost, but also can shorten the shooting time and improve the recognition efficiency.
[0065] Figure 3 It is a flowchart of the material recognition method according to the embodiments of the present disclosure, and this embodiment is further optimized on the basis of the above embodiments. As Figure 3 shown, the method specifically includes the following:
[0066] S301. Extract the target material image from the currently captured video stream.
[0067] S302. Extract any entity from the target material image and determine the inclination angle of the any entity.
[0068] S303. Adjust the angle of the target material image according to the inclination angle of the any entity.
[0069] Among them, the any entity can be an entity such as text or a graphic, and the inclination angle can be determined according to the camera coordinate system of the photographing device. Adjusting the angle of the target material image upright can not only improve the accuracy of subsequent text recognition, but also facilitate the archiving of the material image and improve the quality of the archived image. Regarding the method for entity extraction, any image recognition method in the prior art can be used to implement the embodiments of the present disclosure, and details are not described herein.
[0070] S304. Classify the target material image according to the image features of the target material image to determine the target type of the target material.
[0071] Among them, the image features can be the image features extracted by using any image recognition method when recognizing the target material image, or the image features extracted during the process of detecting the target when extracting the target material image from the video stream. A feature library of multiple materials to be recognized can be configured in advance, and then the image features of the target material image are compared with the features in the feature library to determine the target type of the target material. Of course, other image classification methods in the prior art can also be used to classify the target material image, and the embodiments of the present disclosure do not make any limitation thereto.
[0072] S305. Crop the target material image according to the background area in the target material image.
[0073] Use the image recognition method to recognize the foreground and background of the target material image, or after target detection, the area other than the target is the background area. After cropping, a more accurate material image can be obtained, which is beneficial to subsequent archiving and improves the accuracy of text recognition, avoiding interference from the background area.
[0074] S306. Select a target text recognition model corresponding to the target type to perform text recognition on the target material image to obtain the structured target data of the target material.
[0075] S307. Determine whether the target material image is the same as the image of the recognized material according to the structured target data of the target material. If they are the same, execute S308; otherwise, execute S309.
[0076] S308. Play a preset third warning prompt sound.
[0077] S309. Play a preset recognition success prompt sound.
[0078] Specifically, during the process of the operator feeding in materials, in order to avoid duplicate feeding or duplicate recognition caused by other reasons, the embodiments of the present disclosure can, after identifying the structured target data, determine whether there is a duplicate with the images of the materials that have been recognized. For example, compare the currently identified structured target data with the structured data extracted from the images of the recognized materials. If they are the same, it means that the current target material image is a duplicate. Then, a third warning prompt sound (such as "Do not scan repeatedly") can be played to promptly notify the operator to make adjustments. Otherwise, a recognition success prompt sound, such as a short "beep" sound, can be played. In this way, the operator can know that the current material has been recognized, without the need for the operator to manually check the recognition status of the current material, and only need to focus on the feeding operation, thereby achieving non-stop feeding, scanning, and recognition, improving the overall recognition efficiency and the speed of business processing.
[0079] In addition, to further improve the interactivity between the operator and the device, in the embodiments of the present disclosure, the structured target data of the target material can also be displayed. Moreover, a list of materials to be recognized in the current task can be displayed, and the recognized and unrecognized materials in the material list can be distinguished and displayed. That is to say, when handling a business, after determining the materials to be recognized for the current business, the device can display the list of materials to be recognized in the current task. Then, during the process of feeding and recognizing, each time a material is recognized, the material in the list can be highlighted in any way to distinguish the unrecognized and recognized materials. Whether it is the operator or the user handling the business, it can be clearly seen at a glance.
[0080] In one implementation manner, the processes of video stream processing, document classification, and text recognition can be implemented on the device side. In another implementation manner, according to the actual situation, the text recognition process can also be implemented in the cloud. For example, after the device determines the target type of the current target material, the target material image and the target type are sent to the cloud server together. The cloud server selects the corresponding text recognition model according to the target type to perform text recognition on the target material image, and then feeds back the extracted structured data to the device side. At the same time, the material image can also be stored in the cloud server. The embodiments of the present disclosure do not make any limitations on the above two implementation manners.
[0081] In addition, to further improve the speed of business processing, the method according to the embodiments of the present disclosure may also display the review results of structured target data. Specifically, the cloud server is used to review the structured data of the materials required for the current business processing in real time, and the review results are fed back from the cloud to the device side and displayed simultaneously. If there are any abnormalities in the review results, the operator can make timely adjustments, achieving a "one-stop processing" and "only one trip" processing flow where users submit materials, business personnel receive and review the materials, and problems are clarified on the spot, reducing the probability of reprocessing the business. The back-office department can also only review the materials, reducing unnecessary workload and overall improving the speed of material processing and work efficiency.
[0082] The technical solution of the embodiments of the present disclosure can not only classify the material images into documents, automatically identify the material types, and automatically select the corresponding text recognition model for the materials according to the type, avoiding business personnel from manually selecting the recognition model for the materials, realizing non-stop document scanning and recognition, and improving the recognition efficiency. At the same time, based on the interactive feedback mechanism of the device, the business processing efficiency can be further improved, making the material processing flow of document submission, scanning, and recognition smoother, also reducing the operation threshold of the device and the labor cost.
[0083] Figure 4 is a flowchart of the material recognition method according to the embodiments of the present disclosure, and this embodiment is further optimized on the basis of the above embodiments. As Figure 4 shown, exemplarily, in the embodiments of the present disclosure, the entities of the materials required for business processing include certificates, Material 1 and Material 2 of the material category, and also include Material 3 of the bill category. They are scanned by a machine, and the machine performs video stream processing and automatic picture classification. Then, the machine sends the document category and the corresponding pictures to the cloud, and the cloud performs picture text recognition processing. For Material 1 and Material 2, the models corresponding to the certificate and the material category are selected for recognition, and for Material 3, the model corresponding to the bill category can be selected for recognition. Structured data is obtained through KV (Key-Value) extraction.
[0084] Figure 5 is a scenario diagram of a material recognition method that can implement the embodiments of the present disclosure. Figure 5A list of materials required for business processing is shown. In this list, the names of the materials to be currently recognized are sequentially shown in the left area, and the right area is initially blank, indicating that they have not been recognized yet. Then, after being scanned and recognized by the device, the list of materials required for business processing will be automatically refreshed, and the originally blank area on the right side will display the currently recognized material images and corresponding structured data for business personnel to view in a timely manner. At the same time, for the materials that have been recognized, they are distinguished by "check marks" and colors in the left area of the list, so it is clear at a glance which materials have not been recognized yet. It should be noted that this list of materials required for business processing can be displayed on any device connected to the device (such as a PC or a mobile device). For a device that itself has a display device, it can also be directly displayed on the display device of the device itself. The embodiments of the present disclosure do not make any limitations in this regard.
[0085] Figure 6 is another scenario diagram that can implement the material recognition method of the embodiments of the present disclosure. Through Figure 6 As can be seen from the shown content, this scenario is a financial business processing scenario where customers submit materials and business clerks use devices to scan and recognize the materials. Among them, any device (such as a PC or a mobile device) can be bound to the device. During the process of scanning and recognizing by the device, the device can locate problems when problems occur and display corresponding prompt messages through the device bound to the device, such as "Please align the image to be scanned", "Please reposition the image" or "Please get closer", etc., without the business clerk having to locate the problems by themselves, which provides convenience for the business clerk. At the same time, the structured data is audited by the cloud and the audit results are timely fed back to the device and displayed through the device connected to the device, which is convenient for the business clerk to adjust in a timely manner and improves the efficiency of business processing. For the business clerk, they can perform in-piece scanning without stopping and process it in a timely manner after the audit results are returned by the cloud. For the customers who submit materials, they only need to make one trip, avoiding repeated business processing.
[0086] Figure 7 is a schematic structural diagram of a material recognition device according to the embodiments of the present disclosure. This embodiment is applicable to the situation of scanning and recognizing multiple materials, such as the situation where financial network personnel use financial devices such as high-definition cameras to scan and recognize materials, which involves the field of computer technology, especially artificial intelligence, NLP, computer vision, deep learning, and intelligent finance technologies. This device can implement the material recognition method described in any embodiment of the present disclosure. As Figure 7 shown, this device 700 specifically includes:
[0087] An image extraction module 701, configured to extract a target material image from the currently captured video stream;
[0088] The document classification module 702 is used to classify the target material image to determine the target type of the target material;
[0089] The character recognition module 703 is used to select a target character recognition model corresponding to the target type to perform character recognition on the target material image, and obtain the structured target data of the target material.
[0090] Optionally, the target material is any one of at least one material to be recognized.
[0091] Optionally, the image extraction module 701 includes:
[0092] The frame extraction unit is used to extract frames from the currently captured video stream to obtain the current frame image;
[0093] The target detection unit is used to perform target detection on the current frame image for the at least one material to be recognized, and obtain the current target material on the current frame image;
[0094] The target material image determination unit is used to use the current frame image as the target material image in response to the current target material not being occluded.
[0095] Optionally, the image extraction module 701 further includes:
[0096] The loop unit is used to, in response to the current target material being occluded, re-execute the operation of extracting frames from the currently captured video stream until the current target material is not occluded.
[0097] Optionally, the device further includes:
[0098] The first warning prompt sound playing module is used to play a preset first warning prompt sound if the non-occluded current target material cannot be obtained within a preset time period.
[0099] Optionally, the image extraction module 701 further includes an occlusion determination unit, specifically used for:
[0100] If, through the target detection, a target object different from the current target material is detected on the current frame image, determine the overlapping area of the detection frames of the current target material and the target object;
[0101] If the overlapping area of the detection frames is higher than a preset threshold, determine that the current target material is occluded.
[0102] The first warning prompt sound playing module is used to play a preset first warning prompt sound if the non-occluded current target material cannot be obtained within a preset time period.
[0103] Optionally, the target material image determination unit is specifically configured to:
[0104] In response to the current target material having no occlusion and the exposure parameter of the current frame image meeting a preset standard, use the current frame image as the target material image.
[0105] Optionally, the device further includes:
[0106] A second warning prompt sound playing module, configured to play a preset second warning prompt sound if the exposure parameter of the current frame image does not meet the preset standard.
[0107] Optionally, the device further includes an angle adjustment module, specifically configured to:
[0108] Before the document classification module 702 classifies the target material image, extract any entity from the target material image and determine the inclination angle of the any entity;
[0109] Adjust the angle of the target material image according to the inclination angle of the any entity.
[0110] Optionally, the document classification module 702 is specifically configured to:
[0111] Classify the target material image according to the image features of the target material image to determine the target type of the target material.
[0112] Optionally, the device further includes a cropping module, specifically configured to:
[0113] Before the text recognition module 703 selects a target text recognition model corresponding to the target type to perform text recognition on the target material image, crop the target material image according to the background area in the target material image.
[0114] Optionally, the device further includes a duplicate judgment module, specifically configured to:
[0115] After the text recognition module 703 obtains the structured target data of the target material, judge whether the target material image is the same as the image of the already recognized material according to the structured target data of the target material;
[0116] If the judgment is the same, play a preset third warning prompt sound;
[0117] If the judgment is different, play a preset recognition success prompt sound.
[0118] Optionally, the device further includes:
[0119] The first display module is configured to display the structured target data of the target material.
[0120] Optionally, the device further includes:
[0121] The second display module is configured to display the list of materials to be recognized in the current task, and distinguish and display the recognized and unrecognized materials in the material list.
[0122] Optionally, the device further includes:
[0123] The third display module is configured to display the review result of the structured target data.
[0124] The above product can execute the method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0125] Figure 8 It is a schematic structural diagram of a financial interaction device according to an embodiment of the present disclosure. The financial interaction device in the embodiment of the present disclosure can be, for example, a device on financial machines such as a high-speed document scanner, and can scan and recognize materials when a user handles business. As shown in the figure, the financial interaction device 800 includes:
[0126] The shooting device 801 is configured to shoot a video stream of the material to be recognized;
[0127] The video stream processing device 802 is configured to process the video stream and extract the target material image from the video stream;
[0128] The material recognition device 803 is configured to perform document classification on the target material image, determine the target type of the target material, select a target character recognition model corresponding to the target type to perform character recognition on the target material image, and obtain the structured target data of the target material.
[0129] In addition, the financial interaction device may further include a communication device (not shown in the figure) for communicating with a cloud server and sending information such as the material image and the type of the material to the cloud server. The financial interaction device may further include a display device for displaying necessary information such as a list of materials required for handling business.
[0130] The financial interaction device in this embodiment can execute the material recognition method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method. The specific execution process of the method has been described in the above embodiments and will not be repeated here.
[0131] It should be noted that in the process of using existing financial machines to identify materials, due to the diverse types and different sizes of materials, the operator needs to place them on bases with certain fixed positions and sizes one by one for shooting. In addition, factors such as lighting during shooting also affect the imaging clarity. The entire process requires the operator to manually adjust and repeatedly confirm in the video screen before pressing the shooting button. Since the types and quantities of materials are relatively large, this process is time-consuming and prone to repetition, and it also depends on the operator's experience. After scanning, the operator needs to manually select the recognition model according to the type of the material. Since the types and quantities of materials are relatively large, this step needs to be performed for each material, and there is no room for compressing the duration to improve efficiency. In addition, there is no appropriate interaction during the scanning process, and the operator cannot know whether the operation result is correct or correct the operation based on the feedback prompt.
[0132] The important improvement of the financial interaction device in the embodiments of the present disclosure compared with existing market products is that it introduces processing strategies including intelligent processing of video streams and intelligent classification of pictures based on eigenvalue. Based on this processing strategy, it can effectively reduce the operation time of material shooting, improve the quality of material pictures, and automatically identify and determine the picture type in different business processes, making the financial business more efficient. In addition, based on this intelligent processing strategy, the picture is continuous and does not freeze during the material scanning and input process, the accuracy is high, and the non-stop incoming interaction process is smoother and more convenient, which also improves the experience of the operator using the financial machine.
[0133] In addition, the interaction method of the financial interaction device in the embodiments of the present disclosure has a low threshold, which is convenient for early judgment of problems existing in the materials. It can be applied to the self-service pre-review of the materials prepared by customers coming to the network before handling business, avoiding the discovery of primary problems during the business handling process, and can greatly reduce the unnecessary waiting time and negative emotions of customers.
[0134] Therefore, the existing processing algorithms and technologies are integrated on the financial interaction device in the embodiments of the present disclosure, and are flexibly combined and applied to the machine, so that the business handling process and interaction experience based on the machine have been greatly improved. As a result, the existing business processes in the network can be optimized, and the original process of customers submitting materials, business personnel receiving materials, back-office departments reviewing materials, and reprocessing after problems are found, which takes several days to 1 week, can be simplified into a "one-stop processing" and "only need to make one trip" process where customers pre-review and submit materials in the network, business personnel receive and review materials, and problems are clearly defined on the spot, reducing the probability of reprocessing. The back-office department can also only review the materials, reducing unnecessary workload. The technical solution of the embodiments of the present disclosure is a revolutionary improvement to the current interaction method of such financial interaction devices, and will greatly improve the experience in terms of both effect and efficiency for both business personnel and customers coming to handle business.
[0135] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0136] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0137] Figure 9 A schematic block diagram of an exemplary electronic device 900 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0138] As Figure 9 shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0139] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0140] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the material identification method. For example, in some embodiments, the material identification method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the material identification method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the material identification method in any other suitable manner (e.g., by means of firmware).
[0141] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0142] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0143] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0144] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0145] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0146] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server may also be a server of a distributed system or a server combined with blockchain.
[0147] Artificial intelligence is a discipline that studies to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.
[0148] Cloud computing refers to a technical system that accesses an elastic and scalable shared physical or virtual resource pool through a network. The resources can include servers, operating systems, networks, software, applications, and storage devices, etc., and the resources can be deployed and managed in a on-demand and self-service manner. Through cloud computing technology, it can provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.
[0149] It should be understood that various forms of processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recorded in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided in this disclosure can be achieved, and no limitations are imposed herein.
[0150] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for material identification, comprising: extracting a target material image from a currently captured video stream; performing document classification on the target material image to determine the target type of the target material; selecting a target text recognition model corresponding to the target type to perform text recognition on the target material image to obtain the structured target data of the target material; wherein, the extracting a target material image from a currently captured video stream includes: extracting frames from the currently captured video stream to obtain a current frame image; wherein, at least one stacked material is included in a single frame of the video stream; performing target detection on the current frame image to obtain the current target material on the current frame image; wherein, the current target material is the material located at the topmost layer; in response to the current target material not being occluded, taking the current frame image as the target material image; wherein, the process of determining whether the current target material is occluded includes: if, through the target detection, a target object different from the current target material is detected on the current frame image, determining the overlapping area of the detection frames of the current target material and the target object; if the overlapping area of the detection frames is higher than a preset threshold, determining that the current target material is occluded.
2. The method according to claim 1, wherein, the extracting a target material image from a currently captured video stream further includes: in response to the current target material being occluded, re - executing the operation of extracting frames from the currently captured video stream until the current target material is not occluded.
3. The method according to claim 1, further comprising: if the non - occluded current target material cannot be obtained within a preset time period, playing a preset first warning prompt sound.
4. The method according to claim 1, wherein, the taking the current frame image as the target material image in response to the current target material not being occluded includes: in response to the current target material not being occluded and the exposure parameter of the current frame image meeting a preset standard, taking the current frame image as the target material image.
5. The method according to claim 4, further comprising: if the exposure parameter of the current frame image does not meet the preset standard, playing a preset second warning prompt sound.
6. The method according to claim 1, wherein, before performing document classification on the target material image, the method further includes: extracting any entity from the target material image and judging the tilt angle of the any entity; adjusting the angle of the target material image according to the tilt angle of the any entity.
7. The method according to claim 1, wherein, the performing document classification on the target material image to determine the target type of the target material includes: performing document classification on the target material image according to the image features of the target material image to determine the target type of the target material.
8. The method according to claim 1, wherein, Before performing text recognition on the target material image using the target text recognition model corresponding to the target type, the method further includes: Cropping the target material image according to the background region in the target material image.
9. According to the method of claim 1, after obtaining the structured target data of the target material, the method further includes: Judging whether the target material image is the same as the image of the recognized material according to the structured target data of the target material; If the judgment is the same, play a preset third warning prompt sound; If the judgment is different, play a preset recognition success prompt sound.
10. According to the method of claim 1, the method further includes: Displaying the structured target data of the target material.
11. According to the method of claim 1, the method further includes: Displaying a list of materials to be recognized in the current task, and distinguishing and displaying the recognized and unrecognized materials in the material list.
12. According to the method of claim 1, the method further includes: Displaying the review result of the structured target data.
13. According to the method of claim 1, wherein, The target material is any one of the materials to be recognized, and the number of the materials to be recognized is two or more; In terms of the recognition sequence, compared with any other material in the materials to be recognized, the recognition sequence of the target material includes prior recognition or subsequent recognition.
14. A material recognition device, including: An image extraction module, configured to extract a target material image from a currently captured video stream; A document classification module, configured to perform document classification on the target material image to determine the target type of the target material; A text recognition module, configured to select a target text recognition model corresponding to the target type to perform text recognition on the target material image to obtain the structured target data of the target material; Optionally, the image extraction module includes: A frame extraction unit, configured to extract frames from the currently captured video stream to obtain a current frame image; wherein, at least one stacked material is included in a single frame of the video stream; A target detection unit, configured to perform target detection on the current frame image for the at least one material to be recognized to obtain the current target material on the current frame image; A target material image determination unit, configured to, in response to the current target material not being occluded, use the current frame image as the target material image; wherein, the current target material is the material located at the top layer; An occlusion judgment unit, configured to, if a target object different from the current target material is detected on the current frame image through the target detection, determine the overlapping area of the detection frames of the current target material and the target object; if the overlapping area of the detection frames is higher than a preset threshold, judge that the current target material is occluded.
15. An electronic device, including: A photographing device, configured to photograph a video stream of a material to be recognized; At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the material identification method according to any one of claims 1-13.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, the computer instructions are used to cause a computer to execute the material identification method according to any one of claims 1-13.
17. A computer program product, comprising a computer program which, when executed by a processor, implements the material identification method according to any one of claims 1-13.
18. A financial interaction device, comprising: a photographing device for photographing a video stream of a material to be identified; a video stream processing device for processing the video stream and extracting a target material image from the video stream; a material identification device for classifying the target material image into a document category, determining a target type of the target material, selecting a target character recognition model corresponding to the target type to perform character recognition on the target material image, and obtaining structured target data of the target material; wherein the video stream processing device is specifically configured to: extract a frame from the currently photographed video stream to obtain a current frame image; wherein at least one stacked material is included in a single frame of the video stream; perform target detection on the current frame image to obtain a current target material on the current frame image; wherein the current target material is the material located at the top layer; in response to the current target material not being occluded, use the current frame image as the target material image; wherein the process of determining whether the current target material is occluded includes: if, through the target detection, a target object different from the current target material is detected on the current frame image, determine an overlapping area of the detection frames of the current target material and the target object; if the overlapping area of the detection frames is higher than a preset threshold, determine that the current target material is occluded.
Citation Information
Patent Citations
Video frame extraction method and system based on deep learning
CN113792600A
Document generation method, device and platform, electronic equipment and storage medium
CN113971810A
Object recognition method and device and computer readable storage medium
CN114220045A