Information identification method, device and terminal equipment

By acquiring multiple frames of waybill images through terminal devices, selecting clear target waybill images, and parsing waybill fields, the problem of low efficiency and low accuracy in express waybill recognition is solved, achieving fast and accurate recognition of key waybill information, and improving last-mile delivery efficiency and user experience.

CN115393860BActive Publication Date: 2026-07-24HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
Filing Date
2022-08-30
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, identifying key information on express delivery waybills is inefficient, inaccurate, and results in a poor user experience, making it impossible to complete last-mile delivery quickly and accurately.

Method used

By acquiring multiple frames of waybill images, selecting the target waybill image and parsing the waybill fields, the terminal device is used to identify key information of the waybill in real time, including waybill detection, field recognition and anomaly detection, to ensure image clarity and accuracy.

Benefits of technology

It enables rapid and accurate identification of key information on express delivery waybills, improving last-mile delivery efficiency, reducing application costs, and enhancing anti-interference capabilities and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393860B_ABST
    Figure CN115393860B_ABST
Patent Text Reader

Abstract

The application provides an information identification method, device and terminal equipment. The method comprises: acquiring multiple express delivery sheet images for an express delivery sheet; selecting a target express delivery sheet image from the multiple express delivery sheet images; wherein the displacement between the center point of a barcode in the target express delivery sheet image and the center point of a barcode in a candidate express delivery sheet image is less than a preset displacement threshold, and the candidate express delivery sheet image is a previous express delivery sheet image of the target express delivery sheet image; parsing a target express delivery sheet field from the target express delivery sheet image; and determining express delivery sheet key information in the target express delivery sheet image based on the target express delivery sheet field. Through the technical solution of the application, the express delivery sheet key information in the express delivery sheet can be identified based on a clear target express delivery sheet image, and the identification efficiency, accuracy and speed of the express delivery sheet key information are high, and the user experience is good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technology, and in particular to an information identification method, apparatus and terminal equipment. Background Technology

[0002] With the rapid development of e-commerce, the logistics industry has also experienced rapid growth, with express delivery volume increasing exponentially. Faced with the demand for handling such a large volume of express deliveries, higher requirements have been placed on logistics delivery speed and efficiency, making the rapid and accurate completion of last-mile delivery a primary need.

[0003] To complete last-mile delivery quickly and accurately, it is necessary to be able to identify key information on the waybill (such as name, mobile phone number, address, etc.). However, in the relevant technologies, the key information on the waybill is manually checked by the user, which has problems such as low efficiency, low accuracy, and poor user experience, and cannot accurately obtain the key information on the waybill. Summary of the Invention

[0004] This application provides an information identification method applied to a terminal device, the method comprising:

[0005] Acquire multiple frames of images of the express delivery waybill;

[0006] A target label image is selected from the multiple label images; wherein the displacement between the center point of the barcode in the target label image and the center point of the barcode in the candidate label image is less than a preset displacement threshold, and the candidate label image is the label image preceding the target label image.

[0007] Parse the target shipping label fields from the target shipping label image;

[0008] Based on the target waybill field, determine the key information of the waybill in the target waybill image.

[0009] This application provides an information identification device applied to a terminal device, the device comprising:

[0010] The acquisition module is used to acquire multiple frames of images of the express delivery waybill.

[0011] The processing module is used to select a target label image from the multi-frame label images; wherein the displacement between the center point of the barcode in the target label image and the center point of the barcode in the candidate label image is less than a preset displacement threshold, and the candidate label image is the label image preceding the target label image.

[0012] The determination module is used to parse the target waybill field from the target waybill image and determine the key information of the waybill in the target waybill image based on the target waybill field.

[0013] This application provides a terminal device, including: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the information recognition method of the example above in this application.

[0014] As can be seen from the above technical solutions, in this embodiment, a target waybill image can be selected from multiple frames of waybill images. This target waybill image is clear, not blurry, enabling the identification of key information on the waybill based on the clear target image. This allows for rapid identification of key information, facilitating quick and accurate last-mile delivery. The identification efficiency, accuracy, and speed of key information are high, and the user experience is good. The identification process can be completed in real-time by the terminal device, resulting in low application cost, wide application range, minimal device resource consumption, and independence from network and platform requirements. This significantly improves express delivery efficiency and effectively identifies key information, increasing the accuracy of the identification results. The identification process is highly versatile, exhibiting strong resistance to user behavior and slow or inaccurate camera focusing. It also has minimal constraints on user behavior and low requirements for terminal device imaging. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings of the embodiments of this application.

[0016] Figure 1 This is a flowchart illustrating an information identification method according to one embodiment of this application;

[0017] Figure 2 This is a flowchart illustrating an information identification method according to one embodiment of this application;

[0018] Figure 3 This is a schematic diagram illustrating the determination of a target shipping label image in one embodiment of this application;

[0019] Figure 4 This is a schematic diagram of the waybill detection process in one embodiment of this application;

[0020] Figure 5 This is a schematic diagram of the label detection process of the target label image in one embodiment of this application;

[0021] Figure 6 This is a schematic diagram of the field identification process in one embodiment of this application;

[0022] Figure 7A and Figure 7B This is a schematic diagram of the network model in one embodiment of this application;

[0023] Figure 8 This is a schematic diagram of the structure of an information identification device according to one embodiment of this application;

[0024] Figure 9 This is a hardware structure diagram of a terminal device according to one embodiment of this application. Detailed Implementation

[0025] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “the,” and “the” as used in this application and claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations comprising one or more of the associated listed items.

[0026] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when," "when," or "in response to a determination."

[0027] This application proposes an information identification method, which can be applied to terminal devices. See [link to relevant documentation]. Figure 1 The diagram shown is a flowchart of the information recognition method, which may include:

[0028] Step 101: Obtain multiple frames of the waybill image for the express delivery waybill.

[0029] Step 102: Select the target label image from the multiple label images; wherein the displacement between the center point of the barcode in the target label image and the center point of the barcode in the candidate label image is less than a preset displacement threshold, and the candidate label image is the label image preceding the target label image.

[0030] For example, if the barcode in the current frame of the label image meets the barcode verification rules, and the barcode in the previous frame of the label image also meets the barcode verification rules, then the coordinates of the first center point can be determined based on the coordinates of the four vertices of the barcode in the current frame of the label image, and the coordinates of the second center point can be determined based on the coordinates of the four vertices of the barcode in the previous frame of the label image. If the displacement between the first center point coordinates and the second center point coordinates is less than a preset displacement threshold (which can be configured empirically), then the current frame of the label image can be selected as the target label image.

[0031] Step 103: Parse the target waybill fields from the target waybill image.

[0032] For example, the target waybill image can be subjected to waybill detection to obtain the four corner points and their order. Based on the four corner points, an initial sub-image of the waybill region can be extracted from the target waybill image. Based on the order of the four corner points, the initial sub-image of the waybill region is rotated and corrected to obtain the rotated and corrected target sub-image of the waybill region. The target waybill fields are then parsed from the target sub-image of the waybill region.

[0033] In one possible implementation, parsing the target waybill field from the target sub-image of the waybill area may include, but is not limited to: performing text detection on the target sub-image of the waybill area to obtain the coordinate position corresponding to the waybill field, and extracting the text region corresponding to the coordinate position from the target sub-image of the waybill area; then, performing text recognition on the text region to obtain the target waybill field.

[0034] Step 104: Determine the key information of the target waybill in the target waybill image based on the target waybill field.

[0035] For example, after parsing the target label field from the target label image, it is also possible to detect whether there are any anomalies in the target label field; if not, the operation of determining the key information of the label in the target label image based on the target label field is performed; if so, a new target label image can be selected from multiple label images, and the operation of parsing the target label field from the new target label image is returned.

[0036] In one possible implementation, detecting whether a target form field is abnormal may include, but is not limited to: inputting the target form field into a trained target network model to obtain the confidence score and ambiguity of each character in the target form field; wherein the target network model includes a target backbone subnetwork, a target recognition subnetwork, and a target classification subnetwork, the target backbone subnetwork is used to determine target features based on the target form field, the target recognition subnetwork is used to determine the confidence score of each character in the target form field based on the target features, and the target classification subnetwork is used to determine the ambiguity of the target form field based on the target features. Based on the confidence score and / or the ambiguity of each character in the target form field, it is determined whether the target form field is abnormal; wherein, if the confidence score of any character is less than a preset confidence threshold, the target form field is abnormal; if the ambiguity indicates that the target form field is ambiguous, the target form field is abnormal.

[0037] In one possible implementation, the initial network model may include an initial backbone subnetwork, an initial recognition subnetwork, and an initial classification subnetwork. The training process of the target network model may include, but is not limited to: inputting first sample data from a first training dataset into the initial backbone subnetwork to obtain first sample features, and inputting the first sample features into the initial recognition subnetwork to obtain the recognition result of the first sample data; optimizing the initial backbone subnetwork and the initial recognition subnetwork based on the recognition result and the label information of the first sample data to obtain a target backbone subnetwork and a target recognition subnetwork; inputting second sample data from a second training dataset into the target backbone subnetwork to obtain second sample features, and inputting the second sample features into the initial classification subnetwork to obtain the classification result of the second sample data; optimizing the initial classification subnetwork based on the classification result and the label information of the second sample data to obtain a target classification subnetwork; and generating a target network model based on the target backbone subnetwork, the target recognition subnetwork, and the target classification subnetwork, i.e., combining these subnetworks into a target network model.

[0038] As can be seen from the above technical solutions, in this embodiment, a target waybill image can be selected from multiple frames of waybill images. This target waybill image is clear, not blurry, enabling the identification of key information on the waybill based on the clear target image. This allows for rapid identification of key information, facilitating quick and accurate last-mile delivery. The identification efficiency, accuracy, and speed of key information recognition are high, providing a good user experience and accurate access to key information. The identification process can be completed in real-time by the terminal device, resulting in low application cost, wide application range, minimal equipment resource consumption, and independence from network and platform requirements. This significantly improves express delivery efficiency and effectively identifies key information on the waybill, enhancing the accuracy of the identification results.

[0039] When there are anomalies on the waybill, such as see-through defects, incompleteness, soiling, or wrinkles, or when there are anomalies such as slow or inaccurate camera focusing, after obtaining the target waybill fields, the presence of anomalies can be detected. If anomalies are found, the information recognition is re-triggered until no anomalies are found, thus obtaining an accurate recognition result. This avoids outputting target waybill fields under abnormal conditions. In other words, the recognition process has strong versatility (i.e., by detecting whether there are anomalies in the target waybill fields, accurate recognition results can be obtained in any situation), strong resistance to user behavior (i.e., no additional operations such as tidying up the waybill or removing soiling or wrinkles are required for accurate recognition results), strong resistance to slow or inaccurate camera focusing (accurate recognition results can still be obtained even with slow or inaccurate camera focusing), few constraints on user behavior, and low requirements for terminal device imaging (accurate recognition results can still be obtained even if the terminal device cannot capture high-quality images).

[0040] The information recognition method of this application embodiment will be described below in conjunction with specific application scenarios.

[0041] To identify key information on express delivery waybills, users can manually review the information, which is inefficient, inaccurate, and results in a poor user experience. Another possible approach is to use specialized equipment to scan the barcode on the waybill and retrieve the key information from a database (the database records the mapping between barcode information and key information). However, this method requires specialized equipment to scan the barcode, leading to high application costs and limited application scope. Furthermore, the specialized equipment needs to send the barcode information to a server, which then queries the database to obtain the key information and sends it back to the specialized equipment. This reliance on network and platform resources for key information identification limits network requirements and application scope. Due to the complexity of waybill formats (dozens of express delivery companies, hundreds of waybill styles) and significant differences in user behavior, scanning results in large variations in images, making it difficult to effectively identify key information in complex situations.

[0042] In response to the above findings, this application proposes an information recognition method. The recognition process for key information on waybills can be completed in real time by the terminal device. This method is low-cost, widely applicable, requires minimal terminal device resources, and is independent of networks and platforms, significantly improving express delivery efficiency. It effectively identifies key information on waybills and increases the accuracy of the recognition results. The recognition process is highly versatile, exhibits strong resistance to user behavior interference, and is highly resistant to slow or inaccurate camera focusing. It has minimal constraints on user behavior and low requirements for terminal device imaging.

[0043] The information recognition method in this application embodiment can be applied to terminal devices, which may include, but are not limited to, high-speed scanners, PDAs, mobile phones, etc., and there is no limitation on the type of terminal device.

[0044] The information identification method in this application embodiment is described in [reference]. Figure 2 As shown, the process may involve the image acquisition process of the shipping label, the inspection process of the shipping label, the field recognition process, and the field anomaly detection process. The following describes the image acquisition process of the shipping label, the inspection process of the shipping label, the field recognition process, and the field anomaly detection process.

[0045] First, the waybill image acquisition process. During the waybill image acquisition process, a clear waybill image (denoted as the target waybill image) can be obtained. For example, considering factors such as user behavior, the waybill image acquired by the terminal device may be a blurry image. Therefore, the stability and clarity of the current frame waybill image can be determined by factors such as barcode effect and barcode displacement. If it is, the current frame waybill image is used as the target waybill image; otherwise, the next frame waybill image is acquired and analyzed.

[0046] For example, the terminal device can scan the express delivery waybill in real time to obtain multiple frames of waybill images. Barcode detection and recognition are performed on each frame. If the barcode detection results of two consecutive frames are the same (meeting the barcode verification rules), and the center point displacement of these two frames is less than a preset displacement threshold, then the second frame of the two consecutive frames is determined to be a clear waybill image, and this image can be used as the target waybill image for subsequent processing.

[0047] In one possible implementation, multiple frames of waybill images can be acquired for the express delivery waybill, and a target waybill image can be selected from these frames. The displacement between the center point of the barcode in the target waybill image and the center point of the barcode in the candidate waybill image is less than a preset displacement threshold, and the candidate waybill image is the previous frame of the target waybill image. For example, if the barcode in the current frame of the waybill image satisfies the barcode verification rules, and the barcode in the previous frame of the waybill image also satisfies the barcode verification rules, then the coordinates of the first center point can be determined based on the coordinates of the four vertices of the barcode in the current frame of the waybill image, and the coordinates of the second center point can be determined based on the coordinates of the four vertices of the barcode in the previous frame of the waybill image. If the displacement between the first center point coordinates and the second center point coordinates is less than a preset displacement threshold, then the current frame of the waybill image can be selected as the target waybill image.

[0048] See Figure 3 The diagram illustrates the process of determining a target surface image, which may include:

[0049] Step 301: Scan the current frame image of the express waybill using the terminal device.

[0050] Step 302: Determine whether the barcode in the current frame image meets the barcode verification rules.

[0051] For example, after obtaining the current frame waybill image, the current frame waybill image can be parsed to obtain the barcode in the current frame waybill image. The barcode is a graphic identifier that uses multiple black bars and spaces of different widths arranged according to certain encoding rules to express a set of information. This barcode can be understood as the unique identifier of the express waybill, that is, different express waybills correspond to different barcodes.

[0052] After obtaining the barcode in the current frame's single image, it can be determined whether the barcode meets the barcode verification rules. In this embodiment, the verification process for these barcode verification rules is not restricted.

[0053] For example, if the barcode in the current frame single image does not meet the barcode verification rules, then step 303 is executed; if the barcode in the current frame single image meets the barcode verification rules, then step 304 is executed.

[0054] Step 303: Scan the next frame of the waybill image for the current frame of the waybill using the terminal device, use the next frame of the waybill image as the current frame of the waybill image, and return to step 302.

[0055] Step 304: Determine whether the barcode in the previous frame of the shipping label image meets the barcode verification rules. If not, proceed to step 303; if yes, proceed to step 305.

[0056] For example, if the barcode in the current frame of the form image satisfies the barcode verification rules, it is also necessary to determine whether the barcode in the previous frame of the form image also satisfies the barcode verification rules. If yes, it means that the barcodes in two consecutive frames of form images satisfy the barcode verification rules, and the current frame of the form image may become the target form image. Further analysis based on subsequent steps is needed to determine whether the current frame of the form image should be used as the target form image. If no, it means that the current frame of the form image is a form image that satisfies the barcode verification rules, and the current frame of the form image can be used as a reference to continue scanning the next frame of the form image.

[0057] Step 305: If the barcode in the current frame of the order form meets the barcode verification rules, and the barcode in the previous frame of the order form also meets the barcode verification rules, then determine the coordinates of the first center point based on the coordinates of the four vertices of the barcode in the current frame of the order form, and determine the coordinates of the second center point based on the coordinates of the four vertices of the barcode in the previous frame of the order form.

[0058] For example, after obtaining the current frame label image, the current frame label image can be parsed to obtain the coordinates of the four vertices of the barcode in the current frame label image, namely the coordinates of the upper left corner, the upper right corner, the lower right corner, and the lower left corner. Based on the coordinates of the four vertices of the barcode, the coordinates of the center point corresponding to the barcode in the current frame label image can be determined, and denoted as the first center point coordinates.

[0059] Similarly, when the previous frame of the label image is obtained, the previous frame can be parsed to obtain the coordinates of the four vertices of the barcode in the previous frame, namely the coordinates of the top left corner, the top right corner, the bottom right corner, and the bottom left corner. Based on the coordinates of the four vertices of the barcode, the coordinates of the center point corresponding to the barcode in the previous frame can be determined, and denoted as the second center point coordinates.

[0060] Step 306: Determine the displacement between the barcode center point in the current frame image and the barcode center point in the previous frame image based on the coordinates of the first center point and the second center point.

[0061] For example, the first center point coordinates are the coordinates of the barcode center point in the current frame's label image, and the second center point coordinates are the coordinates of the barcode center point in the previous frame's label image. Clearly, based on the first and second center point coordinates, the displacement between these two center points can be determined. For instance, assuming the first center point coordinates are (M+1.x, M+1.y), the second center point coordinates are (Mx, My), and the displacement between these two center points is (Diff.x, Diff.y), then the displacement between these two center points can be determined using the following formulas: Diff.x = M+1.x – Mx; Diff.y = M+1.y – My.

[0062] Step 307: Determine whether the displacement between the two center points is less than the preset displacement threshold.

[0063] If no, proceed to step 303. If yes, proceed to step 308.

[0064] Step 308: Select the current frame label image as the target label image.

[0065] For example, if the displacement between the center point of the barcode in the current frame of the label image and the center point of the barcode in the previous frame of the label image is less than a preset displacement threshold, then the current frame of the label image is determined to be a clear label image, and the current frame of the label image is selected as the target label image, while the previous frame of the label image is selected as the candidate label image. Clearly, the barcodes in both the target and candidate label images satisfy the barcode verification rules, and the displacement between the center point of the barcode in the target label image and the center point of the barcode in the candidate label image is less than the preset displacement threshold. Therefore, the target label image is successfully selected.

[0066] If the displacement between the center point of the barcode in the current frame of the label image and the center point of the barcode in the previous frame of the label image is not less than the preset displacement threshold, the current frame of the label image will not be used as the target label image. Instead, the current frame of the label image can be used as a subsequent reference to continue scanning the next frame of the label image and return to step 302.

[0067] Second, the shipping label detection process. In the shipping label detection process, the target shipping label image can be detected, and a sub-image of the shipping label region can be extracted from the target shipping label image (that is, the target shipping label image is processed to obtain the sub-image of the shipping label region), and then the sub-image of the shipping label region is rotated and corrected.

[0068] See Figure 4The diagram shown illustrates the label inspection process, which may include:

[0069] Step 401: Perform waybill detection on the target waybill image to obtain the four corner points (such as the top left corner, top right corner, bottom right corner, and bottom left corner) and the order of the four corner points.

[0070] For example, after obtaining the target waybill image, waybill detection can be performed on it. The purpose of waybill detection is to find the four corner points and their order within the target waybill image. There are no restrictions on the waybill detection process, as long as the four corner points and their order can be found. For instance, the target waybill image can be input into a waybill detection model, which processes the image to obtain the four corner points and their order, and then outputs the order of the four corner points.

[0071] For example, the waybill detection model can be a deep learning model or a neural network model. There are no restrictions on the type of waybill detection model, as long as it can achieve the function of waybill detection, that is, outputting the four corner points and their order on the waybill. For instance, the waybill detection model can include, but is not limited to, deep learning models based on East, YOLO, DB, etc.

[0072] For example, for the four corner points corresponding to the express delivery waybill, these four corner points can be denoted as corner point A1, corner point A2, corner point A3 and corner point A4. After performing waybill detection on the target waybill image, the coordinates (x1, y1) and the order of corner point A1 can be obtained (as shown in ①), the coordinates (x2, y2) and the order of corner point A2 can be obtained (as shown in ③), the coordinates (x3, y3) and the order of corner point A3 can be obtained (as shown in ④), and the coordinates (x4, y4) and the order of corner point A4 can be obtained (as shown in ②).

[0073] In this sequence, ① represents the top left corner of the express delivery slip, ② represents the top right corner of the express delivery slip, ③ represents the bottom right corner of the express delivery slip, and ④ represents the bottom left corner of the express delivery slip.

[0074] Step 402: Extract the initial sub-image of the shipping label region from the target shipping label image based on the four corner points.

[0075] For example, after obtaining the four corner points corresponding to the express delivery waybill, that is, after obtaining the coordinates of corner point A1, corner point A2, corner point A3 and corner point A4, these four corner points can be found in the target waybill image, namely corner point A1, corner point A2, corner point A3 and corner point A4. Then, based on these four corner points, the target waybill image is processed to obtain the sub-image of the waybill area (that is, the sub-image where the express delivery waybill is located). For ease of distinction, this sub-image is called the initial sub-image of the waybill area.

[0076] Step 403: Based on the order of the four corner points, perform rotation correction on the initial sub-image of the single-area region (i.e., perform a rotation operation on the image) to obtain the rotated and corrected target sub-image of the single-area region.

[0077] For example, after extracting the initial sub-image of the shipping label area from the target shipping label image, the initial sub-image of the shipping label area corresponds to four corner points, and the order of these four corner points is known, that is, the order of corner points A1, A2, A3, and A4 is known. Based on this, the initial sub-image of the shipping label area can be rotated and corrected according to the order of the four corner points. For example, corner point A1 in order ① is rotated and corrected to the upper left corner point, corner point A4 in order ② is rotated and corrected to the upper right corner point, corner point A2 in order ③ is rotated and corrected to the lower right corner point, and corner point A3 in order ④ is rotated and corrected to the lower left corner point. After the initial sub-image of the shipping label area undergoes the above rotation correction, the rotated and corrected target sub-image of the shipping label area can be obtained.

[0078] See Figure 5 As shown, the waybill detection process of the target waybill image is illustrated. The left side shows the four corner points and their order corresponding to the waybill, while the right side shows the target sub-image of the waybill area.

[0079] Clearly, by extracting an initial sub-image of the shipping label area from the target shipping label image—that is, processing only the shipping label area—the influence of the external background on the shipping label can be removed. By performing rotation correction on the initial sub-image of the shipping label area, the image at any angle can be rotated to the correct angle, ensuring that the shipping label area is upright.

[0080] Third, the field recognition process. In the field recognition process, field detection and recognition can be performed on the target sub-image of the label area to obtain the target label fields in the label area target sub-image.

[0081] See Figure 6 The diagram shown illustrates the field recognition process, which may include:

[0082] Step 601: Perform text detection on the target sub-image of the face sheet region to obtain the coordinate position corresponding to the face sheet field. The coordinate position can be the coordinates of the four vertices of the smallest bounding rectangle of the face sheet field, such as the coordinates of the top left vertex, the top right vertex, the bottom right vertex, and the bottom left vertex.

[0083] For example, after obtaining the target sub-image of the shipping label area, text detection can be performed on the target sub-image of the shipping label area. The purpose of text detection is to find the coordinate positions corresponding to the fields on the shipping label. There are no restrictions on the text detection process, as long as the coordinate positions corresponding to the fields on the shipping label can be found. For instance, the target sub-image of the shipping label area can be input into a text detection model, which will process the target sub-image of the shipping label area, obtain the coordinate positions corresponding to the fields on the shipping label, and output the coordinate positions.

[0084] For example, the text detection model can be a deep learning model or a neural network model. There are no restrictions on the type of text detection model, as long as it can achieve the function of text detection, that is, the text detection model can output the coordinate positions corresponding to the fields on the waybill. For instance, the text detection model can include, but is not limited to, deep learning models based on East, YOLOv3, DB, etc.

[0085] For example, if there is name information on the express delivery waybill, when performing text detection on the target sub-image of the waybill area, the coordinate position B1 corresponding to the name field (i.e., the name field as a waybill field) can be obtained, and the coordinate position B1 can be the coordinates of the four vertices of the smallest bounding rectangle of the name field.

[0086] If the waybill contains a mobile phone number, then when performing text detection on the target sub-image of the waybill area, the coordinate position B2 corresponding to the mobile phone number field (i.e., the mobile phone number field is used as a waybill field) can be obtained, and the coordinate position B2 can be the coordinates of the four vertices of the smallest bounding rectangle of the mobile phone number field.

[0087] If the waybill contains address information (i.e., the recipient's address), then when performing text detection on the target sub-image of the waybill area, the coordinate position B3 corresponding to the address field (i.e., the address field as a waybill field) can be obtained, and the coordinate position B3 can be the coordinates of the four vertices of the smallest bounding rectangle of the address field.

[0088] If the express waybill contains three-segment code information, then when performing text detection on the target sub-image of the waybill area, the coordinate position B4 corresponding to the three-segment code field (i.e., the three-segment code field as a waybill field) can be obtained, and the coordinate position B4 can be the coordinates of the four vertices of the smallest bounding rectangle of the three-segment code field.

[0089] Step 602: Extract the text region corresponding to the coordinate position from the target sub-image of the face sheet area.

[0090] For example, after obtaining the coordinates of the label field, that is, after obtaining the coordinates of the four vertices of the minimum bounding rectangle of the label field, these four vertex coordinates can be found in the target sub-image of the label area. Then, based on these four vertex coordinates, the label area target sub-image is processed to obtain the text area corresponding to the coordinate position, that is, the text area corresponding to the label field.

[0091] For example, based on the coordinate position B1 corresponding to the name field, the text region C1 corresponding to coordinate position B1 can be extracted from the target sub-image of the shipping label area. The text region C1 includes the name information.

[0092] Based on the coordinate position B2 corresponding to the mobile phone number field, the text region C2 corresponding to the coordinate position B2 can be extracted from the target sub-image of the waybill area. The text region C2 includes the mobile phone number information.

[0093] Based on the coordinate position B3 corresponding to the address field, the text region C3 corresponding to the coordinate position B3 can be extracted from the target sub-image of the shipping label area. The text region C3 includes the address information.

[0094] Based on the coordinate position B4 corresponding to the three-segment code field, the text region C4 corresponding to the coordinate position B4 can be extracted from the target sub-image of the shipping label area. The text region C4 includes the three-segment code information.

[0095] Step 603: Perform text recognition on the text area to obtain the target form field.

[0096] For example, after obtaining the text region, text recognition can be performed on it. The purpose of text recognition is to find the target label field, that is, to identify the content in the text region. The content in the text region is the target label field. For example, the target label field can be a name, mobile phone number, address, three-segment code, etc. The text recognition process is not limited, as long as the content in the text region can be recognized. For instance, the text region can be input into a text recognition model, which processes the text region to obtain the target label field and outputs it.

[0097] For example, the text recognition model can be a deep learning model or a neural network model. There is no restriction on the type of text recognition model, as long as it can achieve the function of text recognition, that is, the text recognition model can output the target label field (i.e., the content in the text area). For example, the text recognition model can include, but is not limited to: deep learning models based on CTC, LSTM, Attention, etc.

[0098] For example, text region C1 can be input into the text recognition model, which will then recognize the content within it. This content is denoted as target waybill field D1, which can be a name field, such as Zhang San or Li Si. Similarly, text region C2 can be input into the model, and its content will be recognized. This content is denoted as target waybill field D2, which can be a phone number field, such as 13811111111. Text region C3 can be input into the model, and its content will be recognized. This content is denoted as target waybill field D3, which can be an address field, such as XX Street, XX County, Hangzhou City, Zhejiang Province. Finally, text region C4 can be input into the model, and its content will be recognized. This content is denoted as target waybill field D4, which can be a three-segment code field.

[0099] Fourth, the field anomaly detection process. During the field anomaly detection process, anomaly detection can be performed on the target waybill fields, i.e., to check if any of the target waybill fields are abnormal. If an anomaly is found, the user can be prompted to rescan, i.e., return to the waybill image acquisition process and re-acquire the target waybill image. If no anomaly is found, the key information of the waybill in the target waybill image can be determined based on the target waybill fields, and this key information can be output. For example, name information (e.g., Zhang San) can be determined based on target waybill field D1, mobile phone number information (e.g., 13811111111) based on target waybill field D2, address information (e.g., XX Street, XX County, Hangzhou City, Zhejiang Province) based on target waybill field D3, and three-segment code information based on target waybill field D4. The name information, mobile phone number information, address information, and three-segment code information can then form the key information of the waybill, which can then be output.

[0100] For example, although the target waybill image is a clear waybill image, the name field, mobile phone number field, address field, and three-segment code field may have abnormalities such as blurriness, incompleteness, wrinkles, or dirt. These abnormalities will affect the recognition results (i.e., the target waybill fields). Therefore, in order to solve the interference of these abnormalities on the recognition results, anomaly detection can also be performed on the target waybill fields.

[0101] To detect anomalies in the target waybill fields, the following situations may be included, but are not limited to:

[0102] Scenario 1: Identify the phone number information based on the phone number field (i.e., target waybill field D2), and determine if there are any anomalies in the target waybill field based on the phone number information. For example, check if the number of digits conforms to the rules. If the phone number has 11 digits, it conforms to the rules; if the phone number has more than 11 digits, it does not conform to the rules. If the number of digits does not conform to the rules, then the target waybill field is determined to be anomaly. Another example is checking if the first character of the phone number conforms to the rules. For example, if the first character of the phone number is 1, it conforms to the rules; if the first character of the phone number is not 1, it does not conform to the rules. If the first character of the phone number does not conform to the rules, then the target waybill field is determined to be anomaly.

[0103] Scenario 2: Determine the confidence level of each character in the target waybill field. If the confidence level of any character is less than the preset confidence threshold, then the target waybill field is determined to be abnormal. If the confidence level of all characters is not less than the preset confidence threshold, then the target waybill field is determined to be normal.

[0104] For example, assuming the target waybill field is "Zhang San", then the target waybill field includes the characters "Zhang" and "San". We can determine the confidence level for the character "Zhang". Assuming this confidence level is 80%, it means there's an 80% chance this character is "Zhang". Similarly, we can determine the confidence level for the character "San". Assuming this confidence level is 30%, it means there's a 30% chance this character is "San". After obtaining the confidence level for each character, we can determine if any character has a confidence level lower than a preset confidence threshold.

[0105] In one possible implementation, a target network model E1 can be trained. Target network model E1 includes a target backbone subnetwork and a target recognition subnetwork. The target backbone subnetwork is used to determine target features based on the target form field, and the target recognition subnetwork is used to determine the confidence level corresponding to each character in the target form field based on the target features. This embodiment does not limit the structure and function of the target network model E1, as long as the confidence level corresponding to each character in the target form field can be obtained. This embodiment also does not limit the training process of the target network model E1, as long as the target network model E1 can be trained successfully.

[0106] For example, after obtaining the target waybill field, the target waybill field can be input into the target backbone subnetwork of the target network model E1. The target backbone subnetwork extracts features from the target waybill field to obtain the target features corresponding to the target waybill field, and then inputs the target features into the target recognition subnetwork. After obtaining the target features, the target recognition subnetwork can determine the confidence level corresponding to each character in the target waybill field based on the target features. There are no restrictions on this determination process.

[0107] After obtaining the confidence level of each character in the target waybill field, it is possible to determine whether there is an anomaly in the target waybill field based on the confidence level of each character in the target waybill field.

[0108] Scenario 3: Determine the fuzziness level of the target waybill field, and determine whether the target waybill field is abnormal based on the fuzziness level. Specifically, if the fuzziness level of the target waybill field indicates that the field is fuzzy, then the target waybill field is determined to be abnormal. If the fuzziness level of the target waybill field indicates that the field is not fuzzy, then the target waybill field is determined to be normal.

[0109] For example, the fuzziness level of a target waybill field can be a fuzzy score, such as 500. The user can configure a score threshold, such as based on the application scenario or external environment; there are no restrictions on this. If the fuzzy score of the target waybill field is greater than this threshold, it indicates that the target waybill field is fuzzy, meaning there is an anomaly. If the fuzzy score of the target waybill field is not greater than this threshold, it indicates that the target waybill field is not fuzzy, meaning there is no anomaly.

[0110] In one possible implementation, a target network model E2 can be trained. Target network model E2 may include a target backbone subnetwork and a target classification subnetwork. The target backbone subnetwork is used to determine target features based on target label fields, and the target classification subnetwork is used to determine the ambiguity corresponding to the target label fields based on these target features. The structure and function of the target network model E2 are not limited in this embodiment, as long as the ambiguity corresponding to the target label fields can be obtained. The training process of the target network model E2 is not limited in this embodiment, as long as the target network model E2 can be trained successfully.

[0111] For example, after obtaining the target form field, it can be input into the target backbone subnetwork of the target network model E2. The target backbone subnetwork extracts features from the target form field to obtain the target features corresponding to that field, and then inputs these target features into the target classification subnetwork. After obtaining the target features, the target classification subnetwork can determine the ambiguity corresponding to the target form field based on these features. The process of determining this ambiguity is not restricted. After obtaining the ambiguity corresponding to the target form field, it can determine whether the target form field has any anomalies. For example, if the ambiguity indicates that the target form field is ambiguous, then the target form field has an anomaly. If the ambiguity indicates that the target form field is unambiguous, then the target form field does not have an anomaly.

[0112] Scenario 4: Determine the degree of wrinkling corresponding to the target shipping label field, and determine whether the target shipping label field has any anomalies based on the degree of wrinkling. Specifically, if the degree of wrinkling indicates that the target shipping label field has wrinkles, then the target shipping label field is determined to have an anomaly. If the degree of wrinkling indicates that the target shipping label field does not have wrinkles, then the target shipping label field is determined to have no anomalies.

[0113] In one possible implementation, a target network model E3 can be trained. This model determines the degree of wrinkling corresponding to a field on the target form. The structure and training process of this target network model E3 are not limited. After obtaining the target form field, it can be input into the target network model E3, which outputs the degree of wrinkling corresponding to that field. This degree of wrinkling can be a first value, indicating that wrinkles exist in the target form field; or it can be a second value, indicating that wrinkles do not exist in the target form field.

[0114] Scenario 5: Determine the degree of soiling corresponding to the target shipping label field, and based on the degree of soiling, determine whether the target shipping label field is abnormal. Specifically, if the degree of soiling indicates that the target shipping label field is soiled, then the target shipping label field is determined to be abnormal. If the degree of soiling indicates that the target shipping label field is not soiled, then the target shipping label field is determined to be normal.

[0115] In one possible implementation, a target network model E4 can be trained. This model determines the degree of contamination corresponding to a field on the target form. The structure and training process of this target network model E4 are not limited. After obtaining the target form field, it can be input into the target network model E4, which outputs the degree of contamination corresponding to that field. This degree of contamination can be a first value, indicating that the target form field is contaminated, or it can be a second value, indicating that the target form field is not contaminated.

[0116] Case 6: Combine at least two of Cases 1-5 to determine if an anomaly exists in the target waybill field. For example, when combining Case 1 and Case 2, if an anomaly is determined based on the mobile phone number information, and / or based on the confidence level corresponding to each character in the target waybill field, then the target waybill field is determined to be anomaly. If an anomaly is determined based on the mobile phone number information, and also based on the confidence level corresponding to each character in the target waybill field, then the target waybill field is determined to be anomaly. As another example, when combining Case 2 and Case 3, if an anomaly is determined based on the confidence level corresponding to each character in the target waybill field, and / or based on the fuzziness corresponding to the target waybill field, then the target waybill field is determined to be anomaly. If an anomaly is determined based on the confidence level corresponding to each character in the target waybill field, and also based on the fuzziness corresponding to the target waybill field, then the target waybill field is determined to be anomaly. For example, when combining scenarios 1-5, an anomaly is determined to exist in the target waybill field if at least one of the following conditions is met: an anomaly is determined based on the mobile phone number information; an anomaly is determined based on the confidence level of each character in the target waybill field; an anomaly is determined based on the ambiguity of the target waybill field; an anomaly is determined based on the degree of wrinkling of the target waybill field; or an anomaly is determined based on the degree of soiling of the target waybill field. Conversely, if none of the above conditions are met, the target waybill field is determined not to be anomaly.

[0117] In one possible implementation, when combining Case 2 and Case 3, the same target network model can be used to determine the confidence level corresponding to each character in the target form field and the ambiguity corresponding to the target form field. This target network model is denoted as target network model E5. The structure, function, and training process of target network model E5 are described below.

[0118] For example, the training process for the target network model E5 may include the following steps:

[0119] Step S11: Obtain the initial network model, see [link / reference] Figure 7A The diagram shows the structure of the initial network model. The initial network model can include an initial backbone subnetwork, an initial recognition subnetwork, and an initial classification subnetwork. The initial backbone subnetwork is used to extract features from the input data, the initial recognition subnetwork is used to determine the confidence level for each character, and the initial classification subnetwork is used to determine the ambiguity. The initial network model can be a deep learning model or a neural network model; there are no restrictions on its structure.

[0120] Step S12: Obtain the first training dataset and the second training dataset.

[0121] For example, the first training dataset is a field training dataset, which is used to train the initial backbone subnetwork and the initial recognition subnetwork. For ease of distinction, the sample data in the first training dataset is referred to as the first sample data. The first sample data can be a field image (i.e., a string image). The label information of the first sample data is used to represent the field values ​​in the field image, representing the actual content of the field.

[0122] The second training dataset is a character-fuzzy training dataset, used to train the initial classification sub-network. For ease of distinction, the sample data in the second training dataset is referred to as the second sample data. The second sample data can be field images. If the second sample data is a positive sample, it is a sharp field image, and the label information indicates that the field image is sharp, such as the first value. If the second sample data is a negative sample, it is a blurry field image, and the label information indicates that the field image is blurry, such as the second value.

[0123] Step S13: Train the initial backbone subnetwork and the initial recognition subnetwork based on the first sample data in the first training dataset to obtain the trained target backbone subnetwork and target recognition subnetwork.

[0124] For example, the first sample data can be input into the initial backbone subnetwork, which extracts features from the first sample data to obtain the first sample features, and then inputs these features into the initial recognition subnetwork. After obtaining the first sample features, the initial recognition subnetwork can process these features to obtain the recognition result of the first sample data.

[0125] After obtaining the recognition results of the first sample data, the initial backbone subnetwork and initial recognition subnetwork can be optimized based on the recognition results and label information of the first sample data to obtain the target backbone subnetwork and target recognition subnetwork. For example, the target loss value is determined based on the recognition results and label information of the first sample data, the network parameters of the initial backbone subnetwork are adjusted based on the target loss value to obtain the adjusted backbone subnetwork, and the network parameters of the initial recognition subnetwork are adjusted based on the target loss value to obtain the adjusted recognition subnetwork.

[0126] If the adjusted backbone subnetwork and the adjusted recognition subnetwork have converged, the adjusted backbone subnetwork can be used as the target backbone subnetwork, and the adjusted recognition subnetwork can be used as the target recognition subnetwork. Thus, the training process of the target backbone subnetwork and the target recognition subnetwork is completed.

[0127] If the adjusted backbone subnetwork and / or the adjusted recognition subnetwork do not converge, the adjusted backbone subnetwork can be used as the initial backbone subnetwork, and the adjusted recognition subnetwork can be used as the initial recognition subnetwork. The initial backbone subnetwork and the initial recognition subnetwork can continue to be trained based on the first sample data.

[0128] Step S14: Fix the target backbone subnetwork and the target recognition subnetwork, and train the initial classification subnetwork based on the second sample data in the second training dataset to obtain the trained target classification subnetwork.

[0129] For example, the second sample data can be input into the target backbone subnetwork, which extracts features from the second sample data to obtain the second sample features, and then inputs these second sample features into the initial classification subnetwork. After obtaining the second sample features, the initial classification subnetwork can process these features to obtain the classification result of the second sample data.

[0130] The initial classification subnetwork is optimized based on the classification results and label information of the second sample data to obtain the target classification subnetwork. For example, the target loss value is determined based on the classification results and label information of the second sample data. The network parameters of the initial classification subnetwork are adjusted based on the target loss value to obtain the adjusted classification subnetwork. If the adjusted classification subnetwork has converged, it is used as the target classification subnetwork, thus completing the training process of the target classification subnetwork. If the adjusted classification subnetwork has not converged, it is used as the initial classification subnetwork, and training continues based on the second sample data.

[0131] It is important to note that when adjusting the network parameters of the initial classification subnetwork, the target backbone subnetwork and the target recognition subnetwork must be fixed. That is, keep the network parameters of the target backbone subnetwork and the target recognition subnetwork unchanged, and only adjust the network parameters of the initial classification subnetwork.

[0132] Step S15: Generate target network model E5 based on the target backbone subnetwork, target recognition subnetwork, and target classification subnetwork. (See [link]) Figure 7B The diagram shown is a structural schematic of the target network model E5, which can include a target backbone subnetwork, a target recognition subnetwork, and a target classification subnetwork.

[0133] After obtaining the target network model E5, anomalies can be detected in the target form fields based on the target network model E5. The anomaly detection process for the target form fields can include the following steps:

[0134] Step S21: Input the target label field into the target backbone sub-network, and the target backbone sub-network determines the target features based on the target label field, that is, the target backbone sub-network extracts the target features of the target label field.

[0135] Step S22: Input the target features corresponding to the target label field into the target recognition sub-network, and the target recognition sub-network determines the confidence level of each character in the target label field based on the target features.

[0136] Step S23: Input the target features corresponding to the target form field into the target classification sub-network, and the target classification sub-network determines the ambiguity corresponding to the target form field based on the target features.

[0137] Step S24: After obtaining the confidence level and fuzziness level of each character in the target waybill field, it is possible to determine whether the target waybill field is abnormal based on the confidence level and / or fuzziness level. For example, if the confidence level of a character is less than a preset confidence threshold, it is determined that the target waybill field is abnormal; if the fuzziness level indicates that the target waybill field is ambiguous, it is determined that the target waybill field is abnormal.

[0138] In the above embodiments, the target classification subnetwork and the target recognition subnetwork can share the same target backbone subnetwork. The target classification subnetwork determines the ambiguity corresponding to the target form field, and the target recognition subnetwork determines the confidence level corresponding to each character in the target form field. This can accelerate the model training speed without adding extra time to the model training process. The target recognition subnetwork can be trained based on the first training dataset, and the target classification subnetwork can be trained based on the second training dataset. That is, the training datasets are independent, functionally independent, and uncoupled, with strong scalability, allowing for the addition of different samples according to different scenarios.

[0139] As can be seen from the above technical solutions, in this embodiment, a target waybill image can be selected from multiple frames of waybill images. This target waybill image is clear, not blurry, enabling the identification of key information on the waybill based on the clear target image. This allows for rapid identification of key information, facilitating quick and accurate last-mile delivery. The identification efficiency, accuracy, and speed of key information are high, and the user experience is good. The identification process can be completed in real-time by the terminal device, resulting in low application cost, wide application range, minimal device resource consumption, and independence from network and platform requirements. This significantly improves express delivery efficiency and effectively identifies key information, increasing the accuracy of the identification results. The identification process is highly versatile, exhibiting strong resistance to user behavior and slow or inaccurate camera focusing. It also has minimal constraints on user behavior and low requirements for terminal device imaging. It supports arbitrary arrangement of waybills, such as transparent or incomplete displays, and can still determine key information on the waybill even in these cases. Key information on the waybill can include name, address, mobile phone number, and three-segment QR code. After the final output, it can also determine whether the result (i.e., the target waybill field) is reliable, thus outputting only reliable target waybill fields and avoiding user errors caused by abnormal results.

[0140] Based on the same concept as the above method, this application proposes an information identification device for use in terminal devices. See [link to relevant documentation]. Figure 8 The diagram shown is a structural schematic of the device, which includes:

[0141] The acquisition module 81 is used to acquire multiple frames of waybill images for express delivery waybills;

[0142] Processing module 82 is used to select a target label image from the multi-frame label images; wherein the displacement between the center point of the barcode in the target label image and the center point of the barcode in the candidate label image is less than a preset displacement threshold, and the candidate label image is the label image preceding the target label image.

[0143] The determination module 83 is used to parse the target waybill field from the target waybill image and determine the key information of the waybill in the target waybill image based on the target waybill field.

[0144] For example, when the processing module 82 selects a target label image from the multiple label images, it specifically performs the following: if the barcode in the current label image meets the barcode verification rules, and the barcode in the previous label image also meets the barcode verification rules, then the first center point coordinates are determined based on the four vertex coordinates of the barcode in the current label image, and the second center point coordinates are determined based on the four vertex coordinates of the barcode in the previous label image; if the displacement between the first center point coordinates and the second center point coordinates is less than a preset displacement threshold, then the current label image is selected as the target label image.

[0145] For example, when the determining module 83 parses the target waybill field from the target waybill image, it is specifically used to: perform waybill detection on the target waybill image to obtain the four corner points and the order of the four corner points corresponding to the express waybill; extract the initial sub-image of the waybill area from the target waybill image based on the four corner points; perform rotation correction on the initial sub-image of the waybill area based on the order of the four corner points to obtain the target sub-image of the waybill area; and parse the target waybill field from the target sub-image of the waybill area.

[0146] For example, when the determining module 83 parses the target waybill field from the target sub-image of the waybill area, it is specifically used to: perform text detection on the target sub-image of the waybill area to obtain the coordinate position corresponding to the waybill field, extract the text area corresponding to the coordinate position from the target sub-image of the waybill area, and perform text recognition on the text area to obtain the target waybill field.

[0147] For example, the device further includes: a detection module, configured to detect whether there is an anomaly in the target waybill field; if not, the determination module determines the key information of the waybill in the target waybill image based on the target waybill field; if so, the processing module selects a new target waybill image from the multi-frame waybill images, and the determination module parses the target waybill field from the new target waybill image.

[0148] For example, when the detection module detects whether the target form field is abnormal, it specifically performs the following steps: inputting the target form field into a trained target network model to obtain the confidence score corresponding to each character in the target form field and the ambiguity corresponding to the target form field; wherein, the target network model includes a target backbone subnetwork, a target recognition subnetwork, and a target classification subnetwork, the target backbone subnetwork is used to determine target features based on the target form field, the target recognition subnetwork is used to determine the confidence score corresponding to each character in the target form field based on the target features, and the target classification subnetwork is used to determine the ambiguity corresponding to the target form field based on the target features; based on the confidence score corresponding to each character in the target form field and / or the ambiguity corresponding to the target form field, it is determined whether the target form field is abnormal; wherein, if the confidence score corresponding to any character is less than a preset confidence threshold, it is determined that the target form field is abnormal; if the ambiguity indicates that the target form field is ambiguous, it is determined that the target form field is abnormal.

[0149] For example, the initial network model includes an initial backbone subnetwork, an initial recognition subnetwork, and an initial classification subnetwork. The device further includes a training module for training a target network model. Specifically, when training the target network model, the training module performs the following steps: inputting first sample data from a first training dataset into the initial backbone subnetwork to obtain first sample features; inputting the first sample features into the initial recognition subnetwork to obtain the recognition result of the first sample data; optimizing the initial backbone subnetwork and the initial recognition subnetwork based on the recognition result and label information of the first sample data to obtain a target backbone subnetwork and a target recognition subnetwork; inputting second sample data from a second training dataset into the target backbone subnetwork to obtain second sample features; inputting the second sample features into the initial classification subnetwork to obtain the classification result of the second sample data; optimizing the initial classification subnetwork based on the classification result and label information of the second sample data to obtain a target classification subnetwork; and generating the target network model based on the target backbone subnetwork, the target recognition subnetwork, and the target classification subnetwork.

[0150] Based on the same concept as the above method, this application proposes a terminal device, see [link to application]. Figure 9As shown, the terminal device includes a processor 91 and a machine-readable storage medium 92, wherein the machine-readable storage medium 92 stores machine-executable instructions that can be executed by the processor 91; the processor 91 is used to execute the machine-executable instructions to implement the information identification method disclosed in the above example of this application.

[0151] Based on the same concept as the above method, this application also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the information recognition method disclosed in the above examples of this application.

[0152] The aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0153] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0154] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0155] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0156] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0157] Furthermore, these computer program instructions can also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0158] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0159] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. An information identification method, characterized in that, Applied to a terminal device, the method includes: Acquire multiple frames of images of the express delivery waybill; A target label image is selected from the multiple label images; wherein the displacement between the center point of the barcode in the target label image and the center point of the barcode in the candidate label image is less than a preset displacement threshold, and the candidate label image is the label image preceding the target label image; wherein the barcode in the target label image satisfies the barcode verification rules, and the barcode in the candidate label image satisfies the barcode verification rules, the barcode verification rules including that the barcode detection results of two consecutive label images are the same; Parse the target shipping label fields from the target shipping label image; Based on the target waybill field, determine the key information of the waybill in the target waybill image.

2. The method according to claim 1, characterized in that, The step of selecting the target label image from the multi-frame label images includes: If the barcode in the current frame of the order form satisfies the barcode verification rules, and the barcode in the previous frame of the order form satisfies the barcode verification rules, then the coordinates of the first center point are determined based on the coordinates of the four vertices of the barcode in the current frame of the order form, and the coordinates of the second center point are determined based on the coordinates of the four vertices of the barcode in the previous frame of the order form. If the displacement between the coordinates of the first center point and the coordinates of the second center point is less than a preset displacement threshold, then the current frame single image is selected as the target single image.

3. The method according to claim 1, characterized in that, The step of parsing the target waybill fields from the target waybill image includes: The target waybill image is subjected to waybill detection to obtain the four corner points corresponding to the waybill and the order of the four corner points; based on the four corner points, an initial sub-image of the waybill region is extracted from the target waybill image, and the initial sub-image of the waybill region is rotated and corrected based on the order of the four corner points to obtain the rotated and corrected target sub-image of the waybill region; The target label field is parsed from the target sub-image of the label area.

4. The method according to claim 3, characterized in that, The step of parsing the target waybill field from the target sub-image of the waybill area includes: Text detection is performed on the target sub-image of the shipping label area to obtain the coordinate positions corresponding to the shipping label fields, and the text area corresponding to the coordinate positions is extracted from the target sub-image of the shipping label area. The text region is subjected to text recognition to obtain the target form field.

5. The method according to any one of claims 1-4, characterized in that, After parsing the target waybill fields from the target waybill image, the method further includes: Detect whether there are any anomalies in the target form fields; If not, then perform the operation of determining the key information of the shipping label in the target shipping label image based on the target shipping label field; if yes, then select a new target shipping label image from the multi-frame shipping label images and return to perform the operation of parsing the target shipping label field from the new target shipping label image.

6. The method according to claim 5, characterized in that, The detection of whether there are any anomalies in the target form fields includes: The target form field is input into a trained target network model to obtain the confidence score and ambiguity of each character in the target form field. The target network model includes a target backbone subnetwork, a target recognition subnetwork, and a target classification subnetwork. The target backbone subnetwork is used to determine target features based on the target form field. The target recognition subnetwork is used to determine the confidence score of each character in the target form field based on the target features. The target classification subnetwork is used to determine the ambiguity of the target form field based on the target features. Based on the confidence level of each character in the target waybill field and / or the fuzziness level of the target waybill field, it is determined whether the target waybill field is abnormal; wherein, if the confidence level of any character is less than a preset confidence threshold, the target waybill field is abnormal; if the fuzziness level indicates that the target waybill field is fuzzy, the target waybill field is abnormal.

7. The method according to claim 6, characterized in that, The initial network model includes an initial backbone subnetwork, an initial recognition subnetwork, and an initial classification subnetwork. The training process of the target network model includes: The first sample data in the first training dataset is input into the initial backbone sub-network to obtain the first sample features. The first sample features are then input into the initial recognition sub-network to obtain the recognition result of the first sample data. Based on the recognition result and label information of the first sample data, the initial backbone sub-network and the initial recognition sub-network are optimized to obtain the target backbone sub-network and the target recognition sub-network. The second sample data from the second training dataset is input into the target backbone subnetwork to obtain the second sample features. The second sample features are then input into the initial classification subnetwork to obtain the classification result of the second sample data. Based on the classification result and label information of the second sample data, the initial classification subnetwork is optimized to obtain the target classification subnetwork. The target network model is generated based on the target backbone subnetwork, the target recognition subnetwork, and the target classification subnetwork.

8. An information identification device, characterized in that, Applied to a terminal device, the device includes: The acquisition module is used to acquire multiple frames of images of the express delivery waybill. The processing module is used to select a target label image from the multiple label images; wherein the displacement between the center point of the barcode in the target label image and the center point of the barcode in the candidate label image is less than a preset displacement threshold, and the candidate label image is the label image preceding the target label image; wherein the barcode in the target label image satisfies the barcode verification rules, and the barcode in the candidate label image satisfies the barcode verification rules, the barcode verification rules including that the barcode detection results of two consecutive label images are the same; The determination module is used to parse the target waybill field from the target waybill image and determine the key information of the waybill in the target waybill image based on the target waybill field.

9. The apparatus according to claim 8, Its features are, in, When the processing module selects a target label image from the multi-frame label images, it specifically performs the following steps: If the barcode in the current frame label image meets the barcode verification rules, and the barcode in the previous frame label image also meets the barcode verification rules, then the first center point coordinates are determined based on the four vertex coordinates of the barcode in the current frame label image, and the second center point coordinates are determined based on the four vertex coordinates of the barcode in the previous frame label image; if the displacement between the first center point coordinates and the second center point coordinates is less than a preset displacement threshold, then the current frame label image is selected as the target label image. Specifically, when the determining module parses the target waybill field from the target waybill image, it performs the following steps: performs waybill detection on the target waybill image to obtain the four corner points and their order corresponding to the express waybill; extracts an initial sub-image of the waybill region from the target waybill image based on the four corner points; performs rotation correction on the initial sub-image of the waybill region based on the order of the four corner points to obtain a target sub-image of the waybill region; and parses the target waybill field from the target sub-image of the waybill region. Specifically, when the determining module parses the target waybill field from the target sub-image of the waybill area, it is used to: perform text detection on the target sub-image of the waybill area to obtain the coordinate position corresponding to the waybill field, extract the text area corresponding to the coordinate position from the target sub-image of the waybill area, and perform text recognition on the text area to obtain the target waybill field; The device further includes: a detection module for detecting whether there is an anomaly in the target waybill field; if not, the determination module determines the key information of the waybill in the target waybill image based on the target waybill field; if so, the processing module selects a new target waybill image from the multi-frame waybill images, and the determination module parses the target waybill field from the new target waybill image. Specifically, when the detection module detects whether the target form field is abnormal, it is used to: input the target form field into a trained target network model to obtain the confidence score and ambiguity of each character in the target form field; wherein the target network model includes a target backbone subnetwork, a target recognition subnetwork, and a target classification subnetwork, the target backbone subnetwork is used to determine target features based on the target form field, the target recognition subnetwork is used to determine the confidence score of each character in the target form field based on the target features, and the target classification subnetwork is used to determine the ambiguity of the target form field based on the target features; based on the confidence score and / or the ambiguity of each character in the target form field, it is determined whether the target form field is abnormal; wherein, if the confidence score of any character is less than a preset confidence threshold, it is determined that the target form field is abnormal; if the ambiguity indicates that the target form field is ambiguous, it is determined that the target form field is abnormal. The initial network model includes an initial backbone subnetwork, an initial recognition subnetwork, and an initial classification subnetwork. The device further includes a training module for training a target network model. Specifically, the training module trains the target network model by: inputting first sample data from a first training dataset into the initial backbone subnetwork to obtain first sample features; inputting the first sample features into the initial recognition subnetwork to obtain the recognition result of the first sample data; optimizing the initial backbone subnetwork and the initial recognition subnetwork based on the recognition result and label information of the first sample data to obtain a target backbone subnetwork and a target recognition subnetwork; inputting second sample data from a second training dataset into the target backbone subnetwork to obtain second sample features; inputting the second sample features into the initial classification subnetwork to obtain the classification result of the second sample data; optimizing the initial classification subnetwork based on the classification result and label information of the second sample data to obtain a target classification subnetwork; and generating the target network model based on the target backbone subnetwork, the target recognition subnetwork, and the target classification subnetwork.

10. A terminal device, characterized in that, include: A processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the steps of the method according to any one of claims 1-7.