Bill identification identification method and device, medium and program product
By continuously acquiring ticket images with different exposure times using acquisition equipment, fusing and enhancing image processing, and combining boundary detection and gradient thresholding techniques, the problem of recognition efficiency and accuracy of ticket recognition systems in complex environments has been solved, achieving efficient and accurate ticket identification recognition.
Patent Information
- Application Number
- CN202511070887.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-14
AI Technical Summary
In existing technologies, ticket recognition systems suffer from low recognition efficiency and poor accuracy under complex lighting conditions and physical defects.
Multiple exposure images of the target ticket at different exposure times are continuously acquired by the acquisition device. The image pixel brightness is fused, and image enhancement and boundary detection are performed. The mask area is divided, and the target image is generated using gradient magnitude and dynamic threshold. The ticket identifier is determined by combining the pre-trained image recognition model.
It improves the accuracy and efficiency of document identification, enabling accurate identification of document information in complex environments.
Smart Images

Figure CN120954038A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology, and in particular to a method, device, medium, and program product for recognizing ticket identifiers. Background Technology
[0002] With the development of computer technology, the number of invoices in various fields has surged, making the identification and management of invoices a pressing issue. Invoices, whether invoices, checks, drafts, or various expense reports, all contain crucial business information such as amount, date, parties involved in the transaction, and invoice number. The accurate extraction and processing of this information directly affects the accuracy of financial accounting and the smoothness of business processes. Therefore, accurately identifying invoice numbers is a problem that urgently needs to be solved.
[0003] In existing technologies, typical ticket recognition systems rely on single-exposure imaging to identify ticket serial numbers. However, in real-world applications, the ambient lighting conditions surrounding the tickets are complex and variable, potentially including uneven lighting, excessive brightness, or excessive darkness. Furthermore, tickets themselves may have stains, wrinkles, or tears, which further affect image quality, causing blurry images, broken characters, or distortions, thus increasing the difficulty of recognition. Consequently, single-exposure imaging methods for ticket serial number recognition result in low recognition efficiency and poor accuracy. Summary of the Invention
[0004] This invention provides a method, device, medium, and program product for identifying ticket identifiers, in order to solve the problems of poor identification accuracy and low identification efficiency of ticket identifiers in the prior art.
[0005] According to one aspect of the present invention, a method for identifying a ticket identifier is provided, comprising:
[0006] Multiple exposure images of the target ticket at different exposure times are continuously acquired by the acquisition device. The multiple exposure images are fused according to the pixel brightness of multiple image pixels in the exposure images to obtain a target exposure image corresponding to the target ticket.
[0007] Image enhancement is performed on the target exposure image to obtain a target enhanced image. Target boundary information of the region of interest is determined based on the target enhanced image and the boundary detection model. Target mask region corresponding to the region of interest is determined based on the target boundary information.
[0008] The gradient magnitude of each pixel in the target mask region is obtained, the target mask region is divided into multiple pixel regions, the dynamic threshold of each pixel region is determined based on the gradient magnitude of multiple pixels in each pixel region, the target image is generated based on the multiple dynamic thresholds, and the ticket identifier of the target ticket is determined based on the target image by a pre-trained image recognition model.
[0009] According to one aspect of the present invention, a device for identifying ticket identifiers is provided, comprising:
[0010] The image fusion module is used to continuously acquire multiple exposure images of the target ticket at different exposure times through the acquisition device, and fuse the multiple exposure images according to the pixel brightness of multiple image pixels in the exposure images to obtain a target exposure image corresponding to the target ticket;
[0011] The target mask region determination module is used to perform image enhancement on the target exposure image to obtain a target enhanced image, determine the target boundary information of the region of interest based on the target enhanced image and the boundary detection model, and determine the target mask region corresponding to the region of interest based on the target boundary information.
[0012] The ticket identifier recognition module is used to obtain the gradient magnitude of each pixel in the target mask region, divide the target mask region into multiple pixel regions, determine the dynamic threshold of each pixel region based on the gradient magnitude of multiple pixels in each pixel region, generate a target image based on the multiple dynamic thresholds, and determine the ticket identifier of the target ticket based on the target image using a pre-trained image recognition model.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the ticket identification method according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the ticket identification method according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the ticket identification method according to any embodiment of the present invention.
[0019] The technical solution of this invention involves continuously acquiring multiple exposure images of a target document at different exposure times using an acquisition device. Multiple exposure images are fused based on the pixel brightness of multiple pixels in the exposure images to obtain a target exposure image corresponding to the target document. Image enhancement is performed on the target exposure image to obtain a target enhancement image. Target boundary information of the region of interest is determined based on the target enhancement image and a boundary detection model. A target mask region corresponding to the region of interest is determined based on the target boundary information. The gradient magnitude of each pixel in the target mask region is obtained. The target mask region is divided into multiple pixel regions. A dynamic threshold for each pixel region is determined based on the gradient magnitude of multiple pixels within each pixel region. A target image is generated based on the multiple dynamic thresholds. A pre-trained image recognition model is used to determine the document identifier of the target document based on the target image. This improves the accuracy and efficiency of document identifier recognition on documents.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a method for identifying a ticket identifier according to Embodiment 1 of the present invention;
[0023] Figure 2 This is a flowchart of another method for identifying ticket identifiers provided in Embodiment 2 of the present invention;
[0024] Figure 3 A flowchart illustrating the method for identifying ticket identification numbers provided in an embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of the structure of a ticket identification device according to Embodiment 3 of the present invention;
[0026] Figure 5This is a schematic diagram of the structure of an electronic device that implements the ticket identification method of this invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] Example 1
[0030] Figure 1 This is a flowchart illustrating a method for identifying a ticket identifier according to Embodiment 1 of the present invention. This embodiment is applicable to situations where ticket representations in tickets are identified. This method can be executed by a ticket identifier identification device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method may include:
[0031] S110. Continuously acquire multiple exposure images of the target document at different exposure times using an acquisition device, and fuse the multiple exposure images based on the pixel brightness of multiple image pixels in the exposure images to obtain a target exposure image corresponding to the target document.
[0032] The acquisition device can be a hardware device used to acquire images of the target document, which can be any document that needs to be identified. For example, the target document can be an invoice, receipt, check, bill, or banknote. The exposure time refers to the exposure time of the acquisition device when acquiring images of the target document. Exposure images at different exposure times can be multiple images of the same target document acquired continuously by the acquisition device at different exposure durations. These exposure images include a first exposure image acquired with a short exposure time, a second exposure image acquired with a medium exposure time, and a third exposure image acquired with a long exposure time. By continuously acquiring multiple images of the target document at different exposure durations using the acquisition device, images with varying exposure times can be provided for the identification of the target document, thereby improving the accuracy of the identification.
[0033] In this context, image pixels can be the pixels corresponding to the target image, and can be used to store the color or grayscale information corresponding to the image pixels in the target image. The pixel brightness of an image pixel can be its brightness value or brightness component, used to characterize the brightness of the scene area corresponding to that pixel. The target exposure image can be an image generated by fusing multiple images with different exposures. Generating a target exposure image by fusing multiple images with different exposures can improve image recognition efficiency and accuracy.
[0034] Optionally, multiple exposure images of the target document are continuously acquired by the acquisition device at preset intervals. The multiple exposure images are fused based on the pixel brightness of each image pixel in the exposure images to obtain a target exposure image corresponding to the target document. This may include: determining the mean and standard deviation of pixel brightness based on the pixel brightness of the corresponding multiple image pixels in the first, second, and third exposure images; determining the target weight of each image pixel based on the mean, standard deviation, and pixel brightness; and fusing the first, second, and third exposure images based on the target weight of each image pixel to obtain the target exposure image.
[0035] The mean pixel brightness can be the arithmetic mean of the pixel brightness values of each image pixel determined based on the pixel brightness of the image pixels in the first, second, and third exposure images. The standard deviation of pixel brightness can be the degree of dispersion of the brightness values of each image pixel in different exposure images, determined based on the pixel brightness of the image pixels in the first, second, and third exposure images, and is used to characterize the brightness fluctuation of the image pixels.
[0036] The target weight can be determined based on the mean pixel brightness, standard deviation of pixel brightness, and pixel brightness of the image pixels in the first exposure image, the second exposure image, and the third exposure image, and is the weight value of each image pixel in the image fusion.
[0037] Optionally, the target weight for each image pixel is determined based on the mean pixel brightness, standard deviation of pixel brightness, and pixel brightness using the following formula:
[0038]
[0039] Where W(x,y) is the target weight of the image pixel with x as the horizontal coordinate and y as the vertical coordinate in the exposed image, I(x,y) is the pixel brightness of the image pixel with x as the horizontal coordinate and y as the vertical coordinate in the exposed image, μ is the mean pixel brightness, and σ is the standard deviation of pixel brightness.
[0040] Specifically, multiple exposure images of the target document at different exposure times are continuously acquired by an acquisition device. The brightness of multiple images in the exposure images is fused to obtain a target exposure image corresponding to the target document. This can be achieved by continuously acquiring multiple frames of images for the same target document at preset short, medium, and long exposure times using an acquisition device with adjustable exposure parameters. An image registration algorithm is used to ensure strict alignment of the multiple frames, and the brightness value of each pixel in each exposure image is extracted. Based on the brightness distribution analysis, a target exposure image with a significantly improved dynamic range is generated. This preserves the details of the highlight areas of the image and clearly presents the dark areas, thereby providing high-precision image data for subsequent image recognition and improving the recognition accuracy and efficiency.
[0041] S120. Perform image enhancement on the target exposure image to obtain the target enhanced image. Determine the target boundary information of the region of interest based on the target enhanced image and the boundary detection model. Determine the target mask region corresponding to the region of interest based on the target boundary information.
[0042] The target-enhanced image can be obtained by adjusting or enhancing at least one of the brightness, contrast, sharpness, or color distribution of the target exposed image using a preset image enhancement algorithm. The preset image enhancement algorithm can be a pre-defined algorithm used to enhance the target exposed image. The boundary detection model can be a pre-built model based on deep learning or traditional image processing algorithms used to extract the edge contours of the target object from the target-enhanced image. By enhancing the target exposed image to obtain the target-enhanced image, the difficulty of image recognition can be reduced, and the recognition efficiency and accuracy can be improved.
[0043] The region of interest (ROI) can be a specific area in the enhanced image of the target document that requires further analysis or identification, such as the area corresponding to the document identifier on the target document. The target boundary information can be the set of edge pixel coordinates output by the boundary detection model, used to characterize the boundary of the ROI. The target mask region can be a binary mask image corresponding to the ROI.
[0044] Specifically, the exposed image of the target is enhanced to obtain an enhanced image. The target boundary information of the region of interest is determined based on the enhanced image and the boundary detection model. The target mask region corresponding to the region of interest is then determined based on the target boundary information. This can be achieved by enhancing the exposed image of the target using a preset image enhancement algorithm. The enhanced image is then input into a pre-trained boundary detection model to obtain the boundary coordinate information of the region of interest. Finally, the target mask region corresponding to the region of interest is determined based on the target boundary information, providing high-precision image data for subsequent image recognition and improving image recognition accuracy.
[0045] S130. Obtain the gradient magnitude of each pixel in the target mask region, divide the target mask region into multiple pixel regions, determine the dynamic threshold of each pixel region based on the gradient magnitude of multiple pixels in each pixel region, generate the target image based on the multiple dynamic thresholds, and determine the ticket identifier of the target ticket based on the target image through a pre-trained image recognition model.
[0046] The gradient magnitude is a comprehensive measure of the intensity of grayscale changes in each pixel within the target mask region in both horizontal and vertical directions. It characterizes the salience of the target mask region's edges; a larger gradient magnitude indicates a higher probability that the pixel is located at the edge of the image region. A pixel region can be multiple non-overlapping rectangular sub-regions obtained after dividing the target mask region, and each pixel region can contain multiple pixels from the target mask region. The dynamic threshold is a local threshold dynamically calculated based on the statistical value of the gradient magnitude within each pixel window, used to distinguish edge pixels from background pixels within the target mask region. The target image is the image obtained after binarizing and edge-enhancing the target mask region using the dynamic threshold. By determining the dynamic threshold for each pixel region based on the gradient magnitudes of multiple pixels within each pixel region, and generating the target image based on multiple dynamic thresholds, edge pixels in the target image can be significantly highlighted, background noise can be suppressed, and the recognition accuracy of document identifiers in the target document can be improved.
[0047] The image recognition model can be pre-trained based on deep learning or traditional feature matching algorithms, and is used to extract ticket identification information from the target image. The ticket identification is information that uniquely identifies the target ticket, determining its type, serial number, date, or amount. Examples include invoice numbers, receipt numbers, and banknote identification numbers (unique codes printed in specific locations on banknotes for tracking circulation information).
[0048] Specifically, the target mask region is divided into multiple pixel regions. A dynamic threshold for each pixel region is determined based on the gradient magnitude of multiple pixels within each region. A target image is generated based on these dynamic thresholds. A pre-trained image recognition model is then used to determine the ticket identifier of the target ticket based on the target image. This can be achieved by dividing the target mask region into multiple overlapping pixel windows, calculating an adaptive statistic of the pixel gradient magnitude within each window to generate a dynamic threshold, and then performing binarization or edge enhancement processing on the pixels within the window to construct the target image. This image is then input into a pre-trained image recognition model. The image recognition model, combined with prior knowledge of the ticket, parses out the ticket identifier, achieving high-precision ticket information extraction in different scenarios and improving the recognition accuracy and efficiency of the ticket identifier.
[0049] The technical solution of this invention involves continuously acquiring multiple exposure images of a target document at different exposure times using an acquisition device. Multiple exposure images are fused based on the pixel brightness of multiple pixels in the exposure images to obtain a target exposure image corresponding to the target document. Image enhancement is performed on the target exposure image to obtain a target enhancement image. Target boundary information of the region of interest is determined based on the target enhancement image and a boundary detection model. A target mask region corresponding to the region of interest is determined based on the target boundary information. The gradient magnitude of each pixel in the target mask region is obtained. The target mask region is divided into multiple pixel regions. A dynamic threshold for each pixel region is determined based on the gradient magnitude of multiple pixels within each pixel region. A target image is generated based on the multiple dynamic thresholds. A pre-trained image recognition model is used to determine the document identifier of the target document based on the target image. This improves the accuracy and efficiency of document identifier recognition on documents.
[0050] Example 2
[0051] Figure 2 This is a flowchart of another method for identifying ticket identifiers according to Embodiment 2 of the present invention. Based on the embodiments described above, this embodiment continuously acquires multiple exposure images of the target ticket at different exposure times using an acquisition device. The multiple exposure images are fused based on the pixel brightness of multiple image pixels in the exposure images to obtain a target exposure image corresponding to the target ticket. Image enhancement is performed on the target exposure image to obtain a target enhancement image. Target boundary information of the region of interest is determined based on the target enhancement image and a boundary detection model. A target mask region corresponding to the region of interest is determined based on the target boundary information. The gradient magnitude of each pixel in the target mask region is obtained. The target mask region is divided into multiple pixel regions. A dynamic threshold for each pixel region is determined based on the gradient magnitude of multiple pixels within each pixel region. A target image is generated based on the multiple dynamic thresholds. The method further refines the identification of the target ticket based on the target image using a pre-trained image recognition model. Figure 2 As shown, the method may include:
[0052] S210. Continuously acquire multiple exposure images of the target document at different exposure times using an acquisition device, and fuse the multiple exposure images based on the pixel brightness of multiple image pixels in the exposure images to obtain a target exposure image corresponding to the target document.
[0053] S220. Perform image enhancement on the target exposure image to obtain the target enhanced image. Determine the target boundary information of the region of interest based on the target enhanced image and the boundary detection model. Determine the target mask region corresponding to the region of interest based on the target boundary information.
[0054] Optionally, image enhancement of the target exposure image to obtain a target enhanced image may include: decomposing the target exposure image to obtain an image illumination component and an image reflection component, performing nonlinear stretching on the image reflection component to obtain a target reflection component, and determining the target enhanced image corresponding to the target document based on the target reflection component and the image illumination component.
[0055] The image illumination component can be low-frequency brightness information in the exposed image, determined by the light source intensity, the surface normal direction of the object, and the camera exposure parameters, used to characterize the light and dark distribution of the exposed image. The image reflection component can be high-frequency detail information in the exposed image, determined by the surface material properties of the target document, containing key feature information of the target document.
[0056] The target reflection component can be an enhanced component obtained by nonlinearly stretching the image reflection component. By nonlinearly stretching the image reflection component to obtain the target reflection component, the dynamic range of the reflection component's details can be expanded, improving the contrast between text and background. Nonlinear stretching can be a process of dynamically adjusting the grayscale value of the reflection component using a nonlinear function (such as logarithmic, power, or sigmoid functions) to highlight document details and suppress noise.
[0057] Specifically, the target exposed image is decomposed to obtain the image illumination component and the image reflection component. The image reflection component is then nonlinearly stretched to obtain the target reflection component. Based on the target reflection component and the image illumination component, a target enhancement image corresponding to the target document is determined. This can be achieved by decomposing the target exposed image using a pre-defined decomposition algorithm to obtain the image illumination component and the image reflection component, and then nonlinearly stretching the image reflection component to obtain the target reflection component, which expands the dynamic range of detail and enhances the contrast between text and background. Further determining the target enhancement image corresponding to the target document based on the target reflection component and the image illumination component can highlight the details of the target document in the target exposed image and suppress noise, thereby reducing the difficulty of subsequent image processing and improving image recognition efficiency.
[0058] Optionally, the target boundary information of the region of interest is determined based on the target augmentation image and the boundary detection model, including: scaling the target augmentation image to a preset scaling ratio to obtain a target scaled image, and inputting the target scaled image into the boundary detection model to obtain the target boundary information of the region of interest.
[0059] The target scaled image can be an image obtained by scaling a target enhancement image according to a preset scaling ratio. The preset scaling ratio can be a pre-defined ratio used to adjust the image size of the target enhancement image.
[0060] Specifically, the target enhancement image is scaled to a preset scaling ratio to obtain a scaled target image. This scaled target image is then input into a boundary detection model to obtain the target boundary information of the region of interest. Alternatively, the target enhancement image can be scaled according to the preset scaling ratio to obtain the scaled target image, and then the boundary detection model uses this scaled target image to obtain the target boundary information of the region of interest. By scaling the target enhancement image to a preset scaling ratio, the computational efficiency of the boundary detection model and the accuracy of feature extraction can be balanced, thereby improving both image recognition accuracy and efficiency.
[0061] S230. Obtain the gradient magnitude of each pixel in the target mask region, divide the target mask region into multiple pixel regions, and determine the gradient mean and gradient variance of each pixel region based on the gradient magnitude of multiple pixels in each pixel region.
[0062] The gradient mean can be the arithmetic mean of the gradient magnitudes of all pixels within a single pixel region, used to characterize the edge strength of the pixel region. The gradient variance can be the degree of dispersion of the gradient magnitudes around the gradient mean within a pixel region, used to characterize the degree of pixel gradient variation within the pixel region.
[0063] Specifically, the mean and variance of the gradient for each pixel region are determined based on the gradient magnitudes of multiple pixels within that region. This can be achieved by first determining the mean gradient of the pixel region based on the gradient magnitudes of multiple pixels within that region, and then determining the variance of the gradient for that region based on both the mean and the gradient magnitudes of the multiple pixels. By determining the mean and variance of the gradient for each pixel region based on the gradient magnitudes of the multiple pixels within that region, a data foundation can be provided for subsequent image processing, thereby improving the efficiency and accuracy of image processing.
[0064] S240. Determine the global basic threshold based on the gray-level histogram corresponding to the target mask region using the maximum inter-class variance algorithm. Determine the dynamic threshold of each pixel region in the target mask region based on multiple gradient mean values, gradient variance values, global basic thresholds, and preset empirical coefficients. Generate the target image based on multiple dynamic thresholds. Determine the ticket identifier of the target ticket based on the target image using a pre-trained image recognition model.
[0065] The global baseline threshold can be calculated using the maximum inter-class variance (MOV) algorithm based on the gray-level histogram of the target mask region. This global baseline threshold can be used to distinguish between the target and the background within the target mask region. The MOV algorithm can be an adaptive thresholding algorithm based on the gray-level histogram, which determines the global baseline threshold by maximizing the inter-class variance between the foreground (e.g., the ticket identifier of a target ticket in the target mask region) and the background. The gray-level histogram can be a bar chart that statistically analyzes the pixel frequency of each gray level in the target mask region. The preset empirical coefficients can be weighting parameters used to adjust the dynamic threshold calculation; these preset empirical coefficients can include a first preset empirical coefficient and a second preset empirical coefficient.
[0066] Specifically, the global basic threshold is determined based on the gray-level histogram corresponding to the target mask region by using the maximum inter-class variance algorithm. This can be achieved by calculating the global basic threshold used to distinguish the target from the background in the target mask region based on the gray-level histogram corresponding to the target mask region using the maximum inter-class variance algorithm, thereby providing a data foundation for the identification of the ticket identifier in the target ticket.
[0067] Optionally, the dynamic threshold for each pixel region in the target mask region is determined based on multiple gradient mean values, gradient variance values, a global base threshold, and preset empirical coefficients using the following formula:
[0068] T local (x, y) = T base +α(μ G -βσ G )
[0069] Among them, T local (x, y) represents the dynamic threshold of the pixel region with x-coordinate and y-coordinate in the target mask region, T base The global baseline threshold is defined by α, the first preset empirical coefficient is defined by β, and the second preset empirical coefficient is defined by μ. G Let σ be the gradient mean. G This represents the gradient variance value.
[0070] Optionally, the invoice identifier of the target invoice is determined based on the target image using a pre-trained image recognition model, including: inputting the target image into the image recognition model to obtain a reference invoice identifier of the target invoice and a recognition confidence level corresponding to the reference invoice identifier; and determining the reference invoice identifier as the invoice identifier of the target invoice if the recognition confidence level is greater than or equal to a preset confidence threshold.
[0071] The reference ticket identifier can be a reference identifier for the target ticket obtained by the image recognition model after decoding the input target image. The recognition confidence score corresponding to the reference ticket identifier can be a quantified value of the reliability of the output reference ticket identifier by the image recognition model, used to characterize the matching probability between the reference ticket identifier and the actual ticket identifier. The preset confidence threshold can be a pre-set minimum probability threshold used to determine whether the reference ticket identifier output by the image recognition model is reliable.
[0072] Specifically, under the condition that the recognition confidence level is greater than or equal to the preset confidence threshold, the reference ticket identifier is determined as the ticket identifier of the target ticket. This can be achieved by directly determining the reference ticket identifier as the ticket identifier of the target ticket when the recognition confidence level of the reference ticket identifier output by the image recognition model is greater than or equal to the preset confidence threshold.
[0073] Optionally, determining the ticket identifier of the target ticket based on the target image using a pre-trained image recognition model may further include: inputting the target image into the image recognition model to obtain a reference ticket identifier of the target ticket and a recognition confidence level corresponding to the reference ticket identifier; if the recognition confidence level is less than a preset confidence threshold, determining an anomaly handling strategy based on the recognition confidence level, and determining the ticket identifier of the target ticket based on the reference ticket identifier and the recognition confidence level using the anomaly handling strategy.
[0074] The anomaly handling strategy can be pre-defined, used to process reference ticket identifiers with an identification confidence level lower than a pre-set confidence threshold, and further determine the ticket identifier of the target ticket. When the identification confidence level is lower than the pre-set confidence threshold, a corresponding anomaly handling strategy can be determined based on the specific value of the identification confidence level, and then the ticket identifier of the target ticket can be determined through the corresponding anomaly handling strategy. The anomaly handling strategy may include a re-identification strategy for re-identifying the target image, a manual identification strategy for outputting the target image and reference ticket identifier to the front end for manual identification, and other anomaly handling strategies.
[0075] Specifically, when the recognition confidence level is less than a preset confidence threshold, an anomaly handling strategy is determined based on the recognition confidence level. This strategy, along with the reference document identifier, is then used to determine the target document identifier. Specifically, if the recognition confidence level of the reference document identifier output by the image recognition model is less than the preset confidence threshold, the corresponding anomaly handling strategy is determined based on that level, and the target document identifier is then identified using that strategy. By setting anomaly handling strategies corresponding to different recognition confidence levels, the accuracy of target document recognition can be improved even when the recognition confidence level of the reference document identifier output by the image recognition model is less than the preset confidence threshold.
[0076] For example, in a specific application scenario, the ticket identification method provided by this embodiment of the invention can identify the ticket identification number of a target ticket in a complex scenario through the following steps, solving the problems of poor identification accuracy and low identification efficiency in the process of ticket identification number identification. Figure 3 This is a flowchart illustrating a method for identifying ticket identification numbers provided in an embodiment of the present invention. This method can be used to identify the ticket identification number of a target ticket in complex scenarios. Figure 3 As shown, the method includes multi-exposure HDR (High Dynamic Range) imaging and image alignment, Retinex (an image enhancement theory based on the human visual system, which posits that an image is composed of illumination and reflection components) - wavelet fusion enhancement, Cascade CNN (a cascaded convolutional neural network, a model structure composed of multiple convolutional neural networks connected in series) two-level localization, dynamic adaptive threshold segmentation, CRNN-Transformer OCR (CRNN is a deep learning model combining convolutional neural networks (CNN) and recurrent neural networks (RNN), which can be used for optical character recognition tasks; Transformer is a deep learning model based on a self-attention mechanism; OCR is optical character recognition) recognition, and system-level optimization. For example, the recognition of the ticket identification number of a target ticket in a complex scene can be achieved through the following specific steps:
[0077] Step 1: Multi-exposure HDR imaging and image alignment
[0078] (1) Multi-frame acquisition and exposure control: High-speed industrial cameras (such as 20-megapixel CMOS (complementary metal-oxide-semiconductor) sensors) can be used to continuously capture 3 different exposure images at microsecond intervals during the transmission of the target ticket: short exposure (1 / 1000 sec): preserve the details of the highlight area (such as reflective characters); medium exposure (1 / 500 sec): balance the overall brightness; long exposure (1 / 200 sec): enhance the texture of the dark area (such as dirty areas).
[0079] Synchronous triggering mechanism: The target ticket position is detected by photoelectric sensor to ensure that multiple frames of images are acquired at the same physical location, thus avoiding motion blur.
[0080] (2) HDR Composition and Alignment: Image Registration: SIFT (Scale-Invariant Feature Transform) feature matching and RANSAC (Random Sample Consensus) algorithm are used to eliminate small displacements caused by mechanical vibration; Weight Fusion: Weights are dynamically assigned based on pixel brightness.
[0081]
[0082] Where W(x, y) represents the target weight of the image pixel with x-coordinate and y-coordinate in the exposed image, I(x, y) represents the pixel brightness of the image pixel with x-coordinate and y-coordinate in the exposed image, μ is the mean pixel brightness, and σ is the standard deviation of pixel brightness. By reducing the pixel weights in bright and dark areas, overexposure or underexposure is avoided.
[0083] Tone mapping: The HDR image is compressed to 8-bit grayscale space using the Reinhard global operator (a classic algorithm for mapping high dynamic range (HDR) images to low dynamic range (LDR) display devices) while preserving detail and contrast.
[0084] Step 2: Retinex-Wavelet Fusion Enhancement
[0085] (1) Multiscale Retinex decomposition:
[0086] Perform Gaussian pyramid multi-scale decomposition (scale number = 3) on the HDR image to separate the illumination component L. i (x, y) and reflection component R i (x, y), optional, can be separated into light components L using the following formula. i (x, y) and reflection component R i (x, y):
[0087] log(R i (x, y))=log(Ii (x, y))-log(L i (x, y)*G i (x, y))
[0088] Among them, G i (x, y) is a Gaussian at the i-th scale, with a size of (2... i+1 )X(2 i+1 +1), I i (x, y) represents the pixel value of the image at position (x, y) at the i-th scale.
[0089] Illumination compensation: Nonlinear stretching of the reflection component (Gamma = 0.5) is applied to suppress noise in shadow areas.
[0090] (2) Wavelet domain denoising and reconstruction:
[0091] Wavelet decomposition: The Daubechies-8 wavelet basis is selected to perform a 3-level decomposition on the reflection component to obtain the low-frequency subband (LL) and the high-frequency subband (LH, HL, HH).
[0092] Threshold denoising: Applying an improved BayesShrink soft thresholding method to the high-frequency subbands.
[0093]
[0094] Where T is the threshold, σ is the noise variance, N is the total number of pixels, and high-frequency information of character edges is preserved;
[0095] Image reconstruction: The processed subband is inversely transformed to the spatial domain and fused with the illumination component to output the enhanced image (enhanced target image).
[0096] Step 3: Cascade CNN Two-Level Localization
[0097] (1) Coarse localization stage (improved by YOLOv5 (single-shot multi-frame detector (version 5))):
[0098] Input: Enhanced image scaled to 640×640 resolution;
[0099] Network Structure: Backbone: CSPDarknet53 (53 layers of cross-stage dark network, the backbone network of YOLOv5), with added SE attention module (squeeze-encouragement attention module) to improve the response of CIN region (ticket identification number region); Neck: PANet (Path Aggregation Network) + BiFPN (Bidirectional Feature Pyramid Network) for cross-scale feature fusion; Head: Output bounding box (x,y,w,h) and confidence score; Training Data: Contains 50,000 labeled CIN region images, covering scenes such as wrinkles, dirt, and tilt; Loss Function: CIoU Loss (Complete Intersectionover Union Loss) + Focal Loss (focal loss function) to optimize small object detection capability.
[0100] (2) Refined segmentation stage (U-Net improvement):
[0101] Input: Coarsely located CIN region (ROI (Region of Interest)), zoomed in to 512×512;
[0102] Network structure:
[0103] Encoder: ResNet-34 (34-layer residual network), extracting multi-level features;
[0104] Decoder: Introducing dilated convolution (dilation rate = 2) to expand the receptive field and capture long-range dependencies;
[0105] Skip connections: Add a CBAM attention module (convolutional block attention module) to suppress background interference;
[0106] Output: Binary mask M(x,y), retaining only the CIN character region.
[0107] Step 4: Dynamic Adaptive Threshold Segmentation
[0108] (1) Local gradient statistics:
[0109] Calculate the Sobel gradient magnitude G(x, y) for the masked region (Sobel operator is a discrete differential operator used to detect image edges; it highlights rapidly changing areas in the image by calculating approximate gradient values of pixels in the horizontal and vertical directions). Divide the region into local windows (16×16 pixels). Calculate the mean gradient μ within the window. G With variance σ G Dynamically adjust threshold sensitivity.
[0110] (2) Generation of multi-threshold matrices:
[0111] Basic Threshold: An improved Otsu algorithm (a global automatic thresholding segmentation method based on gray-level histograms, which determines the optimal threshold, i.e., the global basic threshold, by maximizing the inter-class variance (the separation between foreground and background)) is used to solve the global basic threshold T based on the bimodal distribution of the gray-level histogram. base .
[0112] Adaptive correction: The threshold is adjusted based on local gradient information using the following formula:
[0113] T local (x, y) = T base +α(μ G -βσ G )
[0114] Among them, T local (x, y) represents the dynamic threshold of the pixel region with x-coordinate and y-coordinate in the target mask region, T base The global baseline threshold is defined by α, the first preset empirical coefficient is defined by β, and the second preset empirical coefficient is defined by μ. G Let σ be the gradient mean. G This represents the gradient variance. α can be 0.2 and β can be 1.5, adapting to the grayscale fluctuations in the contaminated area.
[0115] Morphological optimization: Close the binarized image (3×3 rectangular kernel) to repair character breaks.
[0116] Step 5: CRNN-Transformer OCR Recognition
[0117] (1) Feature extraction (CRNN part):
[0118] Convolutional layers: 4 stacked Conv-BN-ReLU (convolution-batch normalization-ReLU (corrected linear unit) activation layers), kernel size 3×3, output 64-channel feature maps; BiLSTM (bidirectional long short-term memory network): bidirectional LSTM encodes temporal features, hidden layer dimension 128.
[0119] (2) Transformer (Transformer encoder) timing modeling:
[0120] Encoder structure: 4-layer Transformer Encoder (the core component of the Transformer encoder, used to convert input sequences (such as characters, word vectors, image features, etc.) into high-dimensional feature representations containing global context information), each layer containing 8-head self-attention mechanisms;
[0121] Positional encoding: Incorporating learnable positional embeddings to enhance character order awareness;
[0122] Decoding logic: Candidate sequences are generated through Beam Search (beam width = 5), and easily confused characters (such as "8" and "B") are corrected by combining language model (n-gram statistics).
[0123] (3) End-to-end training:
[0124] Loss function: Joint optimization of CTC Loss (connection-time classification loss function) and Cross-Entropy (cross-entropy loss function);
[0125] Data augmentation: Synthesize randomly corrupted, blurred, and rotated samples to improve generalization;
[0126] Inference acceleration: TensorRT quantization model (a high-performance deep learning inference framework for optimizing model deployment efficiency) is used, with a single frame recognition time of ≤20ms.
[0127] Step Six: System-level Optimization
[0128] Hardware collaboration: FPGA (Field Programmable Gate Array) accelerates HDR synthesis and Retinex computation, reducing latency by 40%; GPU parallelizes Cascade CNN (Cascaded Convolutional Neural Network) inference, supporting real-time processing of 30 frames per second;
[0129] Anomaly handling mechanism: If the OCR confidence level is <90%, trigger a multi-frame voting mechanism (e.g., take the majority result from 3 frames); for unrecognizable CINs, activate the ultraviolet sensor for auxiliary verification.
[0130] The technical solution of this invention involves continuously acquiring multiple exposure images of a target document at different exposure times using an acquisition device. Multiple exposure images are fused based on the pixel brightness of multiple pixels in the exposure images to obtain a target exposure image corresponding to the target document. Image enhancement is performed on the target exposure image to obtain a target enhancement image. Target boundary information of the region of interest is determined based on the target enhancement image and a boundary detection model. A target mask region corresponding to the region of interest is determined based on the target boundary information. The gradient magnitude of each pixel in the target mask region is obtained. The target mask region is divided into multiple pixel regions. A dynamic threshold for each pixel region is determined based on the gradient magnitude of multiple pixels within each pixel region. A target image is generated based on the multiple dynamic thresholds. A pre-trained image recognition model is used to determine the document identifier of the target document based on the target image. This improves the accuracy and efficiency of document identifier recognition on documents.
[0131] Example 3
[0132] Figure 4 This is a schematic diagram of the structure of a ticket identification device provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes: an image fusion module 410, a target mask region determination module 420, and a ticket identification module 430.
[0133] The image fusion module 410 continuously acquires multiple exposure images of the target document at different exposure times using an acquisition device. It fuses these multiple exposure images based on the pixel brightness of multiple pixels in the exposure images to obtain a target exposure image corresponding to the target document. The target mask region determination module 420 enhances the target exposure images to obtain an enhanced target image. It determines the target boundary information of the region of interest based on the enhanced target image and a boundary detection model, and determines the target mask region corresponding to the region of interest based on the target boundary information. The document identification module 430 acquires the gradient magnitude of each pixel in the target mask region, divides the target mask region into multiple pixel regions, determines the dynamic threshold of each pixel region based on the gradient magnitude of multiple pixels within each pixel region, generates a target image based on the multiple dynamic thresholds, and determines the document identification of the target document based on the target image using a pre-trained image recognition model. The exposure images include a first exposure image acquired with a short exposure time, a second exposure image acquired with a medium exposure time, and a third exposure image acquired with a long exposure time.
[0134] Furthermore, the image fusion module 410 is specifically used to: determine the mean and standard deviation of pixel brightness based on the pixel brightness of multiple image pixels corresponding to the first, second, and third exposure images; determine the target weight of each image pixel based on the mean, standard deviation, and pixel brightness; and fuse the first, second, and third exposure images based on the target weight of each image pixel to obtain the target exposure image.
[0135] Furthermore, the image fusion module 410 is specifically used to: determine the target weight of each image pixel based on the mean pixel brightness, the standard deviation of pixel brightness, and the pixel brightness using the following formula:
[0136]
[0137] Where W(x,y) is the target weight of the image pixel with x as the horizontal coordinate and y as the vertical coordinate in the exposed image, I(x,y) is the pixel brightness of the image pixel with x as the horizontal coordinate and y as the vertical coordinate in the exposed image, μ is the mean pixel brightness, and σ is the standard deviation of pixel brightness.
[0138] Furthermore, the ticket identification module 430 is specifically used for: determining the gradient mean and gradient variance of each pixel region based on the gradient magnitudes of multiple pixels within each pixel region; determining a global basic threshold based on the grayscale histogram corresponding to the target mask region using the maximum inter-class variance algorithm; and determining the dynamic threshold of each pixel region in the target mask region based on the multiple gradient mean, gradient variance, global basic threshold, and preset empirical coefficients.
[0139] Furthermore, the ticket identification module 430 is specifically used to: determine the dynamic threshold of each pixel region in the target mask region based on multiple gradient mean values, gradient variance values, global basic thresholds, and preset empirical coefficients using the following formula:
[0140] T local (x, y) = T base +α(μ G -βσ G )
[0141] Among them, T local (x, y) represents the dynamic threshold of the pixel region with x-coordinate and y-coordinate in the target mask region, T base The global baseline threshold is defined by α, the first preset empirical coefficient is defined by β, and the second preset empirical coefficient is defined by μ. G Let σ be the gradient mean. G This represents the gradient variance value.
[0142] Furthermore, the invoice identification module 430 is specifically used to: input the target image into the image recognition model to obtain the reference invoice identifier of the target invoice and the recognition confidence level corresponding to the reference invoice identifier; and determine the reference invoice identifier as the invoice identifier of the target invoice when the recognition confidence level is greater than or equal to a preset confidence threshold.
[0143] Furthermore, the invoice identification module 430 is specifically used for: inputting the target image into the image recognition model to obtain the reference invoice identifier of the target invoice and the corresponding recognition confidence level of the reference invoice identifier; under the condition that the recognition confidence level is less than a preset confidence threshold, determining an anomaly handling strategy based on the recognition confidence level, and determining the invoice identifier of the target invoice based on the reference invoice identifier and the recognition confidence level through the anomaly handling strategy.
[0144] Furthermore, the target mask region determination module 420 is specifically used to: decompose the target exposure image to obtain the image illumination component and the image reflection component, perform nonlinear stretching on the image reflection component to obtain the target reflection component, and determine the target enhancement image corresponding to the target document based on the target reflection component and the image illumination component.
[0145] Furthermore, the target mask region determination module 420 is specifically used to: scale the target enhancement image to a preset scaling ratio to obtain a target scaling image, and input the target scaling image into the boundary detection model to obtain the target boundary information of the region of interest.
[0146] The document identification device provided in this embodiment of the invention can execute the document identification method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0147] Example 4
[0148] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0149] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0150] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0151] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the ticket identification method.
[0152] In some embodiments, the ticket identification method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the ticket identification method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the ticket identification method by any other suitable means (e.g., by means of firmware).
[0153] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0154] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0155] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0156] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0157] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0158] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0159] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0160] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for identifying ticket identifiers, characterized in that, include: Multiple exposure images of the target ticket at different exposure times are continuously acquired by the acquisition device. The multiple exposure images are fused according to the pixel brightness of multiple image pixels in the exposure images to obtain a target exposure image corresponding to the target ticket. Image enhancement is performed on the target exposure image to obtain a target enhanced image. Target boundary information of the region of interest is determined based on the target enhanced image and the boundary detection model. Target mask region corresponding to the region of interest is determined based on the target boundary information. The gradient magnitude of each pixel in the target mask region is obtained, the target mask region is divided into multiple pixel regions, the dynamic threshold of each pixel region is determined based on the gradient magnitude of multiple pixels in each pixel region, the target image is generated based on the multiple dynamic thresholds, and the ticket identifier of the target ticket is determined based on the target image by a pre-trained image recognition model.
2. The method according to claim 1, characterized in that, The exposed images include a first exposed image acquired with a short exposure time, a second exposed image acquired with a medium exposure time, and a third exposed image acquired with a long exposure time; the process of continuously acquiring multiple exposed images of the target document at preset intervals using an acquisition device, and fusing the multiple exposed images based on the pixel brightness of each image pixel in the exposed images to obtain a target exposed image corresponding to the target document, includes: The mean pixel brightness and standard deviation pixel brightness are determined based on the pixel brightness of multiple image pixels corresponding to the first exposure image, the second exposure image, and the third exposure image. The target weight of each image pixel is determined based on the mean pixel brightness, the standard deviation pixel brightness, and the pixel brightness. The first exposure image, the second exposure image, and the third exposure image are fused according to the target weight of each image pixel to obtain the target exposure image.
3. The method according to claim 2, characterized in that, The determination of the target weight for each image pixel based on the mean pixel brightness, the standard deviation of pixel brightness, and pixel brightness is achieved through the following formula: Where W(x,y) is the target weight of the image pixel with x as the horizontal coordinate and y as the vertical coordinate in the exposed image, I(x,y) is the pixel brightness of the image pixel with x as the horizontal coordinate and y as the vertical coordinate in the exposed image, μ is the mean pixel brightness, and σ is the standard deviation of pixel brightness.
4. The method according to claim 1, characterized in that, Determining the dynamic threshold of each pixel region based on the gradient magnitude of multiple pixels within each pixel region includes: The gradient mean and gradient variance of each pixel region are determined based on the gradient magnitude of multiple pixels within each pixel region. The global base threshold is determined based on the gray-level histogram corresponding to the target mask region using the maximum inter-class variance algorithm. The dynamic threshold of each pixel region in the target mask region is determined based on multiple gradient mean values, gradient variance values, the global base threshold, and preset empirical coefficients.
5. The method according to claim 4, characterized in that, The dynamic threshold for each pixel region in the target mask region, determined based on multiple gradient mean values, gradient variance values, a global base threshold, and a preset empirical coefficient, is achieved through the following formula: T local (x,y)=T base +a(m G -bs G ) Among them, T local (x, y) represents the dynamic threshold of the pixel region with x-coordinate and y-coordinate in the target mask region, T base The global baseline threshold is defined by α, the first preset empirical coefficient is defined by β, and the second preset empirical coefficient is defined by μ. G Let σ be the gradient mean. G This represents the gradient variance value.
6. The method according to claim 1, characterized in that, The step of determining the ticket identifier of the target ticket based on the target image using a pre-trained image recognition model includes: The target image is input into the image recognition model to obtain the reference ticket identifier of the target ticket and the recognition confidence level corresponding to the reference ticket identifier; Under the condition that the identification confidence level is greater than or equal to the preset confidence threshold, the reference ticket identifier is determined as the ticket identifier of the target ticket.
7. The method according to claim 6, characterized in that, The step of determining the ticket identifier of the target ticket based on the target image using a pre-trained image recognition model further includes: The target image is input into the image recognition model to obtain the reference ticket identifier of the target ticket and the recognition confidence level corresponding to the reference ticket identifier; When the identification confidence level is less than a preset confidence threshold, an anomaly handling strategy is determined based on the identification confidence level, and the target ticket identifier is determined based on the reference ticket identifier and the identification confidence level using the anomaly handling strategy.
8. The method according to claim 1, characterized in that, The step of enhancing the target exposed image to obtain the target enhanced image includes: The target exposure image is decomposed to obtain the image illumination component and the image reflection component. The image reflection component is nonlinearly stretched to obtain the target reflection component. The target enhancement image corresponding to the target document is determined based on the target reflection component and the image illumination component.
9. The method according to claim 1, characterized in that, The step of determining the target boundary information of the region of interest based on the target enhancement image and the boundary detection model includes: The target enhancement image is scaled to a preset scaling ratio to obtain a target scaled image. The target scaled image is then input into a boundary detection model to obtain target boundary information of the region of interest.
10. A computer program product comprising a computer program / instructions, wherein, When the computer program / instruction is executed by the processor, it implements the ticket identification method according to any one of claims 1-9.
Citation Information
Cited By
Information identification method and information identification system of transformer substation protection device
CN122023407A