Reading identification method and device, network training method and device, equipment and medium

By segmenting digital instrument images, performing channel attention and correlation feature mining, and combining convolutional and self-attention network models, the reliability and accuracy of digital instrument reading recognition are improved, solving the problem of poor recognition in existing technologies.

CN120708206AActive Publication Date: 2025-09-26CHENGDU AIRCRAFT INDUSTRY GROUP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510646531.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-09-26
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The reliability of digital instrument reading recognition in the existing technology is relatively poor, which affects the level of informatization and intelligence of industrial production.

Method used

The initial digital instrument image is segmented, and a dual feature mining method of channel attention and correlation feature mining is used to combine convolutional network and self-attention network models to extract and fuse image features for semantic recognition.

Benefits of technology

The reliability of digital instrument reading recognition is improved, the accuracy and real-time performance of semantic recognition are ensured, and the recognition deficiencies in the existing technology are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708206A_ABST
    Figure CN120708206A_ABST
Patent Text Reader

Abstract

The invention provides a reading recognition method and device, a network training method and device, equipment and a medium, and relates to the technical field of image digital recognition. In the application, firstly, an initial digital instrument image can be segmented to obtain a corresponding target digital instrument image; secondly, performing first feature mining processing on the target digital instrument image to obtain a corresponding intermediate image feature; secondly, performing second feature mining processing on the intermediate image features to obtain corresponding target image features; and finally, semantic recognition processing can be performed on the target image features to obtain a target semantic recognition result corresponding to the target digital instrument. On the basis of the content, the problem that in the prior art, the reading recognition reliability of the digital instrument is relatively poor can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image digital technology, and in particular to a reading recognition method and device, a network training method and device, equipment and medium. Background Art

[0002] Digital instruments are widely used in industrial production to display important parameters such as temperature, pressure, and flow. In the industrial production and manufacturing process, a large amount of real-time production data is required for the purposes of quality control, production monitoring, equipment status analysis, etc. Digital instruments are one of the most common data display instruments. Among them, the accuracy and real-time performance of the recognition of the numbers displayed by digital instruments will directly affect the level of informatization and intelligence of industrial production. The intelligent recognition of digital instruments will effectively solve the risk of misoperation in manual reading of instrument data, and can improve data accuracy and meet the real-time requirements of production monitoring. However, the inventors have found that in the prior art, there is a problem of relatively poor reliability in the recognition of digital instrument readings. Summary of the Invention

[0003] In view of this, the purpose of the present application is to provide a reading recognition method and device, a network training method and device, equipment and medium, so as to improve the problem of relatively poor reliability of digital instrument reading recognition in the prior art.

[0004] To achieve the above objectives, this application adopts the following technical solutions: A digital instrument reading recognition method, comprising: Segmenting an initial digital instrument image to obtain a corresponding target digital instrument image, wherein the initial digital instrument image is formed by performing an image acquisition operation on a target digital instrument, and the target digital instrument image at least includes a digital display area of ​​the target digital instrument; performing a first feature mining process on the target digital instrument image to obtain corresponding intermediate image features, wherein the intermediate image features are used to represent semantic information of digits in the target digital instrument image, and the first feature mining process at least includes a channel attention process for mining semantic information of significant channels in the target digital instrument image; performing a second feature mining process on the intermediate image features to obtain corresponding target image features, wherein the second feature mining process is different from the first feature mining process and the second feature mining process at least includes a correlation feature mining process for mining correlation semantic information between feature columns in the intermediate image features; Performing semantic recognition processing on the target image features to obtain a target semantic recognition result corresponding to the target digital instrument, wherein the target semantic recognition result at least includes a number displayed in a digital display area of ​​the target digital instrument.

[0005] In a preferred embodiment of the present application, in the above-mentioned digital meter reading recognition method, the step of performing first feature mining processing on the target digital meter image to obtain corresponding intermediate image features includes: Using a convolutional network model included in a target digital meter reading recognition network, performing feature extraction processing on the target digital meter image to obtain corresponding initial image features; The channel attention network model included in the target digital meter reading recognition network is used to perform channel attention processing on the initial image features to obtain corresponding intermediate image features.

[0006] In a preferred embodiment of the present application, in the above-mentioned digital meter reading recognition method, the step of performing channel attention processing on the initial image features using the channel attention network model included in the target digital meter reading recognition network to obtain corresponding intermediate image features includes: Using a first pooling unit and a second pooling unit in a channel attention network model included in the target digital meter reading recognition network, respectively performing pooling processing on the initial image features to obtain corresponding first pooled image features and second pooled image features, wherein the first pooling unit and the second pooling unit have different pooling processing modes for capturing different information in the initial image features; Using a multi-layer perceptron in a channel attention network model included in the target digital meter reading recognition network, mapping the first pooled image features and the second pooled image features respectively to obtain corresponding first mapped image features and second mapped image features; Based on the first mapped image feature and the second mapped image feature, an intermediate image feature corresponding to the initial image feature is determined.

[0007] In a preferred embodiment of the present application, in the above-mentioned digital meter reading recognition method, the step of determining the intermediate image feature corresponding to the initial image feature based on the first mapped image feature and the second mapped image feature includes: performing a first feature fusion process on the first mapping image feature and the second mapping image feature to obtain a corresponding fused mapping image feature, and performing an activation process on the fused mapping image feature to obtain a corresponding activated mapping image feature, wherein the activated mapping image feature carries semantic information representing the saliency of the image channel; A second feature fusion process is performed on the activation map image feature and the initial image feature to obtain an intermediate image feature corresponding to the initial image feature.

[0008] In a preferred embodiment of the present application, in the above-mentioned digital meter reading recognition method, the step of performing a second feature mining process on the intermediate image features to obtain corresponding target image features includes: Using a self-attention network model included in a target digital meter reading recognition network, mapping processing is performed on the intermediate image features to obtain corresponding query image features, key image features, and value image features, and based on the query image features, the key image features, and the value image features, determining the self-attention image features corresponding to the intermediate image features; The target digital meter reading recognition network includes a bidirectional feature capture network model, and the self-attention image feature is bidirectionally captured in the contextual semantic information to obtain a corresponding first contextual semantic feature and a second contextual semantic feature. Furthermore, the first contextual semantic feature and the second contextual semantic feature are subjected to feature fusion processing to obtain a target image feature corresponding to the self-attention image feature.

[0009] This application also provides a digital instrument reading recognition network training method, including: Obtaining an initial image segmentation network and an initial digital meter reading recognition network included in an initial neural network, and obtaining a first sample digital meter image, first image label information corresponding to the first sample digital meter image, a second sample digital meter image, and second image label information corresponding to the second sample digital meter image, wherein the first image label information is used to indicate a digital display area in the first sample digital meter image, and the second image label information is used to indicate a digital displayed in the digital display area in the second sample digital meter image; Using the initial image segmentation network, segmenting the first sample digital instrument image to obtain a corresponding target sample digital instrument image; and updating network parameters of the initial image segmentation network based on the target sample digital instrument image and the first image label information to obtain a corresponding target image segmentation network, wherein the target sample digital instrument image at least includes a digital display area of ​​the digital instrument; performing a first feature mining process on the second sample digital meter image using the initial digital meter reading recognition network to obtain corresponding sample intermediate image features, wherein the sample intermediate image features are used to represent semantic information of digits in the second sample digital meter image, and the first feature mining process at least includes a channel attention process for mining semantic information of significant channels in the second sample digital meter image; Using the initial digital meter reading recognition network, performing a second feature mining process on the sample intermediate image features to obtain corresponding sample target image features, wherein the second feature mining process is different from the first feature mining process and the second feature mining process at least includes a correlation feature mining process for mining correlation semantic information between feature columns in the sample intermediate image features; Using the initial digital meter reading recognition network, performing semantic recognition processing on the sample target image features to obtain a sample semantic recognition result corresponding to the digital meter, wherein the sample semantic recognition result at least includes the number displayed in the digital display area of ​​the digital meter; Based on the difference between the sample semantic recognition result and the second image label information, updating the network parameters of the initial digital meter reading recognition network to obtain a corresponding target digital meter reading recognition network; Based on the target image segmentation network and the target digital meter reading recognition network, a corresponding target neural network is determined, wherein the target neural network is used to implement each step in the digital meter reading recognition method described in any one of claims 1-5.

[0010] In a preferred embodiment of the present application, in the above-mentioned digital instrument reading recognition network training method, the steps of using the initial image segmentation network to segment the first sample digital instrument image to obtain a corresponding target sample digital instrument image, and updating the network parameters of the initial image segmentation network based on the target sample digital instrument image and the first image label information to obtain a corresponding target image segmentation network include: Determining a plurality of corresponding bounding boxes based on a plurality of first image label information of a plurality of first sample digital images, wherein each of the bounding boxes corresponds to a piece of the first image label information, and the bounding box is used to indicate a corresponding digital display area; performing clustering processing on the plurality of bounding boxes to obtain corresponding clustering results, and determining at least one target bounding box based on the clustering results; Using the initial image segmentation network, based on the at least one target bounding box, segment the first sample digital instrument image to obtain a corresponding target sample digital instrument image; Based on the target sample digital instrument image and the first image label information, the network parameters of the initial image segmentation network are updated to obtain a corresponding target image segmentation network.

[0011] In a preferred embodiment of the present application, in the above-mentioned digital meter reading recognition network training method, the step of using the initial digital meter reading recognition network to perform first feature mining processing on the second sample digital meter image to obtain corresponding sample intermediate image features includes: Using the convolutional network model included in the initial digital meter reading recognition network, perform feature extraction processing on the second sample digital meter image to obtain corresponding sample image features; Using a first pooling unit and a second pooling unit in a channel attention network model included in the initial digital meter reading recognition network, respectively performing pooling processing on the sample image features to obtain corresponding sample first pooled image features and sample second pooled image features, wherein the first pooling unit and the second pooling unit have different pooling processing modes for capturing different information in the sample image features; Using the multi-layer perceptron in the channel attention network model, mapping processing is performed on the first pooled image features of the sample and the second pooled image features of the sample, respectively, to obtain corresponding first mapped image features of the sample and second mapped image features of the sample; Based on the sample first mapping image feature and the sample second mapping image feature, a sample intermediate image feature corresponding to the sample image feature is determined.

[0012] The present application also provides a digital meter reading recognition device, comprising: an image segmentation module, configured to segment an initial digital instrument image to obtain a corresponding target digital instrument image, wherein the initial digital instrument image is formed by performing an image acquisition operation on a target digital instrument, and the target digital instrument image at least includes a digital display area of ​​the target digital instrument; a first feature mining module configured to perform a first feature mining process on the target digital instrument image to obtain corresponding intermediate image features, wherein the intermediate image features are used to represent semantic information of digits in the target digital instrument image, and the first feature mining process at least includes a channel attention process for mining semantic information of significant channels in the target digital instrument image; a second feature mining module, configured to perform a second feature mining process on the intermediate image features to obtain corresponding target image features, wherein the second feature mining process is different from the first feature mining process and the second feature mining process at least includes a correlation feature mining process for mining correlation semantic information between feature columns in the intermediate image features; The semantic recognition module is used to perform semantic recognition processing on the target image features to obtain a target semantic recognition result corresponding to the target digital instrument, wherein the target semantic recognition result at least includes the number displayed in the digital display area of ​​the target digital instrument.

[0013] The present application also provides a digital instrument reading recognition network training device, comprising: a network acquisition module, configured to acquire an initial image segmentation network and an initial digital meter reading recognition network included in an initial neural network, and to acquire a first sample digital meter image, first image label information corresponding to the first sample digital meter image, a second sample digital meter image, and second image label information corresponding to the second sample digital meter image, wherein the first image label information is used to indicate a digital display area in the first sample digital meter image, and the second image label information is used to indicate a digital displayed in the digital display area in the second sample digital meter image; a first network updating module configured to segment the first sample digital instrument image using the initial image segmentation network to obtain a corresponding target sample digital instrument image, and to update network parameters of the initial image segmentation network based on the target sample digital instrument image and the first image label information to obtain a corresponding target image segmentation network, wherein the target sample digital instrument image at least includes a digital display area of ​​the digital instrument; a first sample feature mining module configured to perform a first feature mining process on the second sample digital meter image using the initial digital meter reading recognition network to obtain corresponding sample intermediate image features, wherein the sample intermediate image features are used to represent semantic information of digits in the second sample digital meter image, and the first feature mining process at least includes a channel attention process for mining semantic information of significant channels in the second sample digital meter image; a second sample feature mining module configured to perform a second feature mining process on the sample intermediate image features using the initial digital meter reading recognition network to obtain corresponding sample target image features, wherein the second feature mining process is different from the first feature mining process and the second feature mining process at least includes a correlation feature mining process for mining correlation semantic information between feature columns in the sample intermediate image features; a sample semantic recognition module, configured to perform semantic recognition processing on the sample target image features using the initial digital meter reading recognition network to obtain a sample semantic recognition result corresponding to the digital meter, wherein the sample semantic recognition result at least includes a number displayed in a digital display area of ​​the digital meter; a second network updating module, configured to update network parameters of the initial digital meter reading recognition network based on a difference between the sample semantic recognition result and the second image label information, to obtain a corresponding target digital meter reading recognition network; The target neural network determination module is used to determine the corresponding target neural network based on the target image segmentation network and the target digital meter reading recognition network.

[0014] Based on the above, the present application further provides an electronic device, including: Memory for storing computer programs; The processor connected to the memory is used to execute the computer program stored in the memory to implement the above-mentioned digital meter reading recognition method, or to implement the above-mentioned digital meter reading recognition network training method.

[0015] Based on the above, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is run, it executes the various steps of the above-mentioned digital instrument reading recognition method or digital instrument reading recognition network training method.

[0016] The reading recognition method and apparatus, network training method and apparatus, device, and medium provided in this application first segment an initial digital instrument image to obtain a corresponding target digital instrument image; secondly, perform a first feature mining process on the target digital instrument image to obtain corresponding intermediate image features; then, perform a second feature mining process on the intermediate image features to obtain corresponding target image features; and finally, perform semantic recognition on the target image features to obtain target semantic recognition results corresponding to the target digital instrument. Based on the above, since the target digital instrument image undergoes two different feature mining processes, the semantic representation capability of the mined target image features is enhanced, thereby improving the reliability of semantic recognition based on the target image features. Furthermore, since the first feature mining process mines the semantic information of salient channels in the target digital instrument image, it can capture detailed information that is more representative of the digital display, thereby further improving the reliability of semantic recognition. Furthermore, since the second feature mining process mines the semantic information associated between feature columns in the intermediate image features, it can further improve the semantic representation capability of the mined target image features. Based on this, by adopting the improved solution of the present application, the reliability of semantic recognition can be fully guaranteed, thereby improving the relatively poor reliability of digital instrument reading recognition in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings.

[0018] Figure 1 This is a structural block diagram of the electronic device provided in an embodiment of the present application.

[0019] Figure 2 A flowchart of a method for identifying digital meter readings provided in an embodiment of the present application.

[0020] Figure 3 A schematic diagram of image segmentation processing provided in an embodiment of the present application.

[0021] Figure 4 A schematic diagram comparing the various channels of the image provided in the embodiments of the present application.

[0022] Figure 5 A schematic diagram of the first feature mining process provided in an embodiment of the present application.

[0023] Figure 6 A flowchart of a digital meter reading recognition network training method provided in an embodiment of the present application.

[0024] Figure 7 A block diagram of a digital meter reading recognition device provided in an embodiment of the present application.

[0025] Figure 8 A block diagram of a digital meter reading recognition network training device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0027] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.

[0028] like Figure 1As shown, an embodiment of the present application provides an electronic device, wherein the electronic device may include a memory, a processor, and a target device, wherein the target device may be a digital meter reading recognition device or a digital meter reading recognition network training device.

[0029] In detail, the memory and the processor are electrically connected directly or indirectly to achieve data transmission or interaction. For example, the memory and the processor can be electrically connected through one or more communication buses or signal lines. The target device includes at least one software function module stored in the memory in the form of software or firmware. The processor is used to execute an executable computer program stored in the memory, for example, the software function module and computer program included in the target device, to implement the target method provided in the embodiment of the present application (such as a digital meter reading recognition method or a digital meter reading recognition network training method).

[0030] Optionally, the memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0031] Furthermore, the processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0032] I understand. Figure 1 The structure shown is only for illustration, and the electronic device may also include Figure 1 More or fewer components than shown, or with Figure 1 The different configurations shown, for example, may also include a communication unit for exchanging information with other devices (such as an image acquisition device, etc.).

[0033] Combine Figure 2 The present application also provides a method for identifying digital meter readings applicable to the electronic device described above, wherein the method steps defined in the process related to the method for identifying digital meter readings can be implemented by the electronic device.

[0034] The following will Figure 2 The specific process shown is explained in detail.

[0035] Step S110 , segmenting the initial digital instrument image to obtain a corresponding target digital instrument image.

[0036] In an embodiment of the present application, the electronic device can perform segmentation processing (also called target detection) on the initial digital instrument image to obtain a corresponding target digital instrument image. The initial digital instrument image is formed by performing an image acquisition operation on the target digital instrument, and the target digital instrument image at least includes the digital display area of ​​the target digital instrument. In other words, the partial image corresponding to the digital display area (such as the image of the target digital instrument) can be detected from the initial digital instrument image. Figure 3 as shown), to avoid interference from images in other areas during subsequent reading and recognition.

[0037] Step S120 : performing a first feature mining process on the target digital instrument image to obtain corresponding intermediate image features.

[0038] In an embodiment of the present application, after obtaining the target digital instrument image, the electronic device may perform a first feature mining process on the target digital instrument image to obtain corresponding intermediate image features. The intermediate image features are used to represent the semantic information of the digits in the target digital instrument image, and the first feature mining process at least includes a channel attention process for mining the semantic information of the significant channels in the target digital instrument image, such as Figure 4 As shown, the inventors of this application have found through research that the characteristics of different channels (color channels) have different importance (significance) for digital display. Therefore, through channel attention processing, important channel information can be focused on. For example, the information of the R channel has a stronger representation ability, so the R channel can be focused on.

[0039] Step S130: performing a second feature mining process on the intermediate image features to obtain corresponding target image features.

[0040] In an embodiment of the present application, after obtaining the intermediate image features, the electronic device may perform a second feature mining process on the intermediate image features to obtain corresponding target image features. The second feature mining process is different from the first feature mining process and includes at least a correlation feature mining process, which is used to mine the semantic information associated with the relationships between the feature columns in the intermediate image features. In other words, the intermediate image features may include multiple feature columns, so that the association information between the feature columns can be mined.

[0041] Step S140 , performing semantic recognition processing on the target image features to obtain a target semantic recognition result corresponding to the target digital instrument.

[0042] In this embodiment of the present application, after obtaining the target image features, the electronic device may perform semantic recognition processing on the target image features (e.g., identifying corresponding numbers) to obtain a target semantic recognition result corresponding to the target digital instrument. The target semantic recognition result includes at least the number displayed in the digital display area of ​​the target digital instrument.

[0043] Based on the above, since the target digital instrument image is subjected to two different feature mining processes, the semantic representation ability of the mined target image features is better, thereby improving the reliability of semantic recognition based on the target image features. In addition, since the first feature mining process can mine the semantic information of the significant channels in the target digital instrument image, it can capture detailed information that is more representative of the digital display, and therefore, the reliability of semantic recognition can be further improved. In addition, since the second feature mining process can mine the correlation semantic information between the feature columns in the intermediate image features, it can further improve the semantic representation ability of the mined target image features. Based on this, by adopting the improved solution of the present application, the reliability of semantic recognition can be fully guaranteed, thereby improving the problem of relatively poor reliability of digital instrument reading recognition in the prior art.

[0044] First, it should be noted that, regarding step S110 , the specific method of segmenting the initial digital instrument image is not limited and can be selected according to actual needs.

[0045] For example, in an alternative embodiment, to improve the efficiency and accuracy of segmenting the initial digital instrument image, a neural network can be used for target detection. The detected digital display area is then segmented to obtain the corresponding target digital instrument image. For example, a YOLO series model (such as YOLO5) can be used for target detection. The core concept of the YOLO series model is to directly predict bounding boxes and categories in the image using a single neural network model, thereby achieving real-time target detection.

[0046] Secondly, it should be noted that for step S120 , the specific method of performing the first feature mining process on the target digital instrument image is not limited and can be selected according to actual needs.

[0047] For example, in an alternative embodiment, in order to improve the reliability of the first feature mining process on the target digital instrument image, the above-mentioned step S120 may further include step S121 and step S122, and the specific content of each step is described as follows.

[0048] Step S121 : Using the convolutional network model included in the target digital meter reading recognition network, perform feature extraction processing on the target digital meter image to obtain corresponding initial image features.

[0049] In an embodiment of the present application, a convolutional network model (such as a convolutional neural network (CNN)) included in a target digital instrument reading recognition network can be used to perform feature extraction processing on the target digital instrument image to obtain corresponding initial image features. In this way, shallow image features can be obtained.

[0050] Step S122: Using the channel attention network model included in the target digital meter reading recognition network, perform channel attention processing on the initial image features to obtain corresponding intermediate image features.

[0051] In an embodiment of the present application, after obtaining the initial image features, the channel attention network model included in the target digital meter reading recognition network can be further used to perform channel attention processing on the initial image features to obtain corresponding intermediate image features.

[0052] It can be understood that, in the above-mentioned step S122, the specific method of performing channel attention processing on the initial image features is not limited and can be selected according to actual needs.

[0053] For example, in an alternative embodiment, in order to make the mined intermediate image features have better representation capabilities, the above-mentioned step S122 may further include step S122a, step S122b and step S122c, and the specific content of each step is described as follows.

[0054] Step S122a, using the first pooling unit and the second pooling unit in the channel attention network model included in the target digital meter reading recognition network, respectively perform pooling processing on the initial image features to obtain corresponding first pooled image features and second pooled image features.

[0055] In the embodiment of this application, combined with Figure 5 , the first pooling unit and the second pooling unit in the channel attention network model included in the target digital meter reading recognition network can be used to perform pooling processing on the initial image features respectively to obtain corresponding first pooled image features and second pooled image features. For example, the initial image features are pooled by the first pooling unit to obtain first pooled image features; and the initial image features are pooled by the second pooling unit to obtain second pooled image features. The first pooling unit and the second pooling unit have different pooling processing methods for capturing different information in the initial image features. Exemplarily, the first pooling unit can be a maximum pooling unit, which can obtain important information of the image and remove interference information; the second pooling unit can be an average pooling unit, which can remove some unimportant detail information and help the model generate smoother output, thereby improving the generalization ability of the model.

[0056] Step S122b: Using the multi-layer perceptron in the channel attention network model included in the target digital meter reading recognition network, the first pooled image features and the second pooled image features are mapped to obtain corresponding first mapped image features and second mapped image features.

[0057] In an embodiment of the present application, after obtaining the first pooled image features and the second pooled image features, a multi-layer perceptron (MLP) in the channel attention network model included in the target digital meter reading recognition network can be used to perform mapping processing on the first pooled image features and the second pooled image features, respectively, to obtain first mapped image features and second mapped image features. For example, the first pooled image features can be processed by the multi-layer perceptron to obtain the first mapped image features, and the second pooled image features can be processed by the multi-layer perceptron to obtain the corresponding second mapped image features.

[0058] Step S122c: determining an intermediate image feature corresponding to the initial image feature based on the first mapped image feature and the second mapped image feature.

[0059] In the embodiment of the present application, after obtaining the first mapped image feature and the second mapped image feature, an intermediate image feature corresponding to the initial image feature may be determined based on the first mapped image feature and the second mapped image feature.

[0060] It should be noted that for the above step S122c, the specific method of determining the intermediate image feature based on the first mapped image feature and the second mapped image feature is not limited and can be selected according to actual needs.

[0061] For example, in an alternative embodiment, in order to ensure that the obtained intermediate image features can fully represent the feature information of each channel to varying degrees, thereby improving the representation capability of the intermediate image features, the above-mentioned step S122c may further include the following: First, the first mapping image feature and the second mapping image feature can be subjected to a first feature fusion process to obtain a corresponding fused mapping image feature, and the fused mapping image feature can be subjected to an activation process to obtain a corresponding activated mapping image feature, wherein the activated mapping image feature carries semantic information representing the significance of the image channel; illustratively, the first mapping image feature and the second mapping image feature can be superimposed to obtain a corresponding fused mapping image feature, and then the fused mapping image feature can be activated by a sigmoid function to obtain a corresponding activated mapping image feature, so that the activated mapping image feature can represent the weight value corresponding to each channel; it should be noted that, in other implementations, the first mapping image feature and the second mapping image feature can also be activated by a sigmoid function to obtain a first activation feature (such as Figure 5 alpha1 shown) and the second activation feature (as Figure 5 Then, the first activation feature and the second activation feature are superimposed to obtain the activation map image feature (as shown in Figure 5 Alpha shown); Secondly, the activation mapping image features and the initial image features can be subjected to a second feature fusion process to obtain intermediate image features corresponding to the initial image features; illustratively, the activation mapping image features and the initial image features can be directly superimposed to obtain corresponding intermediate image features, or the activation mapping image features and the initial image features can be weightedly superimposed to obtain corresponding intermediate image features, wherein the weighting coefficients of the weighted superposition can be used as network parameters of the corresponding neural network, and updated during the training process; based on this, the obtained intermediate image features can pay more attention to relatively more important channel information.

[0062] Thirdly, it should be noted that for step S130, the specific method of performing the second feature mining process on the intermediate image features is not limited and can be selected according to actual needs.

[0063] For example, in an alternative embodiment, in order to further improve the characterization capability of the obtained target image features, the above-mentioned step S130 may further include step S131 and step S132, and the specific content of each step is described as follows.

[0064] Step S131, using the self-attention network model included in the target digital meter reading recognition network, the intermediate image features are mapped to obtain corresponding query image features, key image features and value image features, and based on the query image features, the key image features and the value image features, the self-attention image features corresponding to the intermediate image features are determined.

[0065] In an embodiment of the present application, the self-attention network model included in the target digital meter reading recognition network can be used to first perform mapping processing on the intermediate image features to obtain corresponding query image features, key image features, and value image features, and then determine the self-attention image features corresponding to the intermediate image features based on the query image features, the key image features, and the value image features. Exemplarily, the self-attention network model can include three matrices, such as a query matrix, a key matrix, and a value matrix. Then, based on the query matrix, the key matrix, and the value matrix, the intermediate image features can be mapped (such as multiplied) to obtain corresponding query image features, key image features, and value image features. For example, a similarity matrix between the query image features and the key image features can be calculated, and then the similarity matrix and the value image features can be multiplied (such as by performing a multiplication process according to matrix multiplication) to obtain the corresponding self-attention image features.

[0066] In step S132, the bidirectional feature capture network model included in the target digital meter reading recognition network is used to perform bidirectional contextual semantic information capture processing on the self-attention image feature to obtain the corresponding first contextual semantic feature and second contextual semantic feature, and the first contextual semantic feature and the second contextual semantic feature are subjected to feature fusion processing to obtain the target image feature corresponding to the self-attention image feature.

[0067] In an embodiment of the present application, after obtaining the self-attention image feature, the bidirectional feature capture network model included in the target digital meter reading recognition network can be further utilized (such as a bidirectional long short-term memory network with a forgetting mechanism, which can include two long short-term memory networks in opposite directions, one processing the input sequence in forward order, and the other processing the input sequence in reverse order; the long short-term memory network in each direction will generate an output sequence, and these output sequences will be merged or connected together to provide a comprehensive view of bidirectional information), to perform bidirectional contextual semantic information capture processing on the self-attention image feature to obtain the corresponding first contextual semantic feature and the second contextual semantic feature, and to perform feature fusion processing on the first contextual semantic feature and the second contextual semantic feature to obtain the target image feature corresponding to the self-attention image feature. Fourthly, it should be noted that for step S140 , the specific method of performing semantic recognition processing on the target image features is not limited and can be selected according to actual needs.

[0068] For example, in an alternative embodiment, in order to ensure that the reliability of the semantic recognition processing can be higher and to obtain reliable target semantic recognition results, character prediction output can be performed based on each feature in the target image feature. For example, the character with the highest probability among all the characters corresponding to the feature is used as the output (prediction result) of the feature. In addition, repeated characters and placeholders can be removed to obtain the corresponding text recognition content (numbers). For example, the target image features can be semantically recognized by the transcription layer (Transcription Layer) included in the target digital meter reading recognition network. That is, the transcription layer can map the serialized features to a sequence of text or labels, wherein the transcription layer can be composed of one or more fully connected layers for mapping the target image features to the label space. These fully connected layers can generate the probability distribution of the output sequence through activation functions (such as softmax).

[0069] Combine Figure 6The present application also provides a method for training a digital meter reading recognition network applicable to the electronic device described above. The method steps defined in the process related to the training method for digital meter reading recognition network can be implemented by the electronic device.

[0070] The following will Figure 7 The specific process shown is explained in detail.

[0071] Step S210: Acquire an initial image segmentation network and an initial digital instrument reading recognition network included in the initial neural network, and acquire a first sample digital instrument image, first image label information corresponding to the first sample digital instrument image, a second sample digital instrument image, and second image label information corresponding to the second sample digital instrument image.

[0072] In an embodiment of the present application, the electronic device may first obtain an initial neural network comprising an initial image segmentation network and an initial digital meter reading recognition network, and obtain a first sample digital meter image, first image label information corresponding to the first sample digital meter image, and second image label information corresponding to the second sample digital meter image and the second sample digital meter image. The first image label information indicates the digital display area in the first sample digital meter image (e.g., identified by a rectangular box, etc.), and the second image label information indicates the number displayed in the digital display area in the second sample digital meter image. Furthermore, the first sample digital meter image and the second sample digital meter image may be the same or different. To improve the robustness of the network model, image enhancement (for the first sample digital meter image) may be achieved through methods such as image rotation, cropping, filtering, and histogram equalization, effectively enhancing image contrast and addressing issues such as uneven lighting. The processed image may be annotated using an annotation tool to create a rectangular box corresponding to the "digital display area," retaining the corresponding width, height, and coordinate information. The annotated dataset may then be used to train an initial neural network (e.g., a YOLO5 model), and each image may be segmented based on the annotated rectangular box, thereby saving the image of the digital display area. The segmented digital display area image can be used as input for subsequent models (eg, as a second sample digital instrument image).

[0073] Step S220: Using the initial image segmentation network, the first sample digital instrument image is segmented to obtain a corresponding target sample digital instrument image, and based on the target sample digital instrument image and the first image label information, the network parameters of the initial image segmentation network are updated to obtain a corresponding target image segmentation network.

[0074] In this embodiment of the present application, after obtaining the initial image segmentation network, the first sample digital instrument image, and the corresponding first image label information, the electronic device may utilize the initial image segmentation network to perform segmentation processing (e.g., target area detection) on the first sample digital instrument image to obtain a corresponding target sample digital instrument image. Furthermore, based on (the difference between) the target sample digital instrument image and the first image label information, the electronic device may update the network parameters of the initial image segmentation network to obtain a corresponding target image segmentation network. The target sample digital instrument image includes at least the digital display area of ​​the digital instrument.

[0075] Step S230 : Using the initial digital meter reading recognition network, perform a first feature mining process on the second sample digital meter image to obtain corresponding sample intermediate image features.

[0076] In an embodiment of the present application, after obtaining the initial digital meter reading recognition network, the second sample digital meter image, and the corresponding second image label information, the electronic device can utilize the initial digital meter reading recognition network to perform a first feature mining process on the second sample digital meter image to obtain corresponding sample intermediate image features. The sample intermediate image features are used to characterize the semantic information of the digits in the second sample digital meter image, and the first feature mining process includes at least a channel attention process for mining the semantic information of significant channels in the second sample digital meter image, as previously described.

[0077] Step S240 , using the initial digital meter reading recognition network, performs a second feature mining process on the sample intermediate image features to obtain corresponding sample target image features.

[0078] In an embodiment of the present application, after obtaining the sample intermediate image features, the electronic device may utilize the initial digital meter reading recognition network to perform a second feature mining process on the sample intermediate image features to obtain corresponding sample target image features. The second feature mining process is different from the first feature mining process and includes at least a correlation feature mining process for mining the semantic information of associations between feature columns in the sample intermediate image features, as previously described.

[0079] Step S250 , using the initial digital meter reading recognition network, performs semantic recognition processing on the sample target image features to obtain a sample semantic recognition result corresponding to the digital meter.

[0080] In an embodiment of the present application, after obtaining the sample target image features, the electronic device can use the initial digital meter reading recognition network to perform semantic recognition processing on the sample target image features to obtain a sample semantic recognition result corresponding to the digital meter, wherein the sample semantic recognition result at least includes the numbers displayed in the digital display area of ​​the digital meter, as described above.

[0081] Step S260 : Based on the difference between the sample semantic recognition result and the second image label information, the network parameters of the initial digital meter reading recognition network are updated to obtain a corresponding target digital meter reading recognition network.

[0082] In an embodiment of the present application, after obtaining the sample semantic recognition result, the electronic device may update the network parameters of the initial digital meter reading recognition network based on the difference between the sample semantic recognition result and the second image label information, thereby obtaining a corresponding target digital meter reading recognition network. For example, a CTC loss (Connectionist Temporal Classification Loss) may be calculated between the sample semantic recognition result and the second image label information. CTC loss is a commonly used loss function for image text recognition tasks, particularly for variable-length sequence outputs. It allows the network to freely align input and output sequences without requiring precise alignment. In other implementations, the corresponding loss may be calculated based on other methods (such as edit distance). The network parameters of the initial digital meter reading recognition network may then be updated in a direction that reduces the loss, thereby obtaining a corresponding target digital meter reading recognition network (e.g., when the number of updates reaches a preset number, or when the loss is reduced to a preset loss, or when the loss is reduced by less than a preset loss).

[0083] Step S270: Determine a corresponding target neural network based on the target image segmentation network and the target digital meter reading recognition network.

[0084] In an embodiment of the present application, after obtaining the target image segmentation network and the target digital meter reading recognition network, a corresponding target neural network can be determined based on the target image segmentation network and the target digital meter reading recognition network. For example, the target image segmentation network and the target digital meter reading recognition network can be directly combined to form the target neural network, or a corresponding target neural network can be constructed based on the network parameters of the target image segmentation network and the target digital meter reading recognition network. The target neural network can be used to implement each step of the above-mentioned digital meter reading recognition method.

[0085] First, in the above step S220, the specific method of updating the initial image segmentation network is not limited and can be selected according to actual needs.

[0086] For example, in an alternative embodiment, in order to make the updated target image segmentation network have higher reliability so as to reliably identify and detect the digital display area, the above-mentioned step S220 can further include step S221, step S222, step S223 and step S224, and the specific content of each step is described below.

[0087] Step S221 : determining a plurality of corresponding bounding boxes based on a plurality of first image label information possessed by a plurality of first sample digital images.

[0088] In this embodiment of the present application, after obtaining multiple first sample digital images, multiple corresponding bounding boxes can be determined based on the multiple first image tag information of the multiple first sample digital images. Each of the bounding boxes corresponds to a piece of first image tag information, and the bounding box is used to indicate a corresponding digital display area.

[0089] Step S222 : performing clustering processing on the multiple bounding boxes to obtain corresponding clustering results, and determining at least one target bounding box based on the clustering results.

[0090] In an embodiment of the present application, after obtaining the multiple bounding boxes, clustering processing can be performed on the multiple bounding boxes to obtain corresponding clustering results, and at least one target bounding box can be determined based on the clustering results. For example, a target bounding box (anchor box) can be obtained based on the K-Means clustering algorithm or other clustering algorithms for use in subsequent training. For example, the bounding box belonging to the cluster center can be used as the target bounding box.

[0091] Step S223 : Using the initial image segmentation network, based on the at least one target bounding box, segment the first sample digital instrument image to obtain a corresponding target sample digital instrument image.

[0092] In an embodiment of the present application, after obtaining the at least one target bounding box, the initial image segmentation network can be used to perform segmentation processing on the first sample digital instrument image based on the at least one target bounding box to obtain a corresponding target sample digital instrument image.

[0093] Step S224 : Based on the target sample digital instrument image and the first image label information, the network parameters of the initial image segmentation network are updated to obtain a corresponding target image segmentation network.

[0094] In an embodiment of the present application, after obtaining the target sample digital instrument image, the network parameters of the initial image segmentation network can be updated based on the target sample digital instrument image and the first image label information to obtain a corresponding target image segmentation network. For example, based on the difference between the target sample digital instrument image and the first image label information, a corresponding error (or loss; the specific loss function is not specifically defined herein) can be determined. Then, the network parameters of the initial image segmentation network can be updated in a direction that reduces the error, thereby obtaining a corresponding target image segmentation network. For example, as previously described, the labeled dataset (i.e., multiple labeled images) can be divided into a 7:2:1 ratio, corresponding to a training set, a validation set, and a test set. The training set data can then be used to train a YOLO5 model and update its parameters. The validation set can also be used to evaluate the model's performance during training. Training stops when the accuracy reaches a specified threshold and the loss function no longer changes. The test set is the dataset used for the final evaluation of model performance. Once the model has passed the training and validation phases, evaluating it on an unexplored test set can more objectively assess the model's generalization ability. The purpose of the test set is to simulate new data that the model will encounter in real applications, so as to more accurately evaluate the performance of the model.

[0095] On the second aspect, in the above-mentioned step S230, the specific manner of performing the first feature mining process on the second sample digital instrument image is not limited and can be selected according to actual needs.

[0096] For example, in an alternative embodiment, in order to enable the mined sample intermediate image features to have better characterization capabilities for the second sample digital instrument image, the above-mentioned step S230 may further include step S231, step S232, step S233 and step S234, and the specific contents of each step are described below.

[0097] Step S231 : Using the convolutional network model included in the initial digital meter reading recognition network, perform feature extraction processing on the second sample digital meter image to obtain corresponding sample image features.

[0098] In an embodiment of the present application, after obtaining the second sample digital instrument image, the convolutional network model included in the initial digital instrument reading recognition network can be used to perform feature extraction processing on the second sample digital instrument image to obtain corresponding sample image features, wherein the specific process of feature extraction processing is as described above.

[0099] In step S232, the first pooling unit and the second pooling unit in the channel attention network model included in the initial digital meter reading recognition network are used to perform pooling processing on the sample image features respectively to obtain corresponding sample first pooled image features and sample second pooled image features.

[0100] In an embodiment of the present application, after obtaining the sample image features, the first pooling unit and the second pooling unit in the channel attention network model included in the initial digital meter reading recognition network can be used to perform pooling processing on the sample image features, respectively, to obtain corresponding sample first pooled image features and sample second pooled image features. The first pooling unit and the second pooling unit have different pooling processing modes, which are used to capture different information in the sample image features, as described above.

[0101] Step S233, using the multi-layer perceptron in the channel attention network model, respectively perform mapping processing on the sample first pooled image features and the sample second pooled image features to obtain corresponding sample first mapped image features and sample second mapped image features.

[0102] In an embodiment of the present application, after obtaining the sample first pooled image features and the sample second pooled image features, the multi-layer perceptron in the channel attention network model can be used to perform mapping processing on the sample first pooled image features and the sample second pooled image features, respectively, to obtain the corresponding sample first mapping image features and sample second mapping image features, as described above.

[0103] Step S234 : determining a sample intermediate image feature corresponding to the sample image feature based on the sample first mapping image feature and the sample second mapping image feature.

[0104] In an embodiment of the present application, after obtaining the sample first mapping image feature and the sample second mapping image feature, the sample intermediate image feature corresponding to the sample image feature can be determined based on the sample first mapping image feature and the sample second mapping image feature, as described above.

[0105] Combine Figure 7 The present application also provides a digital meter reading recognition device applicable to the above electronic device. The digital meter reading recognition device may include an image segmentation module, a first feature mining module, a second feature mining module, and a semantic recognition module.

[0106] The image segmentation module can be used to segment the initial digital instrument image to obtain a corresponding target digital instrument image, wherein the initial digital instrument image is formed by performing an image acquisition operation on the target digital instrument, and the target digital instrument image at least includes the digital display area of ​​the target digital instrument. In the embodiment of the present application, the image segmentation module can be used to perform Figure 2 As shown in step S110, for the relevant content of the image segmentation module, reference may be made to the above description of step S110.

[0107] The first feature mining module can be used to perform a first feature mining process on the target digital instrument image to obtain corresponding intermediate image features, wherein the intermediate image features are used to represent the semantic information of the digits in the target digital instrument image, and the first feature mining process at least includes a channel attention process for mining the semantic information of the significant channels in the target digital instrument image. In the embodiment of the present application, the first feature mining module can be used to perform Figure 2 As shown in step S120, for the relevant content of the first feature mining module, please refer to the description of step S120 above.

[0108] The second feature mining module can be used to perform a second feature mining process on the intermediate image features to obtain corresponding target image features, wherein the second feature mining process is different from the first feature mining process, and the second feature mining process at least includes a correlation feature mining process for mining the correlation semantic information between each feature column in the intermediate image features. In the embodiment of the present application, the second feature mining module can be used to perform Figure 2 As shown in step S130, for the relevant content of the second feature mining module, reference can be made to the above description of step S130.

[0109] The semantic recognition module can be used to perform semantic recognition processing on the target image features to obtain the target semantic recognition result corresponding to the target digital instrument, wherein the target semantic recognition result at least includes the number displayed in the digital display area of ​​the target digital instrument. In the embodiment of the present application, the semantic recognition module can be used to perform Figure 2 As shown in step S140, for the relevant content of the semantic recognition module, reference can be made to the above description of step S140.

[0110] Combine Figure 8The present application also provides a network training device for digital meter reading recognition applicable to the aforementioned electronic device. The network training device may include a network acquisition module, a first network update module, a first sample feature mining module, a second sample feature mining module, a sample semantic recognition module, a second network update module, and a target neural network determination module.

[0111] The network acquisition module is used to acquire the initial image segmentation network and the initial digital meter reading recognition network included in the initial neural network, and acquire the first sample digital meter image, the first image label information corresponding to the first sample digital meter image, the second sample digital meter image and the second image label information corresponding to the second sample digital meter image, wherein the first image label information is used to indicate the digital display area in the first sample digital meter image, and the second image label information is used to indicate the number displayed in the digital display area in the second sample digital meter image. In this embodiment of the application, the network acquisition module can be used to perform Figure 6 As shown in step S210, for the relevant content of the network acquisition module, reference can be made to the above description of step S210.

[0112] The first network update module is used to use the initial image segmentation network to segment the first sample digital instrument image to obtain a corresponding target sample digital instrument image, and, based on the target sample digital instrument image and the first image label information, update the network parameters of the initial image segmentation network to obtain a corresponding target image segmentation network, wherein the target sample digital instrument image at least includes the digital display area of ​​the digital instrument. In this embodiment of the present application, the first network update module can be used to perform Figure 6 As shown in step S220, for the relevant content of the first network update module, please refer to the description of step S220 above.

[0113] The first sample feature mining module is used to use the initial digital meter reading recognition network to perform a first feature mining process on the second sample digital meter image to obtain corresponding sample intermediate image features, wherein the sample intermediate image features are used to characterize the semantic information of the digits in the second sample digital meter image, and the first feature mining process at least includes a channel attention process for mining the semantic information of the significant channels in the second sample digital meter image. In this embodiment of the present application, the first sample feature mining module can be used to perform Figure 6 As shown in step S230, for the relevant content of the first sample feature mining module, please refer to the description of step S230 above.

[0114] The second sample feature mining module is used to perform a second feature mining process on the sample intermediate image features using the initial digital meter reading recognition network to obtain corresponding sample target image features, wherein the second feature mining process is different from the first feature mining process, and the second feature mining process at least includes a correlation feature mining process for mining the correlation semantic information between each feature column in the sample intermediate image features. In this embodiment of the application, the second sample feature mining module can be used to perform Figure 6 As shown in step S240, for the relevant content of the second sample feature mining module, please refer to the above description of step S240.

[0115] The sample semantic recognition module is used to use the initial digital meter reading recognition network to perform semantic recognition processing on the sample target image features to obtain a sample semantic recognition result corresponding to the digital meter, wherein the sample semantic recognition result at least includes the number displayed in the digital display area of ​​the digital meter. In the embodiment of the present application, the sample semantic recognition module can be used to perform Figure 6 As shown in step S250, for the relevant content of the sample semantic recognition module, reference can be made to the above description of step S250.

[0116] The second network update module is used to update the network parameters of the initial digital meter reading recognition network based on the difference between the sample semantic recognition result and the second image label information to obtain the corresponding target digital meter reading recognition network. In this embodiment of the application, the second network update module can be used to perform Figure 6 As shown in step S260, for the relevant content of the second network update module, reference can be made to the above description of step S260.

[0117] The target neural network determination module is used to determine the corresponding target neural network based on the target image segmentation network and the target digital meter reading recognition network. In the embodiment of the present application, the target neural network determination module can be used to perform Figure 6 As shown in step S270, for the relevant content of the target neural network determination module, reference can be made to the description of step S270 above.

[0118] In an embodiment of the present application, corresponding to the aforementioned digital meter reading recognition method or digital meter reading recognition network training method applied to the electronic device, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program that, when executed, executes each step of the digital meter reading recognition method and / or digital meter reading recognition network training method. The steps executed when the aforementioned computer program is executed are not described in detail here; reference is made to the aforementioned explanation of the digital meter reading recognition method and / or digital meter reading recognition network training method.

[0119] In summary, the reading recognition method and apparatus, network training method and apparatus, device, and medium provided by this application can first segment an initial digital instrument image to obtain a corresponding target digital instrument image; secondly, perform a first feature mining process on the target digital instrument image to obtain corresponding intermediate image features; then, perform a second feature mining process on the intermediate image features to obtain corresponding target image features; and finally, perform semantic recognition on the target image features to obtain target semantic recognition results corresponding to the target digital instrument. Based on the above, since the target digital instrument image undergoes two different feature mining processes, the semantic representation capability of the mined target image features is enhanced, thereby improving the reliability of semantic recognition based on the target image features. Furthermore, since the first feature mining process can mine the semantic information of salient channels in the target digital instrument image, it can capture detailed information that is more representative of the digital display, thereby further improving the reliability of semantic recognition. Furthermore, since the second feature mining process can mine the semantic information associated between feature columns in the intermediate image features, it can further improve the semantic representation capability of the mined target image features. Based on this, by adopting the improved solution of the present application, the reliability of semantic recognition can be fully guaranteed, thereby improving the relatively poor reliability of digital instrument reading recognition in the prior art.

[0120] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0121] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0122] If the functions are implemented in the form of software modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, electronic device, or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. It should be noted that, in this document, the terms "comprise," "include," or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.

[0123] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for identifying digital instrument readings, characterized in that: include: Segmenting an initial digital instrument image to obtain a corresponding target digital instrument image, wherein the initial digital instrument image is formed by performing an image acquisition operation on a target digital instrument, and the target digital instrument image at least includes a digital display area of ​​the target digital instrument; performing a first feature mining process on the target digital instrument image to obtain corresponding intermediate image features, wherein the intermediate image features are used to represent semantic information of digits in the target digital instrument image, and the first feature mining process at least includes a channel attention process for mining semantic information of significant channels in the target digital instrument image; performing a second feature mining process on the intermediate image features to obtain corresponding target image features, wherein the second feature mining process is different from the first feature mining process and the second feature mining process at least includes a correlation feature mining process for mining correlation semantic information between feature columns in the intermediate image features; Performing semantic recognition processing on the target image features to obtain a target semantic recognition result corresponding to the target digital instrument, wherein the target semantic recognition result at least includes a number displayed in a digital display area of ​​the target digital instrument.

2. The digital meter reading recognition method according to claim 1, characterized in that: The step of performing a first feature mining process on the target digital instrument image to obtain corresponding intermediate image features includes: Using a convolutional network model included in a target digital meter reading recognition network, performing feature extraction processing on the target digital meter image to obtain corresponding initial image features; The channel attention network model included in the target digital meter reading recognition network is used to perform channel attention processing on the initial image features to obtain corresponding intermediate image features.

3. The digital meter reading recognition method according to claim 2, characterized in that: The step of performing channel attention processing on the initial image features using the channel attention network model included in the target digital meter reading recognition network to obtain corresponding intermediate image features includes: Using a first pooling unit and a second pooling unit in a channel attention network model included in the target digital meter reading recognition network, respectively performing pooling processing on the initial image features to obtain corresponding first pooled image features and second pooled image features, wherein the first pooling unit and the second pooling unit have different pooling processing modes for capturing different information in the initial image features; Using a multi-layer perceptron in a channel attention network model included in the target digital meter reading recognition network, mapping the first pooled image features and the second pooled image features respectively to obtain corresponding first mapped image features and second mapped image features; Based on the first mapped image feature and the second mapped image feature, an intermediate image feature corresponding to the initial image feature is determined.

4. The digital meter reading recognition method according to claim 3, characterized in that: The step of determining the intermediate image feature corresponding to the initial image feature based on the first mapped image feature and the second mapped image feature includes: performing a first feature fusion process on the first mapping image feature and the second mapping image feature to obtain a corresponding fused mapping image feature, and performing an activation process on the fused mapping image feature to obtain a corresponding activated mapping image feature, wherein the activated mapping image feature carries semantic information representing the saliency of the image channel; A second feature fusion process is performed on the activation map image feature and the initial image feature to obtain an intermediate image feature corresponding to the initial image feature.

5. The digital meter reading recognition method according to claim 1, characterized in that: The step of performing second feature mining processing on the intermediate image features to obtain corresponding target image features includes: Using a self-attention network model included in a target digital meter reading recognition network, mapping processing is performed on the intermediate image features to obtain corresponding query image features, key image features, and value image features, and based on the query image features, the key image features, and the value image features, determining the self-attention image features corresponding to the intermediate image features; The target digital meter reading recognition network includes a bidirectional feature capture network model, and the self-attention image feature is bidirectionally captured in the contextual semantic information to obtain a corresponding first contextual semantic feature and a second contextual semantic feature. Furthermore, the first contextual semantic feature and the second contextual semantic feature are subjected to feature fusion processing to obtain a target image feature corresponding to the self-attention image feature.

6. A digital instrument reading recognition network training method, characterized in that: include: Obtaining an initial image segmentation network and an initial digital meter reading recognition network included in an initial neural network, and obtaining a first sample digital meter image, first image label information corresponding to the first sample digital meter image, a second sample digital meter image, and second image label information corresponding to the second sample digital meter image, wherein the first image label information is used to indicate a digital display area in the first sample digital meter image, and the second image label information is used to indicate a digital displayed in the digital display area in the second sample digital meter image; Using the initial image segmentation network, segmenting the first sample digital instrument image to obtain a corresponding target sample digital instrument image; and updating network parameters of the initial image segmentation network based on the target sample digital instrument image and the first image label information to obtain a corresponding target image segmentation network, wherein the target sample digital instrument image at least includes a digital display area of ​​the digital instrument; performing a first feature mining process on the second sample digital meter image using the initial digital meter reading recognition network to obtain corresponding sample intermediate image features, wherein the sample intermediate image features are used to represent semantic information of digits in the second sample digital meter image, and the first feature mining process at least includes a channel attention process for mining semantic information of significant channels in the second sample digital meter image; Using the initial digital meter reading recognition network, performing a second feature mining process on the sample intermediate image features to obtain corresponding sample target image features, wherein the second feature mining process is different from the first feature mining process and the second feature mining process at least includes a correlation feature mining process for mining correlation semantic information between feature columns in the sample intermediate image features; Using the initial digital meter reading recognition network, performing semantic recognition processing on the sample target image features to obtain a sample semantic recognition result corresponding to the digital meter, wherein the sample semantic recognition result at least includes the number displayed in the digital display area of ​​the digital meter; Based on the difference between the sample semantic recognition result and the second image label information, updating the network parameters of the initial digital meter reading recognition network to obtain a corresponding target digital meter reading recognition network; Based on the target image segmentation network and the target digital meter reading recognition network, a corresponding target neural network is determined, wherein the target neural network is used to implement each step in the digital meter reading recognition method described in any one of claims 1-5.

7. The digital meter reading recognition network training method according to claim 6, characterized in that: The steps of using the initial image segmentation network to segment the first sample digital instrument image to obtain a corresponding target sample digital instrument image, and updating the network parameters of the initial image segmentation network based on the target sample digital instrument image and the first image label information to obtain a corresponding target image segmentation network include: Determining a plurality of corresponding bounding boxes based on a plurality of first image label information of a plurality of first sample digital images, wherein each of the bounding boxes corresponds to a piece of the first image label information, and the bounding box is used to indicate a corresponding digital display area; performing clustering processing on the plurality of bounding boxes to obtain corresponding clustering results, and determining at least one target bounding box based on the clustering results; Using the initial image segmentation network, based on the at least one target bounding box, segment the first sample digital instrument image to obtain a corresponding target sample digital instrument image; Based on the target sample digital instrument image and the first image label information, the network parameters of the initial image segmentation network are updated to obtain a corresponding target image segmentation network.

8. The digital meter reading recognition network training method according to claim 6 is characterized in that: The step of performing a first feature mining process on the second sample digital meter image using the initial digital meter reading recognition network to obtain corresponding sample intermediate image features includes: Using the convolutional network model included in the initial digital meter reading recognition network, perform feature extraction processing on the second sample digital meter image to obtain corresponding sample image features; Using a first pooling unit and a second pooling unit in a channel attention network model included in the initial digital meter reading recognition network, respectively performing pooling processing on the sample image features to obtain corresponding sample first pooled image features and sample second pooled image features, wherein the first pooling unit and the second pooling unit have different pooling processing modes for capturing different information in the sample image features; Using the multi-layer perceptron in the channel attention network model, mapping processing is performed on the first pooled image features of the sample and the second pooled image features of the sample, respectively, to obtain corresponding first mapped image features of the sample and second mapped image features of the sample; Based on the sample first mapping image feature and the sample second mapping image feature, a sample intermediate image feature corresponding to the sample image feature is determined.

9. A digital instrument reading recognition device, characterized in that: include: an image segmentation module, configured to segment an initial digital instrument image to obtain a corresponding target digital instrument image, wherein the initial digital instrument image is formed by performing an image acquisition operation on a target digital instrument, and the target digital instrument image at least includes a digital display area of ​​the target digital instrument; a first feature mining module configured to perform a first feature mining process on the target digital instrument image to obtain corresponding intermediate image features, wherein the intermediate image features are used to represent semantic information of digits in the target digital instrument image, and the first feature mining process at least includes a channel attention process for mining semantic information of significant channels in the target digital instrument image; a second feature mining module, configured to perform a second feature mining process on the intermediate image features to obtain corresponding target image features, wherein the second feature mining process is different from the first feature mining process and the second feature mining process at least includes a correlation feature mining process for mining correlation semantic information between feature columns in the intermediate image features; The semantic recognition module is used to perform semantic recognition processing on the target image features to obtain a target semantic recognition result corresponding to the target digital instrument, wherein the target semantic recognition result at least includes the number displayed in the digital display area of ​​the target digital instrument.

10. A digital instrument reading recognition network training device, characterized in that: include: a network acquisition module, configured to acquire an initial image segmentation network and an initial digital meter reading recognition network included in an initial neural network, and to acquire a first sample digital meter image, first image label information corresponding to the first sample digital meter image, a second sample digital meter image, and second image label information corresponding to the second sample digital meter image, wherein the first image label information is used to indicate a digital display area in the first sample digital meter image, and the second image label information is used to indicate a digital displayed in the digital display area in the second sample digital meter image; a first network updating module configured to segment the first sample digital instrument image using the initial image segmentation network to obtain a corresponding target sample digital instrument image, and to update network parameters of the initial image segmentation network based on the target sample digital instrument image and the first image label information to obtain a corresponding target image segmentation network, wherein the target sample digital instrument image at least includes a digital display area of ​​the digital instrument; a first sample feature mining module configured to perform a first feature mining process on the second sample digital meter image using the initial digital meter reading recognition network to obtain corresponding sample intermediate image features, wherein the sample intermediate image features are used to represent semantic information of digits in the second sample digital meter image, and the first feature mining process at least includes a channel attention process for mining semantic information of significant channels in the second sample digital meter image; a second sample feature mining module configured to perform a second feature mining process on the sample intermediate image features using the initial digital meter reading recognition network to obtain corresponding sample target image features, wherein the second feature mining process is different from the first feature mining process and the second feature mining process at least includes a correlation feature mining process for mining correlation semantic information between feature columns in the sample intermediate image features; a sample semantic recognition module, configured to perform semantic recognition processing on the sample target image features using the initial digital meter reading recognition network to obtain a sample semantic recognition result corresponding to the digital meter, wherein the sample semantic recognition result at least includes a number displayed in a digital display area of ​​the digital meter; a second network updating module, configured to update network parameters of the initial digital meter reading recognition network based on a difference between the sample semantic recognition result and the second image label information, to obtain a corresponding target digital meter reading recognition network; The target neural network determination module is used to determine the corresponding target neural network based on the target image segmentation network and the target digital meter reading recognition network.

11. An electronic device, characterized in that: include: memory for storing computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the digital meter reading recognition method described in any one of claims 1 to 5, or to implement the digital meter reading recognition network training method described in any one of claims 6 to 8.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when running, executes the digital meter reading recognition method according to any one of claims 1 to 5, or executes the digital meter reading recognition network training method according to any one of claims 6 to 8.

Citation Information

Patent Citations

  • Instrument numerical value identification method and device and related product

    CN114549815A

  • A wheel position indicator

    WO2023094855A1

  • Dinosaur footprint image data set construction method and system based on large language model

    WO2025091919A1