Point reading position identification method, device, point reading device and storage medium
Through image recognition technology, the position of fingers or pen tips is solved, and the problem of existing point reading machines needing supporting hardware equipment is realized, flexible point reading methods are realized, and user experience is improved.
Patent Information
- Application Number
- CN202110488678.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-06
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-05-06
AI Technical Summary
Existing point reading machines require supporting hardware equipment for point reading, which is highly limited and cannot flexibly support multiple point reading methods.
Through image recognition technology, the image to be identified in the point reading area is detected, the position of the finger or pen tip is recognized, specific position information is output, and the finger or pen tip is supported to enrich the point reading method.
Click-to-point reading can be achieved without supporting hardware equipment. Users can flexibly choose fingers or pen tips to reduce the cumbersome operation and improve user experience.
Smart Images

Figure CN113762045B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a point-reading position identification method, apparatus, point-reading device, and storage medium. Background Art
[0002] Reading books is one of the most important ways for teenagers to learn. At the same time, in order to better protect teenagers' eyes, more parents will choose to let their teenagers read paper books. When teenagers encounter unfamiliar words or content that they cannot understand while reading paper books, they need to be answered. With the development of technology, products such as point reading machines have emerged. Point reading machines can support teenagers to point out the content in the book that they are confused about. Point reading machines can identify the location of the point reading, identify the corresponding content in the book, and give a corresponding response. For example, if a teenager points out an unfamiliar word, the point reading machine will provide pronunciation and explanation, etc., or if a teenager points out an unsolvable equation, the point reading machine will provide the solution, etc.
[0003] In related technologies, some point-reading machines must use matching hardware equipment for point reading, such as setting up a magnetic device through customized books and pre-buried point-reading positions, and using a matching point-reading pen to read the corresponding position to trigger the corresponding content. This method has great limitations. Summary of the Invention
[0004] Based on this, it is necessary to provide a point reading position identification method, device, point reading equipment and storage medium that can support a variety of point reading methods to address the above technical problems.
[0005] A point reading position identification method, the method comprising:
[0006] Obtaining the image to be identified in the point reading area;
[0007] Performing target detection on the image to be identified to obtain a target detection result;
[0008] If it is determined according to the target detection result that the image to be identified contains a first preset target or a second preset target, outputting specific position information in the first preset target or the second preset target, where the specific position represents a target position specified in the preset target;
[0009] If it is determined according to the target detection result that the image to be identified includes a first preset target and a second preset target, specific position information of the preset target with a higher priority is output according to the preset priority of the first preset target and the second preset target.
[0010] A point reading position identification device, the device comprising:
[0011] An image acquisition module is used to acquire the image to be identified in the point reading area;
[0012] A target detection module is used to perform target detection on the image to be identified and obtain a target detection result;
[0013] A position information output module is used to output specific position information in the first preset target or the second preset target if it is determined according to the target detection result that the image to be identified contains the first preset target or the second preset target, where the specific position represents the target position specified in the preset target; if it is determined according to the target detection result that the image to be identified contains the first preset target and the second preset target, output specific position information in the preset target with a higher priority according to the priority of the preset first preset target and the second preset target.
[0014] A point reading device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0015] Obtaining the image to be identified in the point reading area;
[0016] Performing target detection on the image to be identified to obtain a target detection result;
[0017] If it is determined according to the target detection result that the image to be identified contains a first preset target or a second preset target, outputting specific position information in the first preset target or the second preset target, where the specific position represents a target position specified in the preset target;
[0018] If it is determined according to the target detection result that the image to be identified includes a first preset target and a second preset target, specific position information of the preset target with a higher priority is output according to the preset priority of the first preset target and the second preset target.
[0019] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0020] Obtaining the image to be identified in the point reading area;
[0021] Performing target detection on the image to be identified to obtain a target detection result;
[0022] If it is determined according to the target detection result that the image to be identified contains a first preset target or a second preset target, outputting specific position information in the first preset target or the second preset target, where the specific position represents a target position specified in the preset target;
[0023] If it is determined according to the target detection result that the image to be identified includes a first preset target and a second preset target, the priorities of the first preset target and the second preset target are read and specific position information of the preset target with a higher priority is output.
[0024] The aforementioned point-reading location recognition method, apparatus, point-reading device, and storage medium capture an image to be identified within a point-reading area and perform target detection on the image to be identified, wherein target detection includes detection of a first preset target and a second preset target. Based on the target detection results, if only one preset target is included, the location of a specific position within that preset target is output; if two preset targets are included, the specific location information within the preset target with a higher priority is output and determined as the point-reading location. The aforementioned method determines the user's indicated location by performing target detection on the image. Point-reading can be achieved without a matching point-reading pen, and supports specifying the point-reading location using two targets, such as finger-pointing or pen-pointing, enriching the supported point-reading methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 1 is a flow chart of a method for identifying a point-reading position in one embodiment;
[0026] Figure 2 1 is a flow chart of a method for identifying a point-reading position in one embodiment;
[0027] Figure 3 A schematic diagram of a process for inputting image features into two or more attribute prediction branches and obtaining attribute prediction results output by each attribute prediction branch in one embodiment;
[0028] Figure 4 A schematic diagram of the network structure of a point-reading position recognition model in a specific embodiment;
[0029] Figure 5 (1) is a schematic diagram of the center position of the preset target (hand) in a specific embodiment;
[0030] Figure 5 (2) is a heat map corresponding to the preset target center position in a specific embodiment;
[0031] Figure 6 (1) is a schematic diagram of a specific position (finger tip) of a preset target in a specific embodiment;
[0032] Figure 6 (2) is a heat map corresponding to the position of the fingertip or pen tip in a specific embodiment;
[0033] FIG7 (1) is a schematic diagram of the corresponding text content recognized and output by the point reading device according to the coordinate position of the point reading target in a specific embodiment;
[0034] FIG7 (2) is a schematic diagram of the corresponding text content identified and output by the point reading device according to the coordinate position of the point reading target in another specific embodiment;
[0035] Figure 8 A structural block diagram of a point reading position identification device in one embodiment;
[0036] Figure 9 FIG. 1 is a diagram showing the internal structure of a point-to-point reading device in one embodiment. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0038] In some embodiments, the point reading position recognition method provided in the present application can be applied to a point reading device, which obtains an image to be recognized in the point reading area and performs target detection on the image to be recognized, wherein the target detection includes the detection of a first preset target and a second preset target. In the result of the target detection, if only one preset target is included, the position of a specific position in the preset target is output; if two preset targets are included, the specific position information in the preset target with a higher priority is output and determined as the point reading position. Subsequently, the text content is recognized based on the position of this specific position and output to achieve the purpose of point reading.
[0039] In other embodiments, the point reading location recognition method provided in this application can be applied to a system including a point reading device and a server. The point reading device communicates with the server via a network. The server obtains an image to be recognized within the point reading area from the point reading device and performs target detection on the image to be recognized, wherein the target detection includes detection of a first preset target and a second preset target. Based on the target detection results, if only one preset target is included, the location of a specific location within the preset target is output; if two preset targets are included, the specific location information within the preset target with a higher priority is output and determined as the point reading location. Finally, the determined point reading location is fed back to the point reading device so that the point reading device can recognize the text content based on the location of this specific location and output it, thereby achieving the purpose of point reading. The server can be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The point reading device can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc. with image acquisition function, but is not limited to these. The terminal and the server can be directly or indirectly connected through wired or wireless communication, and this application does not impose any restrictions here.
[0040] Cloud computing refers to the delivery and usage model of IT infrastructure, enabling on-demand, scalable access to required resources over the internet. In a broader sense, cloud computing refers to the delivery and usage model of services, enabling on-demand, scalable access to required services over the internet. These services can be IT-related, software-related, internet-related, or other services. Cloud computing is the product of the convergence of traditional computer and network technologies, including grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing.
[0041] Cloud computing has rapidly grown, driven by the internet, real-time data streams, the diversification of connected devices, and the growing demand for search services, social networks, mobile commerce, and open collaboration. Unlike previous parallel and distributed computing approaches, the emergence of cloud computing will fundamentally revolutionize the entire internet and enterprise management model.
[0042] In some embodiments of the present application, it involves using computer vision to identify whether a specific target appears in the collected image. Computer vision technology (Computer Vision, CV) Computer vision is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes to identify and measure targets, and further perform graphic processing so that the computer processing becomes an image that is more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0043] In some embodiments of the present application, it is also involved to realize image feature extraction, attribute prediction and the like using neural networks, and neural networks belong to machine learning. Machine Learning (ML) is a multi-disciplinary interdisciplinary subject involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specializes in studying how computers simulate or realize human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and formula-based learning.
[0044] In one embodiment, Figure 1 As shown, a point reading position recognition method is provided, including steps S110 to S140.
[0045] Step S110: Acquire the image to be recognized in the point reading area.
[0046] The "point reading area" refers to the image acquisition area of the point reading device, typically the area covered by the image acquisition module of the point reading device. For example, in one embodiment, when a mobile point reading device is placed on a table, the area covered by the image acquisition module of the point reading device is the point reading area. In this embodiment, the image captured within the point reading area is recorded as the image to be recognized.
[0047] In one embodiment, the above method is applied to a system of a server and a point reading device, and obtaining the image to be identified in the point reading area is that the server obtains the image captured by the image acquisition module of the point reading device from the point reading device; and when the above method is applied to the point reading device, obtaining the image to be identified in the point reading area is that the point reading device obtains the image captured by the image acquisition module.
[0048] Furthermore, in one embodiment, during the operation of the point reading device, an image within the point reading area is acquired at a preset time interval, wherein the preset time period can be set to any length according to actual conditions.
[0049] Step S120: performing target detection on the image to be recognized to obtain a target detection result.
[0050] Object detection, also known as object extraction, is a type of image segmentation based on the geometric and statistical characteristics of the target. In this embodiment, after acquiring the image to be identified, object detection is performed in the image to be identified. Furthermore, in this embodiment, object detection includes detecting whether the image to be identified contains the first preset target and / or the second preset target.
[0051] In one embodiment, target detection of the first preset target and the second preset target in the image to be identified can be achieved in any manner. For example, in one embodiment, two target detection models are trained separately, one of which is used to detect the first preset target and the other is used to detect the second preset target. In another embodiment, a model is trained that includes a feature extraction part and an attribute prediction part, wherein the branch in the attribute prediction part that is used to detect whether the preset target is included includes two channels, one for detecting the first preset target and the other for detecting the second preset target. In other embodiments, it can also be achieved in other ways.
[0052] Step S130 , if it is determined according to the target detection result that the image to be identified contains the first preset target or the second preset target, output specific position information of the first preset target or the second preset target, where the specific position represents the target position specified in the preset target.
[0053] Step S140 , if it is determined according to the target detection result that the image to be identified includes the first preset target and the second preset target, specific position information of the preset target with the higher priority is output according to the priority of the first preset target and the second preset target.
[0054] According to the target detection results, it can be determined whether the image to be identified contains the first preset target or the second preset target. If it is detected that only one preset target is contained, the specific position information of the preset target is output; if it is determined that the image to be identified contains only the first preset target, the specific position information of the first preset target is output; if it is determined that only the second preset target is contained, the specific position information of the second preset target is output.
[0055] If the target detection result determines that the image to be identified contains both the first preset target and the second preset target, the specific location information of the preset target with the higher priority is output based on the preset priority of the first preset target and the second preset target. The priority of the first preset target and the second preset target can be preset according to actual conditions.
[0056] Among them, the specific position in the preset target represents the position set in advance in the preset target. In actual applications, when a user faces text content that he cannot understand, he may point his fingertips, pen tips, etc. to the position where he needs to read, that is, the position that contacts the text to be recognized. When the reading device recognizes the reading position, it is necessary to identify the position of the fingertips, pen tips, etc. In this embodiment, the preset target such as the hand (such as the arm, palm) or pen is first recognized, and then the position of the fingertips, pen tips, etc. is determined and output. In one embodiment, the specific position of the first preset target and the specific position of the second preset target can be set according to actual conditions. For example, if the first preset target is a palm, the specific position of the first preset target is set to a finger, if the second preset target is a pen, the specific position of the second preset target is the pen tip, and so on.
[0057] In another embodiment, if it is determined based on the target detection result that the first preset target or the second preset target is not detected in the image to be identified, it is possible that the current user has not performed a point reading operation, and no position information is output at this time.
[0058] In a specific embodiment, the first preset target represents the user's hand (such as an arm or palm), and the second preset target represents a pen (such as a signature pen, a capacitive pen, or any other pen). Assuming that the priority of the second preset target is higher than that of the first preset target, if only the user's hand is detected in the image to be identified, the specific position information of the hand is output; if only the pen is detected in the image to be identified, the position information of the pen is output; if both the hand and the pen are detected, the specific position information of the pen of the second preset target is output based on the priority. In other embodiments, the preset targets can also be other targets. In this embodiment, the first preset target is a hand and the second preset target is a pen as an example. In the scenario of writing with a pen, the user may point out the text content he wants to understand with the tip of his finger or the tip of the pen. The above method supports both hand and pen reading. In this scenario, the user does not need to deliberately choose a fixed method to read, but can switch between fingertip and pen reading at will, reducing the tedious operation of using a fixed method to read, reducing the user's operating difficulty, and improving the user experience.
[0059] Furthermore, in one embodiment, when outputting the position information of a preset target, the position information of the contact position between the preset target and the text content to be recognized is output. For example, in a specific embodiment, the user points the tip of a finger at the text content to be recognized, that is, the position of the fingertip is the position corresponding to the text content to be recognized that needs to be output; or the user points the tip of a pen at the text content to be recognized, and the position information of the pen tip is output.
[0060] The above-mentioned point reading location recognition method obtains an image to be identified within the point reading area and performs target detection on the image to be identified, wherein the target detection includes detection of a first preset target and a second preset target. Based on the target detection results, if only one preset target is included, the location of a specific position within the preset target is output; if two preset targets are included, the specific location information of the preset target with a higher priority is output and determined as the point reading location. The above-mentioned method determines the user's pointed location by performing target detection on the image. Point reading can be achieved without the supporting point reading pen, and supports specifying the point reading location through two targets, such as finger reading or pen tip reading, enriching the supported point reading methods.
[0061] In one embodiment, target detection is performed on an image to be identified to obtain target detection results, including: extracting image features of the image to be identified; inputting the image features into two or more attribute prediction branches respectively, and obtaining attribute prediction results output by each attribute prediction branch; wherein the attribute prediction branches include: a center position prediction branch of a preset target and a specific position prediction branch within the preset target.
[0062] Image features can be divided into two levels, including low-level visual features and high-level semantic features. Low-level visual features include texture, color, and shape. Semantic features are the relationship between things. Texture feature extraction algorithms include: gray-level co-occurrence matrix method, Fourier power spectrum method; color feature extraction algorithms include: histogram method, cumulative histogram method, color clustering method, etc.; shape feature extraction algorithms include: spatial moment features, etc.; high-level semantic extraction: semantic network, mathematical logic, framework and other methods. In other embodiments, image feature extraction can also be achieved through neural networks.
[0063] In one embodiment, extracting image features of an image to be identified includes: extracting image features of the image to be identified through a feature extraction network; the feature extraction network includes continuous downsampling layers and continuous upsampling layers; wherein the number of downsampling layers is a first preset number, the number of upsampling layers is a second preset number, and the second preset number is less than or equal to the first preset number.
[0064] In machine learning, pattern recognition, and image processing, feature extraction begins with an initial set of measured data and creates derived values (features) designed to be informative and non-redundant, thereby facilitating subsequent learning and generalization steps and, in some cases, leading to better interpretability. In this embodiment, extracting image features of the image to be recognized can be achieved in any of a variety of ways.
[0065] Downsampling, also known as downsampling, has two main purposes: 1. Make the image fit the size of the display area; 2. Generate a thumbnail of the corresponding image. Downsampling principle: For an image I with a size of M*N, it is downsampled s times, that is, a resolution image of (M / s)*(N / s) size is obtained. Of course, s should be a common divisor of M and N. If the image is in matrix form, the image in the s*s window of the original image is converted into a pixel, and the value of this pixel is the mean of all pixels in the window. In a specific embodiment, the image to be recognized is downsampled using continuous downsampling layers, and the downsampling factor used by the downsampling layer is 2.
[0066] Upsampling, also known as image interpolation, is mainly intended to amplify the original image so that it can be displayed on a higher resolution display device. Upsampling principle: Image enlargement almost always adopts the interpolation method, that is, new elements are inserted between pixels based on the original image pixels using a suitable interpolation algorithm. Among them, commonly used methods for image interpolation include traditional image interpolation methods, edge-based image interpolation algorithms, and region-based image interpolation algorithms, etc. In a specific embodiment, continuous upsampling layers are used to upsample the minimum-sized downsampled image features, wherein the sampling multiple of the upsampling layer is 2, that is, the image features output by the upsampling layer are twice the size of the input image features.
[0067] Furthermore, in one embodiment, extracting image features of the image to be identified through a feature extraction network includes: downsampling the image to be identified layer by layer to obtain downsampled image features of different scales; wherein the number of downsampling layers is a first preset number; and upsampling each downsampled image feature layer by layer through a feature pyramid network to obtain image features of the image to be identified; wherein the number of upsampling layers is a second preset number, which is less than or equal to the first preset number. In one specific embodiment, the downsampling layers include five layers and the upsampling layers include three layers, i.e., when the image to be identified is input into the feature extraction network, the size of the output image features is 1 / 4 of the size of the image to be identified.
[0068] In one embodiment, before inputting the image to be identified into the feature extraction network, the process further includes adjusting the size of the image to be identified to a preset size. The preset size can be set to any value based on actual conditions, for example, the preset size is set to 320*320. Adjusting the image size can be achieved in any manner.
[0069] Furthermore, in one embodiment, each downsampled image feature is upsampled layer by layer through a feature pyramid network, including: using the downsampled image feature of the minimum size as the initial upsampled image feature, splicing the downsampled image feature of the minimum size with the initial upsampled image feature and inputting the resultant into the first upsampled layer to obtain the first upsampled image feature; splicing the first upsampled image feature with the downsampled image feature of the same size and inputting the resultant into the second upsampled layer to obtain the second upsampled image feature; splicing the second upsampled image feature with the downsampled image feature of the same size and inputting the resultant into the third upsampled layer to obtain the third upsampled image feature, which is the image feature of the image to be identified.
[0070] In another embodiment, the feature channel of the feature extraction network is N, which can be adjusted according to the device's requirements for model running performance to achieve a balance between accuracy and running speed.
[0071] In the above embodiment, image features are extracted through a pre-trained neural network. After training, the neural network can extract relatively accurate image features for target detection.
[0072] In one embodiment, after extracting image features, the image features are used to detect whether a preset target exists. In this embodiment, attribute prediction is performed on the extracted image features using attribute prediction branches, wherein the attribute prediction branches include at least a branch for predicting the center position of the preset target and a branch for predicting a specific position within the preset target. The output result of the branch for predicting the center position of the preset target is the position information of the center position of the preset target, and the output result of the branch for predicting the specific position within the preset target is the specific position information within the preset target. It can be understood that in this embodiment, the target detection result includes the position information of the center position of the preset target and the specific position information of the preset target.
[0073] In one embodiment, the specific position prediction branch includes a specific position prediction sub-branch and a specific position drift prediction sub-branch; Figure 2 As shown, the image features are input into two or more attribute prediction branches respectively, and the attribute prediction results output by each attribute prediction branch are obtained, including steps S210 to S230.
[0074] In step S210 , the image features are input into the center position prediction branch of the preset target to obtain a heat map corresponding to the center position of the first preset target and a heat map corresponding to the center position of the second preset target.
[0075] A heatmap can use color changes to reflect data information in a two-dimensional matrix or table. It can intuitively represent the size of the data value with a defined color depth. In this embodiment, after obtaining the center position of the preset target, it is converted into the form of a heatmap for representation. In one embodiment, the heatmap is specifically a Gaussian heatmap, and the two-dimensional coordinates are transformed through a Gaussian transformation to obtain the corresponding Gaussian heatmap. Among them, converting the coordinates into a Gaussian heatmap can be achieved by any method.
[0076] The center position of the target refers to the center position of the target bounding box generated when the target is detected. The target bounding box is a box that encloses the target, typically a rectangular box. For example, when an arm is detected, the bounding box encloses the arm; when a pen is detected, the bounding box encloses the pen.
[0077] In one embodiment, the target center prediction branch includes two channels, one of which is used to output a heat map of the center position of the first target, and the other is used to output a heat map of the center position of the second target. In a specific embodiment, the target center prediction branch includes a 3*3 convolution layer and a 1*1 convolution layer, wherein attribute learning is performed through the 3*3 convolution layer, and the number of channels is adjusted to the dimension of the target attribute through the 1*1 convolution layer.
[0078] In step S220 , the image features are input into the specific position prediction sub-branch to obtain a heat map corresponding to the specific position in the first preset target and a heat map corresponding to the specific position in the second preset target.
[0079] Similar to the center position prediction branch, in this step, the specific position in the preset target is also converted into a heat map for representation.
[0080] In one embodiment, the specific location prediction sub-branch includes two channels, one of which is used to output a specific location heat map of a first preset target, and the other channel is used to output a specific location heat map of a second preset target. In a specific embodiment, the specific location prediction sub-branch includes a 3*3 convolution layer and a 1*1 convolution layer, wherein attribute learning is performed through the 3*3 convolution layer, and the number of channels is adjusted to the dimension of the target attribute through the 1*1 convolution.
[0081] In step S230 , the image features are input into the specific position drift prediction sub-branch to obtain the specific position drift when the image to be identified is converted into a heat map.
[0082] Since only integer coordinates of pixels can be represented in a heat map, conversion errors may occur when converting the coordinate positions in the original image to the representation in the heat map. Therefore, in this embodiment, a specific position drift prediction sub-branch is used to predict the conversion error of the coordinate position from float to int when converting the image to be identified to the heat map. In a specific embodiment, the image feature size input into the specific position drift prediction sub-branch is 1 / 4 of the size of the image to be identified. Assuming that the pixel coordinates in the image to be identified are (120, 121), the coordinates of the coordinates in the image feature are (30, 30.25), which are represented as positive integer coordinates (30, 30) in the heat map. In this embodiment, the position drift is (0, 0.25).
[0083] In one embodiment, the specific position drift prediction sub-branch includes two channels; if it is detected that the image to be identified contains both the first preset target and the second preset target, the horizontal coordinate value and the vertical coordinate value of the specific position of the two preset targets need to be output. Therefore, in this embodiment, the four channels are used to output the position drift of the horizontal and vertical coordinate values of the specific positions of the two preset targets, respectively. In a specific embodiment, the specific position drift prediction sub-branch includes a 3*3 convolution layer and a 1*1 convolution layer, wherein attribute learning is performed through the 3*3 convolution layer, and the number of channels is adjusted to the dimension of the target attribute through the 1*1 convolution.
[0084] In this embodiment, three attribute prediction branches are used to predict the center positions of the first preset target and the second preset target, the specific positions in the first preset target and the second preset target, and the position drift of the specific positions in the first preset target and the second preset target respectively; wherein, according to the output result of the center position prediction branch of the preset target, it can be determined whether the first preset target and / or the second preset target are detected. If it is determined that the first preset target and / or the second preset target are detected, the preset target to be output is determined (if only the first preset target is detected, the first preset target is output; if only the second preset target is detected, the second preset target is output; if both the first preset target and the second preset target are detected, the preset target with higher priority is output), and then the specific position in the preset target to be output is determined according to the output results of the specific position prediction sub-branch and the specific position drift prediction sub-branch.
[0085] In one embodiment, Figure 3 As shown, after performing target detection on the image to be recognized and obtaining the target detection result, steps S310 to S330 are also included.
[0086] Step S310 traverses the heat maps corresponding to the center positions output by the center position prediction branch of the preset targets, and determines the predicted position confidence of the first preset target center position and the predicted position confidence of the second preset target center position.
[0087] In one embodiment, by traversing each channel of the center position prediction branch of the preset target, a heat map output for the center positions of the two preset targets can be obtained. In one embodiment, the predicted positions output by any channel of the center position prediction branch of the preset target may include multiple ones, so the confidence corresponding to each predicted position is read separately, that is, in this embodiment, the predicted position confidence of the first preset target center position and the predicted position confidence of the second preset target center position. Confidence is also called reliability, or confidence level, confidence coefficient. The corresponding probability that the estimated value and the overall parameter are within a certain allowable error range is large. This corresponding probability is called confidence. The greater the confidence, the more likely the target value is close to the correct value.
[0088] In step S320 , if the maximum value of the confidence level of the predicted position of the center position of the first preset target is greater than or equal to the preset threshold, it is determined that the image to be identified contains the first preset target.
[0089] In step S330 , if the maximum value of the confidence level of the predicted position of the center position of the second preset target is greater than or equal to the preset threshold, it is determined that the image to be identified contains the second preset target.
[0090] In this embodiment, a preset threshold is set in advance and the preset threshold is compared with the confidence level. Furthermore, in this embodiment, if the confidence level of the center position is greater than the preset threshold, it means that the corresponding preset target is detected in the image to be identified. It is understandable that in another embodiment, if the maximum value of the predicted position confidence level of the center position of the first preset target is less than the preset threshold, it is determined that the image to be identified does not contain the first preset target. If the maximum value of the predicted position confidence level of the center position of the second preset target is less than the preset threshold, it is determined that the image to be identified does not contain the second preset target. The preset threshold can be set according to actual conditions, for example, the preset threshold can be set to 80%, 90%, etc.
[0091] In this embodiment, by comparing the confidence level in the output result of the preset target center position prediction branch with a preset threshold, it is determined whether the image to be identified contains the first preset target and / or the second preset target, thereby improving the accuracy of target detection.
[0092] In one embodiment, the above method also includes: if it is determined according to the target detection result that the image to be identified contains the first preset target and / or the second preset target, the specific position is read according to the thermal map corresponding to the specific position in the preset target, and according to the drift of the specific position, the specific position in the preset target is restored to the same original coordinates as the image to be identified to obtain the specific position information corresponding to the preset target.
[0093] In this embodiment, after determining that the image to be identified contains the first preset target and / or the second preset target, the coordinate information of the specific position corresponding to the preset target needs to be output. According to the specific position prediction sub-branch, the coordinates of the specific position in the heat map can be known. According to the specific position drift prediction sub-branch, the position drift of the specific position when converted to the heat map can be known. Then, the coordinates of the specific position in the heat map are restored to the coordinate information of the same size as the image to be identified, thereby determining the position of the position in the image to be identified, and subsequently identifying the text content at the position to achieve the purpose of point reading.
[0094] In one embodiment, the attribute prediction branch also includes: a size prediction branch of a preset target, and a distance prediction branch between the center position of the preset target and a specific position; when training the feature extraction network and each attribute prediction branch, the parameters of the feature extraction network and each attribute prediction branch are adjusted based on the sample prediction results output by each attribute prediction branch.
[0095] In one embodiment, the target size prediction branch is used to output the size of the target's bounding box; further, the target size prediction branch is used to output the width and height of the target's bounding box. The target center-to-specific location prediction branch is used to output the distance between the target's center and the specific location; for example, the distance between a finger and the center of an arm's bounding box, or the distance between a pen tip and the center of a pen's bounding box.
[0096] Among them, training the neural network can be achieved in any way. When training the feature extraction network and each attribute prediction branch, training is performed based on sample data carrying labeled data, and the sample data is input into the preset neural network framework (including the feature extraction network and each attribute prediction branch), and the sample prediction results output by the five attribute prediction branches are output, including the corresponding heat map of the center position of the preset target of the sample image, the corresponding heat map of the specific position in the preset target of the sample image, the drift of the specific position in the preset target of the sample image, the bounding box size of the preset target of the sample image, and the distance between the specific position and the center position of the preset target of the sample image. Then, based on the sample prediction results output by each attribute prediction branch, the parameters of the feature extraction network and each attribute prediction branch are adjusted, and the training is stopped when the termination condition is reached, to obtain a neural network determined by training, including the feature extraction network and each attribute prediction branch.
[0097] In one embodiment, the sample data for training the neural network includes at least: sample data containing only the first preset target, sample images containing only the second preset target, and sample images containing both the first preset target and the second preset target.
[0098] In a specific embodiment, the structures of the size prediction branch of the preset target and the distance prediction branch between the preset target center position and the specific position are similar, both including a 3*3 convolution layer and a 1*1 convolution layer, wherein attribute learning is performed through the 3*3 convolution layer, and the number of channels is adjusted to the dimension of the target attribute through the 1*1 convolution.
[0099] In this embodiment, when training the network, the output results of the preset target size prediction branch and the distance prediction branch between the preset target center position and the specific position are also combined for training. These are effective supervision information in the training stage, which is beneficial to the training of the overall task.
[0100] In one embodiment, features of the image to be identified are extracted using a pre-trained feature extraction network, and attributes of the image features are predicted using a pre-trained attribute prediction branch network. In one specific embodiment, the feature extraction network uses MobileNetV2, and the attribute prediction branch uses a convolutional network to predict different attributes. In other embodiments, other networks may be used for the feature extraction network and the attribute prediction branch network.
[0101] The present application also provides an application scenario, which applies the above-mentioned point reading position recognition method.
[0102] Specifically, the application of the point reading position recognition method in this application scenario is as follows:
[0103] In this embodiment, the feature extraction network and each output attribute prediction branch are collectively referred to as a point reading position recognition model. Figure 4 Shown is a schematic diagram of the network structure of a point-reading position recognition model in a specific embodiment.
[0104] 1. Obtain the image to be recognized within the point reading area, resize the image to be recognized to 320*320, and use it as the input of the model.
[0105] 2. The point reading location recognition model is divided into two parts: a) Backbone network feature extraction part <backbonefeature>, b) Attribute prediction part<Attribute Predict> .
[0106] a) The backbone network feature extraction part uses MobileNetV2 as the backbone. The last three layers are then upsampled using FPN (feature pyramid networks) and feature fusion is performed to obtain image features. The size of the image features is 1 / 4 the size of the input image to be recognized (80*80), and the number of feature channels is N. This can be adjusted to achieve a balance between accuracy and speed based on the device's performance requirements for the model.
[0107] b) The attribute prediction part uses a convolutional network to predict different attributes, focusing on detecting two pre-set targets: hands and pens. It also regresses the coordinates of the fingertips and pen tips. From the feature map extracted in step 1, each attribute branch first learns the attribute through a 3x3 convolution, and then uses a 1x1 convolution to adjust the number of channels to the dimension of the target attribute. The attributes include the following:
[0108] i. Center (N = 2, channel 2), predict the position heatmap of the preset target center point. For the bounding box (BoundingBox) of the annotated data hand and pen, calculate the coordinates of its center point and convert it into a Gaussian heatmap with the center point coordinates as the center. As shown in Figure 5 (1), it is a schematic diagram of the center position of the preset target (hand) in a specific embodiment, and as shown in Figure 5 (2), it is a heatmap corresponding to the center position of the preset target in a specific embodiment.
[0109] ii. W / H (N=2, channel 2), the height and width of the BoundingBox of the hand or pen, and the width and height of the preset target if one exists for each point.
[0110] iii. Relation (N=4, channel 4), the coordinate difference between the fingertip or pen tip coordinates and the center point coordinates, that is, the coordinate value of the fingertip or pen tip relative to the center of the hand or pen, used to control the relative relationship between the target point and the center point
[0111] iv.Keypoints (N=2, channel 2), coordinate heat map of the fingertip or pen tip (same as Center operation), as shown in Figure 6 (1), a schematic diagram of the specific position (fingertip) of the preset target in a specific embodiment, and as shown in Figure 6 (2), a heat map corresponding to the position of the fingertip or pen tip in a specific embodiment.
[0112] v.Offset (N=4, 4 channels), fingertip or pen tip coordinate drift. Offset is a floating-point number after 1 / 4 downsampling from the original image coordinates. However, it can only represent integer coordinates on the heat map. Offset is used to learn the conversion error from float to int.
[0113] When training the point-reading location recognition model, two types of data are collected for model training: a) fingertip gesture data alone, excluding pen data, and b) pen-holding hand data, including both pen-holding and non-pen-holding hand data. During the training phase, object detection utilizes both data from the ab and bb libraries, while training target heatmap data is supplemented with empty heatmaps to enhance the semantic distinction between fingertip gestures, standard gestures, and pen-holding gestures.
[0114] 3. After the model is trained and integrated into the SDK (Software Development Kit), the following point reading coordinate solution method is used. During the model operation phase, only the Center, Keypoints, and Offset branches are used. The steps are as follows:
[0115] a) Traverse the two channels of Center, where channel 0 represents the hand target and channel 1 represents the pen target. Find the position with the maximum value in each channel. If the value is greater than the set threshold A, it means there is a valid hand or pen at that position. Otherwise, there is no preset target in the current frame.
[0116] b) Traverse the two channels of Keypoints, where channel 0 represents the position of the fingertip and channel 1 represents the position of the pen tip, and find the position with the largest value in each channel. If the value is greater than the set threshold B, it means that there is a valid fingertip or pen tip at that position. Otherwise, there is no specific position of the preset target in the current frame.
[0117] c) The coordinate position found in b is converted to the original coordinate by the value of Offset.
[0118] d) Based on the results of a) and b), if only the pen tip exists, the position coordinates of the pen tip are returned; if only the fingertip exists, the position coordinates of the fingertip are returned; if both the fingertip and the pen tip exist, the position coordinates of the pen tip with higher priority are returned based on the priority of the fingertip and the pen tip.
[0119] Furthermore, after the point reading target coordinate position is output in the above embodiment, the corresponding text content is identified based on the point reading target coordinate position to achieve the purpose of point reading; as shown in Figures 7(1) and 7(2), there are schematic diagrams of the corresponding text content identified and output by the point reading device based on the point reading target coordinate position.
[0120] The above-mentioned point-reading position recognition method uses a multi-task learning method of a deep learning model to simultaneously detect the hand and pen with target detection as the main body, and simultaneously regresses the coordinates of the fingertip and pen tip, realizing the fusion recognition of the fingertip and pen tip in a single model, and can effectively solve the semantic confusion problem between ordinary gestures and point-reading gestures; and by setting the pen tip priority, it can simultaneously support fingertip recognition and pen tip recognition. When both the fingertip and pen tip are present, the pen tip is given priority by setting the pen tip to have a high priority. The above-mentioned method can realize a seamless interactive process in the point-reading scenario, and there is no interruption in the operation between the user writing with a pen and reading with a point, which greatly improves the learning efficiency and immersive experience. Without changing the camera-based point-reading device, a point-reading interaction method that integrates the pen tip and fingertip is provided, realizing product functions that can be used for both fingertip and pen tip reading, and greatly improving the user experience.
[0121] In one embodiment, the present application also provides a point reading method, comprising the steps of: acquiring an image to be identified in a point reading area, identifying corresponding text content based on specific position information of a preset target in the image to be identified, and displaying the text content on a display screen. The determination of the text content corresponding to the specific position information of the preset target in the image to be identified comprises the steps of: performing target detection on the image to be identified to obtain a target detection result; if it is determined according to the target detection result that the image to be identified contains a first preset target or a second preset target, outputting specific position information in the first preset target or the second preset target, wherein the specific position represents a target position specified in the preset target; if it is determined according to the target detection result that the image to be identified contains a first preset target and a second preset target, outputting specific position information in the preset target with a higher priority based on the preset priority of the first preset target and the second preset target.
[0122] For the specific embodiments of the above-mentioned point reading method, please refer to the above embodiments of the point reading position recognition method, which will not be repeated here.
[0123] It should be understood that, although each step in each flow chart involved in the above-described embodiment is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless clearly stated herein, the execution of these steps does not have strict order restrictions, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each flow chart involved in the above-described embodiment can include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps.
[0124] In one embodiment, Figure 8 As shown, a point reading position recognition device is provided. The device can be a software module or a hardware module, or a combination of the two to form a part of the point reading device. The device specifically includes: an image acquisition module 810, a target detection module 820 and a position information output module 830, wherein:
[0125] The image acquisition module 810 is used to acquire the image to be recognized in the point reading area;
[0126] The target detection module 820 is used to perform target detection on the image to be identified and obtain a target detection result;
[0127] The position information output module 830 is used to output specific position information of the first preset target or the second preset target if it is determined according to the target detection result that the image to be identified contains the first preset target or the second preset target, and the specific position represents the target position specified in the preset target; if it is determined according to the target detection result that the image to be identified contains the first preset target and the second preset target, the specific position information of the preset target with a higher priority is output according to the priority of the preset first preset target and the second preset target.
[0128] The aforementioned point-reading position recognition device captures an image to be recognized within a point-reading area and performs target detection on the image. Target detection includes detecting a first preset target and a second preset target. If only one preset target is included in the target detection result, the specific location within that preset target is output. If two preset targets are included, the specific location information within the preset target with a higher priority is output and determined as the point-reading position. The aforementioned device determines the user's indicated location by performing target detection on the image. Point-reading can be achieved without a matching point-reading pen, and supports specifying the point-reading position using two targets, such as finger-pointing or pen-tip-pointing, enriching the supported point-reading methods.
[0129] In one embodiment, the target detection module 820 of the above-mentioned device includes: a feature extraction unit, which extracts image features of the image to be identified; an attribute prediction unit, which is used to input the image features into two or more attribute prediction branches respectively, and obtain the attribute prediction results output by each attribute prediction branch; wherein the attribute prediction branch includes: a center position prediction branch of a preset target and a specific position prediction branch in the preset target.
[0130] In one embodiment, the feature extraction unit of the above-mentioned device is also used to: extract image features of the image to be identified through a feature extraction network; the feature extraction network includes continuous downsampling layers and continuous upsampling layers; wherein the number of downsampling layers is a first preset number, the number of upsampling layers is a second preset number, and the second preset number is less than or equal to the first preset number.
[0131] In one embodiment, the specific position prediction branch includes a specific position prediction sub-branch and a specific position drift prediction sub-branch; in this embodiment, the attribute prediction unit of the above-mentioned device includes: a center position heat map prediction sub-unit, which is used to input image features into the center position prediction branch of the preset target to obtain a heat map corresponding to the center position of the first preset target and a heat map corresponding to the center position of the second preset target; a specific position heat map prediction sub-unit, which is used to input image features into the specific position prediction sub-branch to obtain a heat map corresponding to the specific position in the first preset target and a heat map corresponding to the specific position in the second preset target; a specific position drift prediction sub-unit, which is used to input image features into the specific position drift prediction sub-branch to obtain a specific position drift when converting the image to be identified into a heat map.
[0132] In one embodiment, the above-mentioned device also includes a confidence reading module, which is used to traverse the heat map corresponding to each center position output by the center position prediction branch of the preset target, and determine the predicted position confidence of the center position of the first preset target and the predicted position confidence of the center position of the second preset target; in this embodiment, the target detection module 820 is also used to: if the maximum value of the predicted position confidence of the center position of the first preset target is greater than or equal to the preset threshold, determine that the first preset target is contained in the image to be identified; if the maximum value of the predicted position confidence of the center position of the second preset target is greater than or equal to the preset threshold, determine that the second preset target is contained in the image to be identified.
[0133] In one embodiment, the above-mentioned device also includes: a position restoration module, which is used to read the specific position according to the thermal map corresponding to the specific position in the preset target if it is determined according to the target detection result that the image to be identified contains the first preset target and / or the second preset target, and restore the specific position in the preset target to the same original coordinates as the image to be identified according to the drift of the specific position, so as to obtain the specific position information corresponding to the preset target.
[0134] In one embodiment, the attribute prediction branch also includes: a size prediction branch of a preset target, and a distance prediction branch between the center position of the preset target and a specific position; the above-mentioned device also includes a model training module, which is used to adjust the parameters of the feature extraction network and each attribute prediction branch based on the sample prediction results output by each attribute prediction branch when training the feature extraction network and each attribute prediction branch.
[0135] Specific embodiments of the point-reading position identification device can be found in the embodiments of the point-reading position identification method described above and will not be further described here. Each module in the above-described point-reading position identification device can be implemented in whole or in part via software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the point-reading device in hardware form, or can be stored in the memory of the point-reading device in software form, so that the processor can call and execute the corresponding operations of each module.
[0136] In one embodiment, a point reading device is provided, whose internal structure diagram can be as follows: Figure 9 As shown. The point reading device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the point reading device is used to provide computing and control capabilities. The memory of the point reading device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the point reading device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a point reading position recognition method is implemented. The display screen of the point reading device can be a liquid crystal display screen or an electronic ink display screen. The input device of the point reading device can be a touch layer covering the display screen, or a button, trackball or touchpad, image acquisition device (camera) provided on the housing of the point reading device, or an external keyboard, touchpad or mouse.
[0137] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the point reading device to which the solution of the present application is applied. The specific point reading device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0138] In one embodiment, a point reading device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0139] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.
[0140] In one embodiment, a computer program product or computer program is provided. The computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a point reading device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the point reading device to perform the steps of the above-described method embodiments.
[0141] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0142] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0143] The above embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.< / backbonefeature>
Claims
1. A point reading position recognition method, characterized in that: The method comprises: Obtaining the image to be identified in the point reading area; Extracting image features of the image to be identified, inputting the image features into two or more attribute prediction branches respectively, and obtaining attribute prediction results output by each attribute prediction branch; the attribute prediction branch includes a specific position prediction sub-branch and a specific position drift prediction sub-branch; the attribute prediction result of the specific position prediction sub-branch includes a heat map corresponding to the specific position; the attribute prediction result of the specific position drift prediction sub-branch includes a specific position drift when converting the image to be identified to the heat map; the specific position drift represents a conversion error of the coordinate position in the image when converting the image to be identified to the heat map; If it is determined that the image to be identified contains a first preset target or a second preset target according to the attribute prediction results output by each of the attribute prediction branches, the specific position is read according to the heat map corresponding to the specific position, and the specific position in the preset target is restored to the same original coordinates as the image to be identified according to the drift of the specific position, and the specific position information of the first preset target or the second preset target is output, where the specific position represents the target position specified in the preset target; If it is determined that the image to be identified contains a first preset target and a second preset target based on the attribute prediction results output by each of the attribute prediction branches, specific position information of the preset target with a higher priority is output based on the priority of the first preset target and the second preset target.
2. The point reading position recognition method according to claim 1, characterized in that: The extracting the image features of the image to be identified includes: The image features of the image to be identified are extracted through a feature extraction network; the feature extraction network includes continuous downsampling layers and continuous upsampling layers; wherein the number of downsampling layers is a first preset number, the number of upsampling layers is a second preset number, and the second preset number is less than or equal to the first preset number.
3. The point reading position recognition method according to claim 1, characterized in that: The attribute prediction branch also includes a center position prediction branch of a preset target; Inputting the image features into two or more attribute prediction branches respectively, and obtaining attribute prediction results output by each of the attribute prediction branches, including: Inputting the image features into the center position prediction branch of the preset target to obtain a heat map corresponding to the center position of the first preset target and a heat map corresponding to the center position of the second preset target; Inputting the image features into the specific position prediction sub-branch to obtain a heat map corresponding to the specific position in the first preset target and a heat map corresponding to the specific position in the second preset target; The image features are input into the specific position drift prediction sub-branch to obtain the specific position drift when the image to be identified is converted into a heat map.
4. The point reading position recognition method according to claim 3, characterized in that: After inputting the image features into two or more attribute prediction branches respectively and obtaining the attribute prediction results output by each of the attribute prediction branches, the method further includes: Traversing the heat maps corresponding to the center positions output by the center position prediction branch of the preset target, and determining the predicted position confidence of the first preset target center position and the predicted position confidence of the second preset target center position; If the maximum value of the predicted position confidence of the center position of the first preset target is greater than or equal to a preset threshold, it is determined that the image to be identified contains the first preset target; and If the maximum value of the predicted position confidence of the center position of the second preset target is greater than or equal to a preset threshold, it is determined that the image to be identified contains the second preset target.
5. The point reading position recognition method according to any one of claims 2 to 4, characterized in that: The attribute prediction branch also includes: a size prediction branch of a preset target, and a distance prediction branch between the center position of the preset target and a specific position; When training the feature extraction network and each of the attribute prediction branches, the parameters of the feature extraction network and each of the attribute prediction branches are adjusted based on the sample prediction results output by each of the attribute prediction branches.
6. A point reading position recognition device, characterized in that: The device comprises: An image acquisition module is used to acquire the image to be identified in the point reading area; The target detection module includes: a feature extraction unit for extracting the image features of the image to be identified; an attribute prediction unit for inputting the image features into two or more attribute prediction branches respectively to obtain the attribute prediction results output by each attribute prediction branch; the attribute prediction branch includes a specific position prediction sub-branch and a specific position drift prediction sub-branch; the attribute prediction result of the specific position prediction sub-branch includes a heat map corresponding to a specific position; the attribute prediction result of the specific position drift prediction sub-branch includes a specific position drift when the image to be identified is converted to a heat map; the specific position drift represents the conversion error of the coordinate position in the image when the image to be identified is converted to a heat map; a position information output module is used to output the image features according to each attribute prediction branch. The attribute prediction result output by the attribute prediction branch determines that the image to be identified contains the first preset target or the second preset target, reads the specific position according to the heat map corresponding to the specific position, and restores the specific position in the preset target to the same original coordinates as the image to be identified according to the drift of the specific position, and outputs the specific position information in the first preset target or the second preset target, where the specific position represents the target position specified in the preset target; if it is determined that the image to be identified contains the first preset target and the second preset target according to the attribute prediction results output by each of the attribute prediction branches, the specific position information in the preset target with a higher priority is output according to the preset priority of the first preset target and the second preset target.
7. The device according to claim 6, characterized in that The feature extraction unit is further used to: extract image features of the image to be identified through a feature extraction network; the feature extraction network includes continuous downsampling layers and continuous upsampling layers; wherein the number of downsampling layers is a first preset number, the number of upsampling layers is a second preset number, and the second preset number is less than or equal to the first preset number.
8. The device according to claim 6, characterized in that The specific position prediction branch also includes a specific position prediction sub-branch; the attribute prediction unit is further used to: input the image features into the center position prediction branch of the preset target to obtain a heat map corresponding to the center position of the first preset target and a heat map corresponding to the center position of the second preset target; input the image features into the specific position prediction sub-branch to obtain a heat map corresponding to the specific position in the first preset target and a heat map corresponding to the specific position in the second preset target.
9. The device according to claim 6, characterized in that The device further includes a confidence reading module, configured to: traverse the heat maps corresponding to the center positions output by the center position prediction branch of the preset target, and determine the predicted position confidence of the first preset target center position and the predicted position confidence of the second preset target center position; The target detection module is also used to: if the maximum value of the predicted position confidence of the center position of the first preset target is greater than or equal to a preset threshold, determine that the image to be identified contains the first preset target; and if the maximum value of the predicted position confidence of the center position of the second preset target is greater than or equal to the preset threshold, determine that the image to be identified contains the second preset target.
10. The device according to any one of claims 6 to 9, characterized in that The attribute prediction branch also includes: a size prediction branch of a preset target, and a distance prediction branch between the center position of the preset target and a specific position; the device also includes a model training module, which is used to: when training the feature extraction network and each of the attribute prediction branches, adjust the parameters of the feature extraction network and each of the attribute prediction branches based on the sample prediction results output by each of the attribute prediction branches.
11. A point reading device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method and device for user indication-based reading
CN108536287A
Mechanical arm grabbing detection method based on improved CenterNet
CN111523486A