Size measurement method and device based on deep learning

CN115797433BActive Publication Date: 2026-09-04优尼特克斯公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111081218.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-25
Filing Date
2021-09-15
Publication Date
2026-09-04
Estimated Expiration
2041-09-15

AI Technical Summary

Technical Problem

但是,受环境、物料等因素的影响,定位的可重复性及准确性不高,使得根据定位结果得到的尺寸信息的准确性也不高

Benefits of technology

[0014] The embodiments of this disclosure can acquire images of the target object to be measured according to a preset positioning accuracy to obtain an image to be processed, and determine at least one target region from the image to be processed. Each target region includes at least one positioning point. Then, the at least one target region is processed by a pre-trained neural network to obtain the first position information of each positioning point. Based on the positioning accuracy and the first position information of each positioning point, the size information of the target object is determined. In this way, the repeatability and accuracy of the positioning of the target object during the size measurement process can be improved through deep learning, thereby improving the accuracy of the size measurement result (i.e., the size information of the target object).

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115797433B_ABST
    Figure CN115797433B_ABST
Patent Text Reader

Abstract

The present disclosure provides a deep learning-based size measurement method and device, the method comprising: according to a preset positioning accuracy, image acquisition is performed on a target object to be measured to obtain a to-be-processed image, the positioning accuracy being used to indicate the image acquisition accuracy of a positioning point used when the target object is measured; at least one target region is determined from the to-be-processed image, each target region comprising at least one positioning point; the at least one target region is processed by a pre-trained neural network to obtain first position information of each positioning point; and size information of the target object is determined according to the positioning accuracy and the first position information of each positioning point. The embodiments of the present disclosure can improve the repeatability and accuracy of the positioning of the target object in the size measurement process in a deep learning manner, thereby improving the accuracy of the size measurement result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of industrial vision technology, and in particular to a dimension measurement method and apparatus based on deep learning. Background Technology

[0002] In industrial production processes, it is often necessary to quantitatively measure the dimensions of parts or products. In related technologies, after acquiring images of the parts or products to be measured, template matching is typically used to locate the parts or products in the acquired images, obtaining the location results. Based on these location results, the dimensional information of the parts or products is determined. However, due to the influence of environmental and material factors, the repeatability and accuracy of the location are not high, resulting in low accuracy of the dimensional information obtained from the location results. Summary of the Invention

[0003] In view of this, this disclosure proposes a dimension measurement method and apparatus based on deep learning.

[0004] According to one aspect of this disclosure, a deep learning-based size measurement method is provided. The method includes: acquiring an image of a target object to be measured according to a preset positioning accuracy to obtain an image to be processed, wherein the positioning accuracy indicates the image acquisition accuracy of the positioning points used when measuring the target object; determining at least one target region from the image to be processed, each target region including at least one positioning point; processing the at least one target region through a pre-trained neural network to obtain first position information of each positioning point; and determining the size information of the target object according to the positioning accuracy and the first position information of each positioning point.

[0005] In one possible implementation, the neural network includes at least one sub-network corresponding to the target region. The sub-network includes an encoder and a decoder. The process of processing the at least one target region using the pre-trained neural network to obtain the location information of each positioning point includes: for any target region, extracting features from the target region using the encoder in the sub-network corresponding to the target region to obtain a feature map of the target region; and processing the feature map using the decoder in the sub-network corresponding to the target region to obtain the location information of each positioning point in the target region.

[0006] In one possible implementation, determining at least one target region from the image to be processed includes: when the location range of the positioning point is known and the number of positioning points used when measuring the target object is 1, determining the region in the image to be processed corresponding to the location range of the positioning point as the target region, wherein the location range is used to indicate the minimum area where the positioning point appears within the field of view, and the field of view is used to indicate the area range actually captured by the image acquisition module when capturing images of the target object.

[0007] In one possible implementation, determining at least one target region from the image to be processed further includes: when the location range of the positioning points is known and the number of positioning points used when measuring the target object is greater than 1, determining multiple selection methods for the target region based on the location range of each positioning point; determining the total area of ​​the target region under each selection method; determining the selection method with the smallest total area as the target selection method; and determining at least one target region from the image to be processed based on the target selection method.

[0008] In one possible implementation, determining at least one target region from the image to be processed further includes: when the location range of the positioning points is unknown, downsampling the image to be processed according to a preset downsampling ratio to obtain an intermediate image, wherein the intermediate image is the downsampled image of the image to be processed; determining second position information of each positioning point in the intermediate image; determining third position information of each positioning point in the image to be processed based on the second position information, wherein the third position information is used to indicate the position information of the positioning point in the image to be processed corresponding to the second position information; and determining at least one target region from the image to be processed based on the third position information of each positioning point and a preset size.

[0009] In one possible implementation, the step of acquiring an image of the target object to be measured based on a preset positioning accuracy to obtain an image to be processed includes: determining the field of view range when acquiring an image of the target object based on the preset positioning accuracy and the number of acquisition pixels; and acquiring an image of the target object based on the field of view range to obtain an image to be processed.

[0010] In one possible implementation, the method further includes: when multiple positioning points are used when measuring the target object, determining the minimum value of the positioning accuracy of the multiple positioning points as the preset positioning accuracy.

[0011] In one possible implementation, the target object includes at least one of industrial parts or industrial products, and the dimensional information includes at least one of length, width, height, curvature, radianity, or area.

[0012] According to another aspect of this disclosure, a deep learning-based size measurement device is provided. The device includes: an image acquisition module, configured to acquire an image of a target object to be measured according to a preset positioning accuracy to obtain an image to be processed, wherein the positioning accuracy is used to indicate the image acquisition accuracy of the positioning points used when measuring the target object; a target region determination module, configured to determine at least one target region from the image to be processed, each target region including at least one positioning point; a position information determination module, configured to process the at least one target region through a pre-trained neural network to obtain first position information of each positioning point; and a size information determination module, configured to determine the size information of the target object according to the positioning accuracy and the first position information of each positioning point.

[0013] According to another aspect of this disclosure, a deep learning-based size measurement apparatus is provided, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement any of the methods described above when executing the instructions.

[0014] The embodiments of this disclosure can acquire images of the target object to be measured according to a preset positioning accuracy to obtain an image to be processed, and determine at least one target region from the image to be processed. Each target region includes at least one positioning point. Then, the at least one target region is processed by a pre-trained neural network to obtain the first position information of each positioning point. Based on the positioning accuracy and the first position information of each positioning point, the size information of the target object is determined. In this way, the repeatability and accuracy of the positioning of the target object during the size measurement process can be improved through deep learning, thereby improving the accuracy of the size measurement result (i.e., the size information of the target object). Attached Figure Description

[0015] The technical solution and its beneficial effects will become apparent from the following detailed description of specific embodiments of this disclosure, in conjunction with the accompanying drawings.

[0016] Figure 1 A schematic diagram illustrating an application scenario of a deep learning-based dimension measurement method according to an embodiment of the present disclosure is shown.

[0017] Figure 2 A schematic diagram illustrating an application scenario of a deep learning-based dimension measurement method according to an embodiment of the present disclosure is shown.

[0018] Figure 3A flowchart illustrating a deep learning-based dimension measurement method according to an embodiment of the present disclosure is shown.

[0019] Figure 4 A schematic diagram illustrating the processing flow of a deep learning-based dimension measurement method according to an embodiment of the present disclosure is shown.

[0020] Figure 5 A block diagram of a deep learning-based dimensional measurement device according to an embodiment of the present disclosure is shown. Detailed Implementation

[0021] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0022] Currently, the dimensional measurement of parts or products in industrial production processes typically employs template matching. This method involves first acquiring two-dimensional (2D) or three-dimensional (3D) images of the object to be measured (e.g., parts of industrial products such as mobile phones and automobiles). Then, a preset template is used to match the object to be measured within the pixels of the acquired 2D image or the voxels of the 3D image, obtaining a positioning result. Based on this positioning result, the dimensional information of the object to be measured is determined.

[0023] However, in this method, the imaging results (i.e., photos) are often inconsistent due to factors such as the environment (e.g., inconsistent ambient light, inconsistent light sources, dust interference, lens defects, sensor defects, etc.) and materials (e.g., inconsistent material processing technology, inconsistent material materials). The repeatability and accuracy of positioning based on these imaging results are affected. In other words, the template matching method is difficult to be compatible with the diversity of object imaging, resulting in low accuracy of the size information obtained from the positioning results.

[0024] In other technical solutions, the measurement of parts or products is based on manually programmed features for identification. For example, feature points of the object to be measured are defined by manually programming edges and corners; then, images of the object are captured, and the feature points of the object are found in the captured images to determine their positions; then, based on the positions of the feature points, the object is located, and finally, based on the location results, the size information of the object is determined.

[0025] However, this method is not only affected by environmental and material factors, but also heavily reliant on vision engineers. For example, when the object under test uses a new material, vision engineers need to manually write its feature points, which greatly limits the deployment speed of the dimensional measurement function. In addition, when the feature points of the object under test are unclear, this method cannot determine the specific location of the feature points from a global perspective, resulting in low accuracy of the final dimensional information.

[0026] To address the aforementioned technical problems, this disclosure provides a deep learning-based size measurement method. Embodiments of this disclosure can acquire images of the target object to be measured based on a preset positioning accuracy, obtaining an image to be processed. The positioning accuracy indicates the image acquisition accuracy of the positioning points used when measuring the target object. At least one target region is determined from the image to be processed, and each target region includes at least one positioning point. Then, a pre-trained neural network processes the at least one target region to obtain the position information of each positioning point. Based on the positioning accuracy and the first position information of each positioning point, the size information of the target object is determined. This method, through deep learning, improves the repeatability and accuracy of the target object's positioning during size measurement, thereby enhancing the accuracy of the size measurement results (i.e., the size information of the target object).

[0027] Figure 1 A schematic diagram illustrating an application scenario of a deep learning-based size measurement method according to embodiments of the present disclosure is shown. Figure 1 As shown, a deep learning-based size measurement method is applied to an electronic device 100, which includes a processor 110 and an image acquisition module 120.

[0028] Electronic device 100 can be a server, desktop computer, mobile device, or any other type of computing device that includes processor 110 and image acquisition module 120. This disclosure does not limit the specific type of electronic device 100.

[0029] The image acquisition module 120 can be, for example, a camera or webcam that can capture images of the target object 200. This disclosure does not limit the specific implementation of the image acquisition module 120.

[0030] The processor 110 can acquire an image of the target object 200 to be measured through the image acquisition module 120 according to the preset positioning accuracy, obtain an image to be processed, and determine at least one target region from the image to be processed. Each target region includes at least one positioning point. Then, the processor 110 processes the at least one target region through a pre-trained neural network to obtain the position information of each positioning point, and determines the size information of the target object based on the position information of each positioning point.

[0031] Processor 110 can be a general-purpose processor, such as a central processing unit (CPU), or an artificial intelligence processing unit (IPU). For example, an artificial intelligence processor may include one or a combination of graphics processing units (GPUs), neural network processing units (NPUs), digital signal processing units (DSPs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs). This disclosure does not limit the specific type of processor 110.

[0032] Figure 2 A schematic diagram illustrating an application scenario of a deep learning-based size measurement method according to embodiments of the present disclosure is shown. Figure 2 As shown, a deep learning-based size measurement method is applied to an electronic device 100, which includes a processor 110, an image acquisition module 120, and an operation module 130.

[0033] Among them, the image acquisition module 120 and the processor 110 and Figure 1 Similar application scenarios are described above, and will not be repeated here.

[0034] After the processor 110 determines the size information of the target object 200, the processor 110 can also control the operation module 130 (such as a robotic arm) to perform operations such as grasping, position adjustment, and orientation adjustment on the target object based on the size information.

[0035] Figure 3 A flowchart illustrating a deep learning-based dimension measurement method according to an embodiment of the present disclosure is shown. Figure 3 As shown, the method includes:

[0036] Step S310: Based on the preset positioning accuracy, the target object to be measured is imaged to obtain the image to be processed.

[0037] The target object to be measured may include at least one of industrial products (e.g., mobile phones, automobiles, etc.) or industrial parts (e.g., mobile phone parts, automobile parts, etc.) produced during the industrial manufacturing process. This disclosure does not limit the specific type of the target object.

[0038] All spatial measurement problems of the target object can be transformed into localization problems of at least one localization point of the target object. Based on this assumption, before performing spatial measurements on the target object, it is necessary to determine at least one localization point to be used in the measurement.

[0039] For example, when performing a single-point measurement on a target object, that is, when measuring the position of a certain point on the target object, that point can be directly used as the positioning point; when measuring the length of a side on a target object, the two endpoints of the side to be measured can be used as positioning points; when measuring the arc or circle on a target object, at least three points on the arc or circle can be used as positioning points; when measuring the area, perimeter, or other shapes of a target object, all vertices of the shape can be used as positioning points.

[0040] When measuring points, lines, surfaces, and shapes of a target object in three-dimensional space, a similar method as described above can be used to determine the positioning points.

[0041] It should be noted that those skilled in the art can determine the location points of the target object according to the actual situation, and this disclosure does not limit the number of location points of the target object or the method of determination.

[0042] In one possible implementation, the positioning accuracy of each positioning point of the target object can be determined according to the accuracy requirements of the actual measurement task. Positioning accuracy can be used to indicate the image acquisition accuracy of the positioning points used when measuring the target object. The positioning accuracy of each positioning point can be the same or different; this disclosure does not impose any limitation on this.

[0043] In one possible implementation, when only one positioning point is used when measuring the target object, the positioning accuracy of that positioning point can be directly determined as the preset positioning accuracy.

[0044] In one possible implementation, when multiple positioning points are used when measuring the target object, the minimum positioning accuracy among the multiple positioning points can be determined as the preset positioning accuracy.

[0045] For example, if the target object 200 is a cube, when measuring the area of ​​each face of the target object 200, the four vertices of each face can be used as positioning points. Suppose that the four positioning points (i.e. vertices) of a certain face are A1, A2, A3 and A4, and their positioning accuracies are 2um, 3um, 5um and 1um respectively. Since the smaller the positioning accuracy, the higher the positioning accuracy, the minimum value of 1um among 2um, 3um, 5um and 1um can be determined as the preset positioning accuracy when measuring the size of the target object A.

[0046] By setting the minimum positioning accuracy among multiple positioning points as the preset positioning accuracy, the preset positioning accuracy can meet the positioning accuracy requirements of multiple positioning points, thereby improving the positioning accuracy.

[0047] In one possible implementation, after determining the preset positioning accuracy, in step S310, based on the positioning accuracy, the target object to be measured is imaged through an image acquisition module such as a camera or webcam to obtain the image to be processed.

[0048] In one possible implementation, step S310 may include: determining the field of view range for image acquisition of the target object based on a preset positioning accuracy and image acquisition pixels; and acquiring an image of the target object based on the field of view range to obtain an image to be processed.

[0049] Among them, the image acquisition pixels can be used to indicate the pixels of the image acquisition module, such as the pixels of a camera or a webcam; the field of view can be used to indicate the actual area range captured by the image acquisition module when it captures images of the target object.

[0050] In one possible implementation, the field of view when capturing images of the target object can be determined based on the preset positioning accuracy and image capture pixels.

[0051] For example, assuming the image acquisition module is a camera, and the absolute position of the camera is known (i.e., the camera position is fixed), if a single point (i.e., a positioning point) needs to be located with a preset positioning accuracy of r, then, according to the Nyquist sampling theorem, r needs to be covered by multiple pixels. Those skilled in the art can determine the minimum number of pixels covering r based on the actual situation, and this disclosure does not impose any limitations on this.

[0052] Assume that r is covered by at least 4 pixels, and each pixel represents a size of r in the actual space. p ×r p a square (where r) p (representing the side length), then 4×r p ≤r.

[0053] Based on this, assuming the camera's pixels (i.e., image-capturing pixels) are w×h and the field of view is a×b, then the actual spatial distance r corresponding to each pixel of the camera... p The following formula (1) should be satisfied:

[0054] r p =max(a / w,b / h)≤r / 4 (1)

[0055] In formula (1), w represents the number of pixels of the camera in the horizontal direction (i.e., each row), and h represents the number of pixels of the camera in the vertical direction (i.e., each column).

[0056] In w, h, r, r p Given the given information, the maximum values ​​of a and b can be determined using the formula (1) above, and the maximum value of a × b can be used as the field of view when capturing images of the target object to be measured. Here, a represents the length of the field of view, and b represents the width of the field of view.

[0057] In one possible implementation, when measuring the distance (e.g., Euclidean distance) between two points of the target object, and requiring the standard deviation of the distance between the two points to be d, these two points can be used as positioning points, and the positioning accuracy r of each positioning point can be set to...

[0058] With positioning accuracy r Furthermore, when the camera has a pixel count of w×h, the field of view for capturing images of the target object can be determined using the above formula (1).

[0059] In one possible implementation, when multiple positioning points are used when measuring the target object, the minimum value of the positioning accuracy of the multiple positioning points can be determined as the preset positioning accuracy, and the field of view range when imaging the target object can be determined by the above formula (1).

[0060] When measuring the points, lines, surfaces, and shapes of a target object in three-dimensional space, a similar method can be used, that is, the field of view range when capturing images of the target object can be determined by the above formula (1).

[0061] It should be noted that those skilled in the art can determine the field of view in other ways, and this disclosure does not impose any restrictions on this.

[0062] In one possible implementation, after determining the field of view, the image acquisition module can be adjusted according to the field of view, such as adjusting the camera's lens, focal length, and other imaging parameters, and the target object can be captured by the adjusted image acquisition module to obtain the image to be processed.

[0063] By determining the field of view when capturing images of the target object based on the positioning accuracy and the number of pixels captured, and then capturing images of the target object based on the field of view, an image to be processed is obtained. This ensures that the image capturing process of the target object meets the positioning accuracy requirements and improves the accuracy of the image to be processed.

[0064] Step S320: Determine at least one target region from the image to be processed.

[0065] To reduce the computational load of the neural network and improve the accuracy of its processing results, at least one target region can be identified from the image to be processed before determining the location information of each positioning point using the neural network. The target region indicates the area in the image to be processed where the positioning points are located. Each target region includes at least one positioning point; that is, the number of positioning points within each target region is one or more. In other words, the target region is an image region in the image to be processed that includes at least one positioning point.

[0066] In one possible implementation, for any given location point, the smallest area within the field of view where the location point appears can be defined as the location range of that location point. If the smallest area within the field of view where the location point appears can be determined in advance, for example, based on existing experience, then the location range of that location point is considered known; if the smallest area within the field of view where the location point appears cannot be determined in advance, i.e., the location point may appear in any area within the field of view, then the location range of that location point is considered unknown.

[0067] In one possible implementation, when the location range of the positioning point is known (e.g., determined based on existing experience) and the number of positioning points used when measuring the target object is 1, the region in the image to be processed that corresponds to the location range of the positioning point can be determined as the target region.

[0068] For example, suppose the number of positioning points used when measuring the target object is 1, denoted as positioning point B. The location range of positioning point B is known: within the field of view, according to coordinates (x... min ,y min ) and (x max ,y max Given a defined region C, the region in the image to be processed that corresponds to the location range of the positioning point (i.e., region C within the field of view) can be defined as the target region, where x min <x max y min <y max .

[0069] When the location range of the positioning point is known and the number of positioning points used when measuring the target object is 1, the area in the image to be processed that corresponds to the location range of the positioning point is directly determined as the target area. This is simple and fast, thereby improving processing efficiency.

[0070] In one possible implementation, when the location range of the positioning points is known and the number of positioning points used when measuring the target object is greater than 1, multiple selection methods for the target area can be determined based on the location range of each positioning point. Then, the total area of ​​the target area under each selection method is determined, and the selection method with the smallest total area is determined as the target selection method. Then, at least one target area is determined from the image to be processed based on the target selection method.

[0071] Generally, the computational complexity of deep learning-based image processing tasks is directly proportional to the image size; that is, the larger the image size, the greater the computational complexity required to process it. The image size is typically proportional to its area.

[0072] Based on this, in order to reduce the computational load of the neural network and improve the accuracy of the neural network processing results, when the location range of the positioning points is known and the number of positioning points used when measuring the target object is greater than 1, we can first determine multiple selection methods of the target area according to the location range of each positioning point, then determine the total area of ​​the target area under each selection method, and determine the selection method with the smallest total area as the target selection method.

[0073] For example, suppose the location range of the positioning points is known, and the number of positioning points used when measuring the target object is 2, denoted as positioning point D1 and positioning point D2, where the location range of positioning point D1 is: within the field of view, based on coordinates (x... 11 ,y 11 ) and (x 12 ,y 12 The defined region, x 11 <x 12 And y 11 <y 12 The location range of positioning point D2 is: within the field of view, based on coordinates (x... 21 ,y 21 ) and (x 22 ,y 22 The defined region, x 21 <x 22 And y 21 <y 22 .

[0074] Multiple selection methods for the target area can be determined by combining positioning points D1 and D2 in different ways:

[0075] The first method for selecting the target area: Positioning points D1 and D2 are located within the same target area. Based on the positional range of positioning points D1 and D2, the area E1 within the field of view, including positioning points D1 and D2, can be determined. Area E1 is determined based on the coordinates (x′... min ,y′ min ) and (x′ max ,y′ max The region is determined by the first step; then, the region in the image to be processed that corresponds to region E1 within the field of view is determined as the target region F1. Where x′ min =min(x 11 ,x 21 ), x′ max =max(x 12 ,x 22 ), y′ min =min(y 11 ,y 21 ), y′ max =max(y 12 ,y 22 ).

[0076] The second method for selecting the target area: Positioning points D1 and D2 are located in different target areas. The position range of positioning point D1 can be considered as region E2 (within the field of view, based on coordinates (x...)). 11 ,y 11 ) and (x 12 ,y 12 The region corresponding to region E2 within the field of view in the image to be processed is defined as the target region F2; the position range of positioning point D2 is considered as region E3 (within the field of view, based on coordinates (x...)). 21 ,y 21 ) and (x 22 ,y 22 The region corresponding to region E3 within the field of view in the image to be processed is determined as the target region F3.

[0077] Then, the total area of ​​the target region is determined under each selection method: under the first selection method, the total area of ​​the target region is the area S1 of target region F1; under the second selection method, the total area of ​​the target region is the sum of the areas S2 of target region F2 and target region F3.

[0078] By comparing and calculating, the selection method with the smallest area can be determined from multiple selection methods: S = min(S1, S2), and this selection method is determined as the target selection method. Then, based on this target selection method, at least one target region is determined from the image to be processed through image segmentation and other methods.

[0079] It should be noted that the above example only illustrates the method for determining the target area using a target object comprising two positioning points. Those skilled in the art should understand that when the target object comprises multiple positioning points, the method for determining the target area is similar, and will not be repeated here.

[0080] By selecting at least one target region from the image to be processed using the method of minimizing the total area, the total area of ​​the target regions can be reduced, thereby improving the processing efficiency of the subsequent neural network.

[0081] In one possible implementation, when the location range of the target object's positioning points is unknown, the image to be processed can be downsampled according to a preset downsampling ratio to obtain an intermediate image, which is the downsampled version of the image to be processed. Then, through methods such as target localization, the second position information of each positioning point in the intermediate image is determined, and based on the second position information, the third position information of each positioning point in the image to be processed is determined. The third position information indicates the position of the positioning point in the image to be processed corresponding to the second position information. Finally, based on the third position information of each positioning point and a preset size, at least one target region is determined from the image to be processed. The preset size indicates the size of the region in the image to be processed where the positioning point appears.

[0082] For example, suppose the image to be processed is P, and the preset downsampling ratio is C. down When the location range of the target object's positioning point is unknown, it can be determined based on the downsampling ratio C. down The image P to be processed is downsampled in both length and width directions to obtain an intermediate image P′. In other words, the intermediate image P′ is the downsampled image of the image P to be processed.

[0083] The positioning accuracy of the intermediate image P′ is r′=r×C. down The computational complexity of the neural network when processing the intermediate image P′ is O′=O / (C down ×C down O represents the computational load when the neural network processes the image P to be processed.

[0084] Then, the second position information of each positioning point in the intermediate image P′ can be determined through methods such as target localization. Since the localization is performed on the intermediate image P′ (i.e., the image after downsampling from the image to be processed), this localization can be regarded as coarse localization. Compared with the image to be processed P, although this method suffers from a loss of localization accuracy, its localization speed is faster, which is beneficial to improving processing efficiency.

[0085] After obtaining the second position information of each positioning point, the third position information corresponding to the second position information of each positioning point in the image to be processed can be determined based on the positional correspondence between each pixel in the intermediate image P′ and the image to be processed P. The third position information can be regarded as the coarse position information of each positioning point in the image to be processed P.

[0086] Then, based on the third position information of each positioning point and the preset size, at least one target region can be determined from the image P to be processed, so as to accurately locate each positioning point in step S330. The preset size indicates the size of the region in the image to be processed where the positioning point appears.

[0087] For example, suppose the third location information of a certain positioning point is its coordinates (x1, y1) in the image P to be processed, and the preset size is 2×C. down Therefore, from the image P to be processed, based on the coordinates (x1-C) down ,y1-C down ) and (x1+C down ,y1+C down Determine region F′, and define regions greater than or equal to region F′ as the target region including the location point.

[0088] In one possible implementation, after obtaining the target area including each positioning point, at least one target area can be obtained by combining the target areas in a similar manner to the above, so that the total area of ​​each target area is minimized.

[0089] By reducing the size of the image to be processed and performing coarse localization when the location range of the target object's positioning point is unknown, and then determining at least one target region from the image to be processed based on the coarse localization result, the target region including the positioning point can be determined through coarse localization when the positioning point of the target object is unclear, thereby improving the accuracy of the target region.

[0090] Step S330: The at least one target region is processed through a pre-trained neural network to obtain the first position information of each positioning point.

[0091] The pre-trained neural network may include at least one sub-network. Optionally, the number of sub-networks in the neural network is consistent with the number of target regions, that is, each target region corresponds to a sub-network that processes that target region. The sub-networks can be deep learning-based neural networks such as convolutional neural networks (CNNs). This disclosure does not limit the number or specific type of sub-networks.

[0092] In one possible implementation, the sub-network may include an encoder and a decoder. The encoder is used to extract features from the input target region and may include multiple convolutional layers. The encoder can spatially cover the input target region with a sufficiently large field of view, extracting features by reducing the spatial size and increasing the number of channels. The decoder is used to determine the positional information of each localization point and may include multiple convolutional layers.

[0093] In one possible implementation, step S330 may include: for any target region, extracting features from the target region using an encoder in the sub-network corresponding to the target region to obtain a feature map of the target region; and processing the feature map using a decoder in the sub-network corresponding to the target region to obtain first position information of each positioning point in the target region.

[0094] For any target region, assuming its size is w′×h′×c′, where w′ represents the width, h′ represents the height, and c represents the number of channels, if a grayscale camera is used for image acquisition, then c′=1; if a color camera is used, then c′=3. The target object is input into the encoder of its corresponding sub-network for feature extraction, resulting in a feature map of the target region with size w′. e ×h′ e ×c′ e w′ e h′ represents the width of the feature map of the target region. e c′ represents the height of the feature map of the target region. e This represents the number of channels in the feature map of the target region, where w ′e <w′、h′ e <h′、c′ e >c′ and (w′) e ×h′ e ×c′ e )<(w′×h′×c′).

[0095] For any convolutional layer in the encoder, assume that the size of its input feature map is w′. e1 ×h′ e1 ×c′ e1 The size of its output feature map is w′ e2 ×h′ e2 ×c′ e2 , where w′ e1 h′ represents the width of the feature map input to this convolutional layer. e1 c′ represents the height of the input feature map of the convolutional layer. e1w′ represents the number of channels in the input feature map of this convolutional layer. e2 h′ represents the width of the feature map output by the convolutional layer. e2 c′ represents the height of the feature map output by the convolutional layer. e2 Let w′ represent the number of channels in the feature map output by the convolutional layer. e2 ≤w′ e1 And h′ e2 ≤h′ e1 Therefore, the forward propagation process of the encoder can be viewed as a process of compressing the image (i.e., the target region).

[0096] After obtaining the feature map of the target region, it can be processed by the decoder in the sub-network corresponding to the target region to obtain the position information of each localization point in the target region. For example, the feature map of the target region (with size w′) can be processed... e ×h′ e ×c′ e The input feature map is processed by the decoder in the sub-network corresponding to the target region to obtain the location information of each localization point in the target region.

[0097] For any convolutional layer in the decoder, assume that the size of its input feature map is w′. d1 ×h′ d1 ×c′ d1 The size of its output feature map is w′ d2 ×h′ d2 ×c′ d2 , where w′ d1 h′ represents the width of the feature map input to this convolutional layer. d1 c′ represents the height of the input feature map of the convolutional layer. d1 w′ represents the number of channels in the input feature map of this convolutional layer. d2 h′ represents the width of the feature map output by the convolutional layer. d2 c′ represents the height of the feature map output by the convolutional layer. d2 Let w′ represent the number of channels in the feature map output by the convolutional layer. d2 ≤w′ d1 And h′ d2 ≤h′ d1 .

[0098] Determining the first position information of each positioning point by using an encoder and decoder can not only improve the processing effect, but also improve the accuracy of the first position information.

[0099] Step S340: Determine the size information of the target object based on the positioning accuracy and the first position information of each positioning point.

[0100] In one possible implementation, the size information of the target object can be determined based on a preset positioning accuracy and the first position information (e.g., coordinates) of each positioning point. The size information of the target object may include at least one of the following: length, width, height, curvature, radianity, or area.

[0101] It should be noted that the size information of the target object may also include other size-related information that can be measured spatially. Those skilled in the art can determine the specific content of the size information of the target object according to the actual situation, and this disclosure does not impose any limitations on it.

[0102] The embodiments of this disclosure can acquire images of the target object to be measured according to a preset positioning accuracy to obtain an image to be processed, and determine at least one target region from the image to be processed. Each target region includes at least one positioning point. Then, the at least one target region is processed by a pre-trained neural network to obtain the first position information of each positioning point. Based on the positioning accuracy and the first position information of each positioning point, the size information of the target object is determined. In this way, the repeatability and accuracy of the positioning of the target object during the size measurement process can be improved through deep learning, thereby improving the accuracy of the size measurement result (i.e., the size information of the target object).

[0103] In one possible implementation, the method further includes: training the neural network according to a preset training set, wherein the training set includes multiple sample images and annotation information for each sample image, and the annotation information includes reference position information for each positioning point in the sample image.

[0104] In one possible implementation, the neural network may include at least one subnetwork. For any subnetwork, a sample set and multiple sample images corresponding to that subnetwork can be input into the subnetwork for processing to obtain the location information of the localization points in each sample image. The difference between the location information of the localization points in each sample image and the annotation information in each sample image is determined. Then, the network loss of the subnetwork is determined based on the difference, and the parameters of the subnetwork are adjusted according to the network loss. Repeating the above process allows for multiple rounds of training of the subnetwork.

[0105] Training can be terminated when all subnetworks in the neural network meet the preset training termination conditions, resulting in a trained neural network. These termination conditions may include the subnetwork loss converging within a certain threshold, or the subnetwork passing validation on a preset validation set.

[0106] It should be noted that those skilled in the art can set the training termination conditions according to the actual situation, and this disclosure does not impose any restrictions on this.

[0107] Training a neural network with a pre-set training set can improve the accuracy of the neural network, thereby improving the accuracy of the size measurement results (i.e., the size information of the target object).

[0108] Figure 4 A schematic diagram illustrating the processing flow of a deep learning-based size measurement method according to an embodiment of the present disclosure is shown. Figure 4 As shown, when measuring the size of the target object 410, such as the radius measurement, three positioning points are used. First, the target object 410 is imaged according to the preset positioning accuracy to obtain the image to be processed 420. Then, two target regions are determined from the image to be processed 420: target region 431 and target region 432. Target region 431 includes two positioning points, and target region 432 includes one positioning point.

[0109] After the target region is determined, the target regions 431 and 432 can be processed by the pre-trained neural network 440 to obtain the first position information of each positioning point. The neural network 440 includes two sub-networks: sub-network 441 and sub-network 444.

[0110] Sub-network 441 processes target region 431: the encoder 442 extracts features from target region 431 to obtain feature map of target region 431, and then the decoder 443 processes the output of encoder 442 (i.e. feature map of target region 431) to obtain the first position information 451 of two positioning points in target region 431.

[0111] Sub-network 444 processes the target region 432: the encoder 445 extracts features from the target region 432 to obtain the feature map of the target region 432, and then the decoder 446 processes the output of the encoder 445 (i.e. the feature map of the target region 432) to obtain the first position information 452 of a positioning point in the target region 432.

[0112] Then, based on the preset positioning accuracy, the first position information 451 and the first position information 452, the size information 460 of the target object can be determined, such as calculating the curvature of the target object.

[0113] It should be noted that the above only uses two target regions (target region 431 and target region 432) as examples to illustrate the processing procedure of the deep learning-based size measurement method. Those skilled in the art should understand that... Figure 4The number of target areas shown can also be other, such as 3, 5, 10, etc. That is to say, those skilled in the art can set the specific number of target areas according to the actual situation, and this disclosure does not limit it.

[0114] The deep learning-based dimension measurement method described in this disclosure can automatically learn the dimension measurement problem of the target object through deep learning. It has strong robustness to environmental and material factors, and has high repeatability and accuracy in positioning, thereby improving the accuracy of dimension measurement results.

[0115] Furthermore, the deep learning-based dimension measurement method described in this disclosure can determine the target area including the positioning points through coarse positioning when the positioning points of the target object are unclear, and then determine the position information of each positioning point through further positioning of the target area (i.e., precise positioning), thereby improving the accuracy of positioning and thus improving the accuracy of dimension measurement results.

[0116] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further.

[0117] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0118] In addition, this disclosure also provides a deep learning-based dimension measurement device and a computer-readable storage medium, both of which can be used to implement any of the deep learning-based dimension measurement methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the relevant section of the method and will not be repeated here.

[0119] Figure 5 A block diagram of a deep learning-based dimensional measurement device according to an embodiment of the present disclosure is shown. Figure 5 As shown, the device includes:

[0120] The image acquisition module 51 is used to acquire images of the target object to be measured according to a preset positioning accuracy to obtain an image to be processed. The positioning accuracy is used to indicate the image acquisition accuracy of the positioning point used when measuring the target object.

[0121] The target region determination module 52 is used to determine at least one target region from the image to be processed, and each target region includes at least one positioning point;

[0122] The location information determination module 53 is used to process the at least one target area through a pre-trained neural network to obtain the first location information of each positioning point;

[0123] The size information determination module 54 is used to determine the size information of the target object based on the positioning accuracy and the first position information of each positioning point.

[0124] The embodiments of this disclosure can acquire images of the target object to be measured according to a preset positioning accuracy to obtain an image to be processed, and determine at least one target region from the image to be processed. Each target region includes at least one positioning point. Then, the at least one target region is processed by a pre-trained neural network to obtain the first position information of each positioning point. Based on the positioning accuracy and the first position information of each positioning point, the size information of the target object is determined. In this way, the repeatability and accuracy of the positioning of the target object during the size measurement process can be improved through deep learning, thereby improving the accuracy of the size measurement result (i.e., the size information of the target object).

[0125] In one possible implementation, the neural network includes at least one sub-network corresponding to the target region. The sub-network includes an encoder and a decoder. The position information determination module 53 includes: a feature map determination sub-module, which extracts features from the target region for any target region by means of the encoder in the sub-network corresponding to the target region to obtain a feature map of the target region; and a first position information determination sub-module, which processes the feature map by means of the decoder in the sub-network corresponding to the target region to obtain the position information of each positioning point in the target region.

[0126] In one possible implementation, the target region determination module 52 includes: a first region determination submodule, used to determine the region in the image to be processed that corresponds to the location range of the positioning point as the target region when the location range of the positioning point is known and the number of positioning points used when measuring the target object is 1. The location range is used to indicate the minimum area where the positioning point appears within the field of view, and the field of view is used to indicate the area range actually captured by the image acquisition module when capturing images of the target object.

[0127] In one possible implementation, the target region determination module 52 further includes: a first selection method determination submodule, used to determine multiple selection methods of the target region based on the location range of each positioning point when the location range of the positioning points is known and the number of positioning points used when measuring the target object is greater than 1; an area sum determination submodule, used to determine the total area of ​​the target region under each selection method; a second selection method determination submodule, used to determine the selection method with the smallest total area as the target selection method; and a second region determination submodule, used to determine at least one target region from the image to be processed based on the target selection method.

[0128] In one possible implementation, the target region determination module 52 further includes: a sampling submodule, configured to downsample the image to be processed according to a preset downsampling ratio to obtain an intermediate image when the location range of the positioning points is unknown, wherein the intermediate image is the image after downsampling of the image to be processed; a second position information determination submodule, configured to determine the second position information of each positioning point in the intermediate image; a third position information determination submodule, configured to determine the third position information of each positioning point in the image to be processed according to the second position information, wherein the third position information is used to indicate the position information of the positioning point in the image to be processed corresponding to the second position information; and a third region determination submodule, configured to determine at least one target region from the image to be processed according to the third position information of each positioning point and a preset size.

[0129] In one possible implementation, the image acquisition module 51 includes: a field-of-view range determination submodule, used to determine the field-of-view range when acquiring images of the target object based on a preset positioning accuracy and image acquisition pixels; and an image acquisition submodule, used to acquire images of the target object based on the field-of-view range to obtain an image to be processed.

[0130] In one possible implementation, the device further includes a positioning accuracy determination module, used to determine the minimum value of the positioning accuracies of the multiple positioning points as a preset positioning accuracy when multiple positioning points are used when measuring the target object.

[0131] In one possible implementation, the target object includes at least one of industrial parts or industrial products, and the dimensional information includes at least one of length, width, height, curvature, radianity, or area.

[0132] Embodiments of this disclosure also propose a deep learning-based dimension measurement device, the device including a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement any of the above-described deep learning-based dimension measurement methods when executing the instructions.

[0133] Embodiments of this disclosure also provide a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement any of the allocation methods described above. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium.

[0134] The above description is merely an exemplary embodiment of this disclosure and does not limit the scope of patent protection of this disclosure. Any equivalent structural or procedural transformations made using the content of this disclosure and its drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this disclosure.

Claims

1. A dimension measurement method based on deep learning, characterized in that, The method includes: Based on the preset positioning accuracy and image acquisition pixels, the field of view range for image acquisition of the target object is determined; based on the field of view range, the target object is image acquired to obtain an image to be processed; the positioning accuracy is used to indicate the image acquisition accuracy of the positioning point used when measuring the target object. Before determining the location information of each positioning point through a neural network, at least one target region is determined from the image to be processed, and each target region includes at least one positioning point; the at least one target region is processed through a pre-trained neural network to obtain the first location information of each positioning point; Based on the positioning accuracy and the first position information of each positioning point, the size information of the target object is determined; The step of determining at least one target region from the image to be processed includes: When the location range of the positioning points is known, and the number of positioning points used when measuring the target object is greater than 1, multiple selection methods for the target area are determined based on the location range of each positioning point. Determine the total area of ​​the target region under each selection method; The method that minimizes the total area is determined as the target selection method; Based on the target selection method, at least one target region is determined from the image to be processed; The location range is used to indicate the smallest area where the positioning point appears within the field of view, and the field of view is used to indicate the actual area range captured by the image acquisition module when it captures images of the target object. or When the location range of the positioning point is unknown, the image to be processed is downsampled according to a preset downsampling ratio to obtain an intermediate image, which is the downsampled image of the image to be processed. Determine the second position information of each positioning point in the intermediate image; Based on the second location information, the third location information of each positioning point in the image to be processed is determined. The third location information is used to indicate the location information of the positioning point in the image to be processed that corresponds to the second location information. Based on the third position information and preset size of each positioning point, at least one target region is determined from the image to be processed; After obtaining the target area including each positioning point, the target areas are combined to obtain at least one target area that minimizes the total area of ​​all target areas. The location range is used to indicate the smallest area where the positioning point appears within the field of view, and the field of view is used to indicate the area range actually captured by the image acquisition module when it captures images of the target object.

2. The method according to claim 1, characterized in that, The neural network includes at least one sub-network, which corresponds to the target region, and the sub-network includes an encoder and a decoder. The process of processing the at least one target region using a pre-trained neural network to obtain the location information of each positioning point includes: For any target region, feature extraction is performed on the target region by the encoder in the sub-network corresponding to the target region to obtain the feature map of the target region; The feature map is processed by the decoder in the sub-network corresponding to the target region to obtain the location information of each localization point in the target region.

3. The method according to any one of claims 1-2, characterized in that, The method further includes: When multiple positioning points are used to measure the target object, the minimum positioning accuracy among the multiple positioning points is determined as the preset positioning accuracy.

4. The method according to any one of claims 1-2, characterized in that, The target object includes at least one of industrial parts or industrial products, and the dimensional information includes at least one of length, width, height, curvature, radianity, or area.

5. A dimension measurement device based on deep learning, characterized in that, The device includes: The image acquisition module is used to determine the field of view when acquiring images of the target object based on the preset positioning accuracy and image acquisition pixels; and to acquire images of the target object based on the field of view to obtain an image to be processed. The positioning accuracy is used to indicate the image acquisition accuracy of the positioning point used when measuring the target object. The target region determination module is used to determine at least one target region from the image to be processed, and each target region includes at least one positioning point; The location information determination module is used to process the at least one target area through a pre-trained neural network to obtain the first location information of each positioning point; The size information determination module is used to determine the size information of the target object based on the positioning accuracy and the first position information of each positioning point; processor; Memory used to store processor-executable instructions; The processor is configured to implement the method described in any one of claims 1-4 when executing the instructions.

Citation Information

Patent Citations

  • Target detection method and system based on cut area candidate network

    CN110348435A

  • Clothes size measuring method and system based on neural network and electronic device

    CN110349201A

  • Cascaded medical image key point detection method and device

    CN111862047A