Workshop equipment three-dimensional layout diagram acquisition method, electronic equipment and storage medium

By using the SuperPoint network model and descriptor matching technology, a 3D layout diagram of workshop equipment is obtained, which solves the problem of difficulty in accurately obtaining the 3D layout diagram of workshop equipment in the existing technology, and achieves higher accuracy and real-time performance, thereby improving production scheduling efficiency.

CN121120976APending Publication Date: 2025-12-12FUTAIHUA PRECISION ELECTRONICS (ZHENGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511059331.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately obtain three-dimensional layout diagrams of workshop equipment, resulting in low production scheduling efficiency.

Method used

The SuperPoint network model is used to obtain the feature point set of the camera, and a 3D layout map of the workshop equipment is constructed by descriptor matching and stitching techniques.

Benefits of technology

It improves the accuracy and real-time performance of generating 3D layout diagrams of workshop equipment, reduces manual intervention, lowers data errors, and enhances production management and scheduling efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120976A_ABST
    Figure CN121120976A_ABST
Patent Text Reader

Abstract

The invention discloses a method for acquiring a three-dimensional layout diagram of workshop equipment, electronic equipment and a computer readable storage medium. The method comprises the steps of obtaining a first feature point set and a second feature point set by using a preset algorithm, then constructing a feature matching point pair set, and finally splicing images shot by a first camera and a second camera based on the feature matching point pair set to obtain a workshop equipment three-dimensional layout map. In the preset algorithm, the last frame descriptor can be obtained based on the dynamic trajectory amplitude from the first frame feature point to the last frame feature point in the first continuous frame image or the second continuous frame image and the texture complexity of the last frame image; the descriptors in the first feature point set and the second feature point set can be adaptively adjusted in the motion blur region and the texture rich region, so that the feature matching point pair set can be closer to the actual situation of the workshop, and the accuracy of generating the three-dimensional layout map of the workshop equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method for obtaining a three-dimensional layout diagram of workshop equipment, an electronic device, and a computer-readable storage medium. Background Technology

[0002] In modern manufacturing workshops, the scheduling of workshop equipment can positively impact production efficiency. Typically, staff can schedule equipment production based on the workshop layout information. However, accurately obtaining a 3D layout diagram of the workshop equipment is a pressing technical problem that needs to be solved. Summary of the Invention

[0003] In order to overcome the above-mentioned technical defects and at least solve the technical problem of difficulty in accurately obtaining three-dimensional layout drawings of workshop equipment, the embodiments of this application provide a method for obtaining three-dimensional layout drawings of workshop equipment, an electronic device, and a computer-readable storage medium.

[0004] The method for obtaining a three-dimensional layout drawing of workshop equipment according to an embodiment of this application includes:

[0005] Based on a preset algorithm, a first feature point set of the first camera and a second feature point set of the second camera are obtained; wherein the field of view of the first camera and the field of view of the second camera at least partially overlap.

[0006] Descriptor matching is performed on the first feature point set and the second feature point set to construct a set of feature matching point pairs; and

[0007] Based on the set of feature matching point pairs, the images captured by the first camera and the second camera are stitched together to obtain a three-dimensional layout map of the workshop equipment.

[0008] in,

[0009] The step of obtaining the first feature point set of the first camera and the second feature point set of the second camera based on a preset algorithm includes:

[0010] Using the SuperPoint network model, the first frame feature points and first frame descriptor of the first frame image in the target continuous frame image are obtained, as well as the last frame feature points and last frame initial descriptor of the last frame image; wherein, the target continuous frame image includes the first continuous frame image of the first camera in a preset time period or the second continuous frame image of the second camera in the preset time period.

[0011] Based on the dynamic trajectory amplitude from the feature points in the first frame to the feature points in the last frame and the texture complexity of the last frame image, the initial descriptor of the last frame is optimized using the descriptor of the first frame to obtain the descriptor of the last frame.

[0012] A target feature point set is constructed based on the last frame feature points and the last frame descriptor; wherein the target feature point set includes the first feature point set or the second feature point set.

[0013] In some implementations, the step of optimizing the initial descriptor of the last frame using the descriptor of the first frame based on the dynamic trajectory amplitude from the feature points of the first frame to the feature points of the last frame and the texture complexity of the last frame image, to obtain the descriptor of the last frame, includes:

[0014] Using an optical flow algorithm, the dynamic trajectory amplitude weight from the first frame feature point to the last frame feature point is obtained based on the covariance of the trajectory positions of the first frame feature point in the target continuous frame image.

[0015] Based on the grayscale gradient histogram entropy values ​​of the target consecutive frame images, the texture complexity weight of the last frame image is obtained; and

[0016] Based on the dynamic trajectory amplitude weight and the texture complexity weight, a correction weight is obtained. The correction weight and the first frame descriptor are used to correct the last frame initial descriptor, and the last frame descriptor is obtained.

[0017] In some implementations, the step of correcting the initial descriptor of the last frame using the correction weights and the first frame descriptor to obtain the last frame descriptor includes:

[0018] The final frame descriptor D is calculated using the following formula:

[0019]

[0020] Wherein, D1 is the first frame descriptor, D0 is the last frame initial descriptor, w is the correction weight, w1 is the dynamic trajectory amplitude weight, w2 is the texture complexity weight, tr(Cov) is the trace of the covariance of the feature points of the first frame in the target continuous frame image, and h is the gray-level gradient histogram entropy value of the last frame image. max It is the maximum grayscale gradient histogram entropy value in the target continuous frame image.

[0021] In some implementations, after constructing the target feature point set based on the last frame feature points and the last frame descriptor, the method further includes:

[0022] Optical flow algorithm is used to obtain the optical flow feature points of the last frame of the image;

[0023] Determine whether the feature points of the last frame are aligned with the optical flow feature points of the last frame;

[0024] If not, then delete the last frame feature point and the last frame descriptor from the target feature point set, and update the target feature point set.

[0025] In some implementations, the step of performing descriptor matching on the first set of feature points and the second set of feature points to construct a set of feature matching point pairs includes:

[0026] Using a fast nearest neighbor search algorithm, descriptor matching is performed on feature points at the same time point in the first and second feature point sets to construct an initial set of feature matching point pairs; and

[0027] The initial feature matching point pair set is cleaned to remove erroneous feature matching point pairs, thus obtaining the feature matching point pair set.

[0028] In some implementations, the step of stitching together the images captured by the first camera and the second camera based on the set of feature matching point pairs to obtain a three-dimensional layout map of the workshop equipment includes:

[0029] Construct a global coordinate system;

[0030] Based on the coordinates of each feature matching point pair in the feature matching point pair set in the global coordinate system, obtain the transformation matrix corresponding to the feature matching point pair set; wherein, the transformation matrix includes one or more of scaling parameters, horizontal translation parameters, vertical translation parameters, and rotation angle parameters; and

[0031] Based on the transformation matrix, the images captured by the first camera and the second camera are stitched together to obtain a partial or complete three-dimensional layout diagram of the workshop equipment.

[0032] In some embodiments, stitching together the images captured by the first camera and the second camera to obtain a partial or complete three-dimensional layout diagram of the workshop equipment includes:

[0033] Based on the set of feature matching points, a thin plate spline model for image deformation compensation is obtained;

[0034] Based on the thin plate spline model, deformation compensation is performed on the images captured by the first camera and the second camera to obtain the compensated optimized images of the first camera and the second camera;

[0035] The compensated and optimized image is optimized using bilinear interpolation to obtain a difference-optimized image between the first camera and the second camera, so that the coordinate values ​​of all pixels in the compensated and optimized image are integers; and

[0036] Based on the transformation matrix, the difference-optimized images from the first camera and the second camera are stitched together to obtain a partial or complete three-dimensional layout diagram of the workshop equipment.

[0037] In some implementations, after stitching together the images captured by the first camera and the second camera based on the feature matching point pair set to obtain a 3D layout map of the workshop equipment, the method further includes:

[0038] Based on the three-dimensional layout diagram of the workshop equipment, visual statistics are obtained; wherein the visual statistics include the number of workshop equipment and / or the location data of workshop equipment.

[0039] Based on the matching results of the visual statistical data and the actual statistical data of the workshop equipment, the 3D layout diagram of the workshop equipment is optimized.

[0040] The electronic device according to embodiments of this application includes a memory and a processor. The processor is used to execute a computer program stored in the memory to implement the acquisition method as described in any of the above embodiments.

[0041] The computer-readable storage medium of the present application embodiments stores a computer program that, when executed by a processor in an electronic device, implements the acquisition method as described in any of the above embodiments.

[0042] The method, electronic device, and computer-readable storage medium for obtaining a 3D layout diagram of workshop equipment according to embodiments of this application utilize a preset algorithm to obtain a first set of feature points from a first consecutive frame image and a second set of feature points from a second consecutive frame image. Then, descriptor matching is performed on the first and second feature point sets to construct a set of feature-matched point pairs. Finally, based on the set of feature-matched point pairs, the images captured by the first and second cameras are stitched together to obtain a 3D layout diagram of the workshop equipment. Because the preset algorithm can optimize the acquisition of the descriptor for the last frame based on the dynamic trajectory amplitude from the feature points in the first frame to the feature points in the last frame of the first or second consecutive frame image, as well as the texture complexity of the last frame image, the descriptors in the first and second feature point sets can adaptively adjust to motion-blurred and texture-rich regions. This allows the set of feature-matched point pairs to more closely reflect the actual situation of the workshop, improving the accuracy of generating the 3D layout diagram of the workshop equipment.

[0043] Additional aspects and advantages of embodiments of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of embodiments of this application. Attached Figure Description

[0044] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, wherein:

[0045] Figure 1 This is a flowchart illustrating a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application;

[0046] Figure 2 This is a schematic diagram of a production workshop scene according to certain embodiments of this application;

[0047] Figure 3 This is a schematic diagram of a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application;

[0048] Figure 4 This is a schematic diagram of the structure of an electronic device according to certain embodiments of this application;

[0049] Figure 5 This is a flowchart illustrating a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application;

[0050] Figure 6 This is a flowchart illustrating a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application;

[0051] Figure 7 This is a flowchart illustrating a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application;

[0052] Figure 8 This is a flowchart illustrating a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application;

[0053] Figure 9 This is a flowchart illustrating a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application;

[0054] Figure 10 This is a flowchart illustrating a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application;

[0055] Figure 11 This is a flowchart illustrating a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application;

[0056] Figure 12 This is a flowchart illustrating a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application;

[0057] Figure 13 This is a schematic diagram of a method for obtaining a three-dimensional layout diagram of workshop equipment according to certain embodiments of this application.

[0058] Explanation of key component symbols:

[0059] 100 electronic devices, 110 memory, 130 processors, 150 computer programs; 300 manufacturing workshops, 310 workshop equipment, 330 cameras. Detailed Implementation

[0060] The embodiments of this application will be further described below with reference to the accompanying drawings. The same or similar reference numerals in the drawings denote the same or similar elements or elements having the same or similar functions throughout. Furthermore, the embodiments of this application described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of this application, and should not be construed as limiting this application.

[0061] In modern manufacturing workshops, the production scheduling of workshop equipment can positively impact production efficiency. Typically, workers can schedule equipment production based on workshop layout information. However, in related technologies, obtaining workshop equipment layout information mainly relies on on-site photography or measurement by workers, followed by software rendering. This process is time-consuming and labor-intensive, and it's difficult to obtain workshop equipment layout information in real time, thus affecting production scheduling efficiency. To address this issue, this application provides a method for obtaining a three-dimensional layout drawing of workshop equipment. Figure 1 As shown), electronic equipment ( Figure 4 (as shown) and computer-readable storage media.

[0062] Please see Figure 1 The method for obtaining a three-dimensional layout drawing of workshop equipment according to the embodiments of this application includes:

[0063] Step 11: Based on a preset algorithm, obtain the first feature point set of the first camera and the second feature point set of the second camera; wherein the field of view of the first camera and the field of view of the second camera at least partially overlap.

[0064] Step 12: Perform descriptor matching on the first feature point set and the second feature point set to construct a feature matching point pair set.

[0065] Step 13 involves stitching together the images captured by the first and second cameras based on the feature matching point pair set to obtain a 3D layout diagram of the workshop equipment.

[0066] Step 11 includes steps 111 to 113.

[0067] Step 111: Using the SuperPoint network model, obtain the first frame feature points and first frame descriptor of the first frame image in the target continuous frame image, and the last frame feature points and last frame initial descriptor of the last frame image; wherein, the target continuous frame image includes the first continuous frame image of the first camera in a preset time period or the second continuous frame image of the second camera in a preset time period.

[0068] Step 112: Based on the dynamic trajectory amplitude from the feature points in the first frame to the feature points in the last frame and the texture complexity of the last frame image, optimize the initial descriptor of the last frame using the descriptor of the first frame to obtain the descriptor of the last frame.

[0069] Step 113: Construct a target feature point set based on the feature points and descriptors of the last frame; wherein the target feature point set includes a first feature point set or a second feature point set.

[0070] Please combine Figure 2 In some embodiments of this application, the manufacturing workshop 300 may include multiple workshop equipment 310 and at least two cameras 330 (a first camera and a second camera). The cameras 330 are capable of capturing images of the workshop equipment 310, and the fields of view of the at least two cameras 330 can cover part or all of the workshop equipment 310, with some overlap between their fields of view. Depending on the specific production needs of the manufacturing workshop 300, the workshop equipment 310 may include at least one of the following: processing equipment (lathes, milling machines, drilling machines, injection molding machines, bending machines, etc.), automated equipment (AGVs, robotic arms, etc.), and inspection equipment (vision inspection machines, coordinate measuring machines, etc.). The cameras 330 may be CCD cameras, CMOS cameras, infrared cameras, etc.

[0071] Optionally, before the camera 330 takes pictures, it is necessary to calibrate the camera 330 to obtain camera parameters, such as the intrinsic parameters (focal length, distortion) and extrinsic parameters (position, orientation) of the camera 330, so as to align the data of at least two cameras 330 to ensure the normal realization of subsequent image stitching. The camera 330 calibration method can be a traditional calibration method or a self-calibration method, etc., and this application does not limit it.

[0072] Please combine Figure 4 The method for obtaining the three-dimensional layout diagram of workshop equipment described above can be applied to electronic device 100. One embodiment of the electronic device 100 includes a memory 110 and a processor 130. The processor 130 executes the computer program 150 stored in the memory 110 to implement the acquisition method in steps 11 to 13. That is, the processor 130 executes the above steps.

[0073] Specifically, in step 11, the first camera can be any one of at least two cameras 330, such as the camera 330 that covers the largest number of workshop equipment 310 in its field of view. The second camera can be the camera adjacent to the first camera among the multiple cameras 330, and the fields of view of the first camera and the fields of view of the second camera at least partially overlap to facilitate stitching.

[0074] It should be noted that, for the second camera, acquiring the second consecutive frame image and acquiring the first consecutive frame image from the first camera can be done simultaneously to ensure frame time alignment, thereby improving feature matching accuracy and significantly enhancing image stitching quality. For example, the first and second consecutive frame images can be acquired simultaneously using a hardware trigger signal (such as a GPIO signal); or, for example, time can be aligned using software timestamps to achieve simultaneous acquisition of the first and second consecutive frame images.

[0075] Optionally, after acquiring the first consecutive frame image from the first camera or the second consecutive frame image from the second camera, the processor 130 can also perform grayscale normalization preprocessing on the images in the consecutive multi-frame images, such as performing grayscale normalization preprocessing only on the first and last frame images in the multi-frame images. This can simplify the data dimensions, improve processing efficiency, and also improve the stability of the algorithm.

[0076] The acquisition of the first or second consecutive frame image can be achieved by the processor 130 acquiring multiple consecutive frames of images from the first or second camera within a preset time period, and using the first frame image within the preset time period as the first frame image of the multiple consecutive frame image, and the last frame image within the preset time period as the last frame image of the multiple consecutive frame image.

[0077] Optionally, the preset time period can be known data. That is, the preset time period can be data that was set before the electronic device 100 left the factory, or it can be data entered by the user when using the electronic device 100 after it leaves the factory. For example, the preset time period can be 100ms. During this time period, the camera 330 can continuously capture 5 frames of images. The processor 130 can acquire the 5 images captured by the first camera within 100ms, select the starting frame image captured within 100ms as the first frame image, and select the ending frame image captured within 100ms as the last frame image.

[0078] It's important to note that feature points are pixel locations in an image that possess significant local characteristics. They typically correspond to corners, edge intersections, and textured areas within a scene. Feature points exhibit repeatable detection under image transformations (such as rotation, scaling, and changes in lighting). A feature point's position in the image can be represented using coordinates (x, y). A descriptor is a mathematical representation of the local region surrounding the feature point, usually a vector (e.g., 128-dimensional or 256-dimensional). The descriptor uniquely encodes the visual information of that region, enabling the correct matching of the same feature points in different images.

[0079] Optionally, the method for obtaining feature points can be at least one of the following: Harris corner detection algorithm, FAST algorithm, SIFT algorithm, DOG algorithm, SURF algorithm, SuperPoint network model, SuperGlue network, and D2-Net network. The method for obtaining descriptors can be at least one of the following: SIFT algorithm and SuperPoint network model.

[0080] In step 11, a preset algorithm is used. In step 111, the SuperPoint network model is an end-to-end neural network model based on deep learning for feature point detection and descriptor extraction. Through self-supervised training, it jointly optimizes feature point detection and descriptor generation, significantly improving the robustness and efficiency of feature matching. The SuperPoint network model can locate repeatable and stable feature points in an image and generate a highly discriminative descriptive vector (e.g., 256 dimensions) for each feature point.

[0081] Specifically, in step 111, both the first and second consecutive frame images can be used as target consecutive frame images for the preset algorithm. When the first consecutive frame image is used as the target consecutive frame image, the target feature point set output by the preset algorithm is the first feature point set; when the second consecutive frame image is used as the target consecutive frame image, the target feature point set output by the preset algorithm is the second feature point set.

[0082] Specifically, when the first and last frames of a target continuous frame image are input into the SuperPoint network model, the SuperPoint network model can process the first and last frames separately to output the first frame feature points and first frame descriptors (corresponding to the first frame feature points) in the first frame image, and the last frame feature points and last frame initial descriptors (corresponding to the last frame feature points) in the last frame image. It is understood that the number of first frame feature points and first frame descriptors can both be multiple, and there is a one-to-one correspondence between them; similarly, the number of last frame feature points and last frame initial descriptors can both be multiple, and there is a one-to-one correspondence between them.

[0083] In step 112, the motion trajectory amplitude can characterize the intensity of motion of a feature point from the first frame of the first frame to the corresponding feature point in the last frame of the last frame within the target continuous frame images. Methods for obtaining the motion trajectory amplitude may include, but are not limited to, optical flow algorithms and feature matching algorithms. Texture complexity can characterize the richness of texture information in the region surrounding the feature point in the last frame image. Methods for obtaining texture complexity may include, but are not limited to, image gradient statistics (calculating the image gradient of the neighborhood around the feature point in the last frame, and then calculating the variance or average magnitude of the gradient), local entropy, and feature point density.

[0084] Specifically, the processor 130 can optimize the initial descriptor of the last frame using the descriptor of the first frame based on the dynamic trajectory amplitude from the feature points of the first frame to the feature points of the last frame in the target continuous frame image and the texture complexity of the last frame image, and obtain the descriptor of the last frame. This can improve the anti-interference ability of the last frame descriptor, enable the last frame descriptor to continuously adapt to changes in the workshop environment (such as changes in lighting or occlusion), improve the stability and reliability of the last frame descriptor, and help improve the image stitching quality in subsequent steps (such as step 13).

[0085] In step 113, the pixel coordinates of the feature points in the last frame image can be represented as (x... i y i ), the first feature point set P t1 It can be represented as P t1 ={(x1, y1), (x2, y2),...(x n y n The last frame descriptor corresponding to the feature points of the last frame can be represented as d. i d i The descriptor vector is defined by the algorithm. For example, when using the SuperPoint network model to obtain the initial descriptor of the last frame, the dimension of the descriptor vector is 256.

[0086] For example, when the number of feature points in the last frame of the first consecutive frame image of the first camera obtained using the SuperPoint network model is three, the pixel coordinates of the three feature points in the last frame image can be (x1, y1), (x2, y2), and (x3, y3), respectively, and the last frame descriptors corresponding to the three feature points can be d1, d2, and d3, respectively. In this case, the first feature point set P t1 = {(x1, y1), (x2, y2), (x3, y3)}. Processor 130 can repeatedly execute steps 111 to 113 to obtain the second feature point set P of the second camera. t2 The second feature point set P t2 It can be represented as P t2 ={(x1', y1'), (x2', y2'),...(x n ',y n ')}.

[0087] In step 12, descriptor matching refers to the process of finding similar feature points by comparing descriptors in two feature point sets. Descriptor matching methods include, but are not limited to, brute-force matching, fast nearest neighbor search algorithms, K-nearest neighbor algorithms, and deep learning-based matching. In this application, by comparing the descriptors in the feature point sets (such as the first feature point set and the second feature point set) of two adjacent cameras 330, similar feature points in the first and second feature point sets are found, and a set of feature matching point pairs is constructed. The set of feature matching point pairs S can be represented as S = {(M1, M2)}, where M1 and M2 are the first feature point set P, respectively. t1 Second feature point set P t2 Feature points that can be matched with each other.

[0088] Optionally, the feature matching point pair set may contain not only the matching information of the first feature point set and the second feature point set, but also, if there are multiple cameras 330, by repeatedly executing steps 11 to 12, the feature matching point pair set may contain the relevant matching information between other cameras 330. In this case, the feature matching point pair set S may be represented as S = {(M1, M2), (M2, M3), (M3, M4)...}.

[0089] For example, please refer to Figure 3 The manufacturing workshop 300 includes three cameras 330, which are designated as a first camera, a second camera, and a third camera. The fields of view of the first and second cameras overlap, as do the fields of view of the second and third cameras. Figure 3 As shown, the image captured by the first camera is I1, the image captured by the second camera is I2, and the image captured by the third camera is I3. Region A in I1 overlaps with region A' in I2, and region B in I2 overlaps with region B' in I3.

[0090] Set the first camera, second camera, and third camera as the first camera in sequence, and repeat steps 11 to 12 to obtain the first feature point set P of the first camera. t1 The second feature point set P of the second camera t2 The third feature point set P of the third camera t3 Among them, the first feature point set P of the first camera t1 The second feature point set P of the second camera t2 Perform descriptor matching, and perform second feature point set P from the second camera. t2 and the third feature point set P of the third camera t3Descriptor matching is performed to construct a set of feature matching point pairs. Thus, the set of feature matching point pairs not only contains matching information from the first and second feature point sets, but also relevant matching information from the third camera. At this point, the set of feature matching point pairs S can be represented as S = {(M1, M2), (M2, M3)}, where M1 and M2 are the first feature point sets P and P, respectively. t1 Second feature point set P t2 The feature points that can be matched with each other, M2 and M3 are the first feature point set P respectively. t2 Second feature point set P t3 Feature points that can be matched with each other.

[0091] In step 13, the processor 130 can stitch together images captured by at least two cameras 330 according to the feature matching point pair set to obtain a three-dimensional layout map of the workshop equipment 310. In this way, compared with related technologies, on the one hand, it can reduce the difficulty of obtaining the layout information of the workshop equipment 310, reduce human intervention, avoid the introduction of erroneous data due to human error, and improve the accuracy of the three-dimensional layout map; on the other hand, it can improve the real-time performance of the acquisition of the three-dimensional layout map of the workshop equipment 310, provide real-time and accurate information support for production management and scheduling, and improve production efficiency and management level.

[0092] Optionally, step 13 further includes: stitching together images captured by at least two cameras 330 based on a set of feature point matching pairs to obtain a two-dimensional layout diagram of the workshop equipment 310; and using the equipment model to determine a three-dimensional model of the workshop equipment 310, and obtaining a three-dimensional layout diagram of the workshop equipment based on the three-dimensional model and the two-dimensional layout diagram. Thus, compared to the two-dimensional layout diagram, the three-dimensional layout diagram can improve the visualization effect of the workshop equipment layout, which is beneficial for achieving more refined production scheduling and management.

[0093] In the method for obtaining a 3D layout diagram of workshop equipment according to the embodiments of this application, a preset algorithm is used to obtain a first feature point set of a first consecutive frame image and a second feature point set of a second consecutive frame image. Then, descriptor matching is performed on the first feature point set and the second feature point set to construct a feature matching point pair set. Finally, based on the feature matching point pair set, the images captured by the first camera and the second camera are stitched together to obtain a 3D layout diagram of workshop equipment. Because the preset algorithm can optimize the acquisition of the descriptor of the last frame based on the dynamic trajectory amplitude from the feature point of the first frame to the feature point of the last frame in the first or second consecutive frame image and the texture complexity of the last frame image, the descriptors in the first feature point set and the second feature point set can be adaptively adjusted in motion-blurred areas and texture-rich areas, making the feature matching point pair set closer to the actual situation of the workshop and improving the accuracy of generating the 3D layout diagram of workshop equipment.

[0094] Please see Figure 5 In some implementations, step 112 of the preset algorithm includes:

[0095] Step 1121: Using the optical flow algorithm, obtain the dynamic trajectory amplitude weight from the feature point in the first frame to the feature point in the last frame based on the covariance of the trajectory positions of the feature point in the first frame in multiple frames of images.

[0096] Step 1123: Based on the grayscale gradient histogram entropy values ​​of the target consecutive frame images, obtain the texture complexity weight of the last frame image; and

[0097] Step 1125: Obtain the correction weight based on the dynamic trajectory amplitude weight and texture complexity weight, and use the correction weight and the first frame descriptor to correct the initial descriptor of the last frame, and obtain the last frame descriptor.

[0098] Specifically, in step 1121, obtaining the covariance of each trajectory position of the first frame feature point in the target continuous frame image using the optical flow algorithm can be achieved by: using the optical flow algorithm to obtain the trajectory positions of the first frame feature point in adjacent frame images, obtaining the feature point displacement based on the trajectory positions of the first frame feature point in two adjacent frame images, and finally obtaining the covariance Cov of each trajectory position of the first frame feature point in multiple frame images based on the feature point displacements between all two adjacent frame images in the multiple frame images. Specifically, the optical flow algorithm can be the pyramid optical flow algorithm.

[0099] For example, please refer to Figure 6 The image can have 5 frames. In this case, the optical flow algorithm is used to obtain the trajectory position of the feature point in the first frame across the 5 frames. Based on the position of the feature point in the first frame in the first frame and the position of the optical flow feature point in the second frame (the feature point in the second frame tracked by the optical flow algorithm that corresponds to the feature point in the first frame), the displacement 1 of the feature point between the first and second frames is obtained. Based on the trajectory position of the optical flow feature point in the second frame and the trajectory position of the optical flow feature point in the third frame, the displacement 1 of the feature point between the first and third frames is obtained. 2. Based on the trajectory positions of the optical flow feature points in the 3rd and 4th frames of the first frame feature points, obtain the displacement between the 3rd and 4th frames. 3. Based on the trajectory positions of the optical flow feature points in the 4th and 5th frames of the first frame feature points, obtain the displacement between the 4th and 5th frames. 4. Finally, based on displacements 1, 2, 3, and 4, obtain the covariance Cov of each trajectory position of the first frame feature points in the multi-frame images.

[0100] Using the optical flow algorithm, the dynamic trajectory amplitude weight from the feature point in the first frame to the feature point in the last frame can be obtained by calculating the dynamic trajectory amplitude weight using the covariance Cov of each trajectory position of the feature point in the first frame in the target continuous frame image.

[0101] In step 1123, the texture complexity weight of the last frame image can be obtained based on the gray-level gradient histogram entropy value of the target continuous frame images. This can be achieved by calculating the gray-level gradient histogram entropy value of each frame image. For example, the feature points of the target continuous frame images can be obtained using the SuperPoint network model. A preset pixel region (e.g., 5*5) is taken as the center of the feature points in each frame image, and the gray-level gradient histogram entropy value h of each frame image is calculated. Finally, the texture complexity weight of the last frame image is obtained based on the gray-level gradient histogram entropy value h of the multiple frames.

[0102] The gray-level gradient histogram entropy value list for consecutive frames of the target image is h-list = (h1, h2, ..., hk), where k ≥ 3. For example, the image can have 5 frames, and the calculated gray-level gradient histogram entropy values ​​for the 5 frames are h1, h2, h3, h4, and h5, respectively. In this case, the gray-level gradient histogram entropy value list for multiple frames is h-list = (h1, h2, h3, h4, h5).

[0103] In some implementations, step 1125: correcting the initial descriptor of the last frame using the corrected weights and the first frame descriptor to obtain the last frame descriptor includes:

[0104] The final frame descriptor D is calculated using the following formulas:

[0105]

[0106] Where D1 is the first frame descriptor, D0 is the initial target descriptor, w1 is the dynamic trajectory amplitude weight, w2 is the texture complexity weight, tr(Cov) is the trace of the covariance of the feature points in the first frame in the continuous frame images of the target, and h is the gray-level gradient histogram entropy value of the last frame image. max It represents the maximum grayscale gradient histogram entropy value in the target consecutive frames of the image.

[0107] Specifically, in the above embodiment, λ is the attenuation coefficient, and the value of λ can be 0.5. In step 1125, through dynamic weight adjustment, the last frame descriptor can continuously adapt to changes in the workshop environment, improving the stability and reliability of feature points.

[0108] Please see Figure 7 In some implementations, the algorithm further includes the following after step 113 of the preset algorithm:

[0109] Step 1131: Use the optical flow algorithm to obtain the optical flow feature points of the last frame image;

[0110] Step 1133: Determine whether the feature points of the last frame are aligned with the optical flow feature points of the last frame;

[0111] Step 1135: If not, delete the last frame feature points and last frame descriptors from the target feature point set and update the target feature point set.

[0112] Specifically, in step 1131, obtaining the optical flow feature points of the last frame of the last frame image can be achieved by using an optical flow algorithm and the feature points of the first frame to obtain the optical flow feature points of the last frame of the last frame image. For example, the pyramid optical flow algorithm can be used to continuously track the feature points of the first frame to the last frame image to obtain the optical flow feature points of the last frame.

[0113] In step 1133, the method for determining whether the feature points of the last frame are aligned with the optical flow feature points of the last frame can be the Euclidean distance threshold method or the reprojection error method, etc. Specifically, when using the Euclidean distance threshold method, the Euclidean distance between the feature points of the last frame and the optical flow feature points of the last frame can be calculated. If the Euclidean distance is less than a preset distance threshold, the feature points of the last frame are determined to be aligned with the optical flow feature points of the last frame; if the Euclidean distance is greater than the preset distance threshold, the feature points of the last frame are determined to be misaligned with the optical flow feature points of the last frame.

[0114] For example, if the preset distance threshold is 5 pixels, then if the Euclidean distance between the feature point and the optical flow feature point of the last frame is less than 5 pixels, it is determined that the feature point of the last frame is aligned with the optical flow feature point of the last frame; if the Euclidean distance between the feature point and the optical flow feature point of the last frame is greater than 5 pixels, it is determined that the feature point of the last frame is not aligned with the optical flow feature point of the last frame.

[0115] In step 1135, if not, the last frame feature point and the last frame descriptor are deleted from the target feature point set, and the target feature point set is updated. That is, if the Euclidean distance between the last frame feature point and the last frame optical flow feature point is greater than a preset distance threshold, the last frame feature point and its corresponding last frame descriptor are deleted from the target feature point set to update the target feature point set. This can eliminate abnormal feature points caused by occlusion or other reasons, improve the quality of feature points, and help improve the image stitching quality.

[0116] Please see Figure 8 In some implementations, step 12 includes:

[0117] Step 121: Using the Fast Nearest Neighbor Search algorithm, perform descriptor matching on feature points at the same time point in the first and second feature point sets to construct an initial set of feature matching point pairs; and

[0118] Step 123: Clean the initial set of feature matching points, remove erroneous feature matching point pairs, and obtain the set of feature matching point pairs.

[0119] Specifically, in step 121, the Fast Approximate Nearest Neighbor Search (ANN) library algorithms (such as FLANN, FAISS, KD-Tree, etc.) are efficient algorithms for calculating the nearest neighbor to the target point in a dataset, suitable for high-dimensional data (such as image feature descriptors). This application can utilize the Fast Approximate Nearest Neighbor Search library algorithms to match descriptors, that is, to perform descriptor matching on feature points at the same time point in the first and second feature point sets, thereby constructing an initial set of feature matching point pairs. Compared to brute-force matching, the Fast Approximate Nearest Neighbor Search library algorithms can significantly improve matching efficiency while maintaining high matching accuracy during the descriptor matching process through approximation or optimization strategies.

[0120] For example, the first feature point set P t1 ={(x1, y1), (x2, y2),...(x n y n The second feature point set P t2 ={(x1', y1'), (x2', y2'),...(x n ',y n In the process of descriptor matching using the fast nearest neighbor search library algorithm, the descriptors corresponding to the feature points in the first feature point set can be used as a database. For each descriptor corresponding to the feature points in the second feature point set, its nearest neighbor (such as the descriptor with the smallest Euclidean distance) is searched in the database, thereby obtaining the initial set of feature matching point pairs.

[0121] Because two images may contain similar local features, such as repetitive textures (like grass or brick walls), different feature points may have similar descriptors. This could lead to incorrectly matched feature point pairs in the initial feature matching point pair set, affecting the stitching quality of subsequent images. For example, the stitched image may show obvious breaks.

[0122] In step 123, the initial set of feature matching point pairs is cleaned to remove erroneous feature matching point pairs, thus obtaining a complete set of feature matching point pairs. This prevents incorrectly matched feature matching point pairs from affecting subsequent image stitching, thereby improving the image stitching quality. Data cleaning methods include, but are not limited to, Random Sample Consensus (RANSAC), distance ratio testing, cross-validation, and geometric consistency constraints.

[0123] Please see Figure 9 In some implementations, step 13 includes:

[0124] Step 131: Construct a global coordinate system;

[0125] Step 133: Based on the coordinates of each feature matching point pair in the feature matching point pair set in the global coordinate system, obtain the transformation matrix corresponding to the feature matching point pair set; wherein, the transformation matrix includes one or more of the following: scaling parameters, horizontal translation parameters, vertical translation parameters, and rotation angle parameters; and

[0126] Step 135: Based on the transformation matrix, stitch together the images captured by the first camera and the second camera to obtain a partial or complete 3D layout diagram of the workshop equipment.

[0127] Specifically, please combine Figure 2 In step 131, the purpose of constructing a global coordinate system is to unify the images captured by all cameras 330 under the same reference system, enabling them to be correctly aligned and facilitating image stitching. In some embodiments of this application, the coordinate system of any one of the images captured by at least two cameras 330 can be selected as the global coordinate system, and other images can be mapped to this coordinate system through transformation matrices. For example, when stitching three images, and the first and second images have overlapping areas, and the second and third images have overlapping areas, the coordinate system of the second image can be selected as the global coordinate system. This can reduce accumulated errors and improve the image stitching quality in the subsequent stitching process.

[0128] In step 133, the feature matching point pair set can be constructed by descriptor matching of feature point sets from at least two cameras 330 (the two cameras 330 are adjacent and their fields of view overlap). Based on the feature point coordinates of each feature matching point pair in the feature matching point pair set, and the coordinates of each feature matching point pair in the global coordinate system, the transformation matrix corresponding to the feature matching point pair set can be obtained. This allows the images captured by at least two cameras 330 to be aligned to the global coordinate system in subsequent steps using the transformation matrix, thereby achieving image stitching.

[0129] For example, suppose there are two images I1 and I2. The set of feature matching point pairs formed by I1 and I2 contains a feature matching point pair (M1, M2), where the coordinates of M1 are (x1, y1) and the coordinates of M2 are (x2, y2). The transformed coordinates of M1 in the global coordinate system are (x1', y1'), and the transformed coordinates of M2 in the global coordinate system are (x2', y2'). The relationship between (x2, y2) and (x2', y2') satisfies:

[0130]

[0131] H is the transformation matrix. Specifically, when I1's coordinate system is the global coordinate system, (x1', y1') and (x1, y1) are essentially the same, and (x1, y1) and (x2, y2) match. In this case, the value of x2' can be replaced with x1, and the value of y2' can be replaced with y1, thus obtaining the transformation matrix H. The transformation matrix H can be:

[0132]

[0133] Where a = Scosθ, b = -Scosθ, c = T x , d=Ssinθ, e=Scosθ, f=T y .

[0134] Furthermore, please combine Figure 10 Step 133 further includes: obtaining the error between the coordinates of each feature matching point pair in the global coordinate system, wherein the specific formula for calculating the error includes:

[0135] e=(S*x2'*cosθ-S*y2'*sinθ+T x -x1') 2 +(S*x2'*sinθ-S*y2'*cosθ+T y -y1') 2 ;

[0136] Where S is the scaling parameter, θ is the rotation angle parameter, and T x T is the horizontal translation parameter. y Let (x1', y1') and (x2', y2') be the vertical translation parameters, and (x1', y1') and (x2', y2') be the coordinates of a feature matching point pair in the global coordinate system. Obtain the inverse residual projection sum of at least two stitched images. Based on the inverse residual projection sum, use an iterative algorithm to update the rotation angle parameter, horizontal translation parameter, vertical translation parameter, and scaling parameter. Then, based on the updated rotation angle parameter θ' and horizontal translation parameter T... x Vertical translation parameter T y Update the transformation matrix using 'and scaling parameter S'.

[0137] Specifically, assuming there are two images I1 and I2, where I1 is in the global coordinate system, and the set of feature matching point pairs formed by I1 and I2 contains a feature matching point pair (M1, M2), where M1 has coordinates (x1, y1) and M2 has coordinates (x2, y2). The transformed coordinates of M1 in the global coordinate system are (x1', y1'), and the transformed coordinates of M2 in the global coordinate system are (x2', y2'). The error e between (x2', y2') and (x1', y1') satisfies the above calculation formula. Therefore, the sum of the inverse residual projections of at least two image stitched together is E = ∑ i e i .

[0138] Updating rotation angle parameters, horizontal translation parameters, vertical translation parameters, and scaling parameters using iterative algorithms can include: initialization parameters (θ = 0, T...). x =0,T y =0, S=1), choose a suitable learning rate α=0.01, and a maximum number of iterations of 1000; respectively for θ, T x T y Take the partial derivative of S to obtain new parameters, and iterate until the iteration condition is met (the change in parameters is less than the threshold e). -6 If the latter's iteration count is greater than the maximum iteration count, the iteration ends, and the updated rotation angle parameter θ' and horizontal translation parameter T are obtained. x Vertical translation parameter T y 'and scaling parameter S'.

[0139] Please see Figure 11 In some implementations, step 135: stitching together the images captured by the first camera and the second camera to obtain a partial or complete three-dimensional layout diagram of the workshop equipment, including:

[0140] Step 1351: Obtain the thin plate spline model for image deformation compensation based on the feature matching point pair set;

[0141] Step 1353: Based on the thin plate spline model, perform deformation compensation on the images captured by the first camera and the second camera to obtain the compensated and optimized images of the first camera and the second camera;

[0142] Step 1355: Perform bilinear interpolation optimization on the compensated and optimized image to obtain the difference optimized image between the first camera and the second camera, so that the coordinate values ​​of the pixels in the compensated and optimized image are all integers; and

[0143] Step 1357: Based on the transformation matrix, stitch together the difference optimized images from the first camera and the second camera to obtain a partial or complete 3D layout map of the workshop equipment.

[0144] Specifically, in step 1351, the thin plate spline (TPS) model is a non-rigid transformation model used to describe the smooth mapping relationship between two planes. Its core idea is to minimize bending energy so that the feature matching point pairs are aligned while maintaining overall smoothness.

[0145] For example, the set of feature matching point pairs can be Where, p i =(xi, y) i ), qi=(x i ',y i At this point, the steps to obtain the thin plate spline model may include: First, constructing a non-rigid mapping function f(p), where f(p) satisfies: f(p) i )≈q i f(p) can specifically be:

[0146]

[0147] in, For the affine part, A represents the global affine transformation parameters (global rotation, translation, and scaling). For the nonlinear part, φ(r) = r²lnr is the radial basis function, describing non-rigid deformation, w j The weight parameters are those associated with the matching points. Then, the parameters A and W are solved by minimizing the energy function to obtain the thin plate spline model.

[0148] Optionally, the thin plate spline model can be updated periodically. For example, the thin plate spline module can be recalculated every preset number of frames to enable the thin plate spline model to better adapt to changes in the workshop scene (such as changes in lighting).

[0149] In step 1353, using the thin-plate spline model obtained in step 1351, each pixel in the image captured by the first camera is mapped to a new position to obtain a compensated and optimized image of the first camera. Similarly, each pixel in the images captured by the adjacent cameras is mapped to a new position to obtain a compensated and optimized image of the adjacent cameras. This ensures that the feature point positions in the images captured by the first camera and the adjacent cameras are strictly matched, and the transition between non-feature point positions is natural, which helps improve the image stitching quality.

[0150] In step 1355, bilinear interpolation is a method for estimating pixel values ​​at non-integer coordinates on a two-dimensional image or grid. It calculates the pixel value at the target location by weighted averaging of the four nearest integer pixels, ensuring a smooth and continuous interpolation result. Bilinear interpolation can interpolate non-integer coordinates into pixel values ​​at integer coordinates, ensuring the integrity of the pixel grid in the output image (difference-optimized image).

[0151] It should be noted that after deformation compensation using the thin-plate spline model, some pixels may not fall exactly on integer coordinates. Therefore, directly stitching the compensated and optimized image may result in jagged edges and holes in the stitched image, affecting its quality. Thus, in step 1755, bilinear interpolation optimization is performed on the compensated and optimized image. That is, the compensated and optimized image is traversed, and interpolation calculations are performed on each non-integer coordinate to obtain a difference-optimized image. Using the difference-optimized image for image stitching prevents jagged edges and holes in the stitched image, improving overall image quality.

[0152] Alternatively, nearest neighbor interpolation, bicubic interpolation, grid-based interpolation, and inverse distance weighted interpolation can be used to optimize the compensated image so that the coordinate values ​​of the pixels in the compensated image are all integers.

[0153] Please see Figure 2 and Figure 12 In some implementations, after step 13, the acquisition method further includes:

[0154] Step 14: Based on the 3D layout diagram of the workshop equipment, obtain visual statistics; the visual statistics include the quantity of workshop equipment 310 and / or the location data of workshop equipment 310; and

[0155] Step 15: Optimize the 3D layout diagram of the workshop equipment based on the matching results of visual statistical data and actual statistical data of workshop equipment 310.

[0156] Specifically, in step 14, the method for obtaining visual statistical data can be: directly parsing the 3D layout diagram of the workshop equipment based on the 3D model, rendering the 3D layout diagram of the workshop equipment into a 2D diagram (such as a top view), and then using an object detection model for detection. Of course, other methods for obtaining visual statistical data can also be included, which will not be listed here.

[0157] In step 15, taking visual statistical data including the quantity and location data of workshop equipment 310 as an example: If the visual statistical data matches the actual statistical data of workshop equipment 310, that is, the quantity and location data of workshop equipment 310 are consistent with the actual statistical data, then the 3D layout diagram of the workshop equipment is not updated; if the visual statistical data does not match the actual statistical data of workshop equipment 310, such as the quantity and / or location data of workshop equipment 310 being inconsistent with the actual statistical data, then the 3D layout diagram of the workshop equipment is updated.

[0158] For example, please refer to Figure 13 If the 3D layout drawing of the workshop equipment obtained in step 13 is as follows Figure 13As shown in Figure a, and given the discrepancy between the location data of workshop equipment 310 and the actual statistical data, the actual location data of workshop equipment 310 in the visual statistical data is used to correct the location data of workshop equipment 310 in the visual statistical data, thereby updating the 3D layout diagram of the workshop equipment. The updated 3D layout diagram of the workshop equipment is shown below. Figure 13 As shown in Figure b, this improves the accuracy of the 3D layout of workshop equipment, providing accurate and real-time information support for production scheduling and enhancing production efficiency.

[0159] Optionally, please combine Figure 4 The electronic device 100 may also include a prompter. Specifically, if the number of workshop equipment 310 is inconsistent with the actual number of workshop equipment 310 in the statistical data, it indicates that workshop equipment 310 has been missing, moved, or added, or that the camera 330 is obstructed. In this case, the processor 130 can control the prompter to work so that the prompter sends a prompt message to the staff, thereby reminding the staff that the number of workshop equipment 310 has changed. The prompter includes, but is not limited to, indicator lights, speakers, buzzers, displays, and vibration motors.

[0160] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device 100 provided in one embodiment of this application. It is understood that... Figure 4 The illustrated structure is merely an example of electronic device 100 and does not constitute a specific limitation on electronic device 100. Electronic device 100 can be a computer, mobile phone, tablet computer, personal digital assistant (PDA), or other device with applications installed. Electronic device 100 may include more or fewer components than illustrated, or combine certain components, or different components. For example, electronic device 100 may also include input / output devices, network access devices, buses, etc.

[0161] The electronic device 100 provided in this application includes a memory 110 and a processor 130. It should be noted that the electronic device 100 in this embodiment is the same as the electronic device 100 in the above embodiments. Therefore, the explanations and descriptions of the electronic device 100 in the previous embodiments also apply to the electronic device 100 in this embodiment, and similarly, the explanations and descriptions of the electronic device 100 in this embodiment also apply to the electronic device 100 in the previous embodiments.

[0162] Specifically, memory 110 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM).

[0163] The random access memory can be directly read and written by the processor 130. It can be used to store executable programs of other running programs, as well as user and application data. The random access memory can include, but is not limited to, static random-access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), and double data rate synchronous dynamic random access memory (DDR SDRAM).

[0164] Non-volatile memory can also store executable programs and user and application data, and can be pre-loaded into random access memory for direct reading and writing by the processor 130. Non-volatile memory can include disk storage devices and flash memory. For example, flash memory can be Nand Flash.

[0165] The memory 110 stores a computer program 150, which is configured to be executed by the processor 130 to implement the acquisition method in any of the above embodiments. The computer program 150 may include multiple instructions, which, when executed by the processor 130, can implement the method for acquiring the three-dimensional layout drawing of the workshop equipment in any of the above embodiments.

[0166] In other embodiments, the electronic device 100 may also include an external memory interface for connecting to an external memory to expand the storage capacity of the electronic device 100.

[0167] The processor 130 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or a processor, or any conventional processor.

[0168] The processor 130 provides computing and control capabilities. For example, the processor 130 is used to execute the computer program 150 stored in the memory 110 to implement the method for obtaining the three-dimensional layout drawing of the workshop equipment in any of the above embodiments.

[0169] It should be noted that the explanation of the method for obtaining the three-dimensional layout diagram of the workshop equipment in the foregoing embodiments also applies to the electronic device 100 in this embodiment, and will not be elaborated here.

[0170] In the electronic device 100 of this application, a preset algorithm is used to obtain a first feature point set of a first consecutive frame image and a second feature point set of a second consecutive frame image. Then, descriptor matching is performed on the first and second feature point sets to construct a feature matching point pair set. Finally, based on the feature matching point pair set, the images captured by the first and second cameras are stitched together to obtain a three-dimensional layout map of the workshop equipment. Because the preset algorithm can optimize the acquisition of the last frame descriptor based on the dynamic trajectory amplitude from the first frame feature point to the last frame feature point in the first or second consecutive frame image and the texture complexity of the last frame image, the descriptors in the first and second feature point sets can be adaptively adjusted in motion-blurred areas and texture-rich areas. This makes the feature matching point pair set closer to the actual situation of the workshop, improving the accuracy of generating the three-dimensional layout map of the workshop equipment.

[0171] The computer-readable storage medium of the embodiments of this application stores a computer program 150. When the computer program 150 is executed by the processor 130 in the electronic device 100, it implements the method for obtaining the three-dimensional layout drawing of the workshop equipment as described in any of the above embodiments.

[0172] The computer-readable storage medium can be the internal memory 110 of the electronic device 100 in the above embodiments, such as the hard disk or memory of the electronic device 100. The computer-readable storage medium can also be an external storage device of the electronic device 100, such as a plug-in hard disk, smart media card (SMC), flash card, etc. equipped on the electronic device 100.

[0173] In some embodiments, a computer-readable storage medium may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application program required for at least one function, etc.; and the data storage area may store data created based on the use of the electronic device 100, etc.

[0174] It should be noted that the method for obtaining the three-dimensional layout diagram of the workshop equipment and the explanation of the electronic device 100 in the foregoing embodiments are also applicable to the computer-readable storage medium of this embodiment, and will not be described in detail here.

[0175] In the computer-readable storage medium of this application, a preset algorithm is used to obtain a first feature point set of a first consecutive frame image and a second feature point set of a second consecutive frame image. Then, descriptor matching is performed on the first and second feature point sets to construct a feature matching point pair set. Finally, based on the feature matching point pair set, the images captured by the first and second cameras are stitched together to obtain a 3D layout diagram of the workshop equipment. Because the preset algorithm can optimize the acquisition of the last frame descriptor based on the dynamic trajectory amplitude from the first frame feature point to the last frame feature point in the first or second consecutive frame image, as well as the texture complexity of the last frame image, the descriptors in the first and second feature point sets can adaptively adjust to motion-blurred and texture-rich regions. This allows the feature matching point pair set to more closely reflect the actual situation of the workshop, improving the accuracy of the generated 3D layout diagram of the workshop equipment.

[0176] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0177] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.

[0178] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, a computer storage medium can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer storage media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer storage medium could even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in the computer memory.

[0179] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0180] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer storage medium, and when executed, it includes one or a combination of the steps of the method embodiments. Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc.

[0181] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for obtaining a three-dimensional layout drawing of workshop equipment, characterized in that, include: Based on a preset algorithm, a first feature point set of the first camera and a second feature point set of the second camera are obtained; wherein the field of view of the first camera and the field of view of the second camera at least partially overlap. Descriptor matching is performed on the first feature point set and the second feature point set to construct a set of feature matching point pairs; and Based on the set of feature matching point pairs, the images captured by the first camera and the second camera are stitched together to obtain a three-dimensional layout map of the workshop equipment. The step of obtaining the first feature point set of the first camera and the second feature point set of the second camera based on a preset algorithm includes: Using the SuperPoint network model, the first frame feature points and first frame descriptor of the first frame image in the target continuous frame image are obtained, as well as the last frame feature points and last frame initial descriptor of the last frame image; wherein, the target continuous frame image includes the first continuous frame image of the first camera in a preset time period or the second continuous frame image of the second camera in the preset time period. Based on the dynamic trajectory amplitude from the feature points in the first frame to the feature points in the last frame and the texture complexity of the last frame image, the initial descriptor of the last frame is optimized using the descriptor of the first frame to obtain the descriptor of the last frame. A target feature point set is constructed based on the last frame feature points and the last frame descriptor; wherein the target feature point set includes the first feature point set or the second feature point set.

2. The acquisition method according to claim 1, characterized in that, The process of optimizing the initial descriptor of the last frame using the descriptor of the first frame based on the dynamic trajectory amplitude from the feature points of the first frame to the feature points of the last frame and the texture complexity of the last frame image, to obtain the descriptor of the last frame, includes: Using an optical flow algorithm, the dynamic trajectory amplitude weight from the first frame feature point to the last frame feature point is obtained based on the covariance of the trajectory positions of the first frame feature point in the target continuous frame image. Based on the grayscale gradient histogram entropy values ​​of the target consecutive frame images, the texture complexity weight of the last frame image is obtained; and Based on the dynamic trajectory amplitude weight and the texture complexity weight, a correction weight is obtained. The correction weight and the first frame descriptor are used to correct the last frame initial descriptor, and the last frame descriptor is obtained.

3. The acquisition method according to claim 2, characterized in that, The step of correcting the initial descriptor of the last frame using the corrected weights and the first frame descriptor to obtain the last frame descriptor includes: The final frame descriptor D is calculated using the following formula: Wherein, D1 is the first frame descriptor, D0 is the last frame initial descriptor, w is the correction weight, w1 is the dynamic trajectory amplitude weight, w2 is the texture complexity weight, tr(Cov) is the trace of the covariance of the feature points of the first frame in the target continuous frame image, and h is the gray-level gradient histogram entropy value of the last frame image. max It is the maximum grayscale gradient histogram entropy value in the target continuous frame image.

4. The acquisition method according to claim 1, characterized in that, After constructing the target feature point set based on the last frame feature points and the last frame descriptor, the method further includes: Optical flow algorithm is used to obtain the optical flow feature points of the last frame of the image; Determine whether the feature points of the last frame are aligned with the optical flow feature points of the last frame; If not, then delete the last frame feature point and the last frame descriptor from the target feature point set, and update the target feature point set.

5. The acquisition method according to claim 1, characterized in that, The step of performing descriptor matching on the first feature point set and the second feature point set to construct a feature matching point pair set includes: Using a fast nearest neighbor search algorithm, descriptor matching is performed on feature points at the same time point in the first and second feature point sets to construct an initial set of feature matching point pairs; and The initial feature matching point pair set is cleaned to remove erroneous feature matching point pairs, thus obtaining the feature matching point pair set.

6. The acquisition method according to claim 1, characterized in that, The step of stitching together the images captured by the first camera and the second camera based on the feature matching point pair set to obtain a 3D layout map of the workshop equipment includes: Construct a global coordinate system; Based on the coordinates of each feature matching point pair in the feature matching point pair set in the global coordinate system, obtain the transformation matrix corresponding to the feature matching point pair set; wherein, the transformation matrix includes one or more of scaling parameters, horizontal translation parameters, vertical translation parameters, and rotation angle parameters; and Based on the transformation matrix, the images captured by the first camera and the second camera are stitched together to obtain a partial or complete three-dimensional layout diagram of the workshop equipment.

7. The method for obtaining according to claim 6, characterized in that, The step of stitching together the images captured by the first camera and the second camera to obtain a partial or complete three-dimensional layout diagram of the workshop equipment includes: Based on the set of feature matching points, a thin plate spline model for image deformation compensation is obtained; Based on the thin plate spline model, deformation compensation is performed on the images captured by the first camera and the second camera to obtain the compensated optimized images of the first camera and the second camera; The compensated and optimized image is optimized using bilinear interpolation to obtain a difference-optimized image between the first camera and the second camera, so that the coordinate values ​​of all pixels in the compensated and optimized image are integers; and Based on the transformation matrix, the difference-optimized images from the first camera and the second camera are stitched together to obtain a partial or complete three-dimensional layout diagram of the workshop equipment.

8. The acquisition method according to claim 1, characterized in that, After stitching together the images captured by the first camera and the second camera based on the feature matching point pair set to obtain a 3D layout map of the workshop equipment, the method further includes: Based on the three-dimensional layout diagram of the workshop equipment, visual statistics are obtained; wherein the visual statistics include the number of workshop equipment and / or the location data of workshop equipment. Based on the matching results of the visual statistical data and the actual statistical data of the workshop equipment, the 3D layout diagram of the workshop equipment is optimized.

9. An electronic device, characterized in that, The electronic device includes: Memory; and A processor for executing a computer program stored in the memory to implement the acquisition method as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor in an electronic device, implements the acquisition method as described in any one of claims 1-8.