Automated recognition of reconstruction of integrated decisions, systems, methods, electronic devices, and media
By using semantic recognition and proportion judgment of image sets, the system automatically identifies image scene patterns and calculates initialization parameters, solving the problem that existing technologies cannot identify image scene patterns and achieving automation and simplification of automatic 3D reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-02
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies cannot identify the scene pattern of an image and then calculate the initialization parameters of the image based on the scene pattern.
By acquiring an image set and performing semantic recognition, the proportion of semantic labels for each pixel is determined. Initialization parameters are obtained using incremental and global spatial calculations, including the determination of the proportion of small and large semantic labels. The image scene type is then determined and corresponding calculations are performed.
It enables automatic 3D reconstruction even when user input parameters are incomplete or missing, reducing the professional requirements and entry barriers for reconstruction personnel.
Smart Images

Figure CN116310082B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional reconstruction technology, and more specifically to a reconstruction system, method, electronic device, and medium for automatic identification and comprehensive decision-making. Background Technology
[0002] 3D reconstruction refers to the process of reconstructing 3D information from single-view or multi-view images. It requires first acquiring 2D images of 3D objects with a camera, then establishing an effective imaging model through camera calibration, solving for the camera's intrinsic and extrinsic parameters, and combining the image matching results to obtain the coordinates of 3D points in space, thereby achieving the purpose of 3D reconstruction.
[0003] In existing technologies, the initialization parameters of an image are calculated based on the scene pattern of the image after the scene pattern of the image cannot be identified. Summary of the Invention
[0004] The purpose of this invention is to provide an automatic identification and comprehensive decision-making reconstruction system, method, electronic device, and medium to solve the problem in the prior art that it is impossible to calculate the initialization parameters of an image based on the scene pattern of the image.
[0005] To achieve the above objectives, embodiments of the present invention provide an automatic identification and comprehensive decision reconstruction method, the method specifically including:
[0006] A set of images for 3D reconstruction is acquired, and semantic recognition is performed on the images in the set to obtain semantic labels for each image pixel. The semantic labels include large semantic labels and small semantic labels.
[0007] Determine the first proportion of the pixels with the small semantic tags in the total number of pixels, and determine whether the first proportion is within a first threshold range. If so, perform incremental spatial calculation on the image to obtain the initialization parameters corresponding to the image.
[0008] Determine the second proportion of pixels with the large semantic tag in the total number of pixels, and determine whether the second proportion is within the range of a second threshold. If so, determine whether the initialization parameters of the images in the image set are missing. If so, perform incremental spatial calculation on the image set to obtain the initialization parameters corresponding to the images.
[0009] Based on the above technical solution, the present invention can be further improved as follows:
[0010] Further, determining the first proportion of pixels with the small semantic tag in the total number of pixels, and judging whether the first proportion is within a first threshold range, includes:
[0011] When the first proportion value is within the first threshold range, the image is determined to be a small object scene, wherein the first threshold range is [0.3, 0.8].
[0012] Further, the incremental spatial calculation of the image to obtain the initialization parameters corresponding to the image includes:
[0013] Based on two matching images in the image set, the pose data of the two images are obtained by calculating the fundamental matrix and / or the essential matrix;
[0014] After obtaining the pose data, one image is added each time. The pose data of the newly added image is obtained by calculating the fundamental matrix and / or the essential matrix, thereby obtaining the initialization parameters corresponding to each image in the image set.
[0015] Further, determining the second proportion of pixels with the large semantic tag in the total number of pixels, and judging whether the second proportion is within the range of the second threshold, includes:
[0016] When the second proportion value is within the range of the second threshold, the image is determined to be a large city scene, wherein the range of the second threshold is [0.5, 0.9].
[0017] Further, determining whether the initialization parameters of the images in the image set are missing includes:
[0018] Determine the camera mode. When the camera mode is a multi-camera mode, store the images corresponding to the same camera in the same image set.
[0019] Further, determining whether the initialization parameters of the images in the image set are missing includes:
[0020] When the camera mode is multi-camera mode, a separate camera model is created for each camera.
[0021] Further, the step of determining whether the initialization parameters of the images in the image set are missing, and if so, performing incremental spatial calculations on the image set to obtain the initialization parameters corresponding to the images, includes:
[0022] When the initialization parameters of the images in the image set are complete, a global spatial calculation is performed on the image set to obtain the initialization parameters corresponding to the images.
[0023] An automatic identification and comprehensive decision-making reconstruction system includes:
[0024] The acquisition module is used to acquire the image set for 3D reconstruction;
[0025] The semantic recognition module is used to perform semantic recognition on the images in the image set to obtain semantic labels for each image pixel, wherein the semantic labels include large semantic labels and small semantic labels;
[0026] The determining module is used to determine the first proportion of the pixels with the small semantic tags in the total number of pixels;
[0027] Also used for:
[0028] Determine the second proportion of the pixels with the large semantic tag in the total number of pixels;
[0029] The judgment module is used to determine whether the first proportion value is within the first threshold range. If so, it performs incremental spatial calculation on the image to obtain the initialization parameters corresponding to the image.
[0030] Also used for:
[0031] Determine whether the second proportion value is within the range of the second threshold. If so, determine whether the initialization parameters of the images in the image set are missing. If so, perform incremental spatial calculation on the image set to obtain the initialization parameters corresponding to the images.
[0032] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the steps of the method described herein.
[0033] A non-transitory computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method.
[0034] The embodiments of the present invention have the following advantages:
[0035] The reconstruction method for automatic identification and comprehensive decision-making in this invention acquires an image set for 3D reconstruction, performs semantic recognition on the images in the image set to obtain semantic labels for each image pixel, wherein the semantic labels include large semantic labels and small semantic labels; determines a first proportion value of pixels with the small semantic labels in the total number of pixels, and determines whether the first proportion value is within a first threshold range. If so, incremental spatial calculation is performed on the image to obtain the initialization parameters corresponding to the image; determines a second proportion value of pixels with the large semantic labels in the total number of pixels, and determines whether the second proportion value is within a second threshold range. If so, it determines whether the initialization parameters of the images in the image set are missing. If so, incremental spatial calculation is performed on the image set to obtain the initialization parameters corresponding to the image. This solves the problem in the prior art of calculating image initialization parameters based on image scene patterns after failing to identify the scene patterns of the image. Attached Figure Description
[0036] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.
[0037] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.
[0038] Figure 1 This is a flowchart of the reconstruction method for automatic identification and comprehensive decision-making according to the present invention;
[0039] Figure 2 This is an architecture diagram of the reconstruction system for automatic identification and comprehensive decision-making according to the present invention;
[0040] Figure 3 This is a schematic diagram of the physical structure of the electronic device provided by the present invention.
[0041] The attached figures are labeled as follows:
[0042] Acquisition module 10, semantic recognition module 20, determination module 30, judgment module 40, electronic device 50, processor 501, memory 502, bus 503. Detailed Implementation
[0043] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] Example
[0045] Figure 1 This is a flowchart of an embodiment of the reconstruction method for automatic identification and comprehensive decision-making of the present invention, as shown below. Figure 1 As shown, the automatic identification and comprehensive decision reconstruction method provided by this embodiment of the invention includes the following steps:
[0046] S101, acquire the image set for 3D reconstruction, perform semantic recognition on the images in the image set, and obtain the semantic labels of each image pixel, wherein the semantic labels include large semantic labels and small semantic labels;
[0047] Specifically, deep learning technology is used for "automatic scene recognition" to identify the semantic information of images, which can identify large semantic tags such as "buildings", "roads" and "toys" as well as small semantic tags in the images.
[0048] S102, determine the first proportion of the pixels with small semantic labels in the total number of pixels, determine whether the first proportion is within the first threshold range, if so, perform incremental spatial calculation on the image to obtain the initialization parameters corresponding to the image;
[0049] Specifically, PN toy <PT toy PN represents the first proportion of pixels with small semantic tags in the total number of pixels, and PT is the threshold setting of PN. When the first proportion is within the first threshold range, the image is determined to be a small object scene, wherein the first threshold range is [0.3, 0.8].
[0050] Based on two matching images in the image set, the pose data of the two images are obtained by calculating the fundamental matrix and / or the essential matrix;
[0051] After obtaining the pose data, one image is added each time. The pose data of the newly added image is obtained by calculating the fundamental matrix and / or the essential matrix, thereby obtaining the initialization parameters corresponding to each image in the image set.
[0052] Incremental, the pose and position information of each image is initially unknown;
[0053] Calculate the camera's extrinsic matrix using Formula 1;
[0054]
[0055] In the formula, z c t is the scale factor, K is the intrinsic parameter matrix, R is the rotation matrix, and t is the translation matrix.
[0056] S103, determine the second proportion of pixels with large semantic labels in the total number of pixels, determine whether the second proportion is within the range of the second threshold, if so, determine whether the initialization parameters of the images in the image set are missing, if so, perform incremental spatial calculation on the image set to obtain the initialization parameters corresponding to the images.
[0057] Specifically, PN city <PT city (PT cityThe threshold range is [0.5, 0.9]), PN represents the second proportion of pixels with large semantic labels in the total number of pixels, and PT is the threshold setting of PN. When the second proportion is within the second threshold range, the image is determined to be a large city scene, wherein the second threshold range is [0.5, 0.9].
[0058] Determine the camera mode; when the camera mode is a multi-camera mode, store the images corresponding to the same camera in the same image set.
[0059] Determine whether the user has input multiple cameras. If the user has input multiple cameras, store the images corresponding to different cameras in different image sets. When the camera mode is multi-camera mode, build a separate camera model for each camera.
[0060] If the user has not set up multiple cameras, the system will use single-camera mode, meaning all input images will come from the same camera.
[0061] It determines whether the user has entered the focal length. If the user has not entered the focal length information, it automatically calculates the focal length and estimates the focal length and other information of the image set.
[0062] Determine whether the user has input POS (position and orientation system) data. If the user outputs POS data and the initialization parameters of the images in the image set are complete, prioritize global spatial calculation of the image set to obtain the initialization parameters corresponding to the images; otherwise, use incremental spatial calculation.
[0063] This invention provides an automatic recognition and comprehensive decision-making reconstruction method. It acquires an image set for 3D reconstruction, performs semantic recognition on the images in the image set to obtain semantic labels for each image pixel, wherein the semantic labels include large semantic labels and small semantic labels; determines a first proportion value of pixels with small semantic labels in the total number of pixels, and determines whether the first proportion value is within a first threshold range. If so, it performs incremental spatial calculation on the image to obtain the initialization parameters corresponding to the image; determines a second proportion value of pixels with large semantic labels in the total number of pixels, and determines whether the second proportion value is within a second threshold range. If so, it determines whether the initialization parameters of the images in the image set are missing. If so, it performs incremental spatial calculation on the image set to obtain the initialization parameters corresponding to the image. This method solves the problem in the prior art of calculating image initialization parameters based on image scene patterns after failing to recognize the scene patterns of the image.
[0064] The reconstruction method based on automatic identification and comprehensive decision-making can automatically perform 3D reconstruction even when user input parameters are incomplete or nonexistent. This effectively solves the problem of manually inputting many professional parameters in the initial stage of 3D reconstruction, lowering the entry barrier and professional requirements for reconstruction personnel.
[0065] Figure 2 This is a flowchart of an embodiment of the reconstruction system for automatic identification and comprehensive decision-making according to the present invention; as follows: Figure 2 As shown in the figure, an automatic identification and comprehensive decision-making reconstruction system provided by an embodiment of the present invention includes the following steps:
[0066] Module 10 is used to acquire an image set for 3D reconstruction;
[0067] The semantic recognition module 20 is used to perform semantic recognition on the images in the image set to obtain semantic labels for each image pixel, wherein the semantic labels include large semantic labels and small semantic labels;
[0068] The determining module 30 is used to determine the first proportion value of the pixel with the small semantic tag in the total number of pixels;
[0069] Also used for:
[0070] Determine the second proportion of the pixels with the large semantic tag in the total number of pixels;
[0071] Determine the camera mode. When the camera mode is a multi-camera mode, store the images corresponding to the same camera in the same image set. When the camera mode is a multi-camera mode, create a separate camera model for each camera.
[0072] The judgment module 40 is used to determine whether the first proportion value is within the first threshold range. If so, it performs incremental spatial calculation on the image to obtain the initialization parameters corresponding to the image.
[0073] Also used for:
[0074] Determine whether the second proportion value is within the range of the second threshold. If so, determine whether the initialization parameters of the images in the image set are missing. If so, perform incremental spatial calculation on the image set to obtain the initialization parameters corresponding to the images.
[0075] When the first proportion value is within the first threshold range, the image is determined to be a small object scene, wherein the first threshold range is [0.3, 0.8].
[0076] When the second proportion value is within the range of the second threshold, the image is determined to be a large city scene, wherein the range of the second threshold is [0.5, 0.9].
[0077] The calculation module is used to obtain the pose data of two matching images in the image set by calculating the fundamental matrix and / or the essential matrix; after obtaining the pose data, each time an image is added, the pose data of the newly added image is obtained by calculating the fundamental matrix and / or the essential matrix, thereby obtaining the initialization parameters corresponding to each image in the image set.
[0078] It is also used to: when the initialization parameters of the images in the image set are complete, perform global spatial calculations on the image set to obtain the initialization parameters corresponding to the images.
[0079] An automatic recognition and comprehensive decision-making reconstruction system of the present invention acquires an image set for 3D reconstruction through an acquisition module 10; performs semantic recognition on the images in the image set through a semantic recognition module 20 to obtain semantic labels for each image pixel, wherein the semantic labels include large semantic labels and small semantic labels; determines a first proportion value of pixels with small semantic labels in the total number of pixels through a determination module 30; determines a second proportion value of pixels with large semantic labels in the total number of pixels; and determines whether the first proportion value is within a first threshold range. If so, incremental spatial calculation is performed on the image to obtain the initialization parameters corresponding to the image. It also determines whether the second proportion value is within a second threshold range. If so, it determines whether the initialization parameters of the images in the image set are missing. If so, incremental spatial calculation is performed on the image set to obtain the initialization parameters corresponding to the image.
[0080] Figure 3 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention, such as... Figure 3 As shown, the electronic device 50 includes: a processor 501, a memory 502, and a bus 503;
[0081] The processor 501 and the memory 502 communicate with each other via the bus 503.
[0082] The processor 501 is used to call program instructions in the memory 502 to execute the methods provided in the above-described method embodiments, such as: acquiring an image set for three-dimensional reconstruction; performing semantic recognition on the images in the image set to obtain semantic labels for each image pixel, wherein the semantic labels include large semantic labels and small semantic labels; determining a first proportion value of pixels with the small semantic labels in the total number of pixels; determining whether the first proportion value is within a first threshold range; if so, performing incremental spatial calculation on the image to obtain initialization parameters corresponding to the image; determining a second proportion value of pixels with the large semantic labels in the total number of pixels; determining whether the second proportion value is within a second threshold range; if so, determining whether initialization parameters for the images in the image set are missing; if so, performing incremental spatial calculation on the image set to obtain initialization parameters corresponding to the image.
[0083] This embodiment provides a non-transitory computer-readable medium storing computer instructions that cause a computer to execute the methods provided in the above-described method embodiments. For example, the instructions include: acquiring an image set for 3D reconstruction; performing semantic recognition on the images in the image set to obtain semantic labels for each image pixel, wherein the semantic labels include large semantic labels and small semantic labels; determining a first proportion of pixels with small semantic labels in the total number of pixels; determining whether the first proportion is within a first threshold range; if so, performing incremental spatial calculation on the image to obtain initialization parameters corresponding to the image; determining a second proportion of pixels with large semantic labels in the total number of pixels; determining whether the second proportion is within a second threshold range; if so, determining whether initialization parameters for the images in the image set are missing; if so, performing incremental spatial calculation on the image set to obtain initialization parameters corresponding to the image.
[0084] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned medium includes various media that can store program code, such as ROM, RAM, magnetic disk, or optical disk.
[0085] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0086] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0087] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A method for automatic identification of reconstruction of integrated decisions, characterized in that, The method specifically comprises: acquiring an image set for three-dimensional reconstruction, performing semantic recognition on images in the image set to obtain semantic labels of pixel points of each image, wherein the semantic labels comprise large semantic labels and small semantic labels; determining a first proportion value of pixel points with the small semantic labels in overall pixel points, judging whether the first proportion value is within a first threshold range, if yes, performing incremental spatial calculation on the image to obtain initialization parameters corresponding to the image; determining a second proportion value of pixel points with the large semantic labels in overall pixel points, judging whether the second proportion value is within a second threshold range, if yes, judging whether initialization parameters of images in the image set are missing, if yes, performing incremental spatial calculation on the image set to obtain initialization parameters corresponding to the image.
2. The reconstruction method for automatic identification and comprehensive decision-making according to claim 1, characterized in that, The determination of the first proportion value of the pixel points with the small semantic labels in the overall pixel points and the judgment of whether the first proportion value is within the first threshold range comprise: when the first proportion value is within the first threshold range, it is determined that the image is a small object scene, wherein the first threshold range is [0.3, 0.8].
3. The method of claim 1, wherein the automatic identification of the reconstruction of the integrated decision is based on a plurality of factors. The incremental spatial calculation on the image to obtain the initialization parameters corresponding to the image comprises: based on two matching images in the image set, pose data of the two images is obtained by calculating a fundamental matrix and / or an essential matrix; after the pose data is obtained, each time an image is added, pose data of the added image is obtained by calculating a fundamental matrix and / or an essential matrix, so that initialization parameters corresponding to each image in the image set are obtained.
4. The method of claim 1, wherein the automatic identification of the reconstruction of the integrated decision is based on a plurality of factors. The determination of the second proportion value of the pixel points with the large semantic labels in the overall pixel points and the judgment of whether the second proportion value is within the second threshold range comprise: when the second proportion value is within the second threshold range, it is determined that the image is a large city scene, wherein the second threshold range is [0.5, 0.9].
5. The method of claim 1, wherein the automatic identification of the reconstruction of the integrated decision is based on a plurality of factors. The judgment of whether the initialization parameters of the images in the image set are missing comprises: determining a camera mode, and when the camera mode is a multi-camera mode, images corresponding to the same camera are stored in the same image set.
6. The method of claim 5, wherein the automatic identification of the reconstruction of the integrated decision is based on a comparison of the integrated decision with a plurality of stored decisions. The judgment of whether the initialization parameters of the images in the image set are missing comprises: when the camera mode is a multi-camera mode, a camera model is established for each camera separately.
7. The method of claim 1, wherein the automatic identification of the reconstruction of the integrated decision is based on a plurality of factors. The judgment of whether the initialization parameters of the images in the image set are missing, if yes, the incremental spatial calculation on the image set to obtain the initialization parameters corresponding to the image comprises: when the initialization parameters of the images in the image set are complete, global spatial calculation is performed on the image set to obtain the initialization parameters corresponding to the image.
8. A system for automatically identifying reconsolidation of integrated decisions, characterized by, comprise: an acquisition module configured to acquire an image set for three-dimensional reconstruction; a semantic recognition module configured to perform semantic recognition on images in the image set to obtain semantic labels of pixel points of each image, wherein the semantic labels comprise large semantic labels and small semantic labels; a determination module configured to determine a first proportion value of pixel points with the small semantic labels in overall pixel points; and further configured to determine a second proportion value of the pixel points with the large semantic label in the whole pixel points; determine whether the first proportion value is within a first threshold range, and if so, perform incremental spatial calculation on the image to obtain the initialization parameter corresponding to the image; further configured to: determine whether the second proportion value is within a second threshold range, and if so, determine whether the initialization parameter of the image in the image set is missing, and if so, perform incremental spatial calculation on the image set to obtain the initialization parameter corresponding to the image.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1 to 7.
10. A non-transitory computer readable medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Semantic analysis method and device based on multi-intention recognition, equipment and storage medium
CN113723114A
3D point cloud data semi-automatic labeling method and device based on pseudo labels
CN113901991A