Image processing method, device and system

By acquiring sparse image sequences at the edge and offloading image processing to the edge, and utilizing the computing power pooling technology of the edge device, the problems of complex deployment and insufficient computing power of the image acquisition end are solved, thereby improving the quality and efficiency of 3D reconstruction.

CN121767540APending Publication Date: 2026-03-31HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, the complex deployment and difficulty in synchronizing image acquisition terminals lead to inaccurate image sequences acquired by the server, resulting in low quality of 3D reconstruction.

Method used

By acquiring sparse image sequences at the edge and offloading image processing functions to the edge, image processing and 3D reconstruction are performed using the computing power pooling technology of the edge device, reducing the deployment complexity and data volume of the edge and improving the quality of 3D reconstruction.

Benefits of technology

It achieves extremely simplified deployment of end-side devices, reduces deployment costs and latency, and improves data transmission rate and 3D reconstruction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767540A_ABST
    Figure CN121767540A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method, device and system, and relates to the field of artificial intelligence. The end side device obtains a sparse image sequence of multiple visual angles in the first scene and outputs the sparse image sequence of the multiple visual angles; the side device carries out image processing on the sparse image sequence of the multiple visual angles, three-dimensional reconstruction is carried out according to the processed image, a reconstruction model of the first scene is obtained, and image processing comprises signal image processing and intelligent processing. The image processing function of the end side is unloaded to the edge side, a sparse image sequence is collected at the end side, a dense image sequence does not need to be collected, and the end side extremely-simplified deployment form is achieved. Besides, since the data volume of the sparse image sequence acquired by the end side is small, image processing does not need to be performed on the sparse image sequence, so that the end side transmits the sparse image sequence to the edge side in time, image processing and three-dimensional reconstruction are performed on the edge side, and the problem of small computing power of the end side is solved through a computing power centralized pooling technology of an edge side device. And the three-dimensional reconstruction quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to an image processing method, apparatus and system. Background Technology

[0002] Currently, a reconstructed model of a scene is obtained by acquiring dense image sequences for 3D reconstruction. This process requires image processing of the dense image sequences, followed by data transmission to a server for 3D reconstruction. However, the complex deployment and synchronization difficulties at the image acquisition terminals lead to inaccurate image sequences received by the server, resulting in low-quality 3D reconstruction. Summary of the Invention

[0003] This application provides an image processing method, apparatus, and system, thereby improving the quality of 3D reconstruction.

[0004] In a first aspect, an image processing method is provided, which is applied to an image processing system. The image processing system includes an end-side device and a side-side device. The method includes: the end-side device acquiring a sparse image sequence of multiple views in a first scene and outputting the sparse image sequence of multiple views; the side-side device performing image processing on the sparse image sequence of multiple views, performing three-dimensional reconstruction based on the processed image, and obtaining a reconstruction model of the first scene. The image processing includes signal image processing and intelligent processing.

[0005] In this way, image processing functions are offloaded from the edge to the device side. By acquiring sparse image sequences at the edge, instead of dense image sequences, a highly simplified deployment is achieved. Furthermore, since the amount of data in the sparse image sequences acquired at the edge is small, there is no need for image processing. This allows the edge to transmit the sparse image sequences to the device side in a timely manner for image processing and 3D reconstruction. The computational power pooling technique used on the edge device addresses the issue of limited computing power at the edge, improving the quality of the 3D reconstruction.

[0006] In one possible implementation, the end-side device includes a small number of image acquisition modules, such as cameras, for acquiring sparse image sequences.

[0007] Therefore, by deploying a small number of image acquisition modules on the edge device, the ease of installation and deployment on the edge is improved, solving the problem of excessive waste in edge deployment and saving on edge deployment costs. This enables the small number of image acquisition modules deployed on the edge to work together efficiently, transmitting sparse image sequences to the edge in a timely manner for image processing and 3D reconstruction. The computational power pooling technology on the edge device addresses the issue of limited computing power on the edge, improving the quality of 3D reconstruction.

[0008] In another possible implementation, the end device does not perform image processing on the sparse image sequence; image processing includes signal image processing and intelligent processing.

[0009] In another possible implementation, the method further includes: the end device performing shallow compression encoding on the sparse image sequence of multiple viewpoints to obtain a bitstream of the sparse image sequence of multiple viewpoints; outputting the sparse image sequence of multiple viewpoints, including: the end device outputting the bitstream of the sparse image sequence of multiple viewpoints.

[0010] Shallow compression coding is used to solve the problem of long image coding time, that is, to shorten the coding time of sparse image sequences, reduce the amount of data transmitted, and reduce the loss of sparse image sequences. The sparse image sequences are transmitted to the edge in a timely manner. The problem of low end-to-end computing power is solved by the computing power pooling technology of the edge device, thereby improving the quality of 3D reconstruction.

[0011] In another possible implementation, the method further includes: the end device acquires a sparse image sequence of a first viewpoint in the first scene and outputs the sparse image sequence of the first viewpoint, the first viewpoint being different from multiple viewpoints; the side device updates the reconstruction model of the first scene based on the sparse image sequence of the first viewpoint.

[0012] In another possible implementation, the method further includes: the side device sending instruction information to the end device, the instruction information being used to indicate the adjustment of at least one of a plurality of viewing angles; the end device adjusting at least one viewing angle to a first viewing angle according to the instruction information.

[0013] In another possible implementation, the indication information is used to indicate at least one of a plurality of perspectives determined by defects in the reconstruction model, including at least one of holes, distortions, or deformations.

[0014] By remotely controlling the end-side device through the side-side device, the viewing angle of the sparse image sequence can be adjusted, sparse image sequences from different perspectives can be obtained, the scene reconstruction model can be automatically updated, and the maintainability can be improved.

[0015] Secondly, an image processing method is provided, including: acquiring a sparse image sequence from multiple perspectives in a first scene; performing image processing on the sparse image sequence from multiple perspectives; performing three-dimensional reconstruction based on the processed images to obtain a reconstruction model of the first scene; the image processing includes signal image processing and intelligent processing.

[0016] In one possible implementation, the method further includes: acquiring a sparse image sequence from a first-view perspective; and updating the reconstruction model of the first scene based on the sparse image sequence from the first-view perspective.

[0017] In another possible implementation, the method further includes sending an instruction message that indicates the adjustment of at least one of a plurality of views.

[0018] In another possible implementation, the indication information is used to indicate at least one of a plurality of perspectives determined by defects in the reconstruction model, including at least one of holes, distortions, or deformations.

[0019] Thirdly, an image processing apparatus is provided, comprising modules for performing the methods of the second aspect or any possible design of the second aspect. For example, the image processing apparatus includes a communication module and a processing module.

[0020] The communication module is used to acquire sparse image sequences from multiple perspectives in the first scene; the processing module is used to perform image processing on the sparse image sequences from multiple perspectives, and to perform 3D reconstruction based on the processed images to obtain a reconstruction model of the first scene. The image processing includes signal image processing and intelligent processing.

[0021] In one possible implementation, the communication module is further configured to acquire a sparse image sequence from a first perspective; the processing module is further configured to update the reconstruction model of the first scene based on the sparse image sequence from the first perspective.

[0022] In another possible implementation, the communication module is also used to send instruction information, which indicates the adjustment of at least one of the multiple perspectives.

[0023] In another possible implementation, the indication information is used to indicate at least one of a plurality of perspectives determined by defects in the reconstruction model, including at least one of holes, distortions, or deformations.

[0024] Fourthly, an electronic device is provided, comprising an image acquisition unit, a processor, and a memory; the image acquisition unit is used to acquire a sparse image sequence of multiple views in a first scene and output the sparse image sequence of multiple views; the processor is used to execute instructions stored in the memory to cause the processor to perform the operation steps of the method as described in the first aspect or any possible implementation of the first aspect.

[0025] Fifthly, an image processing system is provided, the image processing system including an end-side device and an edge-side device; the end-side device and the edge-side device are used to perform operational steps of the method as described in the first aspect or any possible implementation of the first aspect, to perform three-dimensional reconstruction of a first scene.

[0026] In one possible implementation, the end-side device includes a lens and a sensor.

[0027] In another possible implementation, the end-side device and the side-side device are connected by optical fiber, which is used to transmit a sparse image sequence of multiple perspectives in a first scene acquired by the end-side device to the side-side device.

[0028] A sixth aspect provides a computer-readable storage medium comprising: computer software instructions; which, when executed in a processor, cause the processor to perform operational steps of the method as described in the first aspect or any possible implementation thereof.

[0029] In a seventh aspect, a computer program product is provided that, when run on a computer, causes the computer to perform the operational steps of the method as described in the first aspect or any possible implementation thereof.

[0030] The technical effects of any of the design methods in aspects two through seven can be found in aspect one or in different design methods in aspect one, and will not be repeated here.

[0031] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0032] Figure 1 A schematic diagram of the architecture of an image processing system provided by the prior art;

[0033] Figure 2 This application provides a schematic diagram of the architecture of an image processing system.

[0034] Figure 3 A flowchart illustrating an image processing method provided in this application;

[0035] Figure 4 A schematic diagram illustrating a hole, twist, or deformation provided for this application;

[0036] Figure 5 A flowchart illustrating another image processing method provided in this application;

[0037] Figure 6 A schematic diagram of the structure of an image processing device provided in this application;

[0038] Figure 7 A schematic diagram of the structure of a computer device provided in this application;

[0039] Figure 8 A schematic diagram of an image processing system provided in this application. Detailed Implementation

[0040] To facilitate understanding, the main terms used in this application will be explained first.

[0041] 3D reconstruction is a computer technology that converts physical entities into digital models. It involves scanning or photographing objects or scenes in the physical world and then using computer vision and image processing techniques to transform two-dimensional images into three-dimensional scene information. 3D reconstruction technology utilizes images captured from multiple perspectives, employing steps such as camera calibration, feature extraction, image registration, and 3D reconstruction, combined with computer algorithms, to transform two-dimensional images into interactive three-dimensional models. In an abstract sense, a three-dimensional model refers to a combination of shape and appearance, rendered as a highly realistic RGB image from different viewpoints.

[0042] In some embodiments, firstly, multiple images of the scene are captured from multiple perspectives, with common parts among the images, for subsequent registration and reconstruction processes. Then, computer vision algorithms are used to analyze the images, solving for transformation parameters between each image, and superimposing and matching multiple images acquired at different times, angles, and illumination levels into a unified coordinate system. Redundant information is eliminated by calculating translation vectors and rotation matrices, generating point cloud data. Finally, based on the point cloud data, a 3D model is reconstructed using 3D modeling techniques, such as voxels and mesh representations.

[0043] 3D reconstruction technology is widely used in digital city modeling, object preservation, virtual reality, game development, autonomous driving, and robot navigation. For example, in digital city construction, realistic 3D reconstruction technology can generate highly realistic digital twin models, providing a realistic 3D model platform for urban information management systems and smart city construction. In object preservation, this technology can be used for non-contact scanning of objects to generate high-precision 3D models, aiding in the protection and inheritance of these objects. In virtual reality and game development, it provides immersive experiences and more interactive game environments.

[0044] In the field of image processing, key components for image acquisition include the lens, image sensor, image signal processor (ISP), and embedded neural network processing unit (NPU). These key components work together to capture and optimize images.

[0045] Figure 1 A schematic diagram of the architecture of an image processing system provided by existing technology. For example... Figure 1 As shown, the image processing system includes an end-side device and a side-side device. The end-side device is used to acquire dense image sequences, perform image processing and deep coding on the dense image sequences, and send the bitstream to the side-side device. The side-side device is used to decode the bitstream and perform 3D reconstruction based on the processed data to obtain a reconstructed model.

[0046] The end-side device 110 includes a lens 111, an image sensor 112, an ISP 113, an NPU 114, and an encoder 115.

[0047] Lens: Used to focus light onto the image sensor. By adjusting parameters such as focal length and aperture, it ensures that light is accurately projected onto the image sensor.

[0048] Image sensor: Converts light rays focused through the lens into electrical signals. This process involves converting the light signals into digital signals, preparing them for subsequent image processing.

[0049] Image Signal Processor (ISP): Used to process images received from image sensors. The ISP uses a series of algorithms to perform operations such as color correction, noise reduction, and brightness adjustment to improve image quality and appearance. The ISP processing flow includes processing the image (e.g., Bayer format) and ultimately outputting an RGB spatial domain image suitable for use by the backend video acquisition unit.

[0050] NPU: Unlike ISP, NPU is an embedded neural network processor with deep learning capabilities. It can perform precise optimization based on different shooting scenarios and continuously improve its optimization capabilities through self-learning. The application of NPU makes image processing more intelligent, such as enabling AI beautification functions, thereby further improving image quality and user experience.

[0051] Encoder: Organizes digital signals according to a specific format for easy storage and transmission. The encoding process includes sampling and quantization, converting continuous sensory data into digital form, encoding pixels using color spaces such as RGB or YUV, and ultimately forming a digital image.

[0052] Thus, from the lens capturing light to the NPU performing intelligent optimization, each step is an indispensable part of the digital image generation and processing process, together ensuring the quality of the final image and the user experience.

[0053] The edge device 120 includes a decoder 121 and a processor 122. The decoder is used to decode the bitstream from the edge device to obtain image data. The processor is used to perform 3D reconstruction based on the image data to obtain a reconstructed model.

[0054] In some embodiments, processor 122 includes one or more processing units. Processor 122 may include, for example, a graphics processing unit (GPU), a central processing unit (CPU), other general-purpose processors, or special-purpose processors. Examples include image signal processors (ISP), neural network processing units (NPU), digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors may be microprocessors or any conventional processor.

[0055] In some embodiments, the end-side device and the edge-side device can be integrated into a single physical device. The image processing system can be a terminal, such as a mobile phone, tablet computer, laptop computer, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, extended reality (ER) device, camera, or in-vehicle terminal, etc.

[0056] In other embodiments, the end-side device and the edge-side device are deployed on two separate physical devices. For example, the end-side device is deployed at the end, and the edge-side device is deployed at the edge. The edge-side device is an edge device (e.g., a box carrying a chip with processing capabilities). The end-side device and the edge-side device are connected via a transmission medium such as optical fiber.

[0057] The edge-side device, including key image acquisition components such as lenses and image sensors, requires specialized personnel for setup. For example, a circular camera deployment is often necessary, requiring the acquisition of dense image sequences, resulting in low image acquisition efficiency and complex, costly deployment. Furthermore, deep compression coding also contributes to low image acquisition efficiency, with significant latency and synchronization difficulties between components in the edge-side device, leading to poor registration and fusion results. Consequently, the side-testing device cannot acquire image sequences in a timely manner, and the resulting image sequences are inaccurate, resulting in low-quality 3D reconstruction. The use of AI computing resources in both the edge-side and side-testing devices also creates computational redundancy.

[0058] To address the issue of low quality in 3D reconstruction, this application provides an image processing method in which an end-side device acquires a sparse image sequence from multiple perspectives in a first scene and outputs the sparse image sequence from multiple perspectives; a side-side device performs image processing on the sparse image sequence from multiple perspectives, and performs 3D reconstruction based on the processed images to obtain a reconstruction model of the first scene. The image processing includes signal image processing and intelligent processing.

[0059] In this way, image processing functions are offloaded from the edge to the device side. By acquiring sparse image sequences at the edge, instead of dense image sequences, a highly simplified deployment is achieved. Furthermore, since the amount of data in the sparse image sequences acquired at the edge is small, there is no need for image processing. This allows the edge to transmit the sparse image sequences to the device side in a timely manner for image processing and 3D reconstruction. The computational power pooling technique used on the edge device addresses the issue of limited computing power at the edge, improving the quality of the 3D reconstruction.

[0060] This application is applicable to mainstream scenarios involving multiple sensing terminals, such as homes, supermarkets, and transportation. The aim of this application is to greatly simplify edge-side image sensing devices and improve algorithm accuracy and performance by offloading image processing to the edge to form a centralized processing platform.

[0061] Figure 2 This is a schematic diagram of the architecture of an image processing system provided in this application. Figure 2 As shown, the image processing system includes an end-side device 210 and a side-side device 220. The end-side device 210 is used to acquire sparse image sequences from multiple perspectives in a first scene and output the sparse image sequences from multiple perspectives. The side-side device 220 is used to perform image processing on the sparse image sequences from multiple perspectives, and to perform 3D reconstruction based on the processed images to obtain a reconstructed model of the first scene. The image processing includes signal image processing and intelligent processing.

[0062] End-side device 210 and edge-side device 220 are connected via a transmission medium such as optical fiber. For example, end-side device 210 transmits a sparse image sequence to edge-side device 220 via optical fiber. This improves the data transmission rate and enables data to be transmitted from the end-side to the edge-side as quickly as possible.

[0063] The end-side device 210 includes multiple image acquisition modules 211 and encoders 212. For example, the image acquisition module 211 includes a camera, radar detector, etc., and the camera includes a lens and an image sensor.

[0064] Side-side device 220 includes decoder 221 and processor 222. In some embodiments, the processor includes one or more processing units. The processor includes, for example, a GPU, CPU, other general-purpose processors, and special-purpose processors. For example, special-purpose processors include ISPs, NPUs, DSPs, ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor, etc.

[0065] Figure 2 The image processing system shown is Figure 1 The difference in the image processing system shown is that the image processing functions on the edge are offloaded to the side. For example, the ISP and NPU included in the edge device 210 are offloaded to the side device 220. The edge device 210 does not include the ISP and NPU and does not need to perform image processing on sparse image sequences. The side device 220 includes the ISP and NPU and performs image processing on sparse image sequences from multiple viewpoints.

[0066] This application does not limit the deployment location of end-side devices and edge-side devices. For example, an end-side device is a mobile robot that moves within its coverage area; an edge-side device is an edge device (e.g., a box carrying a chip with processing capabilities).

[0067] The structures illustrated in the embodiments of this application do not constitute a specific limitation on the image processing system. In practical applications, image processing systems are more... Figure 2 The system may include more or fewer components, combinations of two or more components, or different component configurations. For example, the image processing system may also include components such as a display device and a radar detector. Figure 2 The various components shown are implemented in hardware, software, or a combination of hardware and software, including one or more signal processing or application-specific integrated circuits. The methods in the following embodiments are implemented in an image processing system having the above-described hardware structure.

[0068] The image processing method provided in this application will now be described in detail with reference to the accompanying drawings.

[0069] Figure 3 A flowchart illustrating the image processing method provided in this application. Here, it is shown... Figure 2 The image processing system shown is used to generate a reconstruction model as an example for illustration. Figure 3 As shown, it includes the following steps.

[0070] Step 310: The end device acquires sparse image sequences from multiple perspectives in the first scene and outputs sparse image sequences from multiple perspectives.

[0071] The edge device employs a minimalist deployment; for example, it includes a small number of image acquisition modules deployed with sparse viewpoints. A sparse viewpoint refers to a limited number of angles. Sparse is the opposite of dense. The small number of image acquisition modules deployed with a limited number of viewpoints. For example, multiple viewpoints in a first scene refer to sparse viewpoints in that first scene.

[0072] A small number of image acquisition modules are used to acquire sparse image sequences. For example, sparse image sequences are image data in raw format.

[0073] The edge device acquires sparse image sequences from multiple perspectives in the first scene, i.e., the edge device acquires sparse image sequences from sparse perspectives in the first scene. The edge device outputs the sparse image sequences from multiple perspectives and transmits them to the side device.

[0074] Because the edge device employs a highly simplified deployment, it does not require an ISP or NPU. For example, it eliminates the need for image processing of the acquired images, meaning no signal-image processing or intelligent processing is required. This allows images captured by the edge device to be transmitted to the edge device rapidly. For instance, the edge device can achieve sub-frame level (e.g., less than 5 milliseconds) transmission of sparse image sequences. Furthermore, this improves the ease of installation and deployment of the edge device, saving on deployment costs.

[0075] In some embodiments, the end-side device performs shallow compression encoding on sparse image sequences from multiple viewpoints to obtain a bitstream of sparse image sequences from multiple viewpoints; the end-side device outputs the bitstream of sparse image sequences from multiple viewpoints.

[0076] Shallow compression refers to decorrelation processing of the data to be encoded using intra-frame prediction and transform quantization, followed by entropy coding for information entropy compression to obtain the bitstream. Compared to deep compression, shallow compression generally lacks inter-frame prediction and loop filtering, resulting in lower coding latency, computational complexity, and compression ratio. Existing lossless compression of reference frames, encoding of display interfaces, intra-frame coding in professional fields, and lightweight image coding can all be classified as shallow compression.

[0077] Common shallow compression standards include the following two: Display Stream Compression (DSC) and JPEG-XS (ISO / IEC 21122). Display Stream Compression is an industry standard for video codecs for display interfaces, introduced by the Video Electronics Standard Association (VESA), which compresses video content on video transmission lines. JPEG-XS was standardized in 2019 by the JPEG committee, officially named ISO / IEC SC29 WG1. It is a visually lossless, low-latency, and lightweight image coding standard suitable for video link transmission, VR / AR, drones, and autonomous driving scenarios.

[0078] By performing shallow compression coding on sparse image sequences, the coding time of sparse image sequences is shortened, the amount of data transmitted is reduced, and the loss of sparse image sequences is minimized. Sparse image sequences are then transmitted to the edge devices in a timely manner. The problem of low end-to-end computing power is solved by using the computing power pooling technology of the edge devices, thereby improving the quality of 3D reconstruction.

[0079] Step 320: The side device performs image processing on the sparse image sequence from multiple perspectives, and performs three-dimensional reconstruction based on the processed images to obtain the reconstruction model of the first scene.

[0080] The side-mounted device performs image processing operations such as signal-to-image processing (e.g., ISP processing) and intelligent processing (e.g., NPU processing) on ​​sparse image sequences from multiple viewpoints. As described above... Figure 1 The explanation of ISP and NPU in the text.

[0081] The 3D reconstruction methods provided in this application include those based on sparse images. For example, a processed image is input into a Neural Radiance Fields (NeRF) network, which outputs a reconstructed model of the first scene. Neural Radiance Fields (NeRF) is a neural network-based 3D scene reconstruction method that simulates the radiation amount and density of each point in the scene to recover a realistic 3D scene from 2D image data, achieving high-quality 3D reconstruction. Specific 3D reconstruction methods can be found in the relevant descriptions and will not be elaborated upon here.

[0082] In some embodiments, when a side-side device receives a shallowly compressed encoded bitstream, it first decodes the bitstream and then performs image processing operations such as signal-to-image processing (e.g., ISP processing) and intelligent processing (e.g., NPU processing) on ​​the decoded image. Subsequently, the processed image is input into a NeRF network, which outputs a reconstruction model of the first scene.

[0083] Because the edge has strong computing power, image processing operations are offloaded from the front end to the edge, and image processing and 3D reconstruction are performed on the edge, reducing front-end performance overhead and improving edge resource utilization.

[0084] In other embodiments, the reconstructed model is quality evaluated to identify defective areas such as holes, distortions, and deformations (e.g., Figure 4 As shown in (a), the holes, distortions, and deformations of the reconstructed model are reconstructed, and the reconstructed model is iteratively updated to improve the quality of the reconstructed model.

[0085] Holes refer to missing areas in a reconstructed model. Methods for inspecting holes in reconstructed models mainly include mesh hole region detection and methods for detecting the size of apparent pit defects based on the 3D reconstructed model.

[0086] The specific operations for mesh hole detection include: mesh encoding the reconstructed model to obtain its mesh; identifying boundary edges within the mesh, grouping these edges into ordered, connected 3D closed non-intersecting polygons, and classifying at least one polygon as a hole; determining the area of ​​the hole; and providing feedback on the hole, such as visual feedback. Specifically, it involves performing neighborhood queries on each vertex, edge, and face to check the connectivity of facets. It also checks whether each edge is shared by exactly two faces; if an edge is used by only one face, it is considered a boundary edge, indicating a potential hole.

[0087] A mesh, also known as a static mesh, contains multiple polygons used to describe the boundary surfaces of a volume object. A mesh can contain one or more types of polygons to describe the boundary surfaces of a volume object. For example, a mesh might contain two types of polygons (such as triangles and quadrilaterals) to describe the boundary surfaces of a volume object; or, for instance, a mesh might contain triangles to describe the boundary surfaces of a volume object.

[0088] A polygon comprises multiple vertices and multiple edges. In this embodiment, a polygon is defined by vertices in three-dimensional (3D) space and the way these vertices are connected. The mesh data includes the vertex positions and connections of multiple polygons. Vertex positions indicate the location information of the vertices constituting the polygon in 3D space. For example, vertex positions can be the coordinates of the vertices constituting the polygon in 3D space. If the polygon is a triangle, the vertex positions can include the coordinates of the three vertices of the triangle, such as (0, 2, 1.3, 2.5), representing the positions of the (x, y, z) axes in the coordinate system. Connections indicate the connection relationships between the vertices of the polygons in the mesh.

[0089] Distortion refers to the spiral or bending phenomenon that appears in the reconstructed model. By calculating the normal vector of the reconstructed model, if the normal vector of a certain region suddenly changes significantly, it means that geometric distortion has occurred.

[0090] Deformation refers to phenomena such as scaling, translation, rotation, stretching, or compression that occur in the reconstructed model. Deformation usually causes abnormal changes in edge lengths. Comparing the changes in edge lengths of the reconstructed model's mesh can effectively identify deformed regions. By traversing the faces of the reconstructed model's mesh, extracting the edges between vertices, and then calculating the length of each edge, the length of the edge is compared with the average length of its neighborhood. If the length of a certain edge deviates significantly, it may be a deformed region.

[0091] The aforementioned identification of defects such as holes, distortions, and deformations in the reconstructed model can serve as an initial screening for these defects. In some embodiments, after identifying defects such as holes, distortions, and deformations in the reconstructed model, the images of the initially screened defective areas are input into the identification model for secondary screening to ensure the accuracy of defect identification in the reconstructed model.

[0092] After the above identification, if the reconstructed model has defective areas, it is determined that the quality of the reconstructed model is not up to standard; if the reconstructed model does not have defective areas, it is determined that the quality of the reconstructed model is up to standard.

[0093] Optionally, the quality of the reconstructed model can be evaluated based on model quality conditions. For example, the number of defective regions N in the reconstructed model can be set. If the number of defective regions in the reconstructed model is greater than or equal to N, the quality of the reconstructed model is considered unsatisfactory; if the number of defective regions in the reconstructed model is less than N, the quality of the reconstructed model is considered acceptable.

[0094] If the quality of the reconstructed model is not up to standard, the side device instructs the end device to adjust at least one of the multiple viewpoints to obtain a sparse image sequence of the new viewpoint in the first scene, and iteratively updates the reconstructed model based on the sparse image sequence of the new viewpoint to improve the quality of the reconstructed model.

[0095] like Figure 5 As shown, the image processing method provided in this application also includes the following steps.

[0096] Step 330: The end device acquires a sparse image sequence of the first viewpoint in the first scene and outputs the sparse image sequence of the first viewpoint.

[0097] The side-mounted device performs 3D reconstruction based on the processed data to obtain the first reconstructed model. The quality of the first reconstructed model is then evaluated.

[0098] If the quality of the first reconstruction model is unsatisfactory, the side-side device controls the viewpoint of the end-side device, for example, controlling at least one of multiple viewpoints, so that the end-side device acquires a sparse image sequence of a first viewpoint in the first scene. The first viewpoint is different from the multiple viewpoints. Understandably, the side-side device controls the end-side device to adjust from the current viewpoint to the first viewpoint. The first viewpoint is the viewpoint adjusted by the camera. The sparse image sequence of the first viewpoint is the image sequence captured by the camera after adjusting the viewpoint. The sparse image sequence of the first viewpoint includes an image sequence of the region in the first scene corresponding to the defective region of the reconstruction model.

[0099] For example, the side-side device can control the viewing angle of at least one image acquisition module (e.g., a lens) in the end-side device, adjusting the viewing angle of at least one image acquisition module to a first viewing angle. The first viewing angle includes the adjusted viewing angle of at least one image acquisition module.

[0100] When the side-mounted device can control the viewing angles of two or more image acquisition modules, the adjustment angle and method of the viewing angle of each image acquisition module can be different or the same. That is, the first viewing angle includes the adjusted viewing angles of two or more image acquisition modules, and the adjusted viewing angles of the two or more image acquisition modules can be different or the same.

[0101] In some embodiments, since the side-side device knows the image to which the defect region of the reconstructed model belongs, and the camera to which the image to which the defect region belongs (i.e., the camera that captured the image to which the defect region belongs), the side-side device controls the viewing angle of the camera that captured the image to which the defect region of the reconstructed model belongs. For example, the defect region of the reconstructed model is mainly generated at the boundary of adjacent cameras, and the camera that captured the image to which the defect region belongs is determined by the coordinates of the defect region.

[0102] Optionally, the side device includes a correspondence between cameras and images, which can be used to determine the camera corresponding to the image of the defective area.

[0103] The side-mounted device calculates the median value of the defect area in the reconstructed model and controls the camera to adjust to the position of that median value. For example, as shown... Figure 4 As shown in (b), the first and second cameras are moved towards the center of the hole, so that the first camera changes from a second perspective to a first perspective, allowing it to capture the area in the first scene corresponding to the hole. Similarly, the second camera changes from a second perspective to a first perspective, allowing it to capture the area in the first scene corresponding to the hole. For example, as... Figure 4As shown in (c), the first and second cameras are moved in the opposite direction of the distortion, so that the first camera changes from a second viewpoint to a first viewpoint, allowing its first viewpoint to capture the area in the first scene corresponding to the distortion. Similarly, the second camera changes from a second viewpoint to a first viewpoint, allowing its first viewpoint to capture the area in the first scene corresponding to the distortion. For example, as... Figure 4 As shown in (d), the second viewpoints of the first camera and the second camera are similar, making it impossible to capture a panoramic view of the first scene, potentially leading to deformation of the reconstructed model of the first scene. The first and second cameras are controlled to move in the opposite direction of the deformation, adjusting the first camera from a second viewpoint to a first viewpoint, enabling it to capture the first scene corresponding to the deformation. The second camera is also adjusted from a second viewpoint to a first viewpoint, enabling it to capture the first scene corresponding to the distortion. The first viewpoints of the first camera and the second camera are different. Therefore, the first and second cameras re-capture image sequences of the regions in the first scene corresponding to any one of the holes, distortions, or deformations. The side-mounted device performs 3D reconstruction based on the image sequences of the regions in the first scene corresponding to any one of the holes, distortions, or deformations, to eliminate the holes, distortions, or deformations in the first reconstructed model.

[0104] In some embodiments, the image processing method provided in this application further includes the following steps. For example... Figure 5 As shown, in step 350, the side device sends instruction information to the end device. In step 360, the end device adjusts at least one viewing angle to a first viewing angle according to the instruction information.

[0105] The instruction information is used to indicate the adjustment of at least one of multiple viewing angles. For example, the instruction information includes the camera whose viewing angle is being adjusted and the camera adjustment angle. After receiving the instruction information, the end-side device adjusts the camera's viewing angle according to the camera adjustment angle indicated by the instruction information.

[0106] Thus, the side device adjusts the camera's shooting angle by controlling the end device, and re-acquires a sparse image sequence from a new perspective in the first scene. Based on the sparse image sequence from the new perspective, the reconstruction model is iteratively updated to improve the quality of the reconstruction model.

[0107] Step 340: The side device updates the reconstruction model of the first scene based on the sparse image sequence from the first perspective.

[0108] In some embodiments, the side device replaces the image sequence generating the defect region in the sparse image sequence of multiple perspectives with the sparse image sequence of the first perspective, and performs three-dimensional reconstruction based on the sparse image sequence of the first perspective and other image sequences without defect regions to obtain the updated reconstruction model of the first scene.

[0109] In other embodiments, the side device performs 3D reconstruction using a sparse image sequence from a first perspective to obtain an updated local model of the defect region of the reconstructed model, and updates the reconstructed model of the first scene with the updated local model to obtain an updated reconstructed model of the first scene.

[0110] For example, four images (A, B, C, and D) are captured by cameras 1 and 4, and a reconstruction model of the first scene is generated from these four images. If the model generated based on images A and B has defects, the viewing angles of camera 1 (which captured image A) and camera 2 (which captured image B) are adjusted to acquire new images between images A and B that intersect with image A. Image A or image B is then replaced with the new image, and an updated reconstruction model of the first scene is regenerated.

[0111] Optionally, the edge device evaluates the quality of the updated reconstructed model of the first scene. If the quality of the updated reconstructed model is satisfactory, the updated reconstructed model of the first scene is output; if the quality of the updated reconstructed model is unsatisfactory, the above steps of adjusting the camera perspective and updating the model are repeated. The specific process is described above and will not be repeated here.

[0112] The image processing method provided in this application offloads image processing functions from the edge to the device side. By acquiring sparse image sequences at the edge, it eliminates the need for acquiring dense image sequences, achieving a highly simplified deployment at the edge. Deploying a small number of image acquisition modules on the edge device improves ease of installation and deployment, solving the problem of waste in edge deployments and saving costs. This results in high collaborative efficiency among the few image acquisition modules deployed at the edge. Furthermore, since the sparse image sequences acquired at the edge have a smaller data volume, there is no need for image processing. This allows the edge to promptly transmit the sparse image sequences to the edge for image processing and 3D reconstruction. The edge device's centralized computing power pooling technology addresses the issue of limited computing power at the edge, improving the quality of 3D reconstruction. The edge device can remotely control the edge device to adjust the viewing angle for acquiring sparse image sequences, obtaining sparse image sequences from different perspectives and automatically updating the scene reconstruction model, thus improving maintainability.

[0113] It is understood that, in order to achieve the functions in the above embodiments, the image processing system includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0114] Figure 6This is a schematic diagram of the possible image processing apparatus provided in this application. These image processing apparatuses can be used to implement the functions of the side devices in the above method embodiments, and therefore can also achieve the beneficial effects of the above method embodiments. In the embodiments of this application, the image processing apparatus can be as follows: Figure 2 The side device shown can also be a module (such as a chip) applied to the side device.

[0115] like Figure 6 As shown, the image processing apparatus 600 includes a communication module 610 and a processing module 620. The image processing apparatus 600 can be applied to applications such as... Figure 2 The side device shown.

[0116] The communication module 610 is used to acquire sparse image sequences from multiple perspectives in the first scene collected by the end-side device.

[0117] The processing module 620 is used to perform image processing on a sparse image sequence from multiple perspectives, and to perform 3D reconstruction based on the processed images to obtain a reconstructed model of the first scene. The image processing includes signal image processing and intelligent processing. For example, the processing module 620 is used to execute step 320.

[0118] Storage module 630 is used to store sparse image sequences, reconstruction models, and applications required to perform 3D reconstruction.

[0119] For a more detailed description of the communication module 610 and the processing module 620 mentioned above, please refer to [link / reference]. Figure 3 or Figure 5 The relevant descriptions in the method embodiments shown are directly obtained and will not be repeated here.

[0120] Figure 7 This is a schematic diagram of the structure of a computer device 700 provided in this embodiment. As shown in the figure, the computer device 700 includes a processor 710, a bus 720, a memory 730, and a communication interface 740.

[0121] It should be understood that in this embodiment, the processor 710 can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0122] The processor may also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of the program in this application.

[0123] The communication interface 740 is used to enable communication between the computer device 700 and external devices or components. In this embodiment, the computer device 700 is used to implement... Figure 3 or Figure 5 The side device shown functions as follows: the communication interface 740 is used to acquire sparse image sequences. The processor 710 is used to perform image processing on the sparse image sequences from multiple perspectives, and to perform three-dimensional reconstruction based on the processed images to obtain a reconstructed model of the first scene. The image processing includes signal image processing and intelligent processing.

[0124] Bus 720 may include a pathway for transmitting information between the aforementioned components (such as processor 710 and memory 730). In addition to a data bus, bus 720 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus 720 in the figure.

[0125] As an example, computer device 700 may include multiple processors. A processor may be a multi-core (multi-CPU) processor. Here, "processor" can refer to one or more devices, circuits, and / or computing units used to process data (e.g., computer program instructions).

[0126] It is worth noting that, Figure 7 Taking computer device 700 as an example, which includes one processor 710 and one memory 730, the processor 710 and memory 730 are used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to business needs.

[0127] The memory 730 can correspond to the storage medium used to store computer instructions and images in the above method embodiments, such as a disk, like a mechanical hard disk or a solid-state hard disk.

[0128] The aforementioned computer device 700 can be a general-purpose device or a special-purpose device. For example, computer device 700 can be a mobile terminal, tablet computer, laptop computer, VR device, AR device, MR device or ER device, in-vehicle computer device, etc., or it can be an edge device (e.g., a box carrying a chip with processing capabilities).

[0129] It should be understood that the computer device 700 according to this embodiment may correspond to the image processing device 600 in this embodiment, and may correspond to the device executing the image processing device according to this embodiment. Figure 3 or Figure 5 The corresponding subject in any of the methods, and the above and other operations and / or functions of each module in the image processing apparatus 600 are respectively for implementing Figure 3 or Figure 5 For the sake of brevity, the corresponding processes of each method in the code will not be elaborated here.

[0130] Since the modules in the end-side device and the side-side device provided in this application can be distributed and deployed on multiple computers in the same environment or different environments, this application also provides a... Figure 8 The image processing system shown includes multiple computers 800, each computer 800 including a memory 801, a processor 802, a communication interface 803, and a bus 804. The memory 801, processor 802, and communication interface 803 are interconnected via the bus 804.

[0131] The memory 801 can be a read-only memory, a static storage device, a dynamic storage device, or a random access memory. The memory 801 can store computer instructions. When the computer instructions stored in the memory 801 are executed by the processor 802, the processor 802 and the communication interface 803 are used to execute part of the data processing methods of the software system. The memory can also store data such as images and reconstructed models. For example, a portion of the storage resources in the memory 801 is divided into an area for storing images and programs that implement the functions of the embodiments of this application.

[0132] Processor 802 can be a general-purpose CPU, an application-specific integrated circuit (ASIC), a GPU, or any combination thereof. Processor 802 may include one or more chips. Processor 802 may include an AI accelerator, such as an NPU.

[0133] The communication interface 803 uses transceiver modules, such as, but not limited to, transceivers, to enable communication between the computer 800 and other devices or communication networks. For example, the communication interface 803 can acquire sparse image sequences or provide feedback on viewpoint indication information to the computer device.

[0134] Bus 804 may include a pathway for transmitting information between various components of computer 800 (e.g., memory 801, processor 802, communication interface 803).

[0135] Each of the aforementioned computers 800 establishes a communication path through a communication network. Each computer 800 runs any one or more edge devices. Any computer 800 can be a computer in a cloud data center (e.g., a server), a computer in an edge data center, or a terminal computing device.

[0136] Edge functionality can be deployed on each of the 800 computers. For example, GPUs are used to implement edge functionality.

[0137] The method steps in this embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a terminal device. Of course, the processor and storage medium can also exist as discrete components in a network device or terminal device.

[0138] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD). The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image processing method, characterized in that, Applied to an image processing system, the image processing system including an end-side device and an edge-side device, the method includes: The end-side device acquires sparse image sequences from multiple perspectives in the first scene and outputs the sparse image sequences from multiple perspectives. The side-mounted device performs image processing on the sparse image sequence from the multiple perspectives, and performs three-dimensional reconstruction based on the processed image to obtain a reconstruction model of the first scene. The image processing includes signal image processing and intelligent processing.

2. The method according to claim 1, characterized in that, The method further includes: The end-side device performs shallow compression encoding on the sparse image sequences from the multiple viewpoints to obtain the bitstream of the sparse image sequences from the multiple viewpoints. The output of the sparse image sequence from the multiple viewpoints includes: The end-side device outputs a bitstream of the sparse image sequence from the multiple viewpoints.

3. The method according to claim 1 or 2, characterized in that, The method further includes: The end-side device acquires a sparse image sequence of a first viewpoint in the first scene and outputs the sparse image sequence of the first viewpoint, wherein the first viewpoint is different from the plurality of viewpoints; The side-mounted device updates the reconstruction model of the first scene based on the sparse image sequence from the first viewpoint.

4. The method according to claim 3, characterized in that, The method further includes: The side device sends instruction information to the end device, the instruction information being used to instruct adjustment of at least one of the plurality of viewing angles; The end-side device adjusts the at least one viewing angle to the first viewing angle according to the instruction information.

5. The method according to claim 4, characterized in that, The indication information is used to indicate at least one of the plurality of perspectives determined by defects in the reconstruction model, the defects in the reconstruction model including at least one of holes, distortions, or deformations.

6. An image processing method, characterized in that, include: Obtain a sparse image sequence from multiple perspectives in the first scene; Image processing is performed on the sparse image sequence from multiple perspectives, and three-dimensional reconstruction is performed based on the processed images to obtain a reconstruction model of the first scene. The image processing includes signal image processing and intelligent processing.

7. The method according to claim 6, characterized in that, The method further includes: Obtain the sparse image sequence from the first viewpoint; The reconstruction model of the first scene is updated based on the sparse image sequence from the first viewpoint.

8. The method according to claim 7, characterized in that, The method further includes: Send an instruction message, which is used to instruct adjustment of at least one of the plurality of viewpoints.

9. The method according to claim 8, characterized in that, The indication information is used to indicate at least one of the plurality of perspectives determined by defects in the reconstruction model, the defects in the reconstruction model including at least one of holes, distortions, or deformations.

10. An image processing apparatus, characterized in that, include: The communication module is used to acquire sparse image sequences from multiple perspectives in the first scene; The processing module is used to perform image processing on the sparse image sequence from multiple perspectives, and to perform three-dimensional reconstruction based on the processed image to obtain a reconstruction model of the first scene. The image processing includes signal image processing and intelligent processing.

11. An image processing system, characterized in that, The image processing system includes an end-side device and an edge-side device; The end-side device and the side-side device are used to perform the operational steps of the method as described in any one of claims 1 to 5, so as to perform three-dimensional reconstruction of the first scene.

12. The image processing system according to claim 11, characterized in that, The end-side device includes a lens and a sensor.

13. The image processing system according to claim 11 or 12, characterized in that, The end-side device and the side-side device are connected by an optical fiber, which is used to transmit a sparse image sequence of multiple perspectives in the first scene acquired by the end-side device to the side-side device.

14. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device, the computing device performs the operational steps of the method as described in any one of claims 1 to 5.

15. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a computing device, cause the computing device to perform the operational steps of the method as described in any one of claims 1 to 5.