Large-scale three-dimensional image data video generation method and device

Through dynamic segmentation and three-dimensional spatial relationship stitching technology, the problem of low efficiency in processing large-scale object data sets is solved, efficient and clear three-dimensional videos are generated, and the user experience is improved.

CN120602629APending Publication Date: 2025-09-05SHENZHEN BAIHUI XINZHI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510542029.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies suffer from unbalanced computing task loads when processing large-scale object data sets, and bottlenecks exist in resource allocation and task scheduling, resulting in low efficiency, inability to support customized requirements and complex rendering processes, and a poor user experience.

Method used

By receiving the entropy values ​​or gradient intensities of multiple images to be processed, high-complexity areas are determined, the images are dynamically segmented into multiple data blocks, key information is extracted using the maximum intensity projection algorithm, and video frames are spliced ​​through three-dimensional spatial relationships. Error minimization processing is performed to generate clear three-dimensional videos.

Benefits of technology

It significantly improves the processing efficiency of large-scale object data sets, generates clearer and more accurate three-dimensional videos, and improves visual effects and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602629A_ABST
    Figure CN120602629A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a large-scale three-dimensional image data video generation method and device, and the method comprises the steps: determining a high-complexity region in a to-be-processed image according to a plurality of to-be-processed images and the entropy values or gradient intensities of the to-be-processed images, receiving a preset space region parameter or a preset volume feature parameter, and obtaining a large-scale three-dimensional image data video. Adjusting spatial region parameters or preset volume characteristic parameters, dynamically segmenting each to-be-processed image to obtain a plurality of to-be-processed data blocks corresponding to the to-be-processed image, extracting maximum intensity information included in each voxel in each rendering frame according to a maximum intensity projection algorithm, and combining the maximum intensity information into a single video frame corresponding to each to-be-processed image, according to the technical scheme, the three-dimensional space relation among the multiple to-be-processed images is determined, the video frames are spliced according to the three-dimensional space relation, error minimization is carried out on the spliced video frames, the three-dimensional video of the object is obtained, the clearer and more accurate three-dimensional video can be obtained, and the visual effect and the generation efficiency of the three-dimensional video are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to a method and device for generating large-scale three-dimensional image data videos. Background Art

[0002] In the existing technology, with the development of three-dimensional imaging technology, various details of objects can be obtained, resulting in the continuous increase in the amount of data obtained from large object data sets, which can reach the level of several terabytes (TB). Three-dimensional imaging technology can also provide users with intuitive operation methods through a graphical interface.

[0003] However, when processing large-scale object data sets using existing methods, the hardware requirements are high. In addition, when processing large-scale object data sets to obtain corresponding three-dimensional imaging, the existing methods have bottlenecks in resource allocation and task scheduling, resulting in unbalanced load of computing tasks and affecting overall performance. Secondly, the existing graphical interface cannot support users with high customization requirements for three-dimensional imaging, and lacks automated support for complex rendering processes, resulting in low efficiency and poor user experience when users perform batch processing of large-scale object data sets or multi-dimensional analysis. Summary of the Invention

[0004] In response to the problems in the existing technology, the present application provides a large-scale three-dimensional image data video generation method and device, which can effectively solve the shortcomings of traditional technology in terms of low efficiency when users perform batch processing of large-scale object data sets or multidimensional analysis, and significantly improve the efficiency of batch processing of large-scale object data sets.

[0005] In order to solve at least one of the above problems, the present application provides the following technical solutions:

[0006] In a first aspect, the present application provides a method for generating large-scale three-dimensional image data video, comprising:

[0007] Receiving multiple images to be processed and determining the entropy value or gradient strength of each image to be processed, wherein the multiple images to be processed are obtained by capturing images of the same object from multiple angles at the same time;

[0008] Based on the entropy value or gradient strength of the image to be processed, a local area of ​​the image to be processed corresponding to an entropy value greater than a preset entropy value threshold or a gradient strength greater than a preset gradient strength is determined as a high complexity area in the image to be processed;

[0009] receiving preset spatial region parameters or preset volume feature parameters, reducing the spatial region parameters or the preset volume feature parameters according to a preset ratio, segmenting the image to be processed according to the spatial region parameters or the preset volume feature parameters in the non-high complexity region, and segmenting the image to be processed according to the reduced spatial region parameters or the preset volume feature parameters in the high complexity region, to obtain a plurality of data blocks to be processed;

[0010] Determining the relative positions of the data blocks to be processed corresponding to each image to be processed in the image to be processed, rendering each data block to be processed according to a preset color bit depth to obtain a single rendered frame corresponding to each data block to be processed, extracting the maximum intensity information of each voxel in each rendered frame based on a maximum intensity projection algorithm, and merging the maximum intensity information corresponding to each data block to be processed into a single video frame corresponding to each image to be processed according to the relative positions;

[0011] Determine the three-dimensional spatial relationship between multiple images to be processed, stitch each video frame according to the three-dimensional spatial relationship, and perform error minimization on the stitched video frames to obtain a three-dimensional video of the object.

[0012] Furthermore, the method further includes: when there is a continuously distributed high-complexity region, recursively segmenting the continuously distributed high-complexity region through an octree structure until the entropy value or gradient strength variance of each data block to be processed obtained by segmentation is less than a preset tolerance, and the continuous distribution is an area where any two high-complexity regions overlap;

[0013] At the junction of the high-complexity area and the non-high-complexity area, a transition data block with a gradually changing size is generated. The size corresponding to the transition data block linearly transitions from the size corresponding to the data block to be processed obtained by dividing the spatial area parameters or preset volume feature parameters after reduction to the size corresponding to the data block to be processed obtained by dividing the spatial area parameters or preset volume feature parameters before reduction.

[0014] Furthermore, the method further includes: rendering each data block to be processed in parallel by computing a unified device architecture, and smoothing voxels in each data block to be processed by a linear interpolation algorithm during the rendering process to obtain a rendering result;

[0015] Perform local histogram equalization on the rendering results to obtain a single rendering frame corresponding to each data block to be processed.

[0016] Furthermore, the method further includes: traversing each voxel in each rendered frame, determining the original intensity value of each voxel, and when the rendered frame is multi-channel data, determining the maximum original intensity value in each channel data as the maximum intensity information;

[0017] Merging the maximum intensity information corresponding to each to-be-processed data into a single video frame corresponding to each to-be-processed image according to each relative position, including:

[0018] Perform anisotropic diffusion filtering on the video frame to smooth the edges of the video frame.

[0019] Furthermore, the method further includes: extracting key points and descriptors of each image to be processed, performing fast feature matching based on the key points and descriptors to obtain a fast feature matching result, and determining a matching confidence corresponding to the fast feature matching result;

[0020] The fast feature matching results corresponding to the matching confidence values ​​less than the preset matching confidence value threshold are determined as mismatching points and the mismatching points are eliminated;

[0021] The fast feature matching results after eliminating the mismatched points are converted into a three-dimensional point cloud to obtain the three-dimensional spatial relationship between multiple images to be processed.

[0022] Furthermore, the method further includes: determining a structural similarity index of the spliced ​​adjacent video frames in the overlapping area, and locating the misaligned area in the adjacent video frames based on the structural similarity index;

[0023] Determining a pixel displacement vector in the misaligned area by using an optical flow method, determining a splicing offset based on the pixel displacement vector, and adjusting the spliced ​​video frames according to the splicing offset;

[0024] The Poisson fusion algorithm is used to process the stitching boundaries in the stitched video frames, and the fused video frames are corrected for global brightness consistency to obtain a three-dimensional video of the object.

[0025] Furthermore, the method further includes: receiving display parameters and selected instructions in a preset format, the preset format including any one of a Lua script file, an HTTP interface, and a Python interface, the selected instruction being used to execute the display parameters on a single video frame;

[0026] The display of the stitched video frames is adjusted according to the display parameters and the selected instruction.

[0027] In a second aspect, the present application provides a large-scale three-dimensional image data video generation device, comprising:

[0028] A receiving module is used to receive multiple images to be processed and determine the entropy value or gradient strength of each image to be processed, wherein the multiple images to be processed are obtained by collecting images of the same object from multiple angles at the same time;

[0029] A first processing module is configured to determine, based on the entropy value or gradient strength of the image to be processed, a local area of ​​the image to be processed corresponding to an entropy value greater than a preset entropy value threshold or a gradient strength greater than a preset gradient strength as a high-complexity area in the image to be processed;

[0030] a second processing module, configured to receive preset spatial region parameters or preset volume feature parameters, reduce the spatial region parameters or preset volume feature parameters according to a preset ratio, segment the image to be processed according to the spatial region parameters or preset volume feature parameters in a non-high complexity region, and segment the image to be processed according to the reduced spatial region parameters or preset volume feature parameters in a high complexity region, to obtain a plurality of data blocks to be processed;

[0031] a third processing module, configured to determine the relative positions of the data blocks to be processed corresponding to each image to be processed in the image to be processed, render each data block to be processed according to a preset color bit depth, obtain a single rendered frame corresponding to each data block to be processed, extract the maximum intensity information of each voxel in each rendered frame based on a maximum intensity projection algorithm, and merge the maximum intensity information corresponding to each data block to be processed into a single video frame corresponding to each image to be processed according to the relative positions;

[0032] The stitching module is used to determine the three-dimensional spatial relationship between multiple images to be processed, stitch each video frame according to the three-dimensional spatial relationship, and perform error minimization on the stitched video frames to obtain a three-dimensional video of the object.

[0033] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the described method when executing the program.

[0034] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the large-scale three-dimensional image data video generation method.

[0035] In a fifth aspect, the present application provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method for generating large-scale three-dimensional image data video.

[0036] It can be seen from the above technical solution that the present application provides a large-scale three-dimensional image data video generation method and device, which innovatively determines the high-complexity area in the image to be processed based on multiple images to be processed and the entropy value or gradient strength of the image to be processed, wherein the multiple images to be processed are obtained by multi-angle image acquisition of the same object at the same time, accepting preset spatial area parameters or preset volume feature parameters, reducing the spatial area parameters or preset volume feature parameters according to a preset ratio, and dynamically segmenting the high-complexity area and non-high-complexity area categories to obtain multiple data blocks to be processed corresponding to the image to be processed, which can reduce the computing load and avoid over-complexity. The redundant calculations introduced by degree segmentation improve image processing speed. The maximum intensity information for each voxel in each rendered frame is extracted using the maximum intensity projection algorithm. The maximum intensity information corresponding to each piece of data to be processed is then merged into a single video frame corresponding to each image to be processed, based on their relative positions. The three-dimensional spatial relationship between the multiple images to be processed is determined, and the video frames are stitched together according to this three-dimensional spatial relationship. Error minimization is then performed on the stitched video frames to produce a three-dimensional video of the object. This method can extract key information using the maximum intensity projection algorithm and stitch the video frames together by minimizing the error, resulting in a clearer and more accurate three-dimensional video, improving both visual quality and the efficiency of three-dimensional video generation. This method effectively addresses the inefficiencies of traditional technologies when users are performing batch processing of large-scale object datasets or multidimensional analysis, significantly improving the efficiency of batch processing of large-scale object datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0038] Figure 1 Schematic diagram of the process of generating large-scale three-dimensional image data video in an embodiment of the present application;

[0039] Figure 2 An object image in a process of rendering a video frame provided by an embodiment of the present application;

[0040] Figure 3 This is a structural diagram of a large-scale three-dimensional image data video generation device in an embodiment of the present application;

[0041] Figure 4 Schematic diagram of the structure of the electronic device in the embodiment of the present application.

[0042] Reference numerals:

[0043] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION

[0044] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0045] The acquisition, storage, use, and processing of data in this application's technical solution comply with relevant national laws and regulations.

[0046] In view of the problems existing in the prior art, the present application provides a large-scale three-dimensional image data video generation method and device, which innovatively determines the high-complexity area in the image to be processed based on multiple images to be processed and the entropy value or gradient strength of the image to be processed, wherein the multiple images to be processed are obtained by multi-angle image acquisition of the same object at the same time, accepting preset spatial area parameters or preset volume feature parameters, reducing the spatial area parameters or preset volume feature parameters according to a preset ratio, and dynamically segmenting the high-complexity area and non-high-complexity area categories to obtain multiple data blocks to be processed corresponding to the image to be processed, which can reduce the computing load and avoid It avoids redundant calculations caused by over-segmentation and improves the image processing speed. It extracts the maximum intensity information of each voxel in each rendering frame according to the maximum intensity projection algorithm, merges the maximum intensity information corresponding to each data to be processed into a single video frame corresponding to each image to be processed according to each relative position, determines the three-dimensional spatial relationship between multiple images to be processed, splices each video frame according to the three-dimensional spatial relationship, and performs error minimization on the spliced ​​video frames to obtain a three-dimensional video of the object. It can extract key information through the maximum intensity projection algorithm and splice video frames by minimizing the error to obtain a clearer and more accurate three-dimensional video, thereby improving the visual effect and the efficiency of three-dimensional video generation.

[0047] In order to effectively solve the shortcomings of traditional technologies in terms of low efficiency when users perform batch processing of large-scale object data sets or multi-dimensional analysis, and significantly improve the efficiency of batch processing of large-scale object data sets, this application provides an embodiment of a large-scale three-dimensional image data video generation method, see Figure 1The large-scale three-dimensional image data video generation method specifically includes the following contents:

[0048] Step S101: receiving a plurality of images to be processed, and determining the entropy value or gradient strength of each image to be processed.

[0049] The multiple images to be processed are obtained by capturing images of the same object from multiple angles at the same time.

[0050] Optionally, this embodiment receives multiple images to be processed, and the multiple images to be processed are obtained by shooting the same object from multiple different angles at the same time, for example, they can be obtained from different spatial angles, evenly distributed viewing angles such as 0°, 45°, 90° and 135°. The image to be processed at each viewing angle contains complete object spatial information. Furthermore, each image to be processed can have a unified time-space reference coordinate system.

[0051] In addition, the entropy value or gradient strength of each image to be processed can be determined according to the received segmentation instruction, wherein the entropy value is an indicator to measure the complexity of the image information of the image to be processed. If the entropy value in the image to be processed is higher, the image to be processed contains more details and changes. The gradient strength is used to indicate the severity of the change in pixel value in the image to be processed. A larger gradient strength is likely to appear at the edge of an object or in an area with rich texture. In the case of a three-dimensional video that requires details and changes, a segmentation instruction for segmenting the image to be processed according to the entropy value can be received, and in the case of a three-dimensional video that needs to focus on the edge of an object or in an area with rich texture, a segmentation instruction for segmenting the image to be processed according to the gradient strength can be received.

[0052] This embodiment achieves segmentation of multiple images to be processed from multiple dimensions, making the segmentation method more flexible and able to maintain consistency in multiple images to be processed.

[0053] Step S102: Based on the entropy value or gradient strength of the image to be processed, a local area of ​​the image to be processed corresponding to an entropy value greater than a preset entropy value threshold or a gradient strength greater than a preset gradient strength is determined as a high complexity area in the image to be processed.

[0054] Optionally, this embodiment can obtain high-complexity regions in the image to be processed based on the calculated entropy value or gradient strength of the image to be processed. After determining the entropy value or gradient strength of the image to be processed, in the case of a method of segmenting the image to be processed based on entropy value, when the entropy value of a local region is greater than a preset entropy value threshold, it is marked as a high-complexity region. In the case of a method of segmenting the image to be processed based on gradient strength, when the gradient strength of a local region is greater than a preset gradient strength, it is marked as a high-complexity region.

[0055] Among them, the high-complexity area includes key features of the object, such as features corresponding to more complex textures, edges or structures.

[0056] This embodiment realizes the division of local areas in the image to be processed into high-complexity areas and non-high-complexity areas, and can perform dynamic segmentation on the image to be processed, thereby improving the efficiency and clarity of segmentation.

[0057] Step S103: Receive preset spatial area parameters or preset volume feature parameters, reduce the spatial area parameters or preset volume feature parameters according to a preset ratio, divide the image to be processed according to the spatial area parameters or preset volume feature parameters in the non-high complexity area, and divide the image to be processed according to the reduced spatial area parameters or preset volume feature parameters in the high complexity area to obtain multiple data blocks to be processed.

[0058] Optionally, this embodiment receives preset spatial area parameters or preset volume feature parameters, wherein, when the image to be processed is segmented by the entropy value of the image to be processed, the preset spatial area parameters can be received, that is, the entropy value segmentation and the preset spatial area parameters correspond to each other, and when the image to be processed is segmented by the gradient intensity of the image to be processed, the preset volume feature parameters can be received, that is, the gradient intensity and the preset volume feature parameters correspond to each other.

[0059] In addition, in order to better process the high-complexity areas and non-high-complexity areas of the image to be processed, the spatial area parameters or the preset volume feature parameters are reduced according to the preset ratio, that is, after receiving any of the above parameters, they are reduced according to the preset ratio, making the segmentation process more detailed.

[0060] In addition, in the process of segmenting the image to be processed, the non-high complexity area is segmented according to the received preset spatial area parameters or preset volume feature parameters, while in the high complexity area, the high complexity area will be segmented according to the preset spatial area parameters after being reduced by a preset reduction ratio or the preset volume feature parameters after being reduced by a preset reduction ratio.

[0061] For example, the preset spatial area parameters can be set to 96×96×96 voxels, the preset reduction ratio is 1.5, and the preset spatial area parameters after reduction are 64×64×64 voxels. The non-high complexity area is segmented to obtain a 96×96×96 voxel data block to be processed, and the high complexity area is segmented to obtain a 64×64×64 voxel data block to be processed.

[0062] This embodiment implements a differentiated dynamic segmentation strategy to ensure that detailed information in the image to be processed is not lost during the segmentation process in high-complexity areas, while improving the efficiency and accuracy of segmentation. In the process of obtaining multiple data blocks to be processed, the processing efficiency is significantly improved while ensuring the processing quality.

[0063] Step S104: Determine the relative position of each data block to be processed corresponding to each image to be processed in the image to be processed, render each data block to be processed according to a preset color bit depth, obtain a single rendered frame corresponding to each data block to be processed, extract the maximum intensity information included in each voxel in each rendered frame based on the maximum intensity projection algorithm, and merge the maximum intensity information corresponding to each data block to be processed into a single video frame corresponding to each image to be processed according to the relative position.

[0064] Optionally, after completing the segmentation of the image to be processed, this embodiment determines the relative position of each data block to be processed in the original image to be processed, so that information of each data block to be processed can be accurately integrated and spliced ​​in subsequent image processing.

[0065] In addition, each data block to be processed is rendered separately according to the preset color bit depth to obtain a rendering frame corresponding to each data block to be processed, wherein the preset color bit depth determines the color richness and detail expression ability of the rendering frame. The preset color bit depth can be 16 bits, that is, each data block to be processed can be rendered in 16 bits to obtain a corresponding rendering frame.

[0066] Furthermore, each rendered frame is processed individually using the Maximum Intensity Projection algorithm to extract the maximum intensity information within each frame. Maximum Intensity Projection is an image processing technique that extracts the maximum intensity information contained in each voxel, the smallest unit in a 3D image. By processing the data block using the Maximum Intensity Projection algorithm to obtain the maximum intensity information, the most representative information can be extracted from each rendered frame, reducing redundant data and improving processing efficiency.

[0067] In addition, according to the relative positions of the data blocks to be processed in the original image, the maximum intensity information corresponding to the data blocks to be processed are merged to generate a single video frame corresponding to each image to be processed.

[0068] This embodiment realizes the extraction of the maximum intensity information of each voxel through the maximum intensity projection algorithm, so as to effectively highlight the key features in the image to be processed, such as the edges and textures of objects, avoids the interference of redundant information, and improves the clarity and detail expression of the video frame.

[0069] Step S105: determining the three-dimensional spatial relationship between the multiple images to be processed, stitching the video frames according to the three-dimensional spatial relationship, and performing error minimization on the stitched video frames to obtain a three-dimensional video of the object.

[0070] Optionally, this embodiment determines the three-dimensional spatial relationship between multiple processed images. The determination of the three-dimensional spatial relationship can be achieved by analyzing information such as the geometric relationship between the images to be processed, the relative positions of objects, and the shooting angles.

[0071] According to the determined three-dimensional spatial relationship, each video frame is spliced ​​in the correct spatial order to generate a preliminary three-dimensional video.

[0072] To further improve the quality of 3D video, error minimization processing is performed on the stitched video frames to obtain a 3D video of the object. Error minimization is an optimization algorithm that can reduce the errors and discontinuities that may occur during the stitching process by adjusting the alignment relationship and fusion method between video frames, thereby generating a smoother and more natural 3D video effect.

[0073] This embodiment realizes the splicing of video frames and performs error minimization processing on them. The error minimization processing can reduce the errors and discontinuities that may occur during the splicing process, ensure the smoothness and consistency of the three-dimensional video, and enable the three-dimensional video to more accurately reflect the three-dimensional structure and spatial relationship of the object.

[0074] In some embodiments, segmenting the image to be processed in the high-complexity region according to the reduced spatial region parameters or the preset volume feature parameters includes:

[0075] When there are continuously distributed high-complexity regions, the continuously distributed high-complexity regions are recursively segmented using the octree structure until the entropy value or gradient strength variance of each processed data block obtained by segmentation is less than the preset tolerance. Continuous distribution refers to the area where any two high-complexity regions overlap.

[0076] At the junction of the high-complexity area and the non-high-complexity area, a transition data block with a gradually changing size is generated. The size corresponding to the transition data block linearly transitions from the size corresponding to the data block to be processed obtained by dividing the spatial area parameters or preset volume feature parameters after reduction to the size corresponding to the data block to be processed obtained by dividing the spatial area parameters or preset volume feature parameters before reduction.

[0077] Optionally, in this embodiment, when it is determined that there are continuously distributed high-complexity areas in the image to be processed, it is recursively segmented through an octree structure, where continuous distribution means that there are overlapping areas between any two high-complexity areas, that is, any two high-complexity areas are spatially connected to each other, rather than isolated.

[0078] Among them, octree is a commonly used three-dimensional space segmentation method that can divide the space into multiple small blocks. Each small block can be further subdivided to achieve fine processing of complex areas.

[0079] During the recursive segmentation process, the entropy value or gradient strength variance of each data block to be processed is continuously monitored. If the entropy value or gradient strength variance of a data block is still greater than the preset tolerance value, the data block to be processed is subdivided until the entropy value or gradient strength variance of all data blocks to be processed is less than the preset tolerance.

[0080] In addition, at the junction of the high-complexity area and the non-high-complexity area, a transition data block with a gradually changing size is generated. The size of the transition data block is a linear transition from the data block to be processed obtained by segmenting the high-complexity area to the data block to be processed obtained by segmenting the non-high-complexity area, that is, a linear transition from the size obtained by segmenting the spatial area parameters or the preset volume feature parameters after reduction to the size obtained by segmenting the spatial area parameters or the preset volume feature parameters before reduction.

[0081] Among them, the size of the transition data block will be adjusted according to its position at the junction. The part close to the high-complexity area will be closer to the segmentation result of the high-complexity area; and the part close to the non-high-complexity area will be closer to the segmentation result of the non-high-complexity area, thereby ensuring that the segmentation result at the junction is more natural and smooth, avoiding obvious segmentation marks.

[0082] This embodiment implements recursive segmentation of continuously distributed high-complexity regions using an octree structure. It can dynamically adjust the degree of surprise in the segmentation based on the entropy value or gradient strength variance to ensure that the complexity of each processed data block obtained by segmentation is within a controllable range, avoid the loss of details due to insufficient segmentation, and avoid the waste of computing resources caused by excessive segmentation. At the same time, at the junction of high-complexity regions and non-high-complexity regions, transitional data blocks with gradually varying sizes are used to alleviate the segmentation differences between high-complexity regions and non-high-complexity regions, avoiding the appearance of obvious boundary traces in the processed image.

[0083] In some embodiments, rendering each data block to be processed according to a preset color bit depth to obtain a single rendered frame corresponding to each data block to be processed includes:

[0084] By computing a unified device architecture, each data block to be processed is rendered in parallel. During the rendering process, the voxels in each data block to be processed are smoothed using a linear interpolation algorithm to obtain a rendering result.

[0085] Perform local histogram equalization on the rendering results to obtain a single rendering frame corresponding to each data block to be processed.

[0086] Optionally, this embodiment can perform parallel rendering of each data block to be processed using the Compute Unified Device Architecture (CUDA). CUDA is a highly efficient computing architecture that can fully utilize the parallel computing capabilities of multi-core processors or graphics processing units (GPUs). CUDA enables multiple data blocks to be rendered simultaneously, thereby improving rendering efficiency.

[0087] During the rendering process, the voxels in each block of data to be processed are smoothed using a linear interpolation algorithm. This algorithm can calculate the value of intermediate voxels based on the values ​​of adjacent voxels, thereby reducing discontinuities between voxels and making the rendering smoother. The linear interpolation algorithm can render the block of data to be processed using a fast interpolation algorithm using shared memory, with mirror filling applied to the voxels at the edge of the block.

[0088] In addition, after obtaining the rendering results, local histogram equalization is performed on the rendering results of each data block to be processed. Histogram equalization is a commonly used image enhancement technology that can adjust the contrast of the image to make the grayscale distribution of the image more uniform. Local histogram equalization is an improved method of histogram equalization, that is, instead of making global adjustments to the entire image, adjustments are made to the local areas of the image. This can better preserve the local details of the image while avoiding over-enhancement or distortion that may be caused by global adjustments.

[0089] Through local histogram equalization, the rendering result of each data block to be processed is further optimized to make it visually clearer and more natural, and a single rendering frame corresponding to each data block to be processed is obtained.

[0090] This embodiment implements parallel rendering of a unified computing device architecture, smoothing of a linear interpolation algorithm, and optimization adjustment of local histogram equalization, thereby ensuring processing efficiency and improving rendering quality when processing complex image data to be processed, and also enabling detail enhancement and visual optimization to generate high-quality rendered frames.

[0091] In some embodiments, extracting maximum intensity information of each voxel in each rendered frame based on a maximum intensity projection algorithm includes:

[0092] Traverse each voxel in each rendered frame and determine the original intensity value of each voxel. When the rendered frame is multi-channel data, determine the maximum original intensity value in each channel data as the maximum intensity information;

[0093] Merging the maximum intensity information corresponding to each to-be-processed data into a single video frame corresponding to each to-be-processed image according to each relative position, including:

[0094] Perform anisotropic diffusion filtering on the video frame to smooth the edges of the video frame.

[0095] Optionally, this embodiment traverses all voxels in each rendering frame, determines the original intensity value of each voxel, determines the data volume of the channel data of the rendering frame, and when the rendering frame is multi-channel data, checks the original intensity value in each channel, compares the original intensity values ​​in each channel, and determines the maximum value therein as the maximum intensity information of the voxel.

[0096] When the rendered frame is single-channel data, the original intensity value of the maximum value in the channel data is determined as the maximum intensity data of the voxel, so that the most representative intensity information can be extracted from each voxel.

[0097] In addition, the extracted maximum intensity information is merged according to the relative position of each data block to be processed in the original image, that is, the maximum intensity information from different data blocks to be processed is recombined according to their positional relationship in the image to be processed to generate a complete video frame.

[0098] After the merging is completed, anisotropic diffusion filtering is performed on the resulting video frames. Anisotropic diffusion filtering is an image processing technique that can smooth image edges while retaining important structures and details. Anisotropic diffusion filtering can identify edges in the image and smooth non-edge areas to avoid blurring important features.

[0099] This embodiment achieves maximum intensity projection algorithm, multi-channel data processing and anisotropic diffusion filtering, which can ensure processing efficiency in the process of processing complex three-dimensional image data, while improving the quality of video frames, achieving efficient information extraction, high-quality video frame generation and optimized visual effects.

[0100] In some embodiments, determining the three-dimensional spatial relationship between the plurality of images to be processed includes:

[0101] Extract key points and descriptors of each image to be processed, perform fast feature matching based on the key points and descriptors, obtain fast feature matching results, and determine the matching confidence corresponding to the fast feature matching results;

[0102] The fast feature matching results corresponding to the matching confidence values ​​less than the preset matching confidence value threshold are determined as mismatching points and the mismatching points are eliminated;

[0103] The fast feature matching results after eliminating the mismatched points are converted into a three-dimensional point cloud to obtain the three-dimensional spatial relationship between multiple images to be processed.

[0104] Optionally, this embodiment extracts key points and descriptors of each image to be processed, where key points refer to points with significant features in the image, such as corner points, edge points, etc. Key points have high gradient changes and can effectively identify the characteristic areas of the image. The descriptor is a quantitative representation of the local features of the key points, which can be a vector used to describe the texture, shape and other feature information around the key points.

[0105] Among them, key points in the image to be processed can be identified through key point detection algorithms, such as the scale-invariant feature transform algorithm (SIFT), the speed up robust features algorithm (SURF) or the image processing algorithm based on feature descriptors (Oriented FAST and Rotated BRIEF, ORB), so as to obtain feature points in the image to be processed that are invariant to changes in illumination, scale and rotation, and ensure the stability and repeatability of the key points under different viewing angles and lighting conditions.

[0106] For each detected key point, a descriptor corresponding to the key point is determined, wherein the descriptor may be calculated by sampling and quantizing pixel values ​​in a neighborhood around the key point to generate a feature vector that can uniquely identify the key point.

[0107] In addition, an efficient feature matching algorithm, such as the Fast Library for Approximate Nearest Neighbors (FLANN) or the Brute-Force Matcher (BFMatcher), can be used to compare the key points and descriptors of different images to be processed, and the similarity between two key points can be determined by calculating the distance between the descriptors (the distance can be determined by Euclidean distance or Hamming distance), and a fast feature matching result can be obtained.

[0108] During the process of obtaining fast feature matching results, a match confidence score is determined for each fast feature matching result. Match confidence is an important indicator of the reliability of a matching pair and can be derived based on factors such as the distance between matching descriptors, uniqueness, and consistency with other matching pairs. A higher match confidence score indicates a more reliable matching pair, while a lower match confidence score may indicate a risk of mismatching.

[0109] In addition, a preset matching confidence threshold is received. The preset matching confidence threshold is used to distinguish reliable matches from false matches. It can be determined according to actual application requirements and the quality of image features of the image to be processed, and can be adjusted based on experiments and experience.

[0110] The fast feature matching results with matching confidence less than the preset matching confidence threshold are determined as mismatch points and removed from the fast feature matching results. This can effectively reduce the impact of mismatches on subsequent three-dimensional spatial relationship calculations and improve the overall matching accuracy and reliability.

[0111] In addition, the key point corresponding to the fast feature matching result from which the false matching points have been eliminated is regarded as a corresponding point in three-dimensional space, and when it is determined that each key point has corresponding position and depth information in the image to be processed, it is converted into three-dimensional coordinates according to its position and depth information in the image to be processed. Multiple three-dimensional coordinate points constitute a three-dimensional point cloud. By analyzing and fitting the three-dimensional coordinates of all fast feature matching points, the three-dimensional spatial relationship between multiple images to be processed is constructed.

[0112] This embodiment achieves key point and descriptor extraction, fast feature matching, mismatch point elimination and three-dimensional point cloud conversion, so that when processing complex images to be processed, matching accuracy and accurate construction of three-dimensional reconstruction are guaranteed, while improving computing efficiency and system robustness.

[0113] In some embodiments, performing error minimization on the stitched video frames to obtain a three-dimensional video of the object includes:

[0114] Determining a structural similarity index of the overlapping region of the spliced ​​adjacent video frames, and locating the misaligned region in the adjacent video frames based on the structural similarity index;

[0115] Determining a pixel displacement vector in the misaligned area by using an optical flow method, determining a splicing offset based on the pixel displacement vector, and adjusting the spliced ​​video frames according to the splicing offset;

[0116] The Poisson fusion algorithm is used to process the stitching boundaries in the stitched video frames, and the fused video frames are corrected for global brightness consistency to obtain a three-dimensional video of the object.

[0117] Optionally, this embodiment determines a structural similarity index (SSIM) of the overlapping area of ​​the spliced ​​adjacent video frames, where SSIM is an indicator for measuring the similarity between two images, and can jointly determine the similarity between the two images from the differences in brightness and contrast and the structural information of the images.

[0118] The SSIM values ​​of the overlapping regions of adjacent video frames are determined, wherein a higher SSIM value indicates that the two overlapping regions are more similar in structure; conversely, a lower SSIM value indicates that the structural difference between the overlapping regions is greater.

[0119] Regions with lower SSIM values ​​were identified as misaligned regions, indicating significant structural differences, which may be caused by inaccurate splicing.

[0120] In addition, after the misaligned area is determined, the pixel displacement vectors of the pixels in the misaligned area can be determined using the optical flow method. The optical flow method is a method for estimating pixel motion in an image sequence. This method assumes that pixels move between adjacent frames and describes the motion by calculating the pixel displacement vectors.

[0121] Optical flow is applied to the misaligned areas to determine the pixel displacement vector for each pixel from one video frame to the other. This vector represents the actual pixel movement. Based on the pixel displacement vector, a stitching offset is determined, and the stitched frames are adjusted based on the stitching offset. The stitching offset is the amount of translation required to align one of the two frames in the overlapping area.

[0122] In addition, the splicing boundaries in the spliced ​​video frames are processed by the Poisson fusion algorithm. Poisson fusion is an image fusion technology based on the Poisson equation, which can smooth the splicing boundaries while retaining image details and reduce splicing traces.

[0123] Furthermore, in order to improve the overall quality of the generated video, the fused video frames are subjected to global brightness consistency correction to obtain a three-dimensional video, thereby ensuring that the brightness and contrast between adjacent video frames remain consistent and avoiding brightness discontinuity caused by splicing.

[0124] This embodiment implements structural similarity index evaluation, optical flow method pixel displacement calculation, Poisson fusion and global brightness consistency correction, and can achieve high-precision stitching, optimized visual effects and improved video quality when processing complex images to be processed.

[0125] In some embodiments, after stitching the video frames according to the three-dimensional spatial relationship, the method further includes:

[0126] receiving display parameters and selected instructions in a preset format, the preset format including any one of a Lua script file, an HTTP interface, and a Python interface, the selected instruction being used to execute the display parameters on a single video frame;

[0127] The display of the stitched video frames is adjusted according to the display parameters and the selected instruction.

[0128] Optionally, this embodiment can access multiple preset formats, including any one of Lua script files, HTTP interfaces, and Python interfaces, and can provide users with multiple options, such as determining display parameters and selecting instructions through preset formats, where display parameters may include but are not limited to brightness, contrast, and color balance, and selected instructions are used to execute display parameters for a single video frame.

[0129] Lua script file: Lua is a lightweight scripting language commonly used in embedded systems and game development. You can define display parameters and selected commands by writing Lua scripts.

[0130] For example, when exporting display parameters and selected instructions, a Lua script file will be generated in the following format:

[0131] set(channel_intensity_ranges,channel_index,frame_index,'1001000'), where channel_intensity_ranges represents the channel intensity range, channel_index represents the index of the target channel, frame_index represents the frame number in the 3D video, and '1001000' represents the intensity range, where the minimum intensity is 100 and the maximum intensity is 1000. The unit of the intensity range depends on the data normalization method. After receiving the Lua script file, you can display the 3D video according to the script file.

[0132] HTTP interface: For network-based applications, display parameters can be received through the HTTP interface. That is, the display effect of video frames can be dynamically adjusted by receiving HTTP requests. This can be used for real-time interactive applications and allows dynamic adjustment of display parameters through third-party methods.

[0133] For example, an HTTP interface can be implemented as follows:

[0134] The path can be expressed as / api / update_display_params, and the request method can be expressed as POST;

[0135] Request parameter (JSON) example: {"data":{"channels":{"0":"range":"1001000","visible":true}}};

[0136] Among them, data represents the root object of all parameters, data.channels represents the channel configuration container, the key is the channel index, data.channels."0" represents the configuration of channel 0, range represents the intensity display range, visible represents the channel visibility, true represents display, and false represents hiding.

[0137] Python interface: Python is a widely used programming language with a rich library and tools. Display parameters can be controlled through Python scripts or programs.

[0138] Furthermore, an embodiment of the present application provides a Python encapsulation library that allows users to render and save large-scale images to be processed completely without an interactive interface. Through the above-mentioned Python library, saved project files can be loaded, all display parameters can be obtained, and display parameters can be modified for each frame. The rendering results can also be saved as 8-bit images.

[0139] Python interface example:

[0140] renderer=LychnisIO.Renderer('data.lyp'); renderer.set_channel_visible(0,False); ren derer.render_to_file('output.tif');

[0141] LychnisIO represents the name of the 3D visualization library, Renderer represents the class, that is, the renderer core class, which is used to load and render the data blocks to be processed, and 'data.lyp' is used to represent the input file path.

[0142] In addition, the received display parameters are parsed, including but not limited to brightness adjustment, contrast adjustment, color correction, sharpening, blurring, etc., and the parsed display parameters are applied to the specified video frame according to the selected instructions. For example, if the selected parameters select a video frame and the display parameters set brightness increase and contrast adjustment, the selected video frame is processed according to the selected parameters and the display parameters.

[0143] Furthermore, real-time feedback can be provided during the adjustment process, so that the adjusted video frames are presented in real time on the mobile terminal, and display parameters and / or selected parameters can be modified and resent.

[0144] This embodiment implements support for preset formats, flexibility in selecting instructions, and real-time interactive capabilities. After splicing video frames, the display effects of the video frames can be flexibly adjusted according to the received display parameters and selected instructions.

[0145] Figure 2 An object image in a process of rendering a video frame provided by an embodiment of the present application, Figure 2 The image shown here is an image with the inverted color effect applied. Figure 2 The boundary is obtained by segmenting and synthesizing the volume feature parameters, and has not been smoothed.

[0146] In order to effectively solve the shortcomings of traditional technologies in terms of low efficiency when users perform batch processing of large-scale object data sets or multi-dimensional analysis, the present application provides an embodiment of a large-scale three-dimensional image data video generation device for implementing all or part of the content of the large-scale three-dimensional image data video generation method, see Figure 3 The large-scale three-dimensional image data video generation device specifically includes the following contents:

[0147] A receiving module 10 is configured to receive a plurality of images to be processed and determine an entropy value or gradient strength of each image to be processed, wherein the plurality of images to be processed are obtained by capturing images of the same object from multiple angles at the same time;

[0148] The first processing module 20 is configured to determine, based on the entropy value or gradient strength of the image to be processed, a local area of ​​the image to be processed corresponding to an entropy value greater than a preset entropy value threshold or a gradient strength greater than a preset gradient strength as a high-complexity area in the image to be processed;

[0149] a second processing module 30 configured to receive a preset spatial region parameter or a preset volume feature parameter, reduce the spatial region parameter or the preset volume feature parameter according to a preset ratio, segment the image to be processed according to the spatial region parameter or the preset volume feature parameter in a non-high complexity region, and segment the image to be processed according to the reduced spatial region parameter or the preset volume feature parameter in a high complexity region, to obtain a plurality of data blocks to be processed;

[0150] The third processing module 40 is configured to determine the relative positions of the data blocks to be processed corresponding to each image to be processed in the image to be processed, render each data block to be processed according to a preset color bit depth to obtain a single rendered frame corresponding to each data block to be processed, extract the maximum intensity information of each voxel in each rendered frame based on a maximum intensity projection algorithm, and merge the maximum intensity information corresponding to each data block to be processed into a single video frame corresponding to each image to be processed according to the relative positions;

[0151] The stitching module 50 is used to determine the three-dimensional spatial relationship between multiple images to be processed, stitch the video frames according to the three-dimensional spatial relationship, and perform error minimization on the stitched video frames to obtain a three-dimensional video of the object.

[0152] From the above description, it can be seen that the large-scale three-dimensional image data video generation device provided in the embodiment of the present application can innovatively determine the high-complexity area in the image to be processed based on multiple images to be processed and the entropy value or gradient strength of the image to be processed, wherein the multiple images to be processed are obtained by multi-angle image acquisition of the same object at the same time, accepting preset spatial area parameters or preset volume feature parameters, reducing the spatial area parameters or preset volume feature parameters according to a preset ratio, and dynamically segmenting the high-complexity area and non-high-complexity area categories to obtain multiple data blocks to be processed corresponding to the image to be processed, which can reduce the computing load and avoid excessive The redundant computations introduced by segmentation speed up image processing. The maximum intensity information for each voxel in each rendered frame is extracted using a maximum intensity projection algorithm. The maximum intensity information corresponding to each piece of data to be processed is then merged into a single video frame corresponding to each image to be processed, based on their relative positions. The three-dimensional spatial relationship between the multiple images to be processed is determined, and the video frames are stitched together according to this relationship. Error minimization is then performed on the stitched video frames to produce a three-dimensional video of the object. This method extracts key information using the maximum intensity projection algorithm and stitches the video frames together using error minimization, resulting in a clearer and more accurate three-dimensional video, improving both visual quality and the efficiency of 3D video generation. This method effectively addresses the inefficiencies of traditional technologies when performing batch processing of large-scale object datasets or multidimensional analysis, significantly improving the efficiency of batch processing of large-scale object datasets.

[0153] From a hardware perspective, in order to effectively address the shortcomings of conventional technologies in terms of low efficiency in batch processing of large-scale object datasets or multidimensional analysis, and to significantly improve the efficiency of batch processing of large-scale object datasets, the present application provides an embodiment of an electronic device for implementing all or part of the content of the large-scale three-dimensional image data video generation method. The electronic device specifically includes the following content:

[0154] A processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to transmit information between the large-scale 3D image data and video generation device and related devices such as core business systems, user terminals, and related databases; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., but this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the large-scale 3D image data and video generation method and the large-scale 3D image data and video generation device in the embodiments, the contents of which are incorporated herein and repeated parts are not repeated.

[0155] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0156] In practical applications, portions of the large-scale 3D image data video generation method can be executed on the electronic device as described above, or all operations can be performed on the client device. The specific selection can be based on the processing capabilities of the client device and the limitations of the user's usage scenario. This application does not impose any restrictions on this. If all operations are performed on the client device, the client device may also include a processor.

[0157] The client device may include a communication module (i.e., a communication unit) that can establish a communication connection with a remote server to implement data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0158] Figure 4 Schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Figure 4 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that the Figure 4 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0159] In one embodiment, the function of the large-scale 3D image data video generation method can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control:

[0160] Step S101: receiving multiple images to be processed and determining the entropy value or gradient strength of each image to be processed;

[0161] Step S102: Based on the entropy value or gradient strength of the image to be processed, a local area of ​​the image to be processed corresponding to an entropy value greater than a preset entropy value threshold or a gradient strength greater than a preset gradient strength is determined as a high complexity area in the image to be processed;

[0162] Step S103: receiving preset spatial region parameters or preset volume feature parameters, reducing the spatial region parameters or the preset volume feature parameters according to a preset ratio, segmenting the image to be processed according to the spatial region parameters or the preset volume feature parameters in the non-high complexity region, and segmenting the image to be processed according to the reduced spatial region parameters or the preset volume feature parameters in the high complexity region, to obtain a plurality of data blocks to be processed;

[0163] Step S104: Determine the relative position of each to-be-processed data block corresponding to each to-be-processed image in the to-be-processed image, render each to-be-processed data block according to a preset color bit depth, obtain a single rendered frame corresponding to each to-be-processed data block, extract the maximum intensity information of each voxel in each rendered frame based on a maximum intensity projection algorithm, and merge the maximum intensity information corresponding to each to-be-processed data block according to the relative position into a single video frame corresponding to each to-be-processed image;

[0164] Step S105: determining the three-dimensional spatial relationship between the multiple images to be processed, stitching the video frames according to the three-dimensional spatial relationship, and performing error minimization on the stitched video frames to obtain a three-dimensional video of the object.

[0165] From the above description, it can be seen that the electronic device provided in the embodiment of the present application innovatively determines the high-complexity area in the image to be processed based on multiple images to be processed and the entropy value or gradient strength of the image to be processed, wherein the multiple images to be processed are obtained by multi-angle image acquisition of the same object at the same time, accepting preset spatial area parameters or preset volume feature parameters, reducing the spatial area parameters or preset volume feature parameters according to a preset ratio, and dynamically segmenting the high-complexity area and the non-high-complexity area categories to obtain multiple data blocks to be processed corresponding to the image to be processed, which can reduce the computing load and avoid redundancy caused by excessive segmentation. The method uses a maximum intensity projection algorithm to extract the maximum intensity information of each voxel in each rendered frame. The maximum intensity information corresponding to each piece of data to be processed is then merged into a single video frame corresponding to each image to be processed according to their relative positions. The three-dimensional spatial relationship between the multiple images to be processed is determined. The video frames are then stitched together according to this three-dimensional spatial relationship, and the error is minimized on the stitched video frames to obtain a three-dimensional video of the object. This method can extract key information through the maximum intensity projection algorithm and stitch the video frames together by minimizing the error, resulting in a clearer and more accurate three-dimensional video, improving both visual quality and the efficiency of three-dimensional video generation. This method effectively addresses the inefficiencies of traditional technologies when users are performing batch processing of large-scale object datasets or multidimensional analysis, significantly improving the efficiency of batch processing of large-scale object datasets.

[0166] In another embodiment, the large-scale three-dimensional image data video generation device can be configured separately from the central processing unit 9100. For example, the large-scale three-dimensional image data video generation device can be configured as a chip connected to the central processing unit 9100, and the function of the large-scale three-dimensional image data video generation method can be implemented under the control of the central processing unit.

[0167] like Figure 4 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Figure 4 In addition, the electronic device 9600 may also include all components shown in Figure 4 For components not shown, reference may be made to the prior art.

[0168] like Figure 4 As shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.

[0169] Memory 9140 can be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It can store the aforementioned failure-related information and also store programs that execute the relevant information. The CPU 9100 can execute the programs stored in memory 9140 to implement information storage or processing.

[0170] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 may be, for example, a keypad or touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display objects such as images and text. The display may be, for example, an LCD display, but is not limited thereto.

[0171] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), or a SIM card. Alternatively, it may be a memory that retains information even when power is off, can be selectively erased, and is provided with more data. Examples of such memory are sometimes referred to as EPROMs. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 by the central processing unit 9100.

[0172] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various driver programs for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0173] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.

[0174] Based on different communication technologies, multiple communication modules 9110 can be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby implementing typical telecommunication functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Furthermore, the audio processor 9130 is also coupled to the central processing unit 9100, enabling local recording via the microphone 9132 and playback of stored audio via the speaker 9131.

[0175] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the method for generating large-scale 3D image data and video in the above-mentioned embodiment, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the computer program implements all steps of the method for generating large-scale 3D image data and video in the above-mentioned embodiment, where the execution subject is a server or a client. For example, when the processor executes the computer program, the following steps are implemented:

[0176] Step S101: receiving multiple images to be processed and determining the entropy value or gradient strength of each image to be processed;

[0177] Step S102: Based on the entropy value or gradient strength of the image to be processed, a local area of ​​the image to be processed corresponding to an entropy value greater than a preset entropy value threshold or a gradient strength greater than a preset gradient strength is determined as a high complexity area in the image to be processed;

[0178] Step S103: receiving preset spatial region parameters or preset volume feature parameters, reducing the spatial region parameters or the preset volume feature parameters according to a preset ratio, segmenting the image to be processed according to the spatial region parameters or the preset volume feature parameters in the non-high complexity region, and segmenting the image to be processed according to the reduced spatial region parameters or the preset volume feature parameters in the high complexity region, to obtain a plurality of data blocks to be processed;

[0179] Step S104: Determine the relative position of each to-be-processed data block corresponding to each to-be-processed image in the to-be-processed image, render each to-be-processed data block according to a preset color bit depth, obtain a single rendered frame corresponding to each to-be-processed data block, extract the maximum intensity information of each voxel in each rendered frame based on a maximum intensity projection algorithm, and merge the maximum intensity information corresponding to each to-be-processed data block according to the relative position into a single video frame corresponding to each to-be-processed image;

[0180] Step S105: determining the three-dimensional spatial relationship between the multiple images to be processed, stitching the video frames according to the three-dimensional spatial relationship, and performing error minimization on the stitched video frames to obtain a three-dimensional video of the object.

[0181] From the above description, it can be seen that the computer-readable storage medium provided in the embodiment of the present application innovatively determines the high-complexity area in the image to be processed based on multiple images to be processed and the entropy value or gradient strength of the image to be processed, wherein the multiple images to be processed are obtained by multi-angle image acquisition of the same object at the same time, accepting preset spatial area parameters or preset volume feature parameters, reducing the spatial area parameters or preset volume feature parameters according to a preset ratio, and dynamically segmenting the high-complexity area and the non-high-complexity area categories to obtain multiple data blocks to be processed corresponding to the image to be processed, which can reduce the computing load and avoid the problems caused by excessive segmentation. This method increases image processing speed by eliminating redundant computations. The maximum intensity information for each voxel in each rendered frame is extracted using a maximum intensity projection algorithm. The maximum intensity information corresponding to each piece of data to be processed is then merged into a single video frame corresponding to each image to be processed, based on their relative positions. The three-dimensional spatial relationship between the multiple images to be processed is determined, and the video frames are stitched together according to this three-dimensional spatial relationship. Error minimization is then performed on the stitched video frames to produce a three-dimensional video of the object. This method extracts key information using the maximum intensity projection algorithm and stitches the video frames together using error minimization, resulting in a clearer and more accurate three-dimensional video, improving both visual quality and the efficiency of three-dimensional video generation. This method effectively addresses the inefficiencies of traditional technologies when performing batch processing of large-scale object datasets or multidimensional analysis, significantly improving the efficiency of batch processing of large-scale object datasets.

[0182] The present application also provides a computer program product capable of implementing all steps of the large-scale 3D image data video generation method described in the above embodiment, where the execution subject is a server or a client. When the computer program / instructions are executed by a processor, the computer program / instructions implement the steps of the large-scale 3D image data video generation method. For example, the computer program / instructions implement the following steps:

[0183] Step S101: receiving multiple images to be processed and determining the entropy value or gradient strength of each image to be processed;

[0184] Step S102: Based on the entropy value or gradient strength of the image to be processed, a local area of ​​the image to be processed corresponding to an entropy value greater than a preset entropy value threshold or a gradient strength greater than a preset gradient strength is determined as a high complexity area in the image to be processed;

[0185] Step S103: receiving preset spatial region parameters or preset volume feature parameters, reducing the spatial region parameters or the preset volume feature parameters according to a preset ratio, segmenting the image to be processed according to the spatial region parameters or the preset volume feature parameters in the non-high complexity region, and segmenting the image to be processed according to the reduced spatial region parameters or the preset volume feature parameters in the high complexity region, to obtain a plurality of data blocks to be processed;

[0186] Step S104: Determine the relative position of each to-be-processed data block corresponding to each to-be-processed image in the to-be-processed image, render each to-be-processed data block according to a preset color bit depth, obtain a single rendered frame corresponding to each to-be-processed data block, extract the maximum intensity information of each voxel in each rendered frame based on a maximum intensity projection algorithm, and merge the maximum intensity information corresponding to each to-be-processed data block according to the relative position into a single video frame corresponding to each to-be-processed image;

[0187] Step S105: determining the three-dimensional spatial relationship between the multiple images to be processed, stitching the video frames according to the three-dimensional spatial relationship, and performing error minimization on the stitched video frames to obtain a three-dimensional video of the object.

[0188] From the above description, it can be seen that the computer program product provided in the embodiment of the present application innovatively determines the high-complexity area in the image to be processed based on multiple images to be processed and the entropy value or gradient strength of the image to be processed, wherein the multiple images to be processed are obtained by multi-angle image acquisition of the same object at the same time, accepting preset spatial area parameters or preset volume feature parameters, reducing the spatial area parameters or preset volume feature parameters according to a preset ratio, and dynamically segmenting the high-complexity area and the non-high-complexity area categories to obtain multiple data blocks to be processed corresponding to the image to be processed, which can reduce the computing load and avoid the problem caused by excessive segmentation. Redundant computation improves image processing speed. Based on the maximum intensity projection algorithm, the maximum intensity information for each voxel in each rendered frame is extracted. The maximum intensity information corresponding to each piece of data to be processed is then merged into a single video frame corresponding to each image to be processed, according to their relative positions. The three-dimensional spatial relationship between the multiple images to be processed is determined, and the video frames are stitched together according to this three-dimensional spatial relationship. Error minimization is then performed on the stitched video frames to produce a three-dimensional video of the object. This method can extract key information through the maximum intensity projection algorithm and stitch the video frames together by minimizing the error, resulting in a clearer and more accurate three-dimensional video, improving both visual quality and the efficiency of 3D video generation. This method effectively addresses the inefficiencies of traditional technologies when performing batch processing of large-scale object datasets or multidimensional analysis, significantly improving the efficiency of batch processing of large-scale object datasets.

[0189] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0190] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as a combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0191] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0192] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0193] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

Claims

1. A method for generating large-scale three-dimensional image data video, characterized in that: The method comprises: Receiving a plurality of images to be processed, and determining an entropy value or a gradient strength of each of the images to be processed, wherein the plurality of images to be processed are obtained by capturing images of the same object from multiple angles at the same time; Based on the entropy value or gradient strength of the image to be processed, determining a local area of ​​the image to be processed corresponding to an entropy value greater than a preset entropy value threshold or a gradient strength greater than a preset gradient strength as a high-complexity area in the image to be processed; receiving a preset spatial region parameter or a preset volume feature parameter, reducing the spatial region parameter or the preset volume feature parameter according to a preset ratio, segmenting the image to be processed according to the spatial region parameter or the preset volume feature parameter in a non-high-complexity region, and segmenting the image to be processed according to the reduced spatial region parameter or the preset volume feature parameter in a high-complexity region, to obtain a plurality of data blocks to be processed; Determining the relative position of each of the to-be-processed data blocks corresponding to each of the to-be-processed images in the to-be-processed image, rendering each of the to-be-processed data blocks according to a preset color bit depth to obtain a single rendered frame corresponding to each of the to-be-processed data blocks, extracting maximum intensity information included in each voxel in each of the rendered frames based on a maximum intensity projection algorithm, and merging the maximum intensity information corresponding to each of the to-be-processed data blocks into a single video frame corresponding to each of the to-be-processed images according to the relative positions; Determine a three-dimensional spatial relationship between the plurality of images to be processed, stitch the video frames together according to the three-dimensional spatial relationship, and perform error minimization on the stitched video frames to obtain a three-dimensional video of the object.

2. The method according to claim 1, characterized in that Segmenting the image to be processed in the high-complexity area according to the reduced spatial area parameter or the preset volume feature parameter includes: When there are continuously distributed high-complexity regions, recursively segment the continuously distributed high-complexity regions using an octree structure until the entropy value or the gradient strength variance of each of the data blocks to be processed obtained by segmentation is less than a preset tolerance, and the continuously distributed region is an area where any two of the high-complexity regions overlap; At the junction of the high-complexity area and the non-high-complexity area, a transition data block with a gradually changing size is generated, and the size corresponding to the transition data block linearly transitions from the size corresponding to the data block to be processed obtained by dividing the spatial area parameters or the preset volume feature parameters after reduction to the size corresponding to the data block to be processed obtained by dividing the spatial area parameters or the preset volume feature parameters before reduction.

3. The method according to claim 1, characterized in that Rendering each of the to-be-processed data blocks according to a preset color bit depth to obtain a single rendering frame corresponding to each of the to-be-processed data blocks includes: Rendering the data blocks to be processed in parallel by computing a unified device architecture, and smoothing the voxels in the data blocks to be processed by a linear interpolation algorithm during the rendering process to obtain a rendering result; Local histogram equalization is performed on the rendering result to obtain a single rendering frame corresponding to each of the data blocks to be processed.

4. The method according to claim 1, wherein The extracting the maximum intensity information of each voxel in each of the rendered frames based on the maximum intensity projection algorithm includes: Traversing each voxel in each of the rendered frames, determining an original intensity value of each of the voxels, and when the rendered frame is multi-channel data, determining the maximum original intensity value in each of the channel data as the maximum intensity information; Merging the maximum intensity information corresponding to each of the to-be-processed data into a single video frame corresponding to each of the to-be-processed images according to each of the relative positions, including: Anisotropic diffusion filtering is performed on the video frame to smooth edges of the video frame.

5. The method according to claim 1, wherein Determining the three-dimensional spatial relationship between the plurality of images to be processed includes: Extracting key points and descriptors of each of the images to be processed, performing fast feature matching based on the key points and the descriptors to obtain a fast feature matching result, and determining a matching confidence corresponding to the fast feature matching result; Determine the fast feature matching result corresponding to the matching confidence less than a preset matching confidence threshold as an erroneous matching point, and eliminate the erroneous matching point; The fast feature matching result after eliminating the mismatched points is converted into a three-dimensional point cloud to obtain the three-dimensional spatial relationship between the multiple images to be processed.

6. The method according to claim 1, characterized in that The step of performing error minimization on the spliced ​​video frames to obtain a three-dimensional video of the object includes: determining a structural similarity index of the spliced ​​adjacent video frames in an overlapping area, and locating a misaligned area in the adjacent video frames based on the structural similarity index; Determining a pixel displacement vector in the misaligned area by using an optical flow method, determining a splicing offset based on the pixel displacement vector, and adjusting the spliced ​​video frames according to the splicing offset; The splicing boundaries in the spliced ​​video frames are processed by a Poisson fusion algorithm, and global brightness consistency correction is performed on the fused video frames to obtain a three-dimensional video of the object.

7. The method according to claim 1, characterized in that After stitching the video frames according to the three-dimensional spatial relationship, the method further includes: receiving display parameters and selected instructions in a preset format, wherein the preset format includes any one of a Lua script file, an HTTP interface, and a Python interface, and the selected instruction is used to execute the display parameters on a single video frame; The display of the spliced ​​video frames is adjusted according to the display parameters and the selected instructions.

8. A large-scale three-dimensional image data video generation device, characterized in that: The device comprises: A receiving module, configured to receive a plurality of images to be processed and determine an entropy value or a gradient strength of each of the images to be processed, wherein the plurality of images to be processed are obtained by capturing images of the same object from multiple angles at the same time; a first processing module, configured to determine, based on the entropy value or gradient strength of the image to be processed, a local area of ​​the image to be processed corresponding to an entropy value greater than a preset entropy value threshold or a gradient strength greater than a preset gradient strength as a high-complexity area in the image to be processed; a second processing module, configured to receive preset spatial region parameters or preset volume feature parameters, reduce the spatial region parameters or the preset volume feature parameters according to a preset ratio, segment the image to be processed according to the spatial region parameters or the preset volume feature parameters in non-high-complexity regions, and segment the image to be processed according to the reduced spatial region parameters or the preset volume feature parameters in the high-complexity regions, to obtain a plurality of data blocks to be processed; a third processing module, configured to determine the relative position of each of the to-be-processed data blocks corresponding to each of the to-be-processed images in the to-be-processed images, render each of the to-be-processed data blocks according to a preset color bit depth to obtain a single rendered frame corresponding to each of the to-be-processed data blocks, extract maximum intensity information of each voxel in each of the rendered frames based on a maximum intensity projection algorithm, and merge the maximum intensity information corresponding to each of the to-be-processed data blocks into a single video frame corresponding to each of the to-be-processed images according to the relative positions; The splicing module is used to determine the three-dimensional spatial relationship between the multiple images to be processed, splice the video frames according to the three-dimensional spatial relationship, and perform error minimization on the spliced ​​video frames to obtain a three-dimensional video of the object.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the large-scale three-dimensional image data video generation method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for generating large-scale three-dimensional image data video according to any one of claims 1 to 7 are implemented.