Method, device and equipment for generating three-dimensional image based on prompt text and medium
By using a text-based 3D image generation method, which automatically processes text embedding vectors and disparity offsets, the problem of low generation efficiency in existing technologies is solved, achieving efficient and reliable 3D image generation with realistic spatial depth and stereoscopic effects.
Patent Information
- Application Number
- CN202511035218.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-25
AI Technical Summary
The existing process for generating 3D images with left and right parallax is cumbersome, consumes a lot of human resources and time, and is easily affected by human intervention, resulting in low generation efficiency.
By using a text-based 3D image generation method, text embedding vectors are processed through an image generation model. Combined with depth estimation and disparity offset models, 3D images with left and right disparities are automatically generated. This includes steps such as feature extraction, fusion, semantic understanding, cropping, and merging, reducing manual operations.
It improves the efficiency and reliability of 3D image generation, reduces generation time, avoids the influence of human intervention, and generates 3D images with realistic spatial depth and stereoscopic effect.
Smart Images

Figure CN120526067B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer vision and artificial intelligence, and in particular to a three-dimensional image generation method based on a prompt text, a three-dimensional image generation device based on a prompt text, a three-dimensional image generation equipment based on a prompt text and a three-dimensional image generation medium based on a prompt text. BACKGROUND
[0002] Three-dimensional images with left-right parallax and three-dimensional images without left-right parallax have significant differences in visual presentation and experience. Three-dimensional images without left-right parallax often only simulate stereoscopic effect through color, shading, etc. It is difficult for the audience to obtain real spatial depth perception. Three-dimensional images with left-right parallax provide slightly different images for the left and right eyes, use the natural parallax mechanism of the human eye, and let the brain integrate a three-dimensional scene with real depth and spatial hierarchy to bring an immersive feeling.
[0003] However, the existing generation process of three-dimensional images with left-right parallax is complicated, which is not conducive to improving the generation efficiency of three-dimensional images with left-right parallax. The reason is that the existing technology mainly adopts a manual operation method to generate three-dimensional images with left-right parallax. The manual operation method requires a technician to use a binocular camera and a stereoscopic shooting support to manually adjust the shooting angle, distance or parameters to simulate the binocular perspective difference when the human eye observes a real scene. This will consume a large amount of human and time resources, increase the generation time of three-dimensional images with left-right parallax, and be easily affected by human intervention. Therefore, it is not conducive to improving the generation efficiency of three-dimensional images with left-right parallax. SUMMARY
[0004] The three-dimensional image generation method based on a prompt text, the three-dimensional image generation device based on a prompt text, the three-dimensional image generation equipment based on a prompt text and the three-dimensional image generation medium based on a prompt text provided by the embodiments of the present application solve the technical problem that the existing generation process of three-dimensional images with left-right parallax is complicated and not conducive to improving the generation efficiency of three-dimensional images with left-right parallax.
[0005] In a first aspect, the embodiments of the present application provide a three-dimensional image generation method based on a prompt text, applied to an electronic device, and the three-dimensional image generation method comprises:
[0006] obtaining a prompt text, performing feature extraction on the prompt text to obtain a text embedding vector, processing the text embedding vector through an image generation model to obtain a picture;
[0007] modifying the current size of the picture to a fixed size to obtain a modified picture, extracting features of different scales from the modified picture, fusing the features of different scales to obtain a fusion graph, and performing optimization processing on the fusion graph to obtain a depth map;
[0008] The size of the depth map is selected as a target size, semantic understanding processing and gradient supervision processing are performed on the fusion map, and a salient subject image is obtained, the salient subject image is divided into a subject range and a non-subject range;
[0009] The current size of the subject range is adjusted to the target size to obtain a modified subject range, an erosion operation is performed on the modified subject range by using a structural element to obtain an inner ring range in the modified subject range, and a plurality of depth values are extracted from the depth map according to the inner ring range;
[0010] The null values or abnormal values in the plurality of depth values are removed to obtain a depth value set, a depth value with the highest frequency of occurrence in the depth value set is selected as a depth value of a zero disparity plane, a disparity offset is generated according to the maximum disparity and a preset disparity offset model, a left image is obtained by shifting the pixel points of the modified image to the right according to the disparity offset, a right image is obtained by shifting the pixel points of the modified image to the left according to the disparity offset, a reference offset of the zero disparity plane is generated according to the depth value of the zero disparity plane, the maximum disparity and a preset reference offset model, a clipping boundary width is generated according to the reference offset of the zero disparity plane and a preset boundary width model, the left boundary of the left image is clipped by using the clipping boundary width to obtain a clipped left image, the right boundary of the right image is clipped by using the clipping boundary width to obtain a clipped right image, the clipped left image and the clipped right image are aligned in subject content by using the zero disparity plane to obtain an aligned left image and an aligned right image, the aligned left image and the aligned right image are filled with holes to obtain a filled left image and a filled right image, and the filled left image and the filled right image are merged to generate a three-dimensional image with left and right disparities.
[0011] In a possible implementation manner of the first aspect, the prompt text is obtained, feature extraction is performed on the prompt text to obtain a text embedding vector, and the text embedding vector is processed by using an image generation model to obtain an image, and the method comprises the following steps.
[0012] The prompt text input by a user is obtained from a text box, feature extraction is performed on the prompt text to obtain a text embedding vector, and the text embedding vector is input into an image generation model, and the text embedding vector is processed by using the image generation model to obtain an image.
[0013] In a possible implementation manner of the first aspect, the prompt text is obtained, feature extraction is performed on the prompt text to obtain a text embedding vector, and the text embedding vector is processed by using an image generation model to obtain an image, and the method comprises the following steps.
[0014] The user voice is acquired, the user voice is converted into prompt text, feature extraction is performed on the prompt text to obtain a text embedding vector, the text embedding vector is input into an image generation model, the text embedding vector is processed by the image generation model, and a picture is obtained.
[0015] In a possible implementation manner of the first aspect, the current size of the picture is modified to a fixed size to obtain a modified picture, different scale features are extracted from the modified picture, the different scale features are fused to obtain a fusion picture, and the fusion picture is subjected to optimization processing to obtain a depth map, including:
[0016] The picture is loaded by using a loading function, the current size of the picture is acquired, the current size of the picture is modified to a fixed size to obtain a modified picture;
[0017] Different scale features are extracted from the modified picture, the different scale features are fused to obtain a fusion picture, the fusion picture is input into a depth estimation model, the depth estimation model is subjected to optimization processing by using a gradient matching loss, and a depth map output by the depth estimation model is obtained, the gradient matching loss being a loss function used to measure the difference between the output result of the depth estimation model and a target result in gradient features.
[0018] In a possible implementation manner of the first aspect, the size of the depth map is selected as a target size, the fusion picture is subjected to semantic understanding processing and gradient supervision processing to obtain a saliency subject image, and the saliency subject image is divided into a subject range and a non-subject range, including:
[0019] The size of the depth map is selected as a target size, the fusion picture is subjected to semantic understanding processing and gradient supervision processing to obtain a saliency subject image, a part of the saliency subject image in which a target is located is marked as a subject range, and a part of the saliency subject image other than the target is marked as a non-subject range.
[0020] In a possible implementation manner of the first aspect, the null values or the abnormal values are removed from the plurality of depth values to obtain a set of depth values, a depth value with a highest frequency of occurrence in the set of depth values is selected as a depth value of the zero disparity surface, a disparity offset is generated according to the maximum disparity and a preset disparity offset model, a left image is obtained by shifting the pixel points of the modified image to the right according to the disparity offset, a right image is obtained by shifting the pixel points of the modified image to the left according to the disparity offset, a reference offset of the zero disparity surface is generated according to the depth value of the zero disparity surface, the maximum disparity and a preset reference offset model, a clipping boundary width is generated according to the reference offset of the zero disparity surface and a preset boundary width model, a left boundary of the left image is clipped using the clipping boundary width to obtain a clipped left image, a right boundary of the right image is clipped using the clipping boundary width to obtain a clipped right image, the clipped left image and the clipped right image are aligned in subject content using the zero disparity surface to obtain an aligned left image and an aligned right image, the aligned left image and the aligned right image are subjected to hole filling to obtain a filled left image and a filled right image, and the filled left image and the filled right image are merged to generate a three-dimensional image with left and right disparities, and the three-dimensional image generation method comprises the following steps of:
[0021] obtaining a display instruction, executing the display instruction, and displaying the three-dimensional image with left and right disparities on the screen.
[0022] In a possible implementation manner of the first aspect, the disparity offset model is a division offset model or a multiplication offset model.
[0023] The division offset model is:
[0024] ;
[0025] is the disparity offset; is a depth value of the pixel point, is the maximum disparity, is a depth offset;
[0026] The maximum disparity refers to a maximum pixel offset of a same object in a horizontal direction in stereovision.
[0027] The reference offset model is:
[0028] ;
[0029] is the reference offset of the zero disparity surface; is the maximum disparity, is the depth value of the zero disparity surface; is the depth offset;
[0030] wherein the boundary width model is:
[0031] ;
[0032] is the clipping boundary width;
[0033] The pixels obtained by the division offset model are batch filled according to the gridding coordinates.
[0034] The multiplication offset model is:
[0035] ;
[0036] is the disparity offset; is the depth value of the pixel point, is the maximum disparity, is the depth value of the zero disparity surface.
[0037] The pixels obtained by the multiplication offset model are filled according to the depth order and the left-right order. The pixel rendering is performed through the depth sorting mechanism (far to near), and the left figure adopts the right-to-left filling strategy, and the right figure adopts the left-to-right filling strategy.
[0038] In a possible implementation of the first aspect, the structural element includes one or a combination of a square structural element, a cross structural element, a circular structural element, an elliptical structural element, a diamond structural element, and a star structural element.
[0039] In a second aspect, an embodiment of the present application provides a three-dimensional image generation device based on prompt text, applied to an electronic device, comprising:
[0040] A first acquisition module is configured to acquire the prompt text, perform feature extraction on the prompt text, obtain a text embedding vector, process the text embedding vector through an image generation model, and obtain a picture.
[0041] A modification module is configured to modify a current size of the picture to a fixed size, obtain a modified picture, extract features of different scales from the modified picture, fuse the features of different scales, obtain a fusion graph, and perform optimization processing on the fusion graph to obtain a depth graph.
[0042] A second acquisition module is configured to select a size of the depth graph as a target size, perform semantic understanding processing and gradient supervision processing on the fusion graph, obtain a saliency subject image, and divide the saliency subject image into a subject range and a non-subject range.
[0043] The extraction module is configured to adjust a current size of the subject range to a target size to obtain a modified subject range, perform an erosion operation on the modified subject range by using a structure element to obtain an inner ring range in the modified subject range, and extract a plurality of depth values from the depth map according to the inner ring range;
[0044] The generation module is configured to remove null values or abnormal values from the plurality of depth values to obtain a set of depth values, select a depth value with a highest frequency of occurrence in the set of depth values as a depth value of a zero disparity plane, generate a disparity offset according to the maximum disparity and a preset disparity offset model, shift pixel points of the modified picture to the right according to the disparity offset to obtain a left picture, shift the pixel points of the modified picture to the left according to the disparity offset to obtain a right picture, generate a reference offset of the zero disparity plane according to the depth value of the zero disparity plane, the maximum disparity, and a preset reference offset model, generate a clipping boundary width according to the reference offset of the zero disparity plane and a preset boundary width model, clip a left boundary of the left picture by using the clipping boundary width to obtain a clipped left picture, clip a right boundary of the right picture by using the clipping boundary width to obtain a clipped right picture, perform subject content alignment on the clipped left picture and the clipped right picture by using the zero disparity plane to obtain an aligned left picture and an aligned right picture, perform hole filling on the aligned left picture and the aligned right picture to obtain a filled left picture and a filled right picture, and merge the filled left picture and the filled right picture to generate a three-dimensional image with left and right disparities.
[0045] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the three-dimensional image generation method of any one of the first aspect when executing the computer program.
[0046] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program is executable on a processor to implement the three-dimensional image generation method of any one of the first aspect.
[0047] In a fifth aspect, a computer program product is provided, which, when executed on an electronic device, causes the electronic device to perform the three-dimensional image generation method of any one of the first aspect.
[0048] The beneficial effects of the embodiment of the present application are twofold. On the one hand, the null values or abnormal values are removed from the plurality of depth values to obtain a depth value set, the depth value with the highest frequency of occurrence in the depth value set is selected as the depth value of the zero disparity surface, the disparity offset is generated according to the maximum disparity and the preset disparity offset model, the pixel points of the modified picture are shifted to the right according to the disparity offset to obtain a left picture, the pixel points of the modified picture are shifted to the left according to the disparity offset to obtain a right picture, the reference offset of the zero disparity surface is generated according to the depth value of the zero disparity surface, the maximum disparity and the preset reference offset model, the clipping boundary width is generated according to the reference offset of the zero disparity surface and the preset boundary width model, the left boundary of the left picture is clipped using the clipping boundary width to obtain a clipped left picture, the right boundary of the right picture is clipped using the clipping boundary width to obtain a clipped right picture, the main content of the clipped left picture and the clipped right picture is aligned using the zero disparity surface to obtain an aligned left picture and an aligned right picture, the aligned left picture and the aligned right picture are filled with holes to obtain a filled left picture and a filled right picture, and the filled left picture and the filled right picture are merged to generate a three-dimensional image with left and right disparities. Since no manual operation is required, the generation time of the three-dimensional image with left and right disparities is reduced, and the generation efficiency of the three-dimensional image with left and right disparities is improved. On the other hand, since the three-dimensional image with left and right disparities is automatically generated, it is not affected by human intervention, and the reliability of the three-dimensional image with left and right disparities is improved. BRIEF DESCRIPTION OF DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0050] Figure 1 The application scenario diagram of the three-dimensional image generation method provided by the embodiment of the present application;
[0051] Figure 2 The flowchart of the three-dimensional image generation method provided by the embodiment of the present application;
[0052] Figure 3 The flowchart of the display list processing result provided by the embodiment of the present application;
[0053] Figure 4 The schematic block diagram of the three-dimensional image generation device provided by the embodiment of the present application;
[0054] Figure 5 The structural schematic diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to make the purposes, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0056] The three-dimensional image generation method provided by the embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), and the like. The embodiments of the present application do not make any limitation on the specific type of electronic device.
[0057] Referring to Figure 1 , Figure 1 The application scenario diagram of the three-dimensional image generation method provided by the embodiments of the present application is described as follows:
[0058] The electronic device obtains the prompt text through a text box or user voice.
[0059] Among them, the input mode of the text box provides a precise editing approach for the user, which facilitates the user to carefully consider the words and correct errors, and is especially suitable for scenes that require rigorous expression or have high requirements for content accuracy.
[0060] Among them, the input mode of the user voice greatly improves the interaction efficiency. When the user is busy with both hands or pursues convenient operation, the user only needs to speak to quickly convey information, so that information entry is more natural and easy.
[0061] In the embodiments of the present application, the electronic device can obtain the prompt text from the text box and the user voice. The two acquisition modes complement each other and can meet the diversified needs of the user in different scenarios, promoting the development of human-computer interaction in a more efficient direction.
[0062] Referring to Figure 2 , Figure 2 The flowchart of the three-dimensional image generation method provided by the embodiments of the present application is shown in FIG. 1, and the method can be applied to an electronic device.
[0063] As Figure 2As shown, the three-dimensional image generation method provided by the embodiment of the application includes the following steps, which are described in detail as follows.
[0064] S201, obtaining prompt text, performing feature extraction on the prompt text to obtain a text embedding vector, processing the text embedding vector through an image generation model to obtain a picture;
[0065] Among them, S201 has two implementation ways, which are described in detail as follows:
[0066] The first way:
[0067] The prompt text input by the user is obtained from the text box, feature extraction is performed on the prompt text to obtain a text embedding vector, the text embedding vector is input into the image generation model, and the text embedding vector is processed through the image generation model to obtain a picture.
[0068] The second way:
[0069] The user's voice is obtained, the user's voice is converted into prompt text, feature extraction is performed on the prompt text to obtain a text embedding vector, the text embedding vector is input into the image generation model, and the text embedding vector is processed through the image generation model to obtain a picture.
[0070] Among them, the prompt text is a natural language instruction for guiding the image generation model to create a picture, and the prompt text functions to convert the user's intention into semantic features that can be understood by the image generation model.
[0071] For ease of illustration, the following examples are given:
[0072] For example, the prompt text is: a cute cartoon panda;
[0073] For example, the prompt text is: a husky pilot wearing goggles.
[0074] S202, modifying the current size of the picture to a fixed size to obtain a modified picture, extracting features of different scales from the modified picture, fusing the features of different scales to obtain a fusion picture, and performing optimization processing on the fusion picture to obtain a depth map;
[0075] Among them, the number of pixels and the spatial resolution contained in pictures of different sizes are different, and large-size pictures may introduce noise due to too many details, and small-size pictures may lead to inaccurate disparity estimation due to insufficient information. Modifying the current size of the picture to a fixed size provides a standardized input for disparity calculation, so that the algorithm can perform pixel matching and disparity calculation at the same spatial scale, effectively reducing the error caused by size difference, and significantly improving the accuracy and reliability of disparity estimation.
[0076] S203, select the size of the depth map as the target size, perform semantic understanding processing and gradient supervision processing on the fusion map, obtain a salient subject image, and divide the salient subject image into a subject range and a non-subject range;
[0077] The salient subject image refers to a subject image generated by automatically identifying and extracting a visual focus area in an image through computer vision algorithms. This technology is based on human visual attention mechanism and determines the most prominent part of the picture by analyzing visual attributes such as color contrast, texture features, and spatial position.
[0078] The subject range focuses on the most critical and representative information in the image, which is the core of the salient subject image. The non-subject range is the relatively secondary background or auxiliary element. After determining the subject range, subsequent algorithm operations and feature extraction operations can be carried out in this key area, avoiding unnecessary processing of the non-subject range, thereby greatly reducing the computational complexity.
[0079] The subject range is extracted from the fusion map according to the boundary of the subject object, which can remove redundant information and focus attention on the subject range.
[0080] S204, adjust the current size of the subject range to the target size to obtain a modified subject range, perform an erosion operation on the modified subject range using a structural element to obtain an inner ring range in the modified subject range, and extract a plurality of depth values from the depth map according to the inner ring range.
[0081] The subject range may contain some small noise points that interfere with accurate analysis and identification of the subject object. Through the erosion operation, the structural element can eliminate the noise points in the subject range. Since noise points are usually isolated and small in area, they are more likely to be completely covered and removed by the structural element, thereby effectively purifying the image and making the outline of the modified subject range clearer.
[0082] The structural element includes one of a square structural element, a cross structural element, a circular structural element, an elliptical structural element, a diamond structural element, a star structural element, or a combination thereof.
[0083] The cross structural element is used for erosion operation in a specific direction, such as preserving horizontal or vertical line features or removing elongated noise.
[0084] The shape of the structural element is selected in combination with specific task requirements.
[0085] The square structural element is simple to calculate and can quickly remove small particle noise, making it suitable for high real-time and regular target scenarios.
[0086] The circular structure element is isotropic, can uniformly erode the main body range, and maximally retains the circular edge of the main body range.
[0087] The elliptical structure element can match the elliptical shape target, accurately erodes along the contour, and is suitable for processing the main body range with specific elliptical features.
[0088] The rhombus structure element is sensitive to specific angle edges, can strengthen the erosion effect of such edges, and is suitable for processing the main body range with sharp angles or specific directional features.
[0089] The star structure element takes into account multi-directional erosion, can effectively remove complex edge noise and retain structures, and is suitable for processing the main body range with radial or complex edges.
[0090] The size of the depth map is selected as the target size, the fusion map is subjected to semantic understanding processing and gradient supervision processing to obtain a salient main body image, the salient main body image is divided into a main body range and a non-main body range, and the method comprises the following steps:
[0091] The size of the depth map is selected as the target size, the fusion map is subjected to semantic understanding processing and gradient supervision processing to obtain a salient main body image, the part where the target is located in the salient main body image is marked as the main body range, and the part other than the target in the salient main body image is marked as the non-main body range.
[0092] The semantic understanding processing is a key means for deeply mining semantic information of the fusion map by relying on advanced artificial intelligence technology.
[0093] The gradient supervision processing is an important technical means in the field of image processing and computer vision, and the core lies in using image gradient information to guide and optimize the learning process of the model. The image gradient can accurately reflect the change of pixel intensity in space, and by calculating the gradient value of each pixel point in the image in different directions, the key local features such as edges and textures of the image can be clearly captured, which plays a crucial role in distinguishing objects from backgrounds and recognizing object contours.
[0094] Through the semantic understanding processing, the range possibly containing the salient main body is preliminarily circled, and the gradient supervision processing can highlight the areas with sharp changes in the image, further refining the boundary of the salient main body and making it more clearly separated from the background. The semantic understanding processing and the gradient supervision processing work together to finally accurately extract the salient main body image from the fusion map.
[0095] S205, the null value or abnormal value is removed in the plurality of depth values, a depth value set is obtained, a depth value with the highest frequency of occurrence in the depth value set is selected as a depth value of a zero disparity surface, a disparity offset is generated according to the maximum disparity and a preset disparity offset model, a left image is obtained by shifting the pixel points of the modified picture to the right according to the disparity offset, a right image is obtained by shifting the pixel points of the modified picture to the left according to the disparity offset, a reference offset of the zero disparity surface is generated according to the depth value of the zero disparity surface, the maximum disparity and a preset reference offset model, a clipping boundary width is generated according to the reference offset of the zero disparity surface and a preset boundary width model, the left boundary of the left image is clipped using the clipping boundary width to obtain a clipped left image, the right boundary of the right image is clipped using the clipping boundary width to obtain a clipped right image, the clipped left image and the clipped right image are aligned in subject content using the zero disparity surface to obtain an aligned left image and an aligned right image, the aligned left image and the aligned right image are filled with holes to obtain a filled left image and a filled right image, and the filled left image and the filled right image are merged to generate a three-dimensional image with left and right disparities.
[0096] In this way, the left boundary of the left image is clipped using the clipping boundary width to obtain a clipped left image, and the right boundary of the right image is clipped using the clipping boundary width to obtain a clipped right image. This clipping method is targeted and only operates on the left boundary of the left image and the right boundary of the right image, so important content is not lost due to excessive processing, the image quality is ensured, and the disparity offset problem is effectively corrected.
[0097] In selecting the zero disparity surface, various comprehensive statistics of the salient subject image are calculated, such as mode, mode frequency, effective depth pixel number, inner circle depth pixel number, mean, median, standard deviation, and extreme value. At the same time, visual analysis is performed in combination with the depth map, the salient subject image, the subject inner circle image, and the subject inner circle depth image by calculating the mean plus or minus one, two, and three times the standard deviation, and sorting the data distribution to verify whether the depth values are normally distributed.
[0098] Based on the current data analysis, in addition to the default mode scheme, a variety of multi-dimensional statistical quantity combination schemes can be flexibly selected according to actual needs when determining the zero disparity surface. Specifically as follows:
[0099] Traditional statistics, such as mean, median, standard deviation, and extreme value, can describe the basic distribution of data and provide a preliminary understanding of the overall situation of the data.
[0100] Advanced statistics, including weighted mean, geometric mean, alpha-trimmed mean, interquartile range, entropy analysis, and kurtosis detection, can enhance the adaptability to complex scenarios and more accurately characterize the characteristics of data in complex situations.
[0101] The Chinese standard translation of alpha-trimmed mean is alpha-trimmed mean, which is a statistical estimation method combining robustness and efficiency. The core logic is: first sort the data by size, remove the extreme values on both ends accounting for α proportion, and then take the arithmetic mean of the remaining middle data. The α proportion is usually 5%-25%.
[0102] Spatial statistics, such as Moran's index and semivariogram, can capture the spatial correlation of deep data and help us understand the distribution and relationship of data in the spatial dimension.
[0103] Robust statistics, such as Qn estimator and Hodges-Lehmann estimator, mainly suppress the interference of outliers on the analysis results, ensuring the stability and reliability of statistical results in the presence of abnormal data.
[0104] The Chinese of Qn estimator is: robust scale estimator, Qn estimator is a discrete degree measurement method based on data quantile, which has strong resistance to outliers, especially suitable for data sets with outliers.
[0105] The Chinese of Hodges-Lehmann estimator is: median estimator, Hodges-Lehmann estimator estimates the location parameter through the median of data pairs, which has better robustness than traditional mean.
[0106] Dynamic scene statistics, such as exponential moving average and KL divergence, can optimize the adaptability to motion scenes and better handle data that changes dynamically over time.
[0107] For different application scenarios, the following specific optimization combination strategies are recommended:
[0108] In the basic analysis scenario, the combination of mode, interquartile range and skewness can quickly locate the core depth interval and provide a basic range reference for subsequent analysis work.
[0109] When there is a need for noise reduction, that is, there may be more outliers in the data affecting the analysis results, the combination of Qn estimator and Hodges-Lehmann estimator can effectively reduce the influence of outliers and improve the accuracy of analysis. If there is a need for spatial analysis, that is, to understand the distribution and correlation of deep data in space, the combination of Moran's index and semivariogram can be used to explore the characteristics of data in the spatial dimension.
[0110] In a dynamic scenario, that is, the data will change over time, the combination of exponential moving average and KL divergence can track deep drift and capture the characteristics of data in the dynamic change process in time.
[0111] The Chinese full name of KL divergence is: Kullback-Leibler divergence. KL divergence is used to quantify the asymmetric difference between two probability distributions.
[0112] Among them, the combination of multi-dimensional statistics can comprehensively analyze the depth distribution characteristics and provide fine decision basis for disparity optimization, significantly improving the quality of stereo image generation.
[0113] Among them, the depth value with the highest frequency in the depth value set is selected as the depth value of the zero disparity surface, which can filter out these interference data and ensure the reliability of the zero disparity surface. Because the depth value with the highest frequency in the depth value set is the mode, the mode is not sensitive to outliers, which can filter out these interference data, so as to ensure the reliability of the zero disparity surface.
[0114] For ease of illustration, the following is given as an example:
[0115] The depth value set has depth value 1, depth value 2, depth value 3, depth value 3, depth value 3, depth value 3, depth value 3, depth value 2, and the depth value 3 has the highest frequency. Select depth value 3 as the depth value of the zero disparity surface.
[0116] Among them, according to the maximum disparity and the preset disparity offset model, the disparity offset is generated, including:
[0117] According to the maximum disparity and the division offset model, the reference offset of the zero disparity surface is generated;
[0118] Among them, the disparity offset model adopts a division offset model or a multiplication offset model.
[0119] The division offset model is:
[0120] ;
[0121] The disparity offset is: is the depth value of the pixel point, Max disparity, Depth offset. It can be adjusted by a certain depth value the size of the corresponding disparity offset, the greater, the smaller, the smaller, the greater.
[0122] It can also be used to prevent the denominator from being zero.
[0123] Wherein, the maximum disparity refers to the maximum pixel offset of the same object in the horizontal direction in stereovision.
[0124] Wherein, the reference offset model is:
[0125] ;
[0126] The reference offset of the zero disparity surface; The depth value of the zero disparity surface; Depth offset. It can be adjusted by a certain depth value the size of the corresponding disparity offset, the greater, the smaller, the smaller, the greater. It can also be used to prevent the denominator from being zero. Wherein, the boundary width model is:
[0127] ;
[0128] The clipping boundary width.
[0129] The pixels obtained by the division offset model are filled in batches according to the gridding coordinates.
[0130] The multiplication offset model is:
[0131] ;
[0132] The disparity offset; The depth value of the pixel point, The maximum disparity, The depth value of the zero disparity surface;
[0133] Wherein, the maximum disparity refers to the maximum pixel offset of the same object in the horizontal direction in stereovision.
[0134] Pixels obtained by multiplication offset model are filled according to depth order and left-right order. Pixel drawing is performed by depth sorting mechanism (far to near), and the left figure adopts right-to-left filling strategy, and the right figure adopts left-to-right filling strategy.
[0135] The depth sorting mechanism is a technique in computer graphics for determining the order of objects or pixels in a three-dimensional scene relative to the viewer, and the core purpose of the depth sorting mechanism is to optimize the rendering process by sorting to avoid unnecessary calculations and visual errors.
[0136] The division offset model is an offset model involving division operation.
[0137] The multiplication offset model is an offset model involving multiplication operation.
[0138] The reasons why the division offset model is used preferentially in most cases of the application are as follows:
[0139] The division offset model: the content of the outer circle of the main body is derived from the background content, the semantics are consistent, and the abnormal phenomenon of image repair can be reduced. The grid batch calculation is fast.
[0140] The multiplication offset model: the content of the outer circle of the main body is not filled by the background content, so image repair is necessary to restore the outer circle of the main body to the background content, but the semantics of the inner and outer main body are inconsistent, and the boundary may be unclear. The pixel offset is done in order, which takes a long time.
[0141] Therefore, the adaptability of the division offset model is higher in most cases.
[0142] The method comprises the following steps: removing the null values or abnormal values in the plurality of depth values to obtain a set of depth values; selecting a depth value with the highest frequency of occurrence in the set of depth values as a depth value of a zero disparity surface; generating a disparity offset according to the maximum disparity and a preset disparity offset model; shifting the pixel points of the modified picture to the right according to the disparity offset to obtain a left picture; shifting the pixel points of the modified picture to the left according to the disparity offset to obtain a right picture; generating a reference offset of the zero disparity surface according to the depth value of the zero disparity surface, the maximum disparity and a preset reference offset model; generating a clipping boundary width according to the reference offset of the zero disparity surface and a preset boundary width model; clipping the left boundary of the left picture by using the clipping boundary width to obtain a clipped left picture; clipping the right boundary of the right picture by using the clipping boundary width to obtain a clipped right picture; aligning the main contents of the clipped left picture and the clipped right picture by using the zero disparity surface to obtain an aligned left picture and an aligned right picture; filling the holes of the aligned left picture and the aligned right picture to obtain a filled left picture and a filled right picture; and merging the filled left picture and the filled right picture to generate a three-dimensional image with left and right disparities.
[0143] The method further comprises the following steps: obtaining a display instruction; and executing the display instruction to display the three-dimensional image with left and right disparities on the screen.
[0144] The method further comprises the following steps: displaying the three-dimensional image with left and right disparities on the screen to present a stereoscopic effect.
[0145] The beneficial effects of the embodiments of the present application are twofold. On the one hand, the null values or abnormal values are removed from the plurality of depth values to obtain a depth value set, the depth value with the highest frequency of occurrence in the depth value set is selected as the depth value of the zero disparity surface, the disparity offset is generated according to the maximum disparity and the preset disparity offset model, the pixel points of the modified picture are shifted to the right according to the disparity offset to obtain a left picture, the pixel points of the modified picture are shifted to the left according to the disparity offset to obtain a right picture, the reference offset of the zero disparity surface is generated according to the depth value of the zero disparity surface, the maximum disparity and the preset reference offset model, the clipping boundary width is generated according to the reference offset of the zero disparity surface and the preset boundary width model, the left boundary of the left picture is clipped using the clipping boundary width to obtain a clipped left picture, the right boundary of the right picture is clipped using the clipping boundary width to obtain a clipped right picture, the main content of the clipped left picture and the clipped right picture is aligned using the zero disparity surface to obtain an aligned left picture and an aligned right picture, the aligned left picture and the aligned right picture are filled with holes to obtain a filled left picture and a filled right picture, and the filled left picture and the filled right picture are merged to generate a three-dimensional image with left and right disparities. Since no manual operation is required, the generation time of the three-dimensional image with left and right disparities is reduced, and the generation efficiency of the three-dimensional image with left and right disparities is improved.
[0146] Please refer to Figure 3 , Figure 3 The flowchart for displaying the list processing result provided by the embodiments of the present application is described as follows:
[0147] S301, loading a picture by loading a function, obtaining the current size of the picture, modifying the current size of the picture to a fixed size to obtain a modified picture.
[0148] S302, extracting features of different scales from the modified picture, fusing the features of different scales to obtain a fusion picture, inputting the fusion picture into a depth estimation model, and optimizing the depth estimation model by gradient matching loss to obtain a depth map output by the depth estimation model. Gradient matching loss is a loss function used to measure the difference between the output result of the depth estimation model and the target result in gradient features.
[0149] Gradient matching loss (Gradient Matching Loss) is an existing and widely used loss function in the field of computer vision and image generation. The specific formula of gradient matching loss is not described here.
[0150] In the embodiment of the present application, the depth estimation model is optimized by gradient matching loss, and the depth map output by the depth estimation model can improve the efficiency of obtaining the depth map.
[0151] Corresponding to the three-dimensional image generation method described in the above embodiment, please refer to Figure 4 , Figure 4 The schematic block diagram of the three-dimensional image generation device provided in the embodiment of the present application is shown in Figure 4 The three-dimensional image generation device 400 shown in Figure 1 The three-dimensional image generation device 400 shown in Figure 4 The three-dimensional image generation device 400 shown in
[0152] The first acquisition module 401 is configured to acquire prompt text, perform feature extraction on the prompt text to obtain a text embedding vector, process the text embedding vector by using an image generation model, and obtain a picture.
[0153] The modification module 402 is configured to modify a current size of the picture to a fixed size to obtain a modified picture, extract features of different scales from the modified picture, fuse the features of different scales to obtain a fusion picture, perform optimization processing on the fusion picture to obtain a depth map.
[0154] The second acquisition module 403 is configured to select a size of the depth map as a target size, perform semantic understanding processing and gradient supervision processing on the fusion picture to obtain a salient subject image, and divide the salient subject image into a subject range and a non-subject range.
[0155] The extraction module 404 is configured to adjust a current size of the subject range to the target size to obtain a modified subject range, perform an erosion operation on the modified subject range by using a structural element to obtain an inner ring range in the modified subject range, and extract a plurality of depth values from the depth map according to the inner ring range.
[0156] The generating module 405 is configured to remove null values or abnormal values from the plurality of depth values to obtain a set of depth values, select a depth value with the highest frequency of occurrence in the set of depth values as a depth value of a zero disparity surface, generate a disparity offset according to the maximum disparity and a preset disparity offset model, shift pixel points of the modified picture to the right according to the disparity offset to obtain a left picture, shift the pixel points of the modified picture to the left according to the disparity offset to obtain a right picture, generate a reference offset of the zero disparity surface according to the depth value of the zero disparity surface, the maximum disparity and a preset reference offset model, generate a clipping boundary width according to the reference offset of the zero disparity surface and a preset boundary width model, clip a left boundary of the left picture using the clipping boundary width to obtain a clipped left picture, clip a right boundary of the right picture using the clipping boundary width to obtain a clipped right picture, perform subject content alignment on the clipped left picture and the clipped right picture using the zero disparity surface to obtain an aligned left picture and an aligned right picture, perform hole filling on the aligned left picture and the aligned right picture to obtain a filled left picture and a filled right picture, and combine the filled left picture and the filled right picture to generate a three-dimensional image with left and right disparities.
[0157] It should be noted that each of the embodiments in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts of each embodiment can be referred to each other.
[0158] The beneficial effects of the embodiments of the present application are twofold. On the one hand, the null values or abnormal values are removed from the plurality of depth values to obtain a depth value set, the depth value with the highest frequency of occurrence in the depth value set is selected as the depth value of the zero disparity surface, the disparity offset is generated according to the maximum disparity and the preset disparity offset model, the pixel points of the modified picture are shifted to the right according to the disparity offset to obtain a left picture, the pixel points of the modified picture are shifted to the left according to the disparity offset to obtain a right picture, the reference offset of the zero disparity surface is generated according to the depth value of the zero disparity surface, the maximum disparity and the preset reference offset model, the clipping boundary width is generated according to the reference offset of the zero disparity surface and the preset boundary width model, the left boundary of the left picture is clipped using the clipping boundary width to obtain a clipped left picture, the right boundary of the right picture is clipped using the clipping boundary width to obtain a clipped right picture, the main content of the clipped left picture and the clipped right picture is aligned using the zero disparity surface to obtain an aligned left picture and an aligned right picture, the aligned left picture and the aligned right picture are filled with holes to obtain a filled left picture and a filled right picture, and the filled left picture and the filled right picture are merged to generate a three-dimensional image with left and right disparities. Since no manual operation is required, the generation time of the three-dimensional image with left and right disparities is reduced, and the generation efficiency of the three-dimensional image with left and right disparities is improved. On the other hand, since the three-dimensional image with left and right disparities is automatically generated, it is not affected by human intervention, and thus the reliability of the three-dimensional image with left and right disparities is improved.
[0159] Please refer to Figure 5 , Figure 5 The structural schematic diagram of the electronic device provided by the embodiments of the present application is shown in FIG. 2.
[0160] As Figure 5 shown, Figure 5 The electronic device 2 comprises at least one processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20. The processor 20 implements the steps in any of the method embodiments described above when executing the computer program 22.
[0161] The electronic device 2 can include, but is not limited to, the processor 20 and the memory 21. Those skilled in the art can understand that Figure 5 The electronic device 2 is only an example and does not constitute a limitation on the electronic device 2, and can include more or fewer components than shown, or combine certain components, or different components, for example, can also include input / output devices, network access devices, etc.
[0162] The processor 20 is configured to run the computer program 22 stored in the memory 21, and implement the following steps when executing the computer program 22:
[0163] obtaining prompt text, performing feature extraction on the prompt text to obtain a text embedding vector, processing the text embedding vector through an image generation model to obtain a picture;
[0164] modifying a current size of the picture to a fixed size to obtain a modified picture, extracting features of different scales from the modified picture, fusing the features of different scales to obtain a fused graph, and performing optimization processing on the fused graph to obtain a depth graph;
[0165] selecting a size of the depth graph as a target size, performing semantic understanding processing and gradient supervision processing on the fused graph to obtain a salient subject image, and dividing the salient subject image into a subject range and a non-subject range;
[0166] adjusting a current size of the subject range to the target size to obtain a modified subject range, performing an erosion operation on the modified subject range by using a structural element to obtain an inner ring range in the modified subject range, and extracting a plurality of depth values from the depth graph according to the inner ring range;
[0167] removing null values or abnormal values from the plurality of depth values to obtain a depth value set, selecting a depth value with the highest frequency of occurrence in the depth value set as a depth value of a zero disparity plane, generating a disparity offset according to a maximum disparity and a preset disparity offset model, shifting pixel points of the modified picture to the right according to the disparity offset to obtain a left image, shifting the pixel points of the modified picture to the left according to the disparity offset to obtain a right image, generating a reference offset of the zero disparity plane according to the depth value of the zero disparity plane, the maximum disparity, and a preset reference offset model, generating a clipping boundary width according to the reference offset of the zero disparity plane and a preset boundary width model, clipping a left boundary of the left image using the clipping boundary width to obtain a clipped left image, clipping a right boundary of the right image using the clipping boundary width to obtain a clipped right image, aligning the subject content of the clipped left image and the clipped right image using the zero disparity plane to obtain an aligned left image and an aligned right image, and filling holes in the aligned left image and the aligned right image to obtain a filled left image and a filled right image, and merging the filled left image and the filled right image to generate a three-dimensional image with left and right disparities.
[0168] In some embodiments, the processor 20 is configured to implement:
[0169] obtaining prompt text input by a user from a text box, performing feature extraction on the prompt text to obtain a text embedding vector, inputting the text embedding vector into an image generation model, processing the text embedding vector through the image generation model to obtain a picture.
[0170] In some embodiments, the processor 20 is configured to implement:
[0171] obtaining a user voice, converting the user voice into prompt text, performing feature extraction on the prompt text to obtain a text embedding vector, inputting the text embedding vector into an image generation model, processing the text embedding vector through the image generation model to obtain a picture.
[0172] In some embodiments, the processor 20 is configured to implement:
[0173] loading the picture through a loading function, obtaining a current size of the picture, modifying the current size of the picture into a fixed size to obtain a modified picture;
[0174] extracting features of different scales from the modified picture, fusing the features of different scales to obtain a fused picture, inputting the fused picture into a depth estimation model, and optimizing the depth estimation model through a gradient matching loss to obtain a depth map output by the depth estimation model, the gradient matching loss being a loss function used to measure the difference between the output result of the depth estimation model and a target result in gradient features.
[0175] In some embodiments, the processor 20 is configured to implement:
[0176] selecting a size of the depth map as a target size, performing semantic understanding processing and gradient supervision processing on the fused picture to obtain a salient subject image, marking a part where a target is located in the salient subject image as a subject range, and marking a part other than the target in the salient subject image as a non-subject range.
[0177] In some embodiments, the processor 20 is configured to implement:
[0178] obtaining a display instruction, executing the display instruction, and displaying a three-dimensional image with left and right parallax on a screen.
[0179] In some embodiments, the processor 20 is configured to implement:
[0180] The structural element includes one of a square structural element, a cross structural element, a circular structural element, an oval structural element, a diamond structural element, a star structural element, or a combination thereof.
[0181] The processor 20 can be a central processing unit (CPU). The processor 20 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits, field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0182] The memory 21 can be an internal storage unit of the electronic device 2 in some embodiments, such as a hard disk or a memory of the electronic device 2. The memory 21 can also be an external storage device of the electronic device 2 in other embodiments, such as a plug-in hard disk, a smart memory card, a Secure Digital (SD) card, a Flash Card, and the like equipped on the electronic device 2. Further, the memory 21 can include both the internal storage unit and the external storage device of the electronic device 2. The memory 21 is used to store an operating system, an application program, a Boot Loader, data, and other programs, such as program codes of the computer program, and the like. The memory 21 can also be used to temporarily store data that has been output or is to be output.
[0183] It should be noted that the information interaction, execution process, and the like between the above apparatuses / units are based on the same concept as the method embodiments of the present application, and specific functions and technical effects brought by the same can be referred to the method embodiments part, which will not be described herein.
[0184] The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in each of the above method embodiments.
[0185] The program code stored in the computer readable storage medium can be called and executed by the processor to implement the three-dimensional image generation method described in the above method embodiments.
[0186] The computer program product provided in the embodiments of the present application, when the computer program product runs on the electronic device, causes the electronic device to execute the three-dimensional image generation method described above.
[0187] The integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium.
[0188] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.
[0189] The above is only the preferred embodiments of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for generating a 3D image based on prompt text, characterized in that, The three-dimensional image generation method, applied to electronic devices, includes: The prompt text is obtained, its features are extracted to obtain a text embedding vector, and the text embedding vector is processed by an image generation model to obtain an image. The current size of the image is modified to a fixed size to obtain the modified image. Features of different scales are extracted from the modified image, and the features of different scales are fused to obtain a fused image. The fused image is then optimized to obtain a depth map. The depth map size is selected as the target size. Semantic understanding and gradient supervision are performed on the fused image to obtain the salient subject image. The salient subject image is then divided into subject range and non-subject range. Adjust the current size of the main body range to the target size to obtain the modified main body range. Use the structuring element to perform an erosion operation on the modified main body range to obtain the inner circle range in the modified main body range. Extract multiple depth values from the depth map based on the inner circle range. Null or outlier values are removed from multiple depth values to obtain a depth value set. The depth value with the highest frequency in the depth value set is selected as the depth value of the zero disparity surface. A disparity offset is generated based on the maximum disparity and a preset disparity offset model. The pixels of the modified image are shifted to the right based on the disparity offset to obtain the left image. The pixels of the modified image are shifted to the left based on the disparity offset to obtain the right image. A reference offset for the zero disparity surface is generated based on the depth value of the zero disparity surface, the maximum disparity, and the preset reference offset model. The reference offset of the zero disparity surface is then used in conjunction with a preset boundary width. The model generates a cropping boundary width. The left edge of the left image is cropped using the cropping boundary width to obtain the cropped left image. The right edge of the right image is cropped using the same cropping boundary width to obtain the cropped right image. The main content of the cropped left and right images is aligned using a zero-parallax plane to obtain aligned left and right images. Holes are filled in the aligned left and right images to obtain filled left and right images. The filled left and right images are then merged to generate a 3D image with left and right parallax.
2. The three-dimensional image generation method according to claim 1, characterized in that, The process of obtaining the prompt text, extracting features from the prompt text to obtain a text embedding vector, and processing the text embedding vector through an image generation model to obtain an image includes: The system retrieves the user-inputted prompt text from the text box, extracts features from the prompt text to obtain a text embedding vector, inputs the text embedding vector into the image generation model, and processes the text embedding vector through the image generation model to obtain the image.
3. The three-dimensional image generation method according to claim 1, characterized in that, The process of obtaining the prompt text, extracting features from the prompt text to obtain a text embedding vector, and processing the text embedding vector through an image generation model to obtain an image includes: The process involves acquiring user speech, converting it into prompt text, extracting features from the prompt text to obtain a text embedding vector, inputting the text embedding vector into an image generation model, and then processing the text embedding vector through the image generation model to obtain an image.
4. The three-dimensional image generation method according to claim 1, characterized in that, The process involves modifying the current size of the image to a fixed size to obtain a modified image, extracting features of different scales from the modified image, fusing the features of different scales to obtain a fused image, and optimizing the fused image to obtain a depth map, including: The image is loaded using a loading function, its current size is obtained, and then the current size is modified to a fixed size to obtain the modified image. Features at different scales are extracted from the modified image, and these features are fused to obtain a fused image. The fused image is then input into a depth estimation model, which is optimized using gradient matching loss to obtain the depth map output by the depth estimation model. Gradient matching loss is a loss function used to measure the difference in gradient features between the output and target results of the depth estimation model.
5. The three-dimensional image generation method according to claim 1, characterized in that, The selected depth map size is used as the target size. Semantic understanding and gradient supervision processing are performed on the fused image to obtain a salient subject image. This salient subject image is then divided into subject regions and non-subject regions, including: The depth map size is selected as the target size. Semantic understanding and gradient supervision are performed on the fused image to obtain a salient subject image. The part of the salient subject image where the target is located is marked as the subject range, and the part of the salient subject image other than the target is marked as the non-subject range.
6. The three-dimensional image generation method according to claim 1, characterized in that, The disparity offset model adopts either a division offset model or a multiplication offset model; The division offset model is as follows: ; This is the parallax offset. The depth value of a pixel. For maximum parallax, This is the depth offset; Among them, the maximum parallax refers to the maximum pixel offset of the same object in the horizontal direction in stereo vision; The reference offset model is as follows: ; The reference offset for the zero parallax plane; For maximum parallax, The depth value for the zero parallax surface; This is the depth offset; The boundary width model is as follows: ; To trim the boundary width; The multiplication offset model is as follows: ; This is the parallax offset. The depth value of a pixel. For maximum parallax, The depth value for the zero parallax surface.
7. The three-dimensional image generation method according to any one of claims 1 to 6, characterized in that, Structural elements include one or a combination of square structural elements, cross structural elements, circular structural elements, elliptical structural elements, rhombus structural elements, and star structural elements.
8. A three-dimensional image generation apparatus based on prompt text, according to the three-dimensional image generation method of any one of claims 1 to 7, characterized in that, Applied to electronic devices, including: The first acquisition module is used to acquire the prompt text, extract features from the prompt text to obtain the text embedding vector, and process the text embedding vector through an image generation model to obtain the image. The modification module is used to modify the current size of the image to a fixed size to obtain the modified image. Features of different scales are extracted from the modified image, and the features of different scales are fused to obtain a fused image. The fused image is then optimized to obtain a depth map. The second acquisition module is used to select the size of the depth map as the target size, perform semantic understanding processing and gradient supervision processing on the fused map to obtain the salient subject image, and distinguish the salient subject image into the subject range and non-subject range; The extraction module is used to adjust the current size of the main body range to the target size to obtain the modified main body range. The erosion operation is performed on the modified main body range using the structuring element to obtain the inner circle range in the modified main body range. Multiple depth values are extracted from the depth map based on the inner circle range. The generation module removes null or outlier values from multiple depth values to obtain a depth value set. It selects the most frequently occurring depth value from this set as the depth value of the zero-disparity surface. Based on the maximum disparity and a preset disparity offset model, it generates a disparity offset. The pixels of the modified image are then shifted to the right based on this disparity offset, resulting in the left image. Similarly, the pixels of the modified image are shifted to the left based on the disparity offset, resulting in the right image. Finally, based on the depth value of the zero-disparity surface, the maximum disparity, and a preset baseline offset model, it generates a baseline offset for the zero-disparity surface. Finally, it uses the baseline offset of the zero-disparity surface and a preset edge... The model generates a cropping boundary width. The left boundary of the left image is cropped using this width, resulting in a cropped left image. The right boundary of the right image is then cropped using the same width, resulting in a cropped right image. A zero-parallax plane is used to align the main content of the cropped left and right images, resulting in aligned left and right images. Holes are filled in the aligned left and right images, resulting in filled left and right images. Finally, the filled left and right images are merged to generate a 3D image with left and right parallax.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the three-dimensional image generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the three-dimensional image generation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method for regulating display depth of three-dimensional image
CN101282492A
Unsupervised stereo matching method
CN115830094A