Image synthesis processing method, related device and medium

By acquiring the basic images of target objects and non-target objects from multiple angles, layer feature correction and local clarity enhancement, the problem of poor image quality in image viewing angle conversion is solved, and the accuracy and quality of the composite image is improved.

CN120374772APending Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510459141.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, in complex lighting scenarios, the synthetic images generated in image viewing angle conversion are prone to overexposed, distorted or blurred, and cannot effectively identify important feature information inside the image, resulting in poor image quality.

Method used

The basic images of the target object and non-target object are obtained from multiple angles, and layer feature information correction and local image clarity enhancement are performed separately. After generating the completed image, the synthesis process is performed.

Benefits of technology

It improves the accuracy of image viewing angle conversion and the quality of synthetic images, overcomes the problems of missing feature information of target object layer and unclear non-target objects, and generates higher quality synthetic images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374772A_ABST
    Figure CN120374772A_ABST
Patent Text Reader

Abstract

The invention provides an image synthesis processing method, a related device and a medium. The method comprises the following steps: respectively acquiring a plurality of basic images of a target object in a background from a plurality of angles; for each basic image, extracting pixel point information of a target object in the basic image; based on the layer feature information of the basic image, carrying out layer adaptation correction on the pixel point information to obtain corrected pixel point information, and based on the corrected pixel point information, carrying out image reconstruction on the background image of the background to obtain a reconstructed image; extracting a local image of a non-target object from the basic image, enhancing the definition of the local image to obtain an enhanced local image, and obtaining a complemented image corresponding to the basic image according to the enhanced local image and the reconstructed image; and synthesizing the complemented images corresponding to the plurality of basic images to obtain a synthesized image. The accuracy of image view angle conversion can be improved. The method can be applied to various scenes such as big data and artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of big data technology, and in particular, to an image synthesis processing method, a related device, and a medium. Background Technique

[0002] Image perspective conversion technology refers to the technology of mapping an image obtained in one image domain to another image domain. An image domain refers to the space where an image is generated. For example, the real-world space is an image domain, and the virtual synthetic image space is another image domain. In a binocular camera of virtual reality (VR) or augmented reality (AR), a camera is provided at each eye of the binocular camera. The camera captures an image in the real-world space. The images in the real-world space captured by the two cameras are processed by the binocular camera and integrated into a visual image in the virtual synthetic image space. This conversion of the image is the image perspective conversion.

[0003] In related technologies, the perspective conversion is often achieved by relying on a multi-plane image generation algorithm. In this algorithm, the left image and the right image respectively captured by the binocular camera are applied to multiple lighting models to obtain multiple derivative left images and multiple derivative right images under multiple lighting models. Based on the first difference image between the derivative left image and the left image and the second difference image between the derivative right image and the right image under each lighting model, the weights of the left image and the right image are respectively determined. Based on the weights of the left image and the right image, the left image and the right image are subjected to weighted synthesis processing to obtain the synthetic image seen by the user through the binocular camera. However, this method depends to a large extent on a lighting model with fixed parameters. In a complex lighting scene, since the related technology adopts the way of overall image comparison, it is impossible to effectively identify the important internal feature information of the image, which will lead to inaccurate first difference images and second difference images, and further lead to overexposure of the weighted synthetic image, resulting in distortion or blurring. Therefore, how to improve the effect of image perspective conversion and generate a synthetic image with better image quality is still an urgent problem to be solved. Summary of the Invention

[0004] Embodiments of the present disclosure provide an image synthesis processing method, a related device, and a medium, which can improve the quality of the synthetic image generated by image perspective conversion and improve the accuracy of image perspective conversion.

[0005] According to one aspect of the present disclosure, there is provided an image synthesis processing method, the method including:

[0006] Obtaining a plurality of basic images of a target object in a background from a plurality of angles, the basic images including the target object and non-target objects;

[0007] Obtaining a background image of the background;

[0008] For each base image, extract the pixel point information of the target object in the base image;

[0009] Based on the layer feature information of the base image, perform layer adaptation correction on the pixel point information to obtain corrected pixel point information, and based on the corrected pixel point information, perform image reconstruction on the background image to obtain a reconstructed image;

[0010] Extract the local image of the non-target object from the base image, enhance the clarity of the local image to obtain an enhanced local image, and obtain the completed image corresponding to the base image according to the enhanced local image and the reconstructed image;

[0011] Perform synthesis processing on the completed images corresponding to the multiple base images to obtain a synthesized image.

[0012] According to one aspect of the present disclosure, there is provided an image synthesis processing apparatus, the apparatus comprising:

[0013] A first acquisition unit for respectively acquiring a plurality of base images of a target object under a background from multiple angles, the base images including the target object and non-target objects;

[0014] A second acquisition unit for acquiring a background image of the background;

[0015] An extraction unit for, for each base image, extracting the pixel point information of the target object in the base image;

[0016] A correction unit for, based on the layer feature information of the base image, performing layer adaptation correction on the pixel point information to obtain corrected pixel point information, and based on the corrected pixel point information, performing image reconstruction on the background image to obtain a reconstructed image;

[0017] A completion unit for extracting the local image of the non-target object from the base image, enhancing the clarity of the local image to obtain an enhanced local image, and obtaining the completed image corresponding to the base image according to the enhanced local image and the reconstructed image;

[0018] A synthesis unit for performing synthesis processing on the completed images corresponding to the multiple base images to obtain a synthesized image.

[0019] Optionally, the synthesis unit includes:

[0020] An acquisition module for acquiring an image mask corresponding to the completed image;

[0021] A determination module, configured to determine an image weight corresponding to the image after completion based on the background image, the image after completion, and an image mask corresponding to the image after completion;

[0022] A first synthesis module, configured to perform weighted synthesis processing on each of the images after completion based on the image weights corresponding to the respective images after completion, to obtain the synthesized image.

[0023] Optionally, the acquisition module is configured to:

[0024] For a single pixel point in a single image after completion among multiple images after completion, determine the respective pixel values of the single pixel point in the multiple images after completion;

[0025] Based on the respective pixel values of the single pixel point in the multiple images after completion, determine a mask of the single pixel point;

[0026] Based on the masks of the respective pixel points in the single image after completion, generate the image mask corresponding to the single image after completion.

[0027] Optionally, the determining a mask of the single pixel point based on the respective pixel values of the single pixel point in the multiple images after completion includes:

[0028] Based on the respective pixel values of the single pixel point in the multiple images after completion, determine a pixel maximum value and a pixel minimum value of the single pixel point;

[0029] Based on the pixel maximum value and the pixel minimum value, determine a pixel range difference of the single pixel point;

[0030] If the pixel range difference is greater than or equal to a preset threshold, determine the mask of the single pixel point as a first value;

[0031] If the pixel range difference is less than the preset threshold, determine the mask of the single pixel point as a second value.

[0032] Optionally, the first synthesis module is configured to:

[0033] For a single pixel point in a single image after completion among multiple images after completion, determine the respective brightness values of the single pixel point in the multiple images after completion;

[0034] Based on the image weights corresponding to the respective images after completion, perform weighted synthesis processing on the respective brightness values of the single pixel point in the multiple images after completion, to obtain a target brightness value of the single pixel point;

[0035] Based on the target brightness values of the respective pixel points, determine the synthesized image.

[0036] Optionally, the determining module includes:

[0037] A mask sub-module, configured to perform mask processing on the complemented image based on the image mask to obtain a masked image;

[0038] A fusion sub-module, configured to perform image fusion on the masked image and the background image to obtain a target fusion image, and compress the target fusion image into a compressed image;

[0039] A first determination sub-module, configured to, for each pixel point in the complemented image, determine a first difference based on the brightness value of the pixel point in the complemented image and the brightness value of the pixel point in the compressed image, and determine the brightness difference corresponding to the complemented image based on the absolute value of the first difference of each pixel point;

[0040] A second determination sub-module, configured to determine the image weight corresponding to the complemented image based on the brightness difference corresponding to the complemented image.

[0041] Optionally, the mask sub-module is configured to:

[0042] For each pixel point in the complemented image, determine the mask value of the pixel point in the image mask, and determine the pixel value of the pixel point in the complemented image;

[0043] Obtain the masked pixel value of the pixel value based on the product of the mask value and the pixel value;

[0044] Determine the masked image based on the masked pixel values of each pixel point.

[0045] Optionally, the fusion sub-module is configured to:

[0046] Generate pyramid features for the masked image and the background image respectively to obtain a plurality of first pyramid features corresponding to the masked image and a plurality of second pyramid features corresponding to the background image;

[0047] Perform feature fusion on the first pyramid features and the second pyramid features at the same scale to obtain pyramid fusion features corresponding to each scale;

[0048] Determine the target fusion image based on the pyramid fusion features of each scale.

[0049] Optionally, the second determination sub-module is configured to:

[0050] For the complemented image, take the reciprocal of the brightness difference to obtain a preliminary weight;

[0051] Sum the preliminary weights of the multiple completed images to obtain a weight sum;

[0052] For the completed image, determine the image weight corresponding to the completed image based on the preliminary weight and the weight sum.

[0053] Optionally, the completion unit is configured to:

[0054] Determine the image sharpness of the local image;

[0055] If the image sharpness meets a first preset condition, use the local image as the enhanced local image;

[0056] If the image sharpness does not meet the first preset condition, perform denoising processing on the local image to obtain a denoised image, and enhance the sharpness of the denoised image to obtain an enhanced local image whose image sharpness meets the first preset condition.

[0057] Optionally, the completion unit is configured to:

[0058] Based on the image size and image position information of the enhanced local image, determine a region to be completed in the reconstructed image;

[0059] Perform encoding processing on the region to be completed to obtain an image encoding region;

[0060] Supplement the enhanced local image to the image encoding region so that the center of the enhanced local image coincides with the center of the image encoding region to obtain the completed image corresponding to the base image.

[0061] Optionally, the pixel point information includes the pixel values of each target pixel point corresponding to the target object in the base image;

[0062] The correction unit is configured to:

[0063] For the pixel value of each target pixel point, determine the channel pixel value of the target pixel point in each color channel;

[0064] For each color channel, based on the channel pixel values of each target pixel point in the color channel, obtain the pixel mean value corresponding to the color channel;

[0065] Extract a layer reference value from the layer feature information, and based on the layer reference value and the pixel mean value, determine the correction coefficient corresponding to the color channel;

[0066] Based on the correction coefficient, adaptively correct the channel pixel values of each target pixel point in the color channel to obtain the corrected channel pixel values of each target pixel point;

[0067] For each target pixel point, merge the corrected channel pixel values of the target pixel point in each color channel to obtain the corrected pixel value;

[0068] Based on the corrected pixel values of each target pixel point, obtain the corrected pixel point information.

[0069] Optionally, the corrected pixel point information includes the pixel point labels, pixel point coordinates, and corrected pixel values of each target pixel point corresponding to the target object;

[0070] The correction unit is configured to:

[0071] For each target pixel point, based on the pixel point coordinates of the target pixel point, determine a mapped pixel point corresponding to the target pixel point in the background image, and the coordinates of the mapped pixel point are the same as the pixel point coordinates;

[0072] If the pixel point label of the target pixel point is the first label value, use the corrected pixel value as the updated pixel value of the mapped pixel point;

[0073] If the pixel point label of the target pixel point is the second label value, obtain the initial pixel value of the mapped pixel point in the background image, perform weighted merging on the initial pixel value and the pixel value of the target pixel point in the base image to obtain the merged pixel value, and use the merged pixel value as the updated pixel value of the mapped pixel point;

[0074] Based on the updated pixel values of each mapped pixel point and the initial pixel values of the non-mapped pixel points in the background image, determine the reconstructed image.

[0075] Optionally, the correction unit is configured to:

[0076] Based on the layer feature information of the base image, determine the pixel points to be corrected among the target pixel points;

[0077] Use the preset reference value as the corrected pixel value of the pixel points to be corrected, and update the pixel point information based on the corrected pixel values of each pixel point to be corrected to obtain the corrected pixel point information.

[0078] Optionally, the corrected pixel point information includes the coordinate parameters and pixel values of each target pixel point corresponding to the target object in the base image, and the corrected pixel values of the pixel points to be corrected among the target pixel points;

[0079] The calibration unit is configured to:

[0080] For the coordinate parameters of the target pixel point, identify background pixel points with the coordinate parameters in the background image, and update the pixel values of the background pixel points with the coordinate parameters based on the pixel value of the target pixel point to obtain an updated background image;

[0081] Locate pixel points with the same coordinate parameters as the pixel points to be calibrated in the updated background image, and replace the pixel values of the located pixel points with the calibration pixel values of the pixel points to be calibrated to obtain the reconstructed image.

[0082] Optionally, the synthesis unit includes:

[0083] A detection module, configured to perform quality detection on the completed image corresponding to each basic image to obtain a quality detection result;

[0084] A second synthesis module, configured to perform synthesis processing on the completed images corresponding to the multiple basic images to obtain the synthesized image after determining that the quality detection results of the completed images corresponding to each basic image meet preset conditions.

[0085] Optionally, the detection module is configured to:

[0086] For each basic image, determine a first pixel mean value and a first pixel standard deviation based on the pixel values of each pixel point in the basic image;

[0087] For the completed image corresponding to the basic image, determine a second pixel mean value and a second pixel standard deviation based on the pixel values of each pixel point in the completed image;

[0088] Determine the pixel covariance between the basic image and the completed image;

[0089] Determine a first value based on the sum of the squares of the first pixel mean value and the second pixel mean value, and the sum of the squares of the first pixel standard deviation and the second pixel standard deviation;

[0090] Determine a second value based on the product of the pixel covariance, the first pixel mean value, and the second pixel mean value;

[0091] Determine the quality detection result of the completed image based on the quotient of the second value and the first value.

[0092] Optionally, the detection module is configured to:

[0093] For each base image, perform feature extraction on the base image to obtain a first image feature, and perform feature extraction on the completed image corresponding to the base image to obtain a second image feature;

[0094] Based on the feature difference between the first image feature and the second image feature, determine the quality detection result of the completed image.

[0095] Optionally, the detection module is configured to:

[0096] For each pixel point in the completed image corresponding to each base image, determine the difference between the pixel value of the pixel point in the base image and the pixel value of the pixel point in the completed image;

[0097] Based on the differences of each pixel point, determine the mean square error of the pixel points;

[0098] Based on the pixel values of each pixel point in the base image, determine the maximum pixel value corresponding to the base image;

[0099] Based on the mean square error of the pixel points and the maximum pixel value, determine the quality detection result of the completed image.

[0100] According to one aspect of the present disclosure, there is provided an electronic device including a memory and a processor, the memory storing a computer program, and the processor implementing the image synthesis processing method as described above when executing the computer program.

[0101] According to one aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program, and the computer program implementing the image synthesis processing method as described above when executed by a processor.

[0102] According to one aspect of the present disclosure, there is provided a computer program product including a computer program, the computer program being read and executed by a processor of an electronic device to cause the electronic device to execute the image synthesis processing method as described above.

[0103] In the embodiments of the present disclosure, multiple basic images of a target object in a background are obtained from multiple perspectives. During the process of generating a synthetic image based on the basic images, the target object (usually a person) in the basic images is likely to lose the original layer feature information during reconstruction, deteriorating the quality of the synthetic image. Moreover, the non-target object (usually an object) is likely to become unclear during the reconstruction process, also deteriorating the quality of the synthetic image. Therefore, instead of comparing the overall image, the target object and the non-target object are separately extracted in the embodiments of the present disclosure, and different processing methods are adopted, and then synthesis is performed finally. Layer adaptation correction is performed on the pixel point information of the target object in the basic images. The clarity of the local image of the non-target object is enhanced. In this way, the problem that the target object is likely to lose the original layer feature information during reconstruction is overcome, and the problem that the non-target object is likely to become unclear during the reconstruction process is also overcome, improving the quality of the generated synthetic image and the accuracy of image perspective conversion.

[0104] Other features and advantages of the present disclosure will be described in the following description, and part of them will become obvious from the description, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained through the structures specifically pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the description. They are used together with the embodiments of the present disclosure to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.

[0106] Figure 1 is a system architecture diagram of the application of the image synthesis processing method according to the embodiments of the present disclosure;

[0107] Figures 2A - 2E shows a schematic diagram of the application of the image synthesis processing method according to the embodiments of the present disclosure in an image synthesis scenario;

[0108] Figure 3 is a flowchart of an image synthesis processing method according to an embodiment of the present disclosure;

[0109] Figure 4 is a flowchart of layer adaptation correction for pixel point information according to an embodiment of the present disclosure;

[0110] Figure 5 is a flowchart of image reconstruction for a background image according to an embodiment of the present disclosure;

[0111] Figure 6 is a schematic diagram of the implementation process of generating a reconstructed image according to an embodiment of the present disclosure;

[0112] Figure 7 It is a flowchart for layer adaptation correction of pixel point information in another embodiment of the present disclosure;

[0113] Figure 8 It is a flowchart for image reconstruction of the background image in another embodiment of the present disclosure;

[0114] Figure 9 It is a schematic diagram of the implementation process for generating a reconstructed image in another embodiment of the present disclosure;

[0115] Figure 10 It is a flowchart for determining the completed image corresponding to the base image in one embodiment of the present disclosure;

[0116] Figures 11A - 11B It is a schematic diagram of the implementation process for determining the completed image corresponding to the base image according to one embodiment of the present disclosure;

[0117] Figure 12 It is a flowchart for performing synthesis processing on the completed image according to one embodiment of the present disclosure;

[0118] Figure 13 It is a flowchart for generating an image mask according to one embodiment of the present disclosure;

[0119] Figure 14 It is a schematic diagram of the implementation process for generating an image mask according to one embodiment of the present disclosure;

[0120] Figure 15 It is a flowchart for determining an image weight according to one embodiment of the present disclosure;

[0121] Figure 16 It is a flowchart for weighted synthesis of the completed image according to one embodiment of the present disclosure;

[0122] Figure 17 It is a flowchart for performing synthesis processing on the completed image according to another embodiment of the present disclosure;

[0123] Figure 18 It is a schematic diagram of the implementation details of an image synthesis method according to one embodiment of the present disclosure;

[0124] Figure 19 It is a module diagram of an image synthesis processing device according to one embodiment of the present disclosure;

[0125] Figure 20 It is a terminal structure diagram of an image synthesis processing method according to one embodiment of the present disclosure;

[0126] Figure 21 It is a server structure diagram of an image synthesis processing method according to one embodiment of the present disclosure. Detailed Implementation Manner

[0127] In order to make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.

[0128] The system architecture and scenarios applied in the embodiments of the present disclosure will be described below.

[0129] Figure 1 It is a system architecture diagram to which the image synthesis processing method according to the embodiments of the present disclosure is applied. It includes an object terminal 140, the Internet 130, a gateway 120, a server 110, an image database 150, etc.

[0130] The object terminal 140 includes various forms such as a desktop computer, a laptop computer, a PDA (Personal Digital Assistant), a tablet computer, a mobile phone, a vehicle-mounted terminal, a home theater terminal, a smart TV, a dedicated terminal, etc. In addition, it can be a single device or a collection composed of multiple devices. The object terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data. The object terminal 140 includes an image processing platform, and the image processing platform is used to synthesize the object images at multiple perspectives selected by the object into an image that conforms to the human eye visual effect.

[0131] The server 110 refers to a computer system that can provide certain services to the object terminal 140. Compared with an ordinary object terminal 140, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, or a cloud server, etc. The server 110 includes various types of services. Among them, the implementation of each service of the server 110 is often associated with some intermediate databases or storage media, etc. The server 110 is used to perform layer adaptation and correction on the image part containing the target object, and enhance the clarity of the image part that does not contain the target object, based on the object images at multiple perspectives selected by the object and the background images extracted from the image database, so as to obtain the complementary images corresponding to each object image, and synthesize the complementary images into an image that conforms to the human eye visual effect. The image database 150 is used to store various images such as the object images at multiple perspectives selected by the object and the background images. Among them, the image database 150 can be set separately or integrated on the server 110 or other electronic devices.

[0132] The gateway 120 is also known as an internetwork connector or protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, or even with completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. Messages sent by the object terminal 140 to the server 110 need to be sent to the corresponding server through the gateway 120. Messages sent by the server 110 to the object terminal 140 also need to be sent to the corresponding object terminal 140 through the gateway 120.

[0133] The embodiments of the present disclosure can be applied in various scenarios, such as Figures 2A - 2E the image synthesis scenario shown, etc.

[0134] As Figure 2A shown, when an object wants to synthesize object images from multiple perspectives into an image that conforms to human eye vision, the object will log in to the image processing platform on the object terminal and enter the image synthesis process. At this time, a prompt field "Please provide the base images from multiple perspectives and the background image:" will be displayed on the page, and there are editing areas for uploading the base images and for selecting the background image. Based on this, the object uploads "left view, right view" in the editing area for uploading the base images; the object uploads "background image 2.jpg" in the editing area for selecting the background image and clicks the "OK" button to determine to generate an image that conforms to human eye vision based on the left view and right view collected by the binocular camera and the background image.

[0135] As Figure 2B shown, after clicking the "OK" button, a prompt field "The following are the pixel point information of the target object in each base image and the partial images of non-target objects:" will be displayed on the page. Specifically, the pixel sequences of the target object in the left view and the right view are composed of the pixel values and pixel coordinates of the pixel points. Among them, the pixel sequence of the target object in the left view is {(150, (1, 2)), (160, (1, 3)), (155, (2, 2))}; the pixel sequence of the target object in the right view is {(153, (1, 2)), (158, (1, 3)), (156, (2, 2))}. In addition, the partial image of the non-target object extracted from the left view is partial image 1.jpg, and the partial image of the non-target object extracted from the right view is partial image 2.jpg.

[0136] As Figure 2CAs shown, after determining the pixel point information and local images corresponding to each base image, a prompt window will be displayed on the page. Among them, the prompt window has a prompt field "The pixel point sequences of each view are being fused into the background image to form a fused image, and each local image is being supplemented into the corresponding fused image to form a completed image. Please wait patiently...", to indicate to the user the generation process of the completed images corresponding to each base image.

[0137] As Figure 2D shown, after generating the completed images corresponding to each perspective, the completed images corresponding to each perspective will be displayed on the page. Specifically, the prompt field "The completed image corresponding to the left view and the completed image corresponding to the right view are as follows:" will be displayed on the page, and the completed image 1 corresponding to the left view and the completed image 2 corresponding to the right view will be shown.

[0138] As Figure 2E shown, after generating the completed image corresponding to the left view and the completed image corresponding to the right view, a prompt window will be displayed on the page. Among them, the prompt window has a prompt field "The synthesis process of multiple completed images is in progress. Please wait patiently...", to indicate to the user the image synthesis process.

[0139] The following provides an overall description of the embodiments of the present disclosure.

[0140] According to an embodiment of the present disclosure, an image synthesis processing method is provided.

[0141] This image synthesis processing method is generally applied to business scenarios that require converting images captured in the real-world space into visual images in the virtual synthesis image space, such as Figures 2A - 2E the image synthesis scenarios shown, etc. The embodiments of the present disclosure provide a solution for layer adaptation correction of the target object and clarity enhancement of non-target objects during perspective conversion, which can improve the quality of the synthesized image generated by image perspective conversion and improve the accuracy of image perspective conversion.

[0142] As Figure 3 shown, according to an embodiment of the present disclosure, the image synthesis processing method can be executed by an electronic device, and the electronic device can be Figure 1 the server or the object terminal shown. According to an embodiment of the present disclosure, the image synthesis processing method can include:

[0143] Step 310: Obtain multiple base images of the target object in the background from multiple angles;

[0144] Step 320: Obtain the background image of the background;

[0145] Step 330: For each base image, extract the pixel point information of the target object in the base image;

[0146] Step 340: Based on the layer feature information of the base image, perform layer adaptation correction on the pixel point information to obtain the corrected pixel point information, and based on the corrected pixel point information, perform image reconstruction on the background image to obtain the reconstructed image;

[0147] Step 350: Extract the local image of the non-target object from the base image, enhance the clarity of the local image to obtain the enhanced local image, and based on the enhanced local image and the reconstructed image, obtain the completed image corresponding to the base image;

[0148] Step 360: Perform synthesis processing on the completed images corresponding to multiple base images to obtain the synthesized image.

[0149] The following gives a detailed description of Steps 310 - 360.

[0150] In Step 310, multiple base images of the target object in the background are obtained from multiple angles respectively.

[0151] The target object refers to the object specified by the user in the base image or the object with the largest pixel area in the base image. For example, the target object is often a person.

[0152] The base image refers to the image captured from different angles by devices such as a monocular camera or a binocular camera.

[0153] Among them, the base image contains the target object and the non-target object. The non-target object is the object other than the target object in the base image. For example, the non-target object is often an object other than a person.

[0154] When this embodiment is specifically implemented, when obtaining the base images from multiple angles respectively, a binocular camera or a multi-camera can be used for real-time shooting to obtain multiple perspective images of the target object in the background, and the multiple perspective images are used as the base images; or multiple perspective images of the target object in the background previously collected and stored can be extracted from the preset image database, and the extracted multiple perspective images are used as the base images.

[0155] For example, the left-eye image and the right-eye image of the target object in the background can be obtained respectively through the binocular cameras set in the VR camera, and the left-eye image and the right-eye image are used as a base image respectively.

[0156] In Step 320, the background image of the background is obtained.

[0157] In the specific implementation of this embodiment, since the background images of the background are often pre-stored in the image database, based on this, with the authorization, the pre-stored background images can be directly extracted from the image database according to the image names or storage paths of the background images of the background.

[0158] In step 330, for each base image, the pixel point information of the target object in the base image is extracted.

[0159] The pixel point information is used to indicate information such as the pixel values and pixel coordinates of the pixel points that make up the target object in the base image.

[0160] In the specific implementation of this embodiment, first, for each base image, the pixel points that make up the target object are determined in the base image; then, for each pixel point that makes up the target object, the pixel value and pixel coordinate of the pixel point are extracted, and the pixel values and pixel coordinates of the multiple pixel points that make up the target object are integrated into the pixel point information of the target object.

[0161] In step 340, based on the layer feature information of the base image, the pixel point information is corrected for layer adaptation to obtain the corrected pixel point information, and based on the corrected pixel point information, an image reconstruction is performed on the background image to obtain a reconstructed image.

[0162] The layer feature information is used to indicate the features in different layers or regions of the base image, where the features include but are not limited to colors, textures, shapes, or edges, etc. in different layers or regions of the base image.

[0163] The corrected pixel point information refers to the information obtained by correcting the pixel values of individual pixel points in the pixel point information.

[0164] The reconstructed image refers to the image formed by fusing the pixel point information into the background image according to the layer feature information of the base image.

[0165] To save space, the specific process of correcting the pixel point information for layer adaptation based on the layer feature information of the base image and the specific process of performing image reconstruction on the background image based on the corrected pixel point information in the embodiments of the present disclosure will be described in detail below and will not be elaborated here.

[0166] In step 350, a partial image of the non-target object is extracted from the base image, the clarity of the partial image is enhanced to obtain an enhanced partial image, and based on the enhanced partial image and the reconstructed image, a completed image corresponding to the base image is obtained.

[0167] The partial image refers to the image part in the base image that represents the non-target object.

[0168] The enhanced local image refers to the image obtained by enhancing the clarity of the local image of the non-target object.

[0169] The completed image refers to the image formed by supplementing the image information of the enhanced local image representing the non-target object into the reconstructed image.

[0170] To save space, the specific process of enhancing the clarity of the local image in the embodiments of the present disclosure, and the specific process of obtaining the completed image corresponding to the base image according to the enhanced local image and the reconstructed image will be described in detail below, and will not be elaborated here.

[0171] In step 360, the completed images corresponding to multiple base images are synthesized to obtain a synthesized image.

[0172] The synthesized image refers to a visual image that conforms to the human eye visual effect based on the base images from multiple perspectives. Among them, the synthesized image and the base image belong to different image domains. For example, the base image is an image of the real-world space, while the synthesized image is a visual image of the virtual synthesized image space.

[0173] To save space, the specific process of synthesizing the completed images corresponding to multiple base images in the embodiments of the present disclosure will be described in detail below, and will not be elaborated here.

[0174] Through the above steps 310-360, the embodiments of the present disclosure respectively obtain multiple base images of the target object in the background from multiple angles. In the process of generating a synthesized image based on the base images, the target object (usually a person) in the base image is likely to lose the original layer feature information during reconstruction, deteriorating the quality of the synthesized image, and the non-target object (usually an object) is likely to become unclear during the reconstruction process, deteriorating the quality of the synthesized image. Therefore, instead of comparing the overall image, the embodiments of the present disclosure separately extract the target object and the non-target object, adopt different processing methods, and then synthesize them finally. The layer adaptation correction is performed on the pixel point information of the target object in the base image. The clarity of the local image of the non-target object is enhanced. In this way, both the problem that the target object is likely to lose the original layer feature information during reconstruction and the problem that the non-target object is likely to become unclear during the reconstruction process are overcome, the quality of the generated synthesized image is improved, and the accuracy of image perspective conversion is improved.

[0175] The above is the overall description of steps 310-360. Since steps 310, 320, and 330 have been described in detail in the above overall description, the specific implementations of steps 340, 350, and 360 will be described in detail below.

[0176] The following provides a detailed description of step 340.

[0177] In step 340, based on the layer feature information of the base image, layer adaptation correction is performed on the pixel point information to obtain corrected pixel point information. Then, based on the corrected pixel point information, image reconstruction is performed on the background image to obtain a reconstructed image.

[0178] In the embodiments of the present disclosure, the pixel point information includes the pixel values of each target pixel point corresponding to the target object in the base image.

[0179] Please refer to Figure 4 , in one embodiment, the specific process of performing layer adaptation correction on the pixel point information in step 340 may include, but is not limited to, the following steps 410-460:

[0180] Step 410: For the pixel value of each target pixel point, determine the channel pixel value of the target pixel point in each color channel;

[0181] Step 420: For each color channel, based on the channel pixel values of each target pixel point in the color channel, obtain the pixel mean value corresponding to the color channel;

[0182] Step 430: Extract the layer reference value from the layer feature information, and based on the layer reference value and the pixel mean value, determine the correction coefficient corresponding to the color channel;

[0183] Step 440: Based on the correction coefficient, perform adaptation correction on the channel pixel values of each target pixel point in the color channel to obtain the corrected channel pixel values of each target pixel point;

[0184] Step 450: For each target pixel point, merge the corrected channel pixel values of the target pixel point in each color channel to obtain the corrected pixel value;

[0185] Step 460: Based on the corrected pixel values of each target pixel point, obtain the corrected pixel point information.

[0186] The following provides a detailed description of steps 410-460.

[0187] In step 410, the color channel refers to the RGB color channel. The channel pixel value is used to indicate the pixel value of the target pixel point in a single color channel.

[0188] In the specific implementation of this embodiment, for the pixel value of each target pixel point, according to the pixel value of the target pixel point in the pixel point information, the channel pixel values of the target pixel point in each color channel are extracted, and the channel pixel value R of the target pixel point in the red channel, the channel pixel value G of the target pixel point in the green channel, and the channel pixel value B of the target pixel point in the blue channel are obtained respectively.

[0189] In step 420, the pixel mean is used to indicate the average pixel value of multiple target pixel points in a certain color channel.

[0190] In the specific implementation of this embodiment, for each color channel, the channel pixel values of multiple target pixel points in this color channel are averaged to obtain the pixel mean corresponding to this color channel.

[0191] In step 430, the layer reference value is used to indicate the color reference value of the base image in the RBG layer. The correction coefficient is used to indicate the degree of correction of the pixel values of the target pixel points in each color channel.

[0192] In the specific implementation of this embodiment, first, the color reference value of the base image in the RBG layer is extracted from the layer feature information to obtain the layer reference value. Then, for each color channel, the layer reference value is divided by the pixel mean of this color channel to obtain the correction coefficient corresponding to this color channel.

[0193] In step 440, the corrected channel pixel value is used to indicate the pixel size obtained after correction of the target pixel point in a single color channel.

[0194] In the specific implementation of this embodiment, when adaptively correcting the channel pixel values of each target pixel point in the color channel, for each color channel, the channel pixel value of each target pixel point in this color channel is multiplied by the correction coefficient corresponding to this color channel to obtain the corrected channel pixel value of the target pixel point in this color channel.

[0195] In step 450, the corrected pixel value is used to indicate the overall pixel size obtained after correction of the target pixel point.

[0196] In the specific implementation of this embodiment, for each target pixel point, the corrected channel pixel values of the target pixel point in each color channel are combined by using pixel weighting or pixel averaging to obtain the corrected pixel value of the target pixel point.

[0197] In step 460, the corrected pixel values of each target pixel point and the pixel point coordinates are integrated to obtain the corrected pixel point information.

[0198] The advantage of this embodiment is that, considering the problem that the target object (usually a person) in the base image is likely to lose the original layer feature information during reconstruction, which may deteriorate the quality of the synthesized image. By calculating the pixel mean of each color channel and determining the correction coefficient based on the layer reference value and the pixel mean, the respective target pixel points of the base image can be adaptively corrected in each color channel, effectively reducing the color deviation in image reconstruction (such as inaccurate white balance, color distortion, etc.). This enables the reconstructed image obtained by reconstructing the image based on the corrected pixel values of the target object to have a more natural and consistent color representation, thereby improving the visual quality of the generated synthesized image.

[0199] In this embodiment, the corrected pixel point information includes the pixel point label, pixel point coordinates, and corrected pixel value of each target pixel point corresponding to the target object.

[0200] The pixel point label is used to indicate the specific method of pixel fusion for the target pixel point during reconstruction.

[0201] The pixel point coordinates are used to indicate the specific position of the target pixel point in the base image and the background image.

[0202] The corrected pixel value is used to indicate the overall pixel size obtained by correcting the target pixel point.

[0203] Please refer to Figure 5 , in this embodiment, the specific process of image reconstruction of the background image in step 340 may include but is not limited to the following steps 510 - 540:

[0204] Step 510: For each target pixel point, based on the pixel point coordinates of the target pixel point, determine the mapped pixel point corresponding to the target pixel point in the background image;

[0205] Step 520: If the pixel point label of the target pixel point is the first label value, use the corrected pixel value as the updated pixel value of the mapped pixel point;

[0206] Step 530: If the pixel point label of the target pixel point is the second label value, obtain the initial pixel value of the mapped pixel point in the background image, perform weighted merging on the initial pixel value and the pixel value of the target pixel point in the base image to obtain the merged pixel value, and use the merged pixel value as the updated pixel value of the mapped pixel point;

[0207] Step 540: Based on the updated pixel values of each mapped pixel point and the initial pixel values of the non - mapped pixel points in the background image, determine the reconstructed image.

[0208] The following will describe steps 510 - 540 in detail.

[0209] In step 510, the coordinates of the mapped pixel points are the same as those of the pixel points.

[0210] In the specific implementation of this embodiment, since the image sizes of the background image and the base image are the same. Based on this, first, for each target pixel point, according to the pixel point coordinates of the target pixel point, find the pixel point with the same pixel point coordinates in the background image, and use this pixel point on the background image as the mapped pixel point corresponding to the target pixel point.

[0211] In step 520, the first tag value is used to indicate that the specific way of pixel fusion for the target pixel point in the reconstruction is to retain the complete pixel value of the target pixel point in the reconstruction. The updated pixel value is used to indicate the pixel value of the mapped pixel point in the reconstructed image (after image reconstruction).

[0212] In the specific implementation of this embodiment, if the pixel point tag of the target pixel point is the first tag value, it indicates that the complete pixel value of the target pixel point needs to be retained in the reconstruction. Therefore, the corrected pixel value is used as the updated pixel value of the mapped pixel point.

[0213] In step 530, the second tag value is used to indicate that the specific way of pixel fusion for the target pixel point in the reconstruction is to perform weighted fusion on the pixel values of the target pixel point and the mapped pixel point. The initial pixel value refers to the pixel value of the mapped pixel point in the background image. The merged pixel value is used to indicate the pixel fusion result of the target pixel point and the mapped pixel point.

[0214] In the specific implementation of this embodiment, if the pixel point tag of the target pixel point is the second tag value, it indicates that the pixel values of the target pixel point and the mapped pixel point are weighted and fused in the reconstruction. Based on this, first, obtain the initial pixel value of the mapped pixel point in the background image. Then, perform weighted merging on the initial pixel value and the pixel value of the target pixel point in the base image according to the preset weight ratio to obtain the merged pixel value. Further, use the merged pixel value as the updated pixel value of the mapped pixel point.

[0215] In step 540, in the background image, replace the initial pixel value with the updated pixel value of each mapped pixel point, and retain the initial pixel value of each non-mapped pixel point in the background image in the background image to obtain the reconstructed image.

[0216] As Figure 6As shown, it is a specific example of image reconstruction based on pixel point information. Specifically, in a base image with an image size of 5×5, there are 13 target pixel points, which are the target pixel point with coordinates (1,4) and pixel value 198, the target pixel point with coordinates (1,5) and pixel value 188, the target pixel point with coordinates (2,3) and pixel value 193, the target pixel point with coordinates (2,4) and pixel value 173, the target pixel point with coordinates (2,5) and pixel value 188, the target pixel point with coordinates (3,3) and pixel value 203, the target pixel point with coordinates (3,4) and pixel value 213, the target pixel point with coordinates (3,5) and pixel value 178, the target pixel point with coordinates (4,3) and pixel value 183, the target pixel point with coordinates (4,4) and pixel value 198, the target pixel point with coordinates (4,5) and pixel value 208, the target pixel point with coordinates (5,4) and pixel value 183, and the target pixel point with coordinates (5,5) and pixel value 193. When the extracted layer reference value is 160, the corrected pixel values of the above target pixel points are 164, 155, 160, 143, 155, 168, 176, 147, 151, 164, 172, 151, 160 in sequence. Based on this, when the pixel values of the corresponding mapped pixel points of the above respective target pixel points in the background image are all 200, and the pixel value labels of the target pixel points with coordinates (1,4), (1,5), (5,4), and (5,5) are the second label value, and the preset weight ratio is 0.5, 0.5, the pixel values of the pixel points with coordinates (1,4), (1,5), (5,4), and (5,5) in the reconstructed image are 199, 194, 191, 196 in sequence. When the pixel value labels of the target pixel values with coordinates (3,3), (3,4), (3,5), (4,3), (4,4), (4,5), (5,3), (5,4), and (5,5) are the first label value, the pixel values of the pixel points with coordinates (3,3), (3,4), (3,5), (4,3), (4,4), (4,5), (5,3), (5,4), and (5,5) in the reconstructed image are 160, 143, 155, 168, 176, 147, 151, 164, 172 in sequence.

[0217] The advantage of this embodiment is that, according to the pixel labels (the first label value or the second label value) of the target pixel points, different processing methods are applied to the pixel points in the background image. The pixel points with the first label value are directly updated using the corrected pixel values, which are suitable for those pixel points that need to be completely replaced or corrected; the pixel points with the second label value are updated by weighted merging, which is suitable for the pixel points that need to retain some of the original information. This method can flexibly handle different types of pixel points and improve the processing accuracy of image reconstruction. In addition, for the pixel points with the second label value, through the weighted merging method, the original details in the background image can be retained, and at the same time, combined with the information of the target pixel points, it can effectively avoid the artifacts and unnatural transition effects that may be brought about by direct replacement, thereby generating a more natural and higher-quality reconstructed image.

[0218] Please refer to Figure 7 , in one embodiment, the specific process of performing layer adaptation correction on the pixel point information in step 340 may include but is not limited to the following steps 710-720:

[0219] Step 710: Based on the layer feature information of the base image, determine the pixel points to be corrected among the target pixel points;

[0220] Step 720: Use the preset reference value as the corrected pixel value of the pixel points to be corrected, and update the pixel point information based on the corrected pixel values of each pixel point to be corrected to obtain the corrected pixel point information.

[0221] The following will describe steps 710-720 in detail.

[0222] In step 710, the pixel points to be corrected refer to some of the target pixel points that need to have their pixel values corrected among the multiple target pixel points.

[0223] In the specific implementation of this embodiment, since the layer feature information of the base image often indicates the image areas that need color correction, based on this, according to the image areas indicated by the layer feature information in the base image, and according to the pixel coordinates or coordinate parameters of each target pixel point, find the target pixel points located in this image area, and determine the found target pixel points as the pixel points to be corrected.

[0224] In step 720, the preset reference value refers to a fixed pixel value preset for the pixel points to be corrected.

[0225] In the specific implementation of this embodiment, first, use the preset reference value as the corrected pixel value of the pixel points to be corrected. Then, update the pixel value of the pixel points to be corrected in the pixel point information to the corrected pixel value to implement the update of the pixel point information and obtain the corrected pixel point information.

[0226] The advantage of this embodiment is that, according to the image area to be corrected indicated by the layer feature information, the pixels to be corrected are identified, and the preset reference value is directly used as the corrected pixel value, which can effectively avoid complex calculation and analysis processes, greatly simplify the correction process, and is conducive to improving the correction efficiency. At the same time, since the pixels to be corrected are all corrected for layer adaptation according to the preset reference value, the overall visual effect of the generated composite image can be effectively improved, making the composite image look more natural and consistent visually, which is conducive to improving the quality of the generated composite image.

[0227] In this embodiment, the corrected pixel point information includes the coordinate parameters and pixel values of each target pixel corresponding to the target object in the base image, and the corrected pixel values of the pixels to be corrected among the target pixels.

[0228] The coordinate parameters are used to indicate the specific position of the target pixel in the base image.

[0229] The pixel value is used to indicate the pixel size of the target pixel in the base image.

[0230] The corrected pixel value is used to indicate the pixel size of the pixel to be corrected.

[0231] Please refer to Figure 8 , in this embodiment, the specific process of image reconstruction on the background image in step 340 may include but is not limited to the following steps 810 - 820:

[0232] Step 810: For the coordinate parameters of the target pixels, identify the background pixels with the coordinate parameters in the background image, and update the pixel values of the background pixels with the coordinate parameters based on the pixel values of the target pixels to obtain the updated background image;

[0233] Step 820: Locate the pixels with the same coordinate parameters as the pixels to be corrected in the updated background image, and replace the pixel values of the located pixels with the corrected pixel values of the pixels to be corrected to obtain the reconstructed image.

[0234] The following is a detailed description of steps 810 - 820.

[0235] In step 810, the updated background image is used to indicate the image formed by fusing the pixel values of all target pixels in the base image into the background image.

[0236] In the specific implementation of this embodiment, the specific process of step 810 is similar to the above-mentioned step 530. For each background pixel point, the pixel value of the background pixel point in the background image and the pixel value of the target pixel point with the same coordinate parameters in the base image are weighted and combined to update the pixel value of the background pixel point, and the updated background image is obtained. For the sake of brevity, it will not be elaborated here.

[0237] In step 820, according to the coordinate parameters of the pixel point to be corrected, the pixel point with the same coordinate parameters as the pixel point to be corrected is located in the updated background image, and the pixel value of the located pixel point in the updated background image is replaced with the corrected pixel value of the pixel point to be corrected to obtain the reconstructed image.

[0238] As Figure 9 shown, it is a specific example of image reconstruction according to pixel point information. Specifically, in the base image with an image size of 5×5, there are 13 target pixel points, which are the target pixel point with coordinates (1,4) and pixel value 198, the target pixel point with coordinates (1,5) and pixel value 188, the target pixel point with coordinates (2,3) and pixel value 193, the target pixel point with coordinates (2,4) and pixel value 173, the target pixel point with coordinates (2,5) and pixel value 188, the target pixel point with coordinates (3,3) and pixel value 203, the target pixel point with coordinates (3,4) and pixel value 213, the target pixel point with coordinates (3,5) and pixel value 178, the target pixel point with coordinates (4,3) and pixel value 183, the target pixel point with coordinates (4,4) and pixel value 198, the target pixel point with coordinates (4,5) and pixel value 208, the target pixel point with coordinates (5,4) and pixel value 183, and the target pixel point with coordinates (5,5) and pixel value 193. Based on this, when the pixel values of the corresponding mapped pixel points of the above-mentioned target pixel points in the background image are all 200, and the pixel points with coordinates (3,3), (3,4), (3,5), (4,3), (4,4), (4,5), (5,3), (5,4), (5,5) are determined as the pixel points to be corrected, and the corrected pixel values of each pixel point to be corrected are 190, 170, 190, 200, 210, 180, 180, 200, 210 in sequence, the pixel values of the pixel points with coordinates (1,4), (1,5), (5,4), (5,5) in the reconstructed image are 199, 194, 191, 196 in sequence, and the pixel points with coordinates (3,3), (3,4), (3,5), (4,3), (4,4), (4,5), (5,3), (5,4), (5,5) are 190, 170, 190, 200, 210, 180, 180, 200, 210 in sequence.

[0239] The advantage of this embodiment is that for each background pixel, the pixel value of the background pixel in the background image and the pixel value of the target pixel with the same coordinate parameters in the base image are weighted and merged to achieve an overall update of the pixel value of the background pixel. Then, the local pixel value of the updated background image is updated according to the corrected pixel value of the pixel to be corrected, so as to obtain a reconstructed image. This method is a step-by-step image reconstruction. First, the pixel value information and the background image are fused at the pixel level, and then the layer adaptation correction situation is propagated to the fused image, which can improve the accuracy of image fusion and the accuracy of layer adaptation correction.

[0240] The following describes step 350 in detail.

[0241] In step 350, a local image of a non-target object is extracted from the base image, the clarity of the local image is enhanced to obtain an enhanced local image, and a completed image corresponding to the base image is obtained according to the enhanced local image and the reconstructed image.

[0242] In the specific implementation of this embodiment, in the base image, other objects except the target object are determined as non-target objects. Then, according to the contour of the non-target object in the base image, the base image is cropped to obtain a local image corresponding to the non-target object.

[0243] Since the local image corresponding to the non-target object in the base image may have the problem of being blurred, it will cause the non-target object to become unclear during reconstruction, resulting in the deterioration of the quality of the synthesized image. Based on this, the embodiment of the present disclosure provides a scheme for enhancing the clarity of the local image, which can overcome the problem that the non-target object is likely to become unclear during the reconstruction process and improve the quality of the generated synthesized image.

[0244] In one embodiment, the specific process of enhancing the clarity of the local image may include but is not limited to the following steps:

[0245] Determine the image clarity of the local image;

[0246] If the image clarity meets the first preset condition, the local image is used as the enhanced local image;

[0247] If the image clarity does not meet the first preset condition, the local image is denoised to obtain a denoised image, and the clarity of the denoised image is enhanced to obtain an enhanced local image whose image clarity meets the first preset condition.

[0248] Among them, image clarity is an important indicator for measuring the image quality of a local image. Image clarity affects the detail presentation of the local image, the readability of text, and the effect of subsequent image processing. The first preset condition is used to measure whether the image clarity of the local image meets the requirements of image completion. The denoised image is used to indicate the result of removing the noise information in the local image.

[0249] Specifically, first, calculate the power spectral density of the local image to obtain the power spectral density of the local image, where the power spectral density is used to indicate the energy distribution of the local image at different frequencies. Since the power spectral density of an image with high clarity is relatively high in the high-frequency region, the image clarity of the local image can be determined according to the power spectral density of the local image. Then, compare the image clarity with the clarity threshold set in the first preset condition. If the image clarity is greater than or equal to the clarity threshold, it is considered that the image clarity of the local image meets the first preset condition, and the local image can be directly used as the enhanced local image. If the image clarity is less than the clarity threshold, it is considered that the image clarity does not meet the first preset condition, and the clarity of the local image needs to be enhanced. Based on this, first, perform denoising processing on the local image using median filtering or non-local means denoising method to remove the noise information in the local image and obtain the denoised image. Then, perform histogram equalization on the denoised image to enhance the contrast of the denoised image and obtain the equalized image, and use the Laplacian operator to sharpen the equalized image to enhance the edge information of the equalized image and obtain the enhanced local image whose image clarity meets the first preset condition.

[0250] The advantage of this embodiment is that it takes into account the problem that the image clarity of the local image of the non-target object directly extracted from the base image may be poor. When it is determined that the image clarity of the local image does not meet the first preset condition, denoising and clarity enhancement are performed on the local image. On the basis of eliminating the noise information of the local image, both the contrast and the image edge of the denoised image are enhanced, which can comprehensively improve the image clarity and obtain the enhanced local image whose image clarity meets the first preset condition. When it is determined that the image clarity meets the first preset condition, the local image is directly used as the enhanced local image, which can improve the image quality of the enhanced local image, overcome the problem that the non-target object is likely to become unclear during the reconstruction process, and thus improve the quality of the generated composite image.

[0251] Please refer to Figure 10 , in this embodiment, step 350 of obtaining the completed image corresponding to the base image according to the enhanced local image and the reconstructed image specifically includes:

[0252] Supplement the enhanced local image to the reconstructed image to obtain the completed image corresponding to the base image;

[0253] Among them, the specific process of supplementing the enhanced local image to the reconstructed image may include but is not limited to the following steps 1010-1030:

[0254] Step 1010: Based on the image size and image position information of the enhanced local image, determine the area to be filled in the reconstructed image;

[0255] Step 1020: Perform encoding processing on the area to be filled to obtain an image encoding area;

[0256] Step 1030: Supplement the enhanced local image to the image encoding area, so that the image center of the enhanced local image coincides with the center of the image encoding area, and obtain the filled image corresponding to the basic image.

[0257] The following is a detailed description of steps 1010-1030.

[0258] In step 1010, the image size is used to indicate the size of the enhanced local image. The image position information is used to indicate the specific position of the local image corresponding to the enhanced local image in the basic image. The area to be filled is used to indicate the image area in the reconstructed image that needs to fill non-target objects.

[0259] It should be noted that the basic image and the reconstructed image have the same size.

[0260] In the specific implementation of this embodiment, first, based on the image position information of the enhanced local image, locate the specific position in the reconstructed image where non-target objects should be filled. Among them, the specific position of the non-target object in the reconstructed image is defined by the endpoint coordinates in the image position information, where the endpoint coordinates include the upper left endpoint, lower left endpoint, upper right endpoint, and lower right endpoint of the non-target object. Then, according to the image size of the enhanced local image and the endpoint coordinates in the image position information, determine the smallest circumscribed rectangle area corresponding to the non-target object in the reconstructed image, and determine the smallest circumscribed rectangle area as the area to be filled.

[0261] In step 1020, the image encoding area is used to mark the area to fill non-target objects in the reconstructed image.

[0262] In the specific implementation of this embodiment, use a preset number or a preset letter to identify each pixel point of the area to be filled to obtain an image encoding area. For example, mark each pixel point in the area to be filled as 1 or 0 to obtain an image encoding area.

[0263] In step 1030, first, determine the image center of the enhanced local image and the center of the image coding region. Then, based on the image center of the enhanced local image and determine the center of the image coding region, supplement the enhanced local image to the image coding region so that the image center of the enhanced local image coincides with the center of the image coding region. Further, if there is still an unsupplemented area in the image coding region after supplementing the enhanced local image to the image coding region, delete the unsupplemented part in the image coding region to obtain the complemented image corresponding to the base image.

[0264] As Figure 11A shown, it is a specific illustration of determining the enhanced local image. Specifically, find the non-target object in the base image and extract the image part corresponding to the non-target object to obtain the local image. Then, enhance the clarity of the local image to obtain the enhanced local image, where the contrast of the enhanced local image is greater than that of the local image.

[0265] As Figure 11B shown, it is a specific illustration of supplementing the enhanced local image to the reconstructed image. Specifically, first, determine a rectangular area of 3*2 as the area to be complemented in the reconstructed image according to the size of the local image corresponding to the enhanced local image and its specific position on the base image, and perform coding processing on the 6 pixel points in the area to be complemented, and label the 6 pixel points as 0. Then, supplement the enhanced local image to the image coding region of the reconstructed image and make the image center of the enhanced local image coincide with the center of the image coding region. Since there are two unsupplemented pixel points in the image coding region after supplementing the enhanced local image to the reconstructed image, at this time, delete the coding values corresponding to these two unsupplemented pixel points to obtain the complemented image corresponding to the base image.

[0266] The advantage of this embodiment is that considering the problems that non-target objects often become unclear or damaged in the reconstructed image, according to the size and image position of the enhanced local image whose clarity of the non-target object meets the requirements, delimit the area to be complemented in the reconstructed image, and perform coding processing on the area to be complemented to clearly distinguish the area to be complemented from other image areas, supplement the enhanced local image to the image coding region, and achieve precise complementation of non-target objects in the reconstructed image, overcoming the problem that non-target objects are prone to becoming unclear during the reconstruction process; at the same time, complementing the reconstructed image during the perspective conversion can reduce the pixel changes between adjacent frame images and the influence of the color, texture, and shape of nearby objects, making the obtained complemented image smoother and more natural, and improving the quality of the complemented images corresponding to each generated base image.

[0267] The following will describe step 360 in detail.

[0268] In step 360, the post-completion images corresponding to multiple base images are subjected to a synthesis process to obtain a synthesized image.

[0269] Please refer to Figure 12 , in one embodiment, step 360 specifically includes but is not limited to the following steps 1210-1230:

[0270] Step 1210: Obtain an image mask corresponding to the post-completion image;

[0271] Step 1220: Based on the background image, the post-completion image, and the image mask corresponding to the post-completion image, determine the image weight corresponding to the post-completion image;

[0272] Step 1230: Based on the image weights corresponding to each post-completion image, perform a weighted synthesis process on each post-completion image to obtain a synthesized image.

[0273] The following will describe steps 1210-1230 in detail.

[0274] In step 1210, the image mask is used to mask some pixel points in the post-completion image.

[0275] To save space, the specific process of obtaining the image mask corresponding to the post-completion image in the embodiments of the present disclosure will be described in detail below and will not be elaborated here.

[0276] In step 1220, the image weight is used to indicate the importance degree of the image feature information of each post-completion image for image synthesis.

[0277] In the specific implementation of this embodiment, first, the post-completion image is masked using the image mask corresponding to the post-completion image to obtain a masked image. Then, the masked image and the background image are fused and compressed to obtain a compressed image. Further, the brightness difference between the pixel points of each post-completion image and the pixel points of the compressed image is calculated, and the image weight corresponding to the post-completion image is determined according to the brightness difference.

[0278] In step 1230, first, for each pixel point in each post-completion image, the pixel value of the pixel point in the post-completion image is weighted using the image weight corresponding to the post-completion image to obtain a weighted pixel value. Then, all the weighted pixel values of the pixel point are added up to obtain the pixel value of the pixel point in the synthesized image. Thus, according to the coordinate positions of each pixel point and the pixel values in the synthesized image, a synthesized image is generated.

[0279] The advantage of this embodiment is that by weighted synthesis of multiple completed images based on an image mask and image weights, the differences between different completed images can be flexibly processed. Different image weights can be assigned according to the quality and importance of each completed image, which can better retain important details in the image while reducing noise and artifacts, thereby generating a synthetic image of higher quality.

[0280] Please refer to Figure 13 , in one embodiment, step 1210 specifically includes but is not limited to the following steps 1310 - 1330:

[0281] Step 1310: For a single pixel point in a single completed image among multiple completed images, determine the respective pixel values of the single pixel point in the multiple completed images;

[0282] Step 1320: Based on the respective pixel values of the single pixel point in the multiple completed images, determine the mask of the single pixel point;

[0283] Step 1330: Based on the masks of each pixel point in a single completed image, generate an image mask corresponding to the single completed image.

[0284] The following will describe steps 1310 - 1330 in detail.

[0285] In step 1310, pixel value extraction is performed for each pixel point in each completed image among multiple completed images to obtain the respective pixel values of each pixel point in the multiple completed images.

[0286] In step 1320, the mask is used to extract a specific region (region of interest) of the completed image and mask other regions in the completed image.

[0287] In the specific implementation of this embodiment, the specific process of determining the mask of a single pixel point based on the respective pixel values of the single pixel point in the multiple completed images may include but is not limited to the following steps:

[0288] Based on the respective pixel values of the single pixel point in the multiple completed images, determine the maximum pixel value and minimum pixel value of the single pixel point;

[0289] Based on the maximum pixel value and minimum pixel value, determine the pixel range of the single pixel point;

[0290] If the pixel range is greater than or equal to a preset threshold, determine the mask of the single pixel point as the first value;

[0291] If the pixel range is less than the preset threshold, determine the mask of the single pixel point as the second value.

[0292] Among them, the pixel maximum value is used to indicate the largest of the multiple pixel values of a pixel point in multiple completed images. The pixel minimum value is used to indicate the smallest of the multiple pixel values of a pixel point in multiple completed images.

[0293] The preset threshold is used to measure whether there is a problem of unclear display of a pixel point in multiple completed images (measuring whether a pixel point is displayed in a part of the completed image but not in another part of the completed image).

[0294] Specifically, first, for each pixel point, compare the pixel values of the pixel point in each completed image in terms of size. Take the largest of this series of pixel values as the pixel maximum value of the pixel point, and take the smallest of this series of pixel values as the pixel minimum value of the pixel point. Then, calculate the difference between the pixel maximum value and the pixel minimum value, subtract the pixel minimum value from the pixel maximum value to obtain the pixel range of the pixel. Further, compare the pixel range with the preset threshold. If the pixel range is greater than or equal to the preset threshold, determine the mask of a single pixel point as the first value, where the first value is 1, and the first value is used to indicate that the pixel point is a pixel point that needs to be retained and concerned in the completed image. If the pixel range is less than the preset threshold, determine the mask of a single pixel point as the second value, where the second value is 0, and the first value is used to indicate that the pixel point is a pixel point that needs to be masked in the completed image.

[0295] In step 1330, represent the masks of each pixel point in the completed image in a binary matrix with the same size as the completed image to obtain an image mask corresponding to a single completed image, where the image mask is a binary image.

[0296] Such as Figure 14As shown, it is a specific illustration of determining the image mask corresponding to the completed image. Specifically, the base image is based on the left-eye image and the right-eye image obtained by the VR binocular camera. Among them, the pixel values of each pixel point in the completed image corresponding to the left-eye image are successively {1, 2, 4, 7, 9, 10, 8, 7, 22}; the pixel values of each pixel point in the completed image corresponding to the right-eye image are successively {1, 5, 5, 9, 6, 10, 8, 7, 23}. Based on this, the difference is taken between the pixel value of each pixel point in the completed image corresponding to the left-eye image and the pixel value of the completed image corresponding to the right-eye image, and the pixel range sequence of the pixel point is obtained as {0, 3, 1, 2, 3, 0, 0, 0, 1}. When the preset threshold is 2, the pixel points with pixel ranges greater than or equal to the preset threshold are the second pixel point, the fourth pixel point, and the fifth pixel point. Therefore, the masks of the second pixel point, the fourth pixel point, and the fifth pixel point are determined to be 1, and the masks of other pixel points except the second pixel point, the fourth pixel point, and the fifth pixel point are determined to be 0, thereby forming an image mask.

[0297] It should be noted that in the embodiments of the present disclosure, the image masks corresponding to each completed image are the same.

[0298] The advantage of this embodiment is that by analyzing the pixel values of the same pixel point in multiple completed images, it is possible to more accurately understand the performance of the pixel point in different images, thereby being able to capture the changes of the pixel point under different conditions. According to the changes and differences of the pixel point under different conditions, the mask of the pixel point is determined, which can improve the accuracy of determining the masks of each pixel point in the completed image, and further improve the accuracy of generating the image mask according to the masks of the pixel points.

[0299] Please refer to Figure 15 , in one embodiment, step 1220 specifically includes but is not limited to the following steps 1510-1540:

[0300] Step 1510: For the completed image, perform mask processing on the completed image based on the image mask to obtain a masked image;

[0301] Step 1520: Perform image fusion on the masked image and the background image to obtain a target fusion image, and compress the target fusion image into a compressed image;

[0302] Step 1530: For each pixel point in the completed image, determine the first difference based on the brightness value of the pixel point in the completed image and the brightness value of the pixel point in the compressed image, and determine the brightness difference corresponding to the completed image based on the absolute value of the first difference of each pixel point;

[0303] Step 1540: Determine the image weight corresponding to the completed image based on the brightness difference corresponding to the completed image.

[0304] The following is a detailed description of steps 1510 - 1540.

[0305] In step 1510, the masked image refers to the image formed after adding an image mask to the completed image.

[0306] In the specific implementation of this embodiment, the specific process of masking the completed image based on the image mask may include but is not limited to the following steps:

[0307] For each pixel point in the completed image, determine the mask value of the pixel point in the image mask, and determine the pixel value of the pixel point in the completed image;

[0308] Based on the product of the mask value and the pixel value, obtain the masked pixel value of the pixel value;

[0309] Based on the masked pixel values of each pixel point, determine the masked image.

[0310] Among them, the masked pixel value refers to the pixel value of the pixel point after masking processing.

[0311] Specifically, first, for each pixel point in the completed image, determine the mask value of the pixel point in the image mask, and determine the pixel value of the pixel point in the completed image. Then, multiply the mask value and the pixel value of the pixel point to obtain the masked pixel value of the pixel value. Further, according to the coordinate positions of each pixel point, fill the masked pixel values of all pixel points into the same image matrix to obtain the masked image.

[0312] In step 1520, the target fusion image is used to indicate the pixel fusion result of each pixel point in the masked image and the background image.

[0313] In the specific implementation of this embodiment, the specific process of image fusion of the masked image and the background image is similar to the specific process of weighted merging of the initial pixel value and the pixel value of the target pixel point in the base image in step 530 above. For the sake of brevity, it will not be elaborated here.

[0314] Furthermore, use a compression algorithm to compress the target fusion image to obtain a compressed image. Among them, the compression algorithm includes but is not limited to lossless compression algorithms such as PNG compression and GIF compression, as well as lossy compression algorithms based on predictive coding or discrete cosine transform.

[0315] In step 1530, the first difference is used to indicate the magnitude of the brightness difference of the pixel point between the completed image and the compressed image. The larger the first difference, the greater the brightness difference of the pixel point between the completed image and the compressed image.

[0316] In the specific implementation of this embodiment, for each pixel point of each completed image, the difference between the brightness value of the pixel point in the completed image and the brightness value of the pixel value in the compressed image is calculated to obtain a first difference, and the absolute value of the first difference of each pixel point is taken to obtain the absolute value of the first difference of each pixel point. Finally, the sum of the absolute values of the first differences of all pixel points in each completed image is calculated to obtain the brightness difference corresponding to the completed image.

[0317] In step 1540, the specific process of determining the image weight corresponding to the completed image based on the brightness difference corresponding to the completed image may include but is not limited to the following steps:

[0318] For the completed image, take the reciprocal of the brightness difference to obtain a preliminary weight;

[0319] Sum the preliminary weights of multiple completed images to obtain a weight sum;

[0320] For the completed image, based on the preliminary weight and the weight sum, determine the image weight corresponding to the completed image.

[0321] Among them, the weight sum is used to indicate the total sum of the preliminary weights of all completed images.

[0322] Specifically, since the larger the brightness difference of the completed image, the greater the impact of the completed image on the image quality of the synthesized image. In order to reduce the impact of the completed image, a smaller image weight should be assigned to the completed image. Based on this, first, for each completed image, take the reciprocal of the brightness difference to obtain the preliminary weight of the completed image. Then, sum the preliminary weights of multiple completed images to obtain a weight sum. Finally, for each completed image, divide the preliminary weight by the weight sum to obtain the image weight corresponding to the completed image.

[0323] The advantage of this embodiment is that by performing mask processing on the completed image through an image mask, specific regions of the completed image can be operated on without affecting other parts, improving the flexibility of local operations; performing image fusion and compression on the mask image and the background image can reduce artifacts such as noise and blur, and integrate the key information of multiple images into the same image to obtain a compressed image; using the difference between the brightness value of the pixel point in the completed image and the brightness value of the pixel point in the compressed image to determine the brightness difference of the completed image, thereby determining the image brightness weight, can improve the accuracy of determining the image weight.

[0324] In one embodiment, the specific process of performing image fusion on the mask image and the background image in step 1520 may include but is not limited to the following steps:

[0325] Generate pyramid features for the mask image and the background image respectively to obtain multiple first pyramid features corresponding to the mask image and multiple second pyramid features corresponding to the background image;

[0326] Perform feature fusion on the first pyramid features and the second pyramid features at the same scale to obtain pyramid fusion features corresponding to each scale;

[0327] Determine the target fusion image based on the pyramid fusion features at each scale.

[0328] Among them, the first pyramid feature is used to indicate the image representation of the mask image at different scales. The second pyramid feature is used to indicate the image representation of the background image at different scales. The pyramid fusion feature is used to indicate the result of fusing the first pyramid feature and the second pyramid feature at the same scale together.

[0329] Specifically, first, generate pyramid features for the mask image and the background image respectively to implement downsampling processing of the mask image and the background image, and obtain the first pyramid features at 1 / 2 scale, 1 / 4 scale, and 1 / 8 scale corresponding to the mask image, and the second pyramid features at 1 / 2 scale, 1 / 4 scale, and 1 / 8 scale corresponding to the background image. Then, during feature fusion, for the 1 / 8 scale, directly use the low-frequency information of the background image and use the second pyramid feature at the 1 / 8 scale as the pyramid fusion feature at the 1 / 8 scale; for the 1 / 2 scale, retain the high-frequency details of the mask image and use the first pyramid feature at the 1 / 2 scale as the pyramid fusion feature at the 1 / 2 scale; for the 1 / 4 scale, use the averaging method to perform feature fusion on the first pyramid feature at the 1 / 4 scale and the second pyramid feature at the 1 / 4 scale, so that the weight ratios of the first pyramid feature at the 1 / 4 scale and the second pyramid feature at the 1 / 4 scale are both 50%, and realize the mixing of the intermediate-frequency details of the mask image and the background image to obtain the pyramid fusion feature at the 1 / 4 scale. Finally, map the pyramid fusion features at each scale to the same feature dimension, and perform feature fusion on the pyramid fusion features in the same feature dimension to obtain the target fusion image.

[0330] The advantage of this embodiment is that it takes into account generating pyramid features for the mask image and the background image respectively, performing layer-by-layer fusion during feature fusion, and reconstructing the target fusion image. A multi-scale fusion method is introduced in the image fusion of the mask image and the background image, which can effectively reduce artifacts in image fusion and improve the image quality of the generated target fusion image.

[0331] Please refer to Figure 16, in one embodiment, step 1230 specifically includes but is not limited to the following steps 1610 - 1630:

[0332] Step 1610: For a single pixel point in a single completed image among multiple completed images, determine the respective brightness values of the single pixel point in each of the multiple completed images;

[0333] Step 1620: Based on the image weights corresponding to each completed image, perform weighted synthesis processing on the respective brightness values of the single pixel point in each of the multiple completed images to obtain the target brightness value of the single pixel point;

[0334] Step 1630: Based on the target brightness values of each pixel point, determine the synthesized image.

[0335] The following is a detailed description of steps 1610 - 1630.

[0336] In step 1610, the brightness value is used to indicate the light and dark degree of the image part corresponding to the pixel point in the completed image.

[0337] In the specific implementation of this embodiment, for a single pixel point in a single completed image among multiple completed images, the pixel value of this pixel point in each completed image is decomposed on the RGB color channels to obtain the pixel component R of the red channel, the pixel component G of the green channel, and the pixel component B of the blue channel of this pixel point. Then, a weighted average is performed on the pixel component R of the red channel, the pixel component G of the green channel, and the pixel component B of the blue channel of this pixel point to obtain the brightness value of this pixel point in each completed image.

[0338] Among them, the brightness value of this pixel point in each completed image can be expressed as the following formula:

[0339] Brightness value = 0.3×R + 0.59×G + 0.11×B;

[0340] In step 1620, the target brightness value is used to indicate the light and dark degree of the image part corresponding to the pixel point in the synthesized image.

[0341] In the specific implementation of this embodiment, for each pixel point, multiply the brightness value of the pixel point in the completed image by the image weight of this completed image to obtain the weighted brightness value, and add the weighted brightness values of the pixel point in multiple completed images to obtain the target brightness value of this pixel point.

[0342] In step 1630, use the target brightness values of each pixel point as the brightness values of each pixel point in the synthesized image, and determine the pixel values of each pixel point in the synthesized image according to the target brightness values, so as to obtain the complete synthesized image.

[0343] The advantage of this embodiment is that, considering the different sensitivities of the human eye's vision to different colors, the pixel value of each pixel point in the completed image is decomposed into pixel components on the RGB color channels, and the pixel components of the three color channels are weighted and averaged to obtain the brightness value corresponding to each pixel point in the completed image, which can improve the accuracy of obtaining the brightness value of each pixel point in the completed image. Further, by weighted combining the brightness values of each pixel point according to the image weight, the accuracy of combining the brightness values of multiple pixel points can be improved, thereby improving the brightness equalization and brightness accuracy of the synthesized image, and further improving the image quality of the synthesized image and the accuracy of image perspective conversion.

[0344] Since not all of the completed images corresponding to the basic images are of high quality during the reconstruction process, often due to various factors such as shooting perspective and lighting conditions, the image quality of some of the completed images corresponding to the basic images is poor. Using low-quality completed images for image synthesis will greatly affect the quality of the generated synthesized image. Based on this, the embodiments of the present disclosure provide an image synthesis scheme based on a quality detection mechanism, which can more accurately perform quality detection on each completed image, and use the completed images that pass the quality detection for image synthesis, which can improve the quality of the generated synthesized image.

[0345] Please refer to Figure 17 , in one embodiment, step 360 specifically includes but is not limited to the following steps 1710-1720:

[0346] Step 1710: Perform quality detection on the completed image corresponding to each basic image to obtain a quality detection result;

[0347] Step 1720: After determining that the quality detection results of the completed images corresponding to each basic image all meet the preset conditions, perform synthesis processing on the completed images corresponding to multiple basic images to obtain a synthesized image.

[0348] The following will describe steps 1710-1730 in detail.

[0349] In step 1710, the quality detection result is used to indicate the quality of the completed image corresponding to the basic image.

[0350] To save space, the specific process of performing quality detection on the completed image corresponding to each basic image in the embodiments of the present disclosure will be described in detail below and will not be elaborated here.

[0351] In step 1720, the preset conditions are used to describe the index requirements that the image quality of the completed image needs to meet.

[0352] In the specific implementation of this embodiment, first, for the completed image corresponding to each base image, the quality detection result of the completed image is compared with the index requirements in the preset conditions. If the quality detection results of the completed images corresponding to each base image meet the preset conditions, the completed images corresponding to multiple base images are subjected to synthesis processing. If there is one or some completed images corresponding to base images whose quality detection results do not meet the preset conditions, the completed images whose direct detection results do not meet the preset conditions are optimized, or the original base images are used to regenerate the completed images again according to the above steps 310 - 350 until the quality detection results of the completed images corresponding to each base image meet the preset conditions, and then the completed images corresponding to multiple base images are subjected to synthesis processing. The specific synthesis processing process is similar to the above steps 1210 - 1230. For the sake of brevity, it will not be elaborated here.

[0353] The advantage of this embodiment is that in the process of synthesizing the completed images corresponding to multiple base images, a quality detection mechanism is introduced. First, the quality of each completed image is detected, and synthesis processing is only carried out after the quality detection results of all completed images meet the preset conditions, which can ensure that the completed images participating in the synthesis processing all have high image quality, thereby improving the quality of the generated synthesized image.

[0354] In one embodiment, the specific process of detecting the quality of the completed image corresponding to each base image may include but is not limited to the following steps:

[0355] For each pixel point in the completed image corresponding to each base image, determine the difference between the pixel value of the pixel point in the base image and the pixel value of the pixel point in the completed image;

[0356] Based on the differences of each pixel point, determine the mean square error of the pixel points;

[0357] Based on the pixel values of each pixel point in the base image, determine the maximum pixel value corresponding to the base image;

[0358] Based on the mean square error of the pixel points and the maximum pixel value, determine the quality detection result of the completed image.

[0359] Specifically, for each pixel in the completed image corresponding to each base image, calculate the difference between the pixel value of the pixel in the base image and the pixel value of the pixel in the completed image to obtain a difference. Then, square the difference to obtain the squared difference of the pixel, and calculate the mean of the squared differences of all pixels to obtain the mean squared error of the pixels. Further, compare the pixel values of each pixel in the base image, and obtain the maximum pixel value as the maximum pixel value corresponding to the base image. Then, divide the squared result of the maximum pixel value by the mean squared error of the pixels to obtain a first result, take the logarithm of the first result to obtain the peak signal-to-noise ratio of the completed image corresponding to the base image, and use the peak signal-to-noise ratio as the quality detection result of the completed image.

[0360] Among them, the peak signal-to-noise ratio of the completed image corresponding to the base image can be expressed as the following formula:

[0361]

[0362] Among them, MSE is the mean squared error of the pixels, PSNR is the peak signal-to-noise ratio, MAX_I is the maximum pixel value, is the first result. I is the base image. Among them, the larger the peak signal-to-noise ratio, the better the image quality. In the preset conditions, it is often limited that the peak signal-to-noise ratio of the completed image should be greater than 30 dB to meet the requirements.

[0363] The advantage of this embodiment is that the ratio of signal to noise between the base image and the completed image is characterized by the difference between the completed image corresponding to the base image, the mean squared error of the pixels of the base image, and the maximum pixel value of the base image, and the image quality of the completed image is measured according to the ratio of signal to noise (peak signal-to-noise ratio). The error is converted into decibels, and the image quality of the completed image can be clearly reflected according to the size of the peak signal-to-noise ratio, improving the accuracy of image quality detection.

[0364] In one embodiment, the specific process of quality detection for each completed image corresponding to each base image may include but is not limited to the following steps:

[0365] For each base image, based on the pixel values of each pixel in the base image, determine the first pixel mean and the first pixel standard deviation;

[0366] For the completed image corresponding to the base image, based on the pixel values of each pixel in the completed image, determine the second pixel mean and the second pixel standard deviation;

[0367] Determine the pixel covariance of the base image and the completed image;

[0368] Determine a first value based on the sum of squares of the first pixel mean and the second pixel mean, and the sum of squares of the first pixel standard deviation and the second pixel standard deviation;

[0369] Determine a second value based on the product of the pixel covariance, the first pixel mean, and the second pixel mean;

[0370] Determine the quality detection result of the completed image based on the quotient of the second value and the first value.

[0371] Specifically, for each base image, calculate the average of the pixel values of each pixel point in the base image to obtain the first pixel mean, and calculate the pixel standard deviation of the pixel points in the base image according to the difference between the pixel value of each pixel point and the first pixel mean to obtain the first pixel standard deviation. Similarly, for the completed image corresponding to the base image, calculate the average of the pixel values of each pixel point in the completed image to obtain the second pixel mean, and calculate the pixel standard deviation of the pixel points in the completed image according to the difference between the pixel value of each pixel point and the second pixel mean to obtain the second pixel standard deviation. Further, for each pixel point, multiply the difference between the pixel value of the pixel point in the base image and the first pixel mean by the difference between the pixel value of the pixel point in the completed image and the second pixel mean to obtain the pixel product result, sum up the pixel product results of all pixel points to obtain the total product result, calculate the quotient between the total product result and the result of subtracting 1 from the total number of pixel points to obtain the pixel covariance between the base image and the completed image. Then, multiply the result of adding the first constant to the sum of squares of the first pixel mean and the second pixel mean by the result of adding the second constant to the sum of squares of the first pixel standard deviation and the second pixel standard deviation to obtain the first value. At the same time, multiply the result of adding twice the pixel covariance to the second constant by the result of adding twice the product of the first pixel mean and the second pixel mean to the first constant to obtain the second value. Finally, divide the second value by the first value to obtain the quotient of the second value and the first value, where the quotient of the second value and the first value is used to indicate the structural similarity between the base image and the corresponding completed image; therefore, use the quotient of the second value and the first value as the quality detection result of the completed image.

[0372] Among them, the structural similarity is used to measure the similarity between the base image and the completed image corresponding to the base image from three aspects: brightness, contrast, and structure. The value range of the structural similarity is [0, 1]. The closer the structural similarity is to 1, the more similar the base image and the completed image are, and the better the image quality of the completed image. In the preset conditions, it is often limited that the structural similarity of the completed image should be greater than a certain value within the range of [0, 1] (such as 0.85) to meet the requirements.

[0373] Among them, the structural similarity (SSIM) between the base image and the corresponding completed image can be expressed as the following formula:

[0374]

[0375] Among them, x represents the base image, y represents the completed image corresponding to the base image x, and μ x represents the first pixel mean of the base image, and μ y represents the second pixel mean of the completed image, and σ x represents the first pixel standard deviation of the base image, and σ y represents the second pixel standard deviation of the completed image. σ xy represents the covariance corresponding to the base image x and the completed image y, C1 represents the first constant, and C2 represents the second constant.

[0376] The advantage of this embodiment is that by using the first pixel mean and the first pixel standard deviation of each pixel point in the base image, the second pixel mean and the second pixel standard deviation of each pixel point in the completed image, and the covariance between the base image and the completed image, the similarity between the base image and the completed image corresponding to the base image is measured from three aspects: brightness, contrast, and structure, which can more comprehensively realize the detection of the image quality of the completed image and improve the comprehensiveness and flexibility of the image quality detection.

[0377] In one embodiment, the specific process of performing quality detection on the completed image corresponding to each base image may include but is not limited to the following steps:

[0378] For each base image, perform feature extraction on the base image to obtain the first image feature, and perform feature extraction on the completed image corresponding to the base image to obtain the second image feature;

[0379] Based on the feature differences between the first image feature and the second image feature, determine the quality detection result of the completed image.

[0380] Specifically, first, for each base image, use a pre-trained convolutional neural network to perform feature extraction on the base image to obtain the first image feature, and use the convolutional neural network to perform feature extraction on the corresponding completed image to obtain the second image feature. Then, calculate the feature mean square difference between the first image feature and the second image feature to obtain the content loss of the completed image, and use the content loss as the quality detection result of the completed image. Among them, the content loss is used to measure the difference between the completed image and the base image in high-level features. The smaller the content loss, the better the image quality of the completed image. The preset conditions often limit that the content loss of the completed image should be less than a certain value to meet the requirements.

[0381] The advantage of this embodiment is that the base image and the completed image corresponding to the base image are input into a pre-trained convolutional neural network for feature extraction, and the perceptual quality of the completed image is measured based on the feature difference between the first image feature corresponding to the base image and the second image feature corresponding to the completed image. Quality detection is performed on the completed image based on the perceptual loss, which can ensure that the perceptual quality of the completed image used for synthesis processing meets the requirements of real vision and is conducive to improving the visual quality of the generated synthetic image.

[0382] The implementation details of the image synthesis processing method according to an embodiment of the present disclosure will be described in detail below.

[0383] As Figure 18 shown, in image synthesis, first, a left-eye view and a right-eye view are obtained by using a VR camera as base images. Then, for each base image, the pixel points constituting the target object are extracted to obtain a pixel sequence, and the specific process is similar to step 330 described above. Further, operations such as pixel-level fusion, correction, propagation, and quality evaluation are performed using the background image and the pixel sequence to obtain the completed image, and the specific process is similar to steps 410-460, steps 510-540, steps 1010-1030, and step 1710 described above. Finally, based on the image mask and the background image, synthesis processing and brightness adjustment are performed on multiple completed images to obtain a synthetic image that conforms to human vision, and the specific process is similar to steps 1210-1230 described above. To save space, it will not be elaborated here.

[0384] The apparatus and device according to the embodiments of the present disclosure will be described below.

[0385] It can be understood that although each step in the above flowcharts is sequentially shown according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0386] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the target object attribute information, the separate permission or separate consent of the target object will be obtained through pop-up windows or by jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data for the normal operation of the embodiments of the present application will be obtained.

[0387] Figure 19 FIG. 1900 is a schematic structural diagram of an image synthesis processing apparatus 1900 provided by an embodiment of the present disclosure. The image synthesis processing apparatus 1900 includes:

[0388] A first acquisition unit 1910, configured to acquire a plurality of basic images of a target object in a background from multiple angles, where the basic images include the target object and non-target objects;

[0389] A second acquisition unit 1920, configured to acquire a background image of the background;

[0390] An extraction unit 1930, configured to extract pixel point information of the target object in each basic image;

[0391] A calibration unit 1940, configured to perform layer adaptation calibration on the pixel point information based on the layer feature information of the basic image to obtain calibrated pixel point information, and perform image reconstruction on the background image based on the calibrated pixel point information to obtain a reconstructed image;

[0392] A completion unit 1950, configured to extract a local image of the non-target object from the basic image, enhance the clarity of the local image to obtain an enhanced local image, and obtain a completed image corresponding to the basic image according to the enhanced local image and the reconstructed image;

[0393] A synthesis unit 1960, configured to perform synthesis processing on the completed images corresponding to the plurality of basic images to obtain a synthesized image.

[0394] Optionally, the synthesis unit 1960 includes:

[0395] An acquisition module (not shown), configured to acquire an image mask corresponding to the completed image;

[0396] A determination module (not shown), configured to determine an image weight corresponding to the completed image based on the background image, the completed image, and the image mask corresponding to the completed image;

[0397] A first synthesis module (not shown) is used to perform weighted synthesis processing on each of the complemented images based on the image weights corresponding to the respective complemented images to obtain a synthesized image.

[0398] Optionally, an acquisition module (not shown) is used for:

[0399] For a single pixel point in a single complemented image among multiple complemented images, determine the respective pixel values of the single pixel point in the multiple complemented images;

[0400] Based on the respective pixel values of the single pixel point in the multiple complemented images, determine the mask of the single pixel point;

[0401] Based on the masks of the respective pixel points in a single complemented image, generate an image mask corresponding to the single complemented image.

[0402] Optionally, determining the mask of a single pixel point based on the respective pixel values of the single pixel point in the multiple complemented images includes:

[0403] Based on the respective pixel values of the single pixel point in the multiple complemented images, determine the maximum pixel value and the minimum pixel value of the single pixel point;

[0404] Based on the maximum pixel value and the minimum pixel value, determine the pixel range of the single pixel point;

[0405] If the pixel range is greater than or equal to a preset threshold, determine the mask of the single pixel point as a first value;

[0406] If the pixel range is less than the preset threshold, determine the mask of the single pixel point as a second value.

[0407] Optionally, a first synthesis module (not shown) is used for:

[0408] For a single pixel point in a single complemented image among multiple complemented images, determine the respective brightness values of the single pixel point in the multiple complemented images;

[0409] Based on the image weights corresponding to the respective complemented images, perform weighted synthesis processing on the respective brightness values of the single pixel point in the multiple complemented images to obtain the target brightness value of the single pixel point;

[0410] Based on the target brightness values of the respective pixel points, determine the synthesized image.

[0411] Optionally, the determination module (not shown) includes:

[0412] A mask sub-module (not shown) is used to perform mask processing on the complemented image based on the image mask for the complemented image to obtain a masked image;

[0413] A fusion sub-module (not shown) is used to perform image fusion on the mask image and the background image to obtain a target fusion image, and compress the target fusion image into a compressed image;

[0414] A first determination sub-module (not shown) is used to, for each pixel point in the complemented image, determine a first difference based on the brightness value of the pixel point in the complemented image and the brightness value of the pixel point in the compressed image, and determine the brightness difference corresponding to the complemented image based on the absolute value of the first difference of each pixel point;

[0415] A second determination sub-module is used to determine the image weight corresponding to the complemented image based on the brightness difference corresponding to the complemented image.

[0416] Optionally, a masking sub-module (not shown) is used for:

[0417] For each pixel point in the complemented image, determine the mask value of the pixel point in the image mask and determine the pixel value of the pixel point in the complemented image;

[0418] Obtain the masked pixel value of the pixel value based on the product of the mask value and the pixel value;

[0419] Determine the mask image based on the masked pixel values of each pixel point.

[0420] Optionally, a fusion sub-module (not shown) is used for:

[0421] Generate pyramid features for the mask image and the background image respectively to obtain a plurality of first pyramid features corresponding to the mask image and a plurality of second pyramid features corresponding to the background image;

[0422] Perform feature fusion on the first pyramid features and the second pyramid features at the same scale to obtain the pyramid fusion features corresponding to each scale;

[0423] Determine the target fusion image based on the pyramid fusion features of each scale.

[0424] Optionally, a second determination sub-module (not shown) is used for:

[0425] For the complemented image, take the reciprocal of the brightness difference to obtain a preliminary weight;

[0426] Sum the preliminary weights of multiple complemented images to obtain a weight sum;

[0427] For the complemented image, determine the image weight corresponding to the complemented image based on the preliminary weight and the weight sum.

[0428] Optionally, the complementing unit 1950 is used for:

[0429] Determine the image sharpness of the local image;

[0430] If the image clarity meets the first preset condition, the partial image is used as the enhanced partial image;

[0431] If the image clarity does not meet the first preset condition, the partial image is denoised to obtain a denoised image, and the clarity of the denoised image is enhanced to obtain an enhanced partial image with image clarity meeting the first preset condition.

[0432] Optionally, the complementing unit 1950 is configured to:

[0433] Determine a region to be complemented in the reconstructed image based on the image size and image position information of the enhanced partial image;

[0434] Perform encoding processing on the region to be complemented to obtain an image encoding region;

[0435] Supplement the enhanced partial image to the image encoding region so that the center of the enhanced partial image coincides with the center of the image encoding region, obtaining a complemented image corresponding to the base image.

[0436] Optionally, the pixel point information includes the pixel values of each target pixel point corresponding to the target object in the base image;

[0437] The correction unit 1940 is configured to:

[0438] For the pixel value of each target pixel point, determine the channel pixel value of the target pixel point in each color channel;

[0439] For each color channel, based on the channel pixel values of each target pixel point in the color channel, obtain the pixel mean value corresponding to the color channel;

[0440] Extract the layer reference value from the layer feature information, and based on the layer reference value and the pixel mean value, determine the correction coefficient corresponding to the color channel;

[0441] Based on the correction coefficient, adaptively correct the channel pixel values of each target pixel point in the color channel to obtain the corrected channel pixel values of each target pixel point;

[0442] For each target pixel point, merge the corrected channel pixel values of the target pixel point in each color channel to obtain the corrected pixel value;

[0443] Based on the corrected pixel values of each target pixel point, obtain the corrected pixel point information.

[0444] Optionally, the corrected pixel point information includes the pixel point label, pixel point coordinates, and corrected pixel value of each target pixel point corresponding to the target object;

[0445] The correction unit 1940 is configured to:

[0446] For each target pixel point, based on the pixel point coordinates of the target pixel point, determine a mapped pixel point corresponding to the target pixel point in the background image, where the coordinates of the mapped pixel point are the same as the pixel point coordinates;

[0447] If the pixel point label of the target pixel point is the first label value, use the corrected pixel value as the updated pixel value of the mapped pixel point;

[0448] If the pixel point label of the target pixel point is the second label value, obtain the initial pixel value of the mapped pixel point in the background image, perform weighted combination on the initial pixel value and the pixel value of the target pixel point in the base image to obtain a combined pixel value, and use the combined pixel value as the updated pixel value of the mapped pixel point;

[0449] Determine a reconstructed image based on the updated pixel values of each mapped pixel point and the initial pixel values of the non-mapped pixel points in the background image.

[0450] Optionally, the correction unit 1940 is configured to:

[0451] Based on the layer feature information of the base image, determine the pixel points to be corrected among the target pixel points;

[0452] Use a preset reference value as the corrected pixel value of the pixel points to be corrected, and update the pixel point information based on the corrected pixel values of each pixel point to be corrected to obtain corrected pixel point information.

[0453] Optionally, the corrected pixel point information includes the coordinate parameters and pixel values of each target pixel point corresponding to the target object in the base image, and the corrected pixel values of the pixel points to be corrected among the target pixel points;

[0454] The correction unit 1940 is configured to:

[0455] For the coordinate parameters of the target pixel point, identify the background pixel points with the coordinate parameters in the background image, and update the pixel values of the background pixel points with the coordinate parameters based on the pixel value of the target pixel point to obtain an updated background image;

[0456] Locate the pixel points in the updated background image that have the same coordinate parameters as the pixel points to be corrected, and replace the pixel values of the located pixel points with the corrected pixel values of the pixel points to be corrected to obtain a reconstructed image.

[0457] Optionally, the synthesis unit 1960 includes:

[0458] A detection module (not shown), configured to perform quality detection on the completed image corresponding to each base image to obtain a quality detection result;

[0459] A second synthesis module (not shown) is configured to perform a synthesis process on the post-completion images corresponding to a plurality of base images to obtain a synthesized image after determining that the quality detection results of the post-completion images corresponding to each base image meet preset conditions.

[0460] Optionally, a detection module (not shown) is configured to:

[0461] For each base image, determine a first pixel mean value and a first pixel standard deviation based on the pixel values of each pixel point in the base image;

[0462] For the post-completion image corresponding to the base image, determine a second pixel mean value and a second pixel standard deviation based on the pixel values of each pixel point in the post-completion image;

[0463] Determine the pixel covariance between the base image and the post-completion image;

[0464] Based on the sum of the squares of the first pixel mean value and the second pixel mean value, and the sum of the squares of the first pixel standard deviation and the second pixel standard deviation, determine a first value;

[0465] Based on the pixel covariance, the product of the first pixel mean value, and the second pixel mean value, determine a second value;

[0466] Based on the quotient of the second value and the first value, determine the quality detection result of the post-completion image.

[0467] Optionally, a detection module (not shown) is configured to:

[0468] For each base image, perform feature extraction on the base image to obtain a first image feature, and perform feature extraction on the post-completion image corresponding to the base image to obtain a second image feature;

[0469] Based on the feature difference between the first image feature and the second image feature, determine the quality detection result of the post-completion image.

[0470] Optionally, a detection module (not shown) is configured to:

[0471] For each pixel point in the post-completion image corresponding to each base image, determine the difference between the pixel value of the pixel point in the base image and the pixel value of the pixel point in the post-completion image;

[0472] Based on the differences of each pixel point, determine the mean square error of the pixel points;

[0473] Based on the pixel values of each pixel point in the base image, determine the maximum pixel value corresponding to the base image;

[0474] Based on the mean square error of the pixel points and the maximum pixel value, determine the quality detection result of the post-completion image.

[0475] Reference Figure 20 , Figure 20 FIG. 1 is a block diagram of a part of a terminal for implementing the image synthesis processing method according to an embodiment of the present disclosure. The terminal may be Figure 1 the object terminal shown in FIG. 1. The terminal includes: a Radio Frequency (RF) circuit 2010, a memory 2015, an input unit 2030, a display unit 2040, a sensor 2050, an audio circuit 2060, a wireless fidelity (WiFi) module 2070, a processor 2080, and a power supply 2090 and other components. Those skilled in the art can understand that Figure 20 the terminal structure shown in FIG. 1 does not limit the mobile phone or computer, and may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0476] The RF circuit 2010 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 2080 for processing; in addition, the uplink data designed is sent to the base station.

[0477] The memory 2015 can be used to store software programs and modules. The processor 2080 executes various functional applications and data processing of the object terminal by running the software programs and modules stored in the memory 2015.

[0478] The input unit 2030 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the object terminal. Specifically, the input unit 2030 may include a touch panel 2031 and other input devices 2032.

[0479] The display unit 2040 can be used to display the input information or provided information and various menus of the object terminal. The display unit 2040 may include a display panel 2041.

[0480] The audio circuit 2060, the speaker 2061, and the microphone 2062 can provide an audio interface.

[0481] In this embodiment, the processor 2080 included in the terminal can execute the image synthesis processing method of the previous embodiment.

[0482] Figure 21 FIG. 2 is a block diagram of a part of a server for implementing the image synthesis processing method according to an embodiment of the present disclosure. The server may be Figure 1The server shown. Servers can vary significantly due to configuration or performance differences and can include one or more central processing units (CPUs) 2122 (e.g., one or more processors) and a memory 2132, and one or more storage media 2130 (e.g., one or more mass storage devices) that store application programs 2142 or data 2144. Among them, the memory 2132 and the storage media 2130 can be transient storage or persistent storage. The programs stored in the storage media 2130 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations on the server. Further, the central processing unit 2122 can be configured to communicate with the storage media 2130 and execute a series of instruction operations in the storage media 2130 on the server.

[0483] The server can also include one or more power supplies 2126, one or more wired or wireless network interfaces 2150, one or more input / output interfaces 2158, and / or one or more operating systems 2141, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0484] The central processing unit 2122 in the server can be used to execute the image synthesis processing method of the embodiments of the present disclosure.

[0485] The embodiments of the present disclosure also provide a computer-readable storage medium for storing a computer program for executing the image synthesis processing method of the foregoing various embodiments.

[0486] The embodiments of the present disclosure also provide a computer program product that includes a computer program. The processor of the electronic device reads and executes the computer program, so that the electronic device executes to implement the above image synthesis processing method.

[0487] In the description of the present disclosure and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0488] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0489] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.

[0490] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0491] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0492] In addition, each functional unit in various embodiments of the present disclosure may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0493] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0494] It should also be understood that the various embodiments provided in the present disclosure can be combined arbitrarily to achieve different technical effects.

[0495] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.

Claims

1. An image synthesis processing method, characterized in that The method includes: Obtaining a plurality of basic images of a target object in a background from multiple perspectives, where the basic images include the target object and non-target objects; Obtaining a background image of the background; For each basic image, extracting pixel point information of the target object in the basic image; Based on the layer feature information of the basic image, performing layer adaptation correction on the pixel point information to obtain corrected pixel point information, and based on the corrected pixel point information, performing image reconstruction on the background image to obtain a reconstructed image; Extracting a local image of the non-target object from the basic image, enhancing the clarity of the local image to obtain an enhanced local image, and obtaining a completed image corresponding to the basic image according to the enhanced local image and the reconstructed image; Performing synthesis processing on the completed images corresponding to the multiple basic images to obtain a synthesized image.

2. The method according to claim 1, characterized in that, The performing synthesis processing on the completed images corresponding to the multiple basic images to obtain a synthesized image includes: Obtaining an image mask corresponding to the completed image; Based on the background image, the completed image, and the image mask corresponding to the completed image, determining an image weight corresponding to the completed image; Based on the image weights corresponding to the respective completed images, performing weighted synthesis processing on the respective completed images to obtain the synthesized image.

3. The method according to claim 2, wherein The obtaining an image mask corresponding to the completed image includes: For a single pixel point in a single completed image among the multiple completed images, determining the respective pixel values of the single pixel point in the multiple completed images; Based on the respective pixel values of the single pixel point in the multiple completed images, determining a mask of the single pixel point; Based on the masks of the respective pixel points in the single completed image, generating the image mask corresponding to the single completed image.

4. The method according to claim 2, wherein The performing weighted synthesis processing on the respective completed images based on the image weights corresponding to the respective completed images to obtain the synthesized image includes: For a single pixel point in a single completed image among the multiple completed images, determining the respective brightness values of the single pixel point in the multiple completed images; Based on the image weights corresponding to the respective completed images, performing weighted synthesis processing on the respective brightness values of the single pixel point in the multiple completed images to obtain a target brightness value of the single pixel point; Based on the target brightness values of the respective pixel points, determining the synthesized image.

5. The method according to claim 2, characterized in that, The determining an image weight corresponding to the completed image based on the background image, the completed image, and the image mask corresponding to the completed image includes: For the completed image, performing mask processing on the completed image based on the image mask to obtain a masked image; Performing image fusion on the masked image and the background image to obtain a target fused image, and compressing the target fused image into a compressed image; For each pixel point in the completed image, determine a first difference based on the brightness value of the pixel point in the completed image and the brightness value of the pixel point in the compressed image, and determine the brightness difference corresponding to the completed image based on the absolute value of the first difference of each pixel point. Determine the image weight corresponding to the completed image based on the brightness difference corresponding to the completed image.

6. The method according to claim 5, wherein The determining the image weight corresponding to the completed image based on the brightness difference corresponding to the completed image includes: For the completed image, take the reciprocal of the brightness difference to obtain a preliminary weight. Sum the preliminary weights of the multiple completed images to obtain a weight sum. For the completed image, determine the image weight corresponding to the completed image based on the preliminary weight and the weight sum.

7. The method according to claim 1, wherein The enhancing the clarity of the local image to obtain an enhanced local image includes: Determine the image clarity of the local image. If the image clarity meets a first preset condition, use the local image as the enhanced local image. If the image clarity does not meet the first preset condition, perform denoising processing on the local image to obtain a denoised image, and enhance the clarity of the denoised image to obtain an enhanced local image whose image clarity meets the first preset condition.

8. The method according to claim 1, wherein The obtaining the completed image corresponding to the base image according to the enhanced local image and the reconstructed image includes: Based on the image size and image position information of the enhanced local image, determine a region to be completed in the reconstructed image. Perform encoding processing on the region to be completed to obtain an image encoding region. Supplement the enhanced local image to the image encoding region so that the center of the image of the enhanced local image coincides with the center of the image encoding region to obtain the completed image corresponding to the base image.

9. The method according to claim 1, wherein The pixel point information includes the pixel values of each target pixel point corresponding to the target object in the base image. The performing layer adaptation correction on the pixel point information based on the layer feature information of the base image to obtain corrected pixel point information includes: For the pixel value of each target pixel point, determine the channel pixel value of the target pixel point in each color channel. For each color channel, obtain the pixel mean value corresponding to the color channel based on the channel pixel values of each target pixel point in the color channel. Extract a layer reference value from the layer feature information, and determine the correction coefficient corresponding to the color channel based on the layer reference value and the pixel mean value. Based on the correction coefficient, perform adaptation correction on the channel pixel values of each target pixel point in the color channel to obtain the corrected channel pixel values of each target pixel point. For each target pixel point, merge the corrected channel pixel values of the target pixel point in each color channel to obtain a corrected pixel value. Based on the corrected pixel values of each target pixel point, obtain the corrected pixel point information.

10. The method according to claim 9, wherein The corrected pixel point information includes the pixel point labels, pixel point coordinates, and corrected pixel values of each target pixel point corresponding to the target object; Based on the corrected pixel point information, performing image reconstruction on the background image to obtain a reconstructed image, including: For each target pixel point, based on the pixel point coordinates of the target pixel point, determining a mapped pixel point corresponding to the target pixel point in the background image, where the coordinates of the mapped pixel point are the same as the pixel point coordinates; If the pixel point label of the target pixel point is the first label value, using the corrected pixel value as the updated pixel value of the mapped pixel point; If the pixel point label of the target pixel point is the second label value, obtaining the initial pixel value of the mapped pixel point in the background image, performing weighted merging on the initial pixel value and the pixel value of the target pixel point in the base image to obtain a merged pixel value, and using the merged pixel value as the updated pixel value of the mapped pixel point; Based on the updated pixel values of each mapped pixel point and the initial pixel values of the non-mapped pixel points in the background image, determining the reconstructed image.

11. The method according to claim 1, wherein Performing synthesis processing on the completed images corresponding to the multiple base images to obtain a synthesized image, including: Performing quality detection on the completed image corresponding to each base image to obtain a quality detection result; After determining that the quality detection results of the completed images corresponding to each base image all meet the preset conditions, performing synthesis processing on the completed images corresponding to the multiple base images to obtain the synthesized image.

12. The method according to claim 11, wherein Performing quality detection on the completed image corresponding to each base image to obtain a quality detection result, including: For each pixel point in the completed image corresponding to each base image, determining the difference between the pixel value of the pixel point in the base image and the pixel value of the pixel point in the completed image; Based on the differences of each pixel point, determining the pixel mean square error; Based on the pixel values of each pixel point in the base image, determining the maximum pixel value corresponding to the base image; Based on the pixel mean square error and the maximum pixel value, determining the quality detection result of the completed image.

13. The method according to claim 11, wherein Performing quality detection on the completed image corresponding to each base image to obtain a quality detection result, including: For each base image, based on the pixel values of each pixel point in the base image, determining the first pixel mean and the first pixel standard deviation; For the completed image corresponding to the base image, based on the pixel values of each pixel point in the completed image, determining the second pixel mean and the second pixel standard deviation; Determining the pixel covariance between the base image and the completed image; Based on the sum of the squares of the first pixel mean and the second pixel mean and the sum of the squares of the first pixel standard deviation and the second pixel standard deviation, determining a first value; Based on the pixel covariance, the product of the first pixel mean, and the second pixel mean, determining a second value; Determine the quality detection result of the completed image based on the quotient of the second value and the first value.

14. The method according to claim 11, wherein The quality detection is performed on the completed image corresponding to each basic image to obtain a quality detection result, including: For each basic image, extract a first image feature from the basic image, and extract a second image feature from the completed image corresponding to the basic image; Determine the quality detection result of the completed image based on the feature difference between the first image feature and the second image feature.

15. An image synthesis processing device, characterized in that, The device includes: A first acquisition unit, configured to respectively acquire a plurality of basic images of a target object in a background from multiple angles, where the basic image includes the target object and a non-target object; A second acquisition unit, configured to acquire a background image of the background; An extraction unit, configured to, for each basic image, extract pixel point information of the target object in the basic image; A correction unit, configured to perform layer adaptation correction on the pixel point information based on the layer feature information of the basic image to obtain corrected pixel point information, and perform image reconstruction on the background image based on the corrected pixel point information to obtain a reconstructed image; A completion unit, configured to extract a local image of the non-target object from the basic image, enhance the clarity of the local image to obtain an enhanced local image, and obtain the completed image corresponding to the basic image according to the enhanced local image and the reconstructed image; A synthesis unit, configured to perform synthesis processing on the completed images corresponding to the plurality of basic images to obtain a synthesized image.

16. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the image synthesis processing method according to any one of claims 1 to 14.

17. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the image synthesis processing method according to any one of claims 1 to 14.

18. A computer program product, the computer program product includes a computer program, the computer program is read and executed by a processor of an electronic device, so that the electronic device executes the image synthesis processing method according to any one of claims 1 to 14.