Image generation methods, chips, electronic devices and storage media

By generating depth images based on real-world parallax distribution, the problem of insufficient binocular depth data quality in existing technologies is solved, resulting in training data that is closer to the real world and reducing costs.

CN115578300BActive Publication Date: 2026-05-26SPREADTRUM COMMUNICATION (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SPREADTRUM COMMUNICATION (SHANGHAI) CO LTD
Filing Date
2022-10-25
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, the binocular depth data used for binocular vision training cannot effectively meet the needs of mobile terminals, resulting in insufficient training data quality.

Method used

By acquiring the original image set, a baseline disparity value is generated based on the disparity distribution of the real world. The original images are processed to generate depth images, including a left depth image, a right depth image, and a disparity map. Images captured by the mobile terminal are combined with images of people, objects, and backgrounds to fuse the target object and generate binocular depth training data that is closer to the real world.

Benefits of technology

It improves the quality of binocular depth training data, making it more consistent with the distribution of depth data in the real world, reducing dependence on complex hardware, and lowering training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115578300B_ABST
    Figure CN115578300B_ABST
Patent Text Reader

Abstract

This application provides an image generation method, chip, electronic device, and storage medium. The method includes: acquiring an original image set, the original image set including multiple original images; processing the original images in the original image set based on a reference disparity value to obtain a depth image; wherein the reference disparity value is generated based on the disparity distribution of the real world, and the depth image includes a left depth image, a right depth image, and a disparity map. The method provided by this application can improve the quality of depth images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning, and more particularly to an image generation method, a chip, an electronic device, and a storage medium. Background Technology

[0002] With the development of artificial intelligence (AI) technology and the continuous improvement of the hardware capabilities of mobile terminals (such as mobile phones and tablets), AI technologies, such as deep learning, are increasingly being applied in mobile terminals, especially in binocular vision technology.

[0003] However, the application of deep learning has also brought some challenges, such as how to train a model with strong generalization capabilities using limited training data, and how to improve the quality of training data. For optimal binocular vision training, a large amount of high-quality binocular depth data is required.

[0004] Currently, the binocular depth data used for binocular vision training is usually synthetic data or data collected by vehicles. These data cannot adequately meet the needs of binocular vision in mobile terminals. Therefore, improving the quality of the aforementioned binocular depth training data is an urgent problem to be solved. Summary of the Invention

[0005] This application provides an image generation method, chip, electronic device, and storage medium that can improve the quality of depth-of-field images.

[0006] In a first aspect, this application provides an image generation method, including:

[0007] Obtain the original image set, which includes multiple original images;

[0008] The original images in the original image set are processed based on the reference disparity to obtain a depth image;

[0009] The reference disparity is generated based on the disparity distribution in the real world, and the depth image includes a left depth image, a right depth image, and a disparity map.

[0010] In this application, depth training data is generated based on the physical meaning of depth data, which makes the generated binocular depth training data more closely match the distribution of depth data in the real world, thereby improving the quality of binocular depth training data.

[0011] In one possible implementation, the multiple original images are obtained by a mobile terminal.

[0012] In one possible implementation, the plurality of original images includes a plurality of person images, a plurality of object images, and a plurality of background images.

[0013] In one possible implementation, processing the original images in the original image set based on the reference disparity to obtain a depth image includes:

[0014] Extract the target person from one or more images of people;

[0015] Extract the target object from one or more object images;

[0016] The target person and the target object are integrated into the first target background image to obtain the left depth image, wherein the first target background image is any one of the plurality of background images;

[0017] The left depth image is processed based on the reference disparity value to obtain the right depth image and the disparity image.

[0018] In one possible implementation, the reference disparity includes person disparity, object disparity, and background disparity. The step of processing the left depth image based on the reference disparity to obtain the right depth image includes:

[0019] A second target background image is generated based on the first target background image and the background disparity.

[0020] The target person is integrated into the second target background image based on the person disparity, and the target object is integrated into the second target background image based on the object disparity, to obtain the depth-of-field right image.

[0021] In one possible implementation, after integrating the target person into the second target background image based on the person disparity and integrating the target object into the second target background image based on the object disparity, the method further includes:

[0022] The target person and the target object integrated into the background image of the second target are offset based on the position offset information;

[0023] The position offset information is determined by the offset of the position of the background in the second target background image relative to the position of the background in the first target background image.

[0024] In one possible implementation, the reference disparity includes person disparity, object disparity, and background disparity. The step of processing the left depth image based on the reference disparity to obtain the right depth image includes:

[0025] The first target background image is stretched to obtain the second target background image;

[0026] The target person is integrated into the second target background image based on the person disparity, and the target object is integrated into the second target background image based on the object disparity, to obtain the depth-of-field right image.

[0027] In one possible implementation, after integrating the target person into the second target background image based on the person disparity and integrating the target object into the second target background image based on the object disparity, the method further includes:

[0028] The target person and the target object integrated into the background image of the second target are offset based on the position offset information;

[0029] The position offset information is determined by the offset and stretching ratio of the position of the background in the second target background image relative to the position of the background in the first target background image.

[0030] In one possible implementation, processing the left depth image based on the reference disparity value to obtain the disparity map includes:

[0031] Obtain the positions of the target person, the target object, and the background in the left depth image;

[0032] A disparity map is generated based on the positions of the target person, the target object, and the background in the left depth image, as well as the reference disparity value.

[0033] In one possible implementation, the reference disparity is randomly generated based on the disparity distribution in the real world.

[0034] Secondly, this application provides a chip including one or more functional modules for performing the image generation method as described in the first aspect.

[0035] Thirdly, this application provides an electronic device, including: a processor and a memory, the memory being used to store a computer program; the processor being used to run the computer program to implement the image generation method as described in the first aspect.

[0036] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to implement the image generation method as described in the first aspect.

[0037] Fifthly, this application provides a computer program that, when run on a processor of an electronic device, causes the electronic device to perform the image generation method described in the first aspect.

[0038] In one possible design, the program in the fifth aspect can be stored wholly or partially on a storage medium packaged with the processor, or it can be stored wholly or partially on a memory not packaged with the processor. Attached Figure Description

[0039] Figure 1 A flowchart illustrating an embodiment of the image generation method provided in this application;

[0040] Figure 2 This is a schematic diagram of the chip structure provided in an embodiment of this application;

[0041] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0042] In this embodiment of the application, unless otherwise stated, the character " / " indicates that the preceding and following objects are in an OR relationship. For example, A / B can represent A or B. "AND / OR" describes the relationship between the associated objects, indicating that three relationships can exist. For example, A AND / OR B can represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0043] It should be noted that the terms "first" and "second" used in the embodiments of this application are used only for distinguishing descriptive purposes and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated, nor should they be construed as indicating or implying order.

[0044] In the embodiments of this application, "at least one" refers to one or more items, and "more than one" refers to two or more items. Furthermore, "at least one of the following" or similar expressions refer to any combination of these items, which may include any combination of a single item or a plurality of items. For example, at least one of A, B, or C can represent: A, B, C, A and B, A and C, B and C, or A, B, and C. Each of A, B, and C can be an element itself or a set containing one or more elements.

[0045] In this application, terms such as "exemplary," "in some embodiments," and "in another embodiment" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.

[0046] In the embodiments of this application, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction, their meanings are consistent. Similarly, in the embodiments of this application, "communication" and "transmission" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction, their meanings are consistent. For example, transmission can include sending and / or receiving, and can be a noun or a verb.

[0047] In the embodiments of this application, the term "equal to" can be used in conjunction with "greater than" to apply to technical solutions employing the condition of "greater than", and can also be used in conjunction with "less than" to apply to technical solutions employing the condition of "less than". It should be noted that when "equal to" is used with "greater than", it cannot be used with "less than"; and when "equal to" is used with "less than", it cannot be used with "greater than".

[0048] The following explanations of some terms used in the embodiments of this application are provided to facilitate understanding by those skilled in the art.

[0049] 1. Parallax. In the embodiments of this application, parallax refers to the change and difference in the position of the object in the field of view when the same object is observed from two different positions.

[0050] 2. Depth of field. In the embodiments of this application, depth of field refers to the range in which the area in front of and behind a point remains in focus when the focal length is pointed at that point.

[0051] With the continuous development of deep learning, its application in mobile terminal (e.g., mobile phones, tablets, etc.) photography is also increasing. Understandably, deep learning is driven by training data, which leads to an increasing demand for training data for related AI models, especially for binocular depth-of-field training data.

[0052] Traditional methods for generating binocular depth training data include two types: hardware- and software-based data synthesis, and purely software-based data synthesis.

[0053] For hardware and software-based data synthesis methods, high-cost and complex hardware such as corresponding optical and imaging systems is required.

[0054] For pure software data synthesis methods, image synthesis is usually performed through common image transformations, such as image flipping, image cropping, image rotation, and image translation. However, these methods cannot change the original disparity map and cannot meet the needs of binocular vision on mobile terminals. They only enhance the robustness of the original data and do not fundamentally improve the quality of the training data.

[0055] Based on the above problems, this application proposes an image generation method that is applied to an electronic device, which may be a personal computer or a server, and this application does not impose any special limitations on it.

[0056] In some alternative embodiments, the electronic device may also be a terminal device, which may be a fixed terminal or a mobile terminal. This application does not impose any special limitations on this.

[0057] Now combined Figure 1 The image generation method provided in the embodiments of this application will be described by way of example. Figure 1 The image generation method shown can be used to generate multiple binocular depth training data. It is understood that the binocular depth training data can be binocular depth images, wherein a binocular depth image can include a left depth image, a right depth image and a disparity image.

[0058] For ease of explanation, this application uses the generation of a single binocular depth image as an example for illustrative purposes. Multiple binocular depth images can be generated in the same manner as the single binocular depth image described above, and will not be repeated below.

[0059] Figure 1 A flowchart illustrating an embodiment of the image generation method provided in this application specifically includes the following steps:

[0060] Step 101: Obtain the original image set.

[0061] Specifically, the original image set may include multiple original images, which may include three types of images, such as images of people, images of objects, and background images.

[0062] It is understood that the original image is the original captured image, which may be unprocessed. This original image can be obtained by a mobile terminal such as a smartphone or tablet, or by other camera devices. This application does not impose any special limitations on the method of capturing the original image. Preferably, the original image can be an image captured by a mobile terminal, thereby better conforming to the depth-of-field data under mobile terminal shooting conditions.

[0063] Furthermore, the original image set can be obtained directly from the shooting device, or from other storage devices or the Internet, and this application embodiment does not impose any special limitations on this.

[0064] Step 102: Extract the target person from one or more person images in the original image set.

[0065] Specifically, one or more images of people can be randomly selected from the original image set, and the target person can be extracted from the one or more images of people randomly selected above.

[0066] The target person can be extracted using a preset portrait segmentation algorithm. For example, the preset portrait segmentation algorithm can separate the outline of a person from an image, and then the image within the outline area is used as the target person. It is understood that the preset portrait segmentation algorithm can be any existing algorithm, and this application embodiment does not impose any special limitations on it. Furthermore, the target person may include one or more persons, and this application embodiment does not impose any special limitations on it.

[0067] Step 103: Extract the target object from one or more object images in the original image set.

[0068] Specifically, one or more object images can be randomly selected from the original image set, and the target object can be extracted from the one or more object images randomly selected above.

[0069] The target object can be extracted using a preset object segmentation algorithm. For example, the preset object segmentation algorithm can separate the outline of the object from the object image, and then the image within the outline region is used as the target object. It is understood that the preset object segmentation algorithm can be an existing algorithm, such as the adaptive binary method or the watershed algorithm, and this application embodiment does not impose any special limitations on it. Furthermore, the target object may include one or more objects, and this application embodiment does not impose any special limitations on it.

[0070] It is understood that step 103 can be executed before step 102, after step 102, or simultaneously with step 102. This application embodiment does not impose any special limitations on this.

[0071] Step 104: Obtain the first target background image, and integrate the target person and target object into the first target background image to obtain the first binocular depth sub-image.

[0072] Specifically, the first target background image can be a background image randomly selected from the original image set.

[0073] Once the first target background image is determined, the target person and target object can be merged into the first target background image to obtain the first binocular depth sub-image.

[0074] Among them, the first binocular depth sub-image can be considered as the left depth image in the binocular depth image.

[0075] Understandably, in order to improve the diversity and volume of training data, when fusing the target person and the target object into the aforementioned first target background image, the target person and the target object can be placed at random positions in the first target background image. This allows a large amount of training data to be generated with less raw data.

[0076] It is understandable that after the target person and target object are merged into the first target background image, the positions of the target person and target object in the first binocular depth sub-image can be recorded so that when constructing the right depth image, the positions of the target person and target object in the right depth image can be determined based on their positions in the left depth image.

[0077] Step 105: Obtain the second binocular depth sub-image based on the first binocular depth sub-image.

[0078] Specifically, after obtaining the first binocular depth sub-image, that is, after obtaining the left depth image, a second binocular depth sub-image can also be obtained, wherein the second binocular depth sub-image can be considered as the right depth image in the binocular depth image.

[0079] For example, the second binocular depth sub-image described above can be obtained in the following two ways.

[0080] Method 1, Fixed parallax method

[0081] First, based on the physical meaning of the depth image, disparity quantities corresponding to people, objects, and backgrounds can be randomly generated according to the disparity distribution in the real world. For example, there are person disparity, object disparity, and background disparity quantities. These can be considered as reference disparity quantities, which conform to the objective laws of human vision. Next, a second target background image can be generated based on the first target background image and the background disparity quantities. It can be understood that the first target background image can be considered the background image of the left depth image, and the second target background image can be considered the background image of the right depth image. The specific method for generating images based on disparity quantities can be found in existing technologies and will not be elaborated here.

[0082] It is understandable that after the second target background image is generated based on the first target background image and the background disparity, the position of the background in the second target background image is offset relative to the position of the background in the first target background image. At this time, the position offset information can be determined based on the offset of the position of the background in the second target background image relative to the position of the background in the first target background image. The position offset information may include the position offset direction and the position offset amount.

[0083] Next, based on the positions of the target person and target object recorded in step 104 in the first target background image, and the aforementioned position offset information, the target person and target object can be fused into the second target background image, thereby obtaining a second binocular depth sub-image. The method for fusion of the target person and target object into the second target background image can be as follows: first, the target person is placed into the second binocular depth sub-image based on their disparity, and the target object is placed into the second binocular depth sub-image based on its disparity; then, the target person and target object already fused into the second binocular depth sub-image are offset based on the aforementioned position offset information, thereby obtaining the second binocular depth sub-image.

[0084] Method 2, Gradual Parallax Method

[0085] First, based on the physical meaning of depth images, disparity values ​​corresponding to people, objects, and backgrounds can be randomly generated according to the disparity distribution in the real world. For example, disparity values ​​for people, objects, and backgrounds can be considered as reference disparity values, which conform to the objective laws of human vision.

[0086] Next, since only horizontal parallax is considered when taking into account parallax, and vertical parallax is not considered, the first target background image can be stretched horizontally to obtain a background image with a gradient parallax effect. Because the size of the stretched background image changes, the stretched first target background image can be cropped to obtain a second target background image, where the size of the second target background image is consistent with the size of the first target background image.

[0087] After obtaining the background image of the second target, the corresponding position offset information can also be obtained.

[0088] Understandably, the parallax effect produced by stretching is gradual, and this gradual effect must be taken into account in the position offset information. For example, the position offset amount can be adjusted according to the stretching ratio.

[0089] Then, the target person and target object can be fused into the second target background image, thereby obtaining the second binocular depth sub-image. The specific method for fusion of the target person and target object into the second target background image can be found in the relevant descriptions in the above embodiments, and will not be repeated here.

[0090] Step 106: Obtain the third binocular depth sub-image based on the first binocular depth sub-image.

[0091] Specifically, after obtaining the first binocular depth sub-image, that is, after obtaining the left depth image, a third binocular depth sub-image can also be obtained. The third binocular depth sub-image can be considered as the disparity map in the binocular depth image.

[0092] The method for obtaining a third binocular depth sub-image based on a first binocular depth sub-image can be as follows: First, obtain the positions of the outlines of people, objects and background in the first binocular depth sub-image. Then, based on the disparity of people, objects and background, as well as the positions of the outlines of people, objects and background, generate the corresponding disparity map.

[0093] It is understood that step 106 can be executed before step 105, after step 105, or simultaneously with step 105. This application embodiment does not impose any special limitations on this.

[0094] Once the first, second, and third binocular depth sub-images are acquired, the generation of a binocular depth image can be considered complete. In other words, this binocular depth image can include the aforementioned first, second, and third binocular depth sub-images. The generation of other binocular depth images can follow the methods described in the above embodiments, and will not be repeated here. This allows for the generation of multiple binocular depth data sets.

[0095] In this embodiment, depth training data is generated based on the physical meaning of depth data, which makes the generated binocular depth training data more closely resemble the distribution of depth data in the real world, thereby improving the quality of the binocular depth training data. Furthermore, the method for generating binocular depth training data in this embodiment does not rely on complex and expensive hardware such as optical devices and imaging devices, thus reducing training costs.

[0096] Figure 2 This is a schematic diagram of the structure of one embodiment of the chip in this application, as shown below. Figure 2 As shown, the chip 20 may include: an acquisition module 21 and an image generation module 22; wherein,

[0097] Acquisition module 21 is used to acquire an original image set, which includes multiple original images;

[0098] Image generation module 22 is used to process the original images in the original image set based on the reference disparity to obtain a depth image;

[0099] The reference disparity is generated based on the disparity distribution in the real world, and the depth image includes a left depth image, a right depth image, and a disparity map.

[0100] In one possible implementation, the multiple original images are obtained by a mobile terminal.

[0101] In one possible implementation, the plurality of original images includes a plurality of person images, a plurality of object images, and a plurality of background images.

[0102] In one possible implementation, the image generation module 22 is specifically used to extract the target person from one or more person images;

[0103] Extract the target object from one or more object images;

[0104] The target person and the target object are integrated into the first target background image to obtain the left depth image, wherein the first target background image is any one of the plurality of background images;

[0105] The left depth image is processed based on the reference disparity value to obtain the right depth image and the disparity image.

[0106] In one possible implementation, the reference disparity includes human disparity, object disparity, and background disparity. The image generation module 22 is specifically used to generate a second target background image based on the first target background image and the background disparity.

[0107] The target person is integrated into the second target background image based on the person disparity, and the target object is integrated into the second target background image based on the object disparity, to obtain the depth-of-field right image.

[0108] In one possible implementation, the chip 20 further includes:

[0109] The offset module is used to offset the target person and the target object that are integrated into the second target background image based on the position offset information;

[0110] The position offset information is determined by the offset of the position of the background in the second target background image relative to the position of the background in the first target background image.

[0111] In one possible implementation, the image generation module 22 is specifically used to stretch the first target background image to obtain the second target background image;

[0112] The target person is integrated into the second target background image based on the person disparity, and the target object is integrated into the second target background image based on the object disparity, to obtain the depth-of-field right image.

[0113] In one possible implementation, the aforementioned offset module is specifically used to offset the target person and the target object that are integrated into the second target background image based on the position offset information.

[0114] The position offset information is determined by the offset and stretching ratio of the position of the background in the second target background image relative to the position of the background in the first target background image.

[0115] In one possible implementation, the image generation module 22 is specifically used to obtain the positions of the target person, the target object and the background in the left depth image;

[0116] A disparity map is generated based on the positions of the target person, the target object, and the background in the left depth image, as well as the reference disparity value.

[0117] In one possible implementation, the reference disparity is randomly generated based on the disparity distribution in the real world.

[0118] Figure 2 The chip 20 provided in the illustrated embodiment can be used to execute the technical solution of the method embodiment shown in this application. Its implementation principle and technical effect can be further referred to the relevant description in the method embodiment.

[0119] It should be understood that the division of the various modules in chip 20 above is merely a logical functional division. In actual implementation, they can be fully or partially integrated onto a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented in software via processing element calls, while others are implemented in hardware. For example, the detection module can be a separate processing element or integrated into a chip within the electronic device. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or implemented independently. During implementation, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0120] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, these modules can be integrated together as a System-On-a-Chip (SOC).

[0121] In the above embodiments, the processor may include, for example, a CPU, DSP, microcontroller, or digital signal processor, and may also include a GPU, embedded neural network processing unit (NPU), and image signal processor (ISP). The processor may also include necessary hardware accelerators or logic processing hardware circuits, such as an ASIC, or one or more integrated circuits for controlling the execution of the program in this application. Furthermore, the processor may have the function of operating one or more software programs, which may be stored in a storage medium.

[0122] The following is combined with Figure 3 The exemplary electronic devices provided in the embodiments of this application are further described. Figure 3 A schematic diagram of the structure of electronic device 300 is shown.

[0123] The aforementioned electronic device 300 may include: at least one processor; and at least one memory communicatively connected to the processor, wherein: the memory stores program instructions executable by the processor, and the processor can execute this application by calling the program instructions. Figure 1 The method provided in the illustrated embodiment.

[0124] Figure 3 A block diagram is shown of an exemplary electronic device 300 suitable for implementing embodiments of this application. Figure 3 The electronic device 300 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0125] like Figure 3As shown, the electronic device 300 is presented in the form of a general-purpose computing device. The components of the electronic device 300 may include, but are not limited to: one or more processors 310, memory 320, communication bus 340 connecting different system components (including memory 320 and processor 310), and communication interface 330.

[0126] Communication bus 340 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MAC) buses, Enhanced ISA buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses.

[0127] Electronic device 300 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, removable and non-removable media.

[0128] Memory 320 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. Although Figure 3 As not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disc drive for reading and writing to a removable non-volatile optical disc (e.g., a compact disc read-only memory (CD-ROM), a digital video disc read-only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to the communication bus 340 via one or more data media interfaces. The memory 320 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this application.

[0129] A program / utility having a set (at least one) of program modules can be stored in memory 320. Such program modules include—but are not limited to—an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules typically perform the functions and / or methods described in the embodiments of this application.

[0130] Electronic device 300 can also communicate with one or more external devices (e.g., keyboard, pointing device, display, etc.), and with one or more devices that enable a user to interact with the electronic device, and / or with any device that enables the electronic device to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed through communication interface 330. Furthermore, electronic device 300 can also communicate through a network adapter (… Figure 3 (Not shown) communicates with one or more networks (e.g., Local Area Network (LAN), Wide Area Network (WAN), and / or public networks, such as the Internet). The aforementioned network adapter can communicate with other modules of the electronic device via communication bus 340. It should be understood that, although... Figure 3 As not shown, other hardware and / or software modules may be used in conjunction with electronic device 300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, Redundant Arrays of Independent Drives (RAID) systems, tape drives, and data backup storage systems.

[0131] The processor 310 executes various functional applications and data processing by running programs stored in the memory 320, such as implementing the methods provided in the embodiments of this application.

[0132] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 300. In other embodiments of this application, the electronic device 300 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0133] In the above embodiments, the processor may include, for example, a CPU, DSP, microcontroller, or digital signal processor, and may also include a GPU, embedded neural network processing unit (NPU), and image signal processor (ISP). The processor may also include necessary hardware accelerators or logic processing hardware circuits, such as an ASIC, or one or more integrated circuits for controlling the execution of the program in this application. Furthermore, the processor may have the function of operating one or more software programs, which may be stored in a storage medium.

[0134] This application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods provided in the embodiments shown in this application.

[0135] This application also provides a computer program product, which includes a computer program that, when run on a computer, causes the computer to perform the methods provided in the embodiments shown in this application.

[0136] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, A and B simultaneously, or B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0137] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0138] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0139] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0140] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application should be determined by the protection scope of the claims.

Claims

1. An image generation method, characterized in that, The method includes: Obtain the original image set, which includes multiple original images; The original images in the original image set are processed based on the reference disparity to obtain a depth image; The reference disparity is generated based on the disparity distribution of the real world, and the depth image includes a left depth image, a right depth image, and a disparity map. The multiple original images include multiple images of people, multiple images of objects, and multiple background images. The process of processing the original images in the original image set based on a reference disparity to obtain a depth image includes: Extract the target person from one or more images of people; Extract the target object from one or more object images; The target person and the target object are integrated into the first target background image to obtain the left depth image, wherein the first target background image is any one of the plurality of background images; The left depth image is processed based on the reference disparity value to obtain the right depth image and the disparity image.

2. The method according to claim 1, characterized in that, The multiple original images were captured by a mobile terminal.

3. The method according to claim 1, characterized in that, The reference disparity includes person disparity, object disparity, and background disparity. The process of processing the left depth-of-field image based on the reference disparity to obtain the right depth-of-field image includes: A second target background image is generated based on the first target background image and the background disparity. The target person is integrated into the second target background image based on the person disparity, and the target object is integrated into the second target background image based on the object disparity, to obtain the depth-of-field right image.

4. The method according to claim 3, characterized in that, After incorporating the target person into the second target background image based on the person disparity and incorporating the target object into the second target background image based on the object disparity, the method further includes: The target person and the target object integrated into the background image of the second target are offset based on the position offset information; The position offset information is determined by the offset of the position of the background in the second target background image relative to the position of the background in the first target background image.

5. The method according to claim 1, characterized in that, The reference disparity includes person disparity, object disparity, and background disparity. The process of processing the left depth-of-field image based on the reference disparity to obtain the right depth-of-field image includes: The first target background image is stretched to obtain the second target background image; The target person is integrated into the second target background image based on the person disparity, and the target object is integrated into the second target background image based on the object disparity, to obtain the depth-of-field right image.

6. The method according to claim 5, characterized in that, After incorporating the target person into the second target background image based on the person disparity and incorporating the target object into the second target background image based on the object disparity, the method further includes: The target person and the target object integrated into the background image of the second target are offset based on the position offset information; The position offset information is determined by the offset and stretching ratio of the position of the background in the second target background image relative to the position of the background in the first target background image.

7. The method according to claim 1, characterized in that, The process of processing the left depth image based on the reference disparity value to obtain the disparity map includes: Obtain the positions of the target person, the target object, and the background in the left depth image; A disparity map is generated based on the positions of the target person, the target object, and the background in the left depth image, as well as the reference disparity value.

8. The method according to any one of claims 1-7, characterized in that, The baseline disparity is randomly generated based on the real-world disparity distribution.

9. A chip, characterized in that, include: The acquisition module is used to acquire a raw image set, which includes multiple raw images; An image generation module is used to process the original images in the original image set based on a reference disparity value to obtain a depth image; The reference disparity is generated based on the disparity distribution of the real world, and the depth image includes a left depth image, a right depth image, and a disparity map. The plurality of original images include a plurality of person images, a plurality of object images and a plurality of background images. The image generation module is further configured to extract a target person from one or more person images; extract a target object from one or more object images; and integrate the target person and the target object into a first target background image to obtain the left depth image, wherein the first target background image is any one of the plurality of background images. The left depth image is processed based on the reference disparity value to obtain the right depth image and the disparity image.

10. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store a computer program; the processor being used to run the computer program to implement the image generation method as described in any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a computer, implements the image generation method as described in any one of claims 1-8.