Data generation method and device, readable medium, electronic equipment and product
By determining the coordinate mapping relationship and identifying the occlusion area based on the spatial depth information and viewpoint position relationship of monocular images, a mask image associated with the first viewpoint is generated, which solves the problems of poor dataset quality and high computational complexity in existing technologies, and achieves efficient and accurate data generation.
Patent Information
- Application Number
- CN202511434569.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-12-23
AI Technical Summary
Existing technologies rely on complex algorithms when constructing new perspective image filling datasets, which leads to error accumulation and high computational complexity, limiting the quality of the dataset and the diversity of acquisition scenarios. Furthermore, the generated masks are prone to misalignment.
By using spatial depth information and viewpoint position relationship based on monocular images, coordinate mapping relationship is determined, and occlusion areas are identified using local rate of change, generating a mask image associated with the first viewpoint as a sample data pair.
It improves the flexibility and quality of sample data pair generation, simplifies the data generation process, increases processing efficiency, and generates more accurate mask images with natural boundaries.
Smart Images

Figure CN121190602A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more specifically, to a data generation method, apparatus, readable medium, electronic device, and product. Background Technology
[0002] New perspective image filling technology is an important technique in the field of computer vision and is widely used in various image processing tasks. The implementation of filling technology usually depends on a dataset containing the source view image, occlusion mask and the new view image after filling. Therefore, the quality of the dataset is of great significance to the implementation of filling technology. Summary of the Invention
[0003] This section is provided to briefly introduce the concepts, which will be described in detail in the Detailed Description section later. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, this disclosure provides a data generation method, the method comprising: Based on the first image from a first-view perspective, determine the spatial depth information corresponding to the first image; Based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint, a first coordinate mapping relationship from the second viewpoint to the first viewpoint is determined; By determining the local rate of change of the first coordinate mapping relationship, the occluded area in the first image is identified; Based on the occlusion area, a mask image corresponding to the first image is generated; The first image and the mask image are identified as a sample data pair associated with the first viewpoint.
[0005] Secondly, this disclosure provides a data generation apparatus, the apparatus comprising: The first determining module is used to determine the spatial depth information corresponding to the first image based on the first image from a first perspective. The second determining module is used to determine a first coordinate mapping relationship from the second viewpoint to the first viewpoint based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint. The third determining module is used to identify the occluded area in the first image by determining the local rate of change of the first coordinate mapping relationship; The generation module is used to generate a mask image corresponding to the first image based on the occluded area; The fourth determining module is used to determine the first image and the mask image as a sample data pair associated with the first viewpoint.
[0006] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect of this disclosure.
[0007] Fourthly, this disclosure provides an electronic device, comprising: A storage device having at least one computer program stored thereon; At least one processing means is configured to execute the at least one computer program in the storage device to implement the steps of the method described in the first aspect of this disclosure.
[0008] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect of this disclosure.
[0009] The above technical solution determines the spatial depth information of the first image from a first-viewpoint. Based on the spatial depth information and the relative positional relationship between the second and first viewpoints, it determines the first coordinate mapping relationship from the second viewpoint to the first viewpoint. By determining the local rate of change of the first coordinate mapping relationship, it identifies occlusion regions in the first image and generates a mask image corresponding to the first image based on these occlusion regions. The first image and the mask image are then used as sample data pairs associated with the first viewpoint. Therefore, regions potentially occluded in the first viewpoint can be determined using the second viewpoint and the coordinate mapping relationship between the two viewpoints. Mask image generation can be achieved based solely on a monocular image, improving the flexibility of sample data pair generation and making it applicable to more data generation scenarios. Furthermore, utilizing the local change characteristics of the coordinate mapping to determine occlusion regions results in more accurate mask images with more natural and smoother boundaries. Additionally, the data generation process does not rely on complex algorithms, improving processing efficiency. Based on this, the generation quality and efficiency of sample data pairs can be effectively improved, providing more and higher-quality usable data for data imputation techniques.
[0010] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1This is a flowchart of a data generation method provided according to one embodiment of the present disclosure; Figure 2 This is a block diagram of a data generation apparatus provided according to one embodiment of the present disclosure; Figure 3 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0012] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0013] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0015] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0016] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0017] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0018] All actions involving the acquisition of signals, information, or data in this disclosure are carried out in accordance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.
[0019] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0020] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0021] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0022] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0023] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0024] As described in the background section, image filling technology relies on a dataset containing an original viewpoint image, an occlusion mask, and a new viewpoint image after filling. In related technologies, constructing such a dataset typically requires acquiring continuous video sequences under camera motion, preprocessing them using algorithms such as SLAM (Simultaneous Localization and Mapping) to obtain camera motion trajectory and scene depth information, and then obtaining image data from the new viewpoint through projection transformation. However, this approach has several shortcomings. First, the algorithms used for preprocessing are prone to accumulating errors, leading to misalignment of the generated new viewpoint image and mask, resulting in poor dataset quality. Second, the algorithms used for preprocessing have high requirements for the input video, limiting the diversity of scenes that can be covered by data acquisition, thus restricting data acquisition. Finally, the algorithms used for preprocessing typically have high computational complexity, resulting in low data processing efficiency.
[0025] To address the aforementioned technical problems, this disclosure provides a data generation method, apparatus, readable medium, electronic device, and product.
[0026] Figure 1 This is a flowchart of a data generation method provided according to one embodiment of this disclosure. For example... Figure 1 As shown, the method provided in this disclosure may include steps 11 to 15.
[0027] In step 11, the spatial depth information corresponding to the first image is determined based on the first image from the first perspective.
[0028] Optionally, the first image can be a monocular image or a frame of monocular video, corresponding to a first viewpoint. The image content of the first image can be considered as an image obtained by the camera capturing the corresponding 3D scene from the first viewpoint.
[0029] Based on the first image, the spatial depth information corresponding to the first image can be determined. This spatial depth information reflects the relative spatial positions of objects in the corresponding 3D scene. The spatial depth information corresponding to the first image can be information that characterizes the spatial depth of each pixel in the first image.
[0030] Optionally, the spatial depth information can be a depth map, which may include the depth value corresponding to each pixel in the first image.
[0031] Optionally, the spatial depth information can be a relative disparity map, which includes the relative disparity value corresponding to each pixel in the first image, wherein the relative disparity value of a pixel can be inversely proportional to the depth of that pixel.
[0032] For example, spatial depth information can be obtained by processing the first image using a pre-generated monocular depth estimation model.
[0033] In step 12, based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint, the first coordinate mapping relationship from the second viewpoint to the first viewpoint is determined.
[0034] The second-person perspective can be selected according to actual needs. The camera position corresponding to the second-person perspective can be based on the position of the camera corresponding to the first-person perspective. x , y , z It is obtained by moving in at least one of the directions. For example, starting from the position of the camera corresponding to the first viewpoint, the position reached after moving the camera horizontally to the left by a certain distance is taken as the camera position of the second viewpoint.
[0035] In one possible implementation, step 12 may include the following steps: Based on spatial depth information and relative positional relationships, determine the second coordinate mapping relationship from the first perspective to the second perspective; Based on the second coordinate mapping relationship, determine the corresponding mapping coordinates of each pixel in the second viewpoint in the first image to obtain the first coordinate mapping relationship.
[0036] The second perspective corresponds to a camera position (i.e., camera coordinates), and the first perspective also corresponds to a camera position. The relative positions mentioned above can be determined by the relative positions between the camera coordinates corresponding to the first and second perspectives, respectively.
[0037] Based on the spatial depth information corresponding to the first image and the relative positional relationship between the second and first viewpoints, a disparity mapping function can be used to determine the positive coordinate mapping relationship from the first to the second viewpoint, which serves as the second coordinate mapping relationship. For each pixel in the first image (i.e., each pixel coordinate), the disparity mapping function can calculate its theoretical corresponding coordinates in the second viewpoint based on the spatial depth information of that pixel and the aforementioned relative positional relationship.
[0038] The second coordinate mapping relationship can be represented by two graphs (or matrices) of the same size as the first image, and may include a new coordinate mapping for storing the coordinates of each pixel in the first image in the second viewpoint. x A graph (or matrix) of coordinates, and a new coordinate system for storing the corresponding pixel in the second viewpoint for each pixel in the first image. y A graph (or matrix) of coordinates. For example, a pixel with coordinates (1,1) in the first viewpoint might have coordinates (1.7,1.2) in the second viewpoint after being calculated by the disparity mapping function.
[0039] Based on the second coordinate mapping relationship, for each pixel in the second viewpoint (i.e., each pixel grid in the second viewpoint), the corresponding mapped coordinates in the first image can be determined. For example, a reverse remapping technique can be used to traverse each pixel grid in the second viewpoint and use an interpolation algorithm to find the corresponding pixel coordinates in the first viewpoint within the second coordinate mapping relationship, thus constructing a reverse mapping relationship from the second viewpoint to the first viewpoint, i.e., the first coordinate mapping relationship. Similar to the second coordinate mapping relationship, the first coordinate mapping relationship can also include the new coordinates corresponding to each pixel in the second viewpoint in the first viewpoint. x A graph (or matrix) of coordinates, and the new coordinates of each pixel in the second viewpoint corresponding to the first viewpoint. y A graph (or matrix) of coordinates.
[0040] In step 13, the occlusion region in the first image is identified by determining the local rate of change of the first coordinate mapping relationship.
[0041] In this disclosure, the local rate of change can be used to quantify the degree of local deformation in the first coordinate mapping relationship. It can characterize the degree of positional change of the mapped coordinates in the first image from the first viewpoint after the image plane of the second viewpoint moves a certain distance. When the camera moves from the first viewpoint to the second viewpoint, a background area occluded by the foreground in the first viewpoint may become visible in the second viewpoint. This change in three-dimensional space causes nearby pixels in the image plane corresponding to the second viewpoint to be mapped to unrelated or even distant coordinates in the first image corresponding to the first viewpoint. This will manifest as a jump in the first coordinate mapping relationship, that is, a sudden increase in the local rate of change. Based on this, by analyzing the local rate of change in the first coordinate mapping relationship and analyzing the abrupt changes in the local rate of change, it is possible to identify a background area that is invisible in the first viewpoint but visible in the second viewpoint due to the change in viewpoint, i.e., the occluded area in the first image.
[0042] In one possible implementation, step 13 may include the following steps: Based on the first coordinate mapping relationship, determine the local rate of change at each pixel in the second viewpoint; Based on the local rate of change at each pixel in the second viewpoint, regions that are invisible in the first viewpoint but visible in the second viewpoint are identified as occlusion regions.
[0043] In one possible implementation, the local rate of change can be determined by multiple rates of change of a pixel along different preset directions when the pixel is mapped from a second viewpoint to a first viewpoint. These multiple rates of change along different preset directions can include those along... x , y The four rates of change in direction, that is, when the second-view pixel is in x When moving in a direction, the corresponding pixel in the first-person view x rate of change of coordinates y The rate of change of coordinates, and when the second-view pixel is in y When moving in a direction, the corresponding pixel in the first-person view x rate of change of coordinates y The rate of change of coordinates. For example, each rate of change can be determined by calculating the partial derivative of the first coordinate mapping relationship in the corresponding direction. After determining the multiple rates of change along different preset directions, since they include rates of change in different directions (i.e., multiple rate of change values), they can be integrated into a single value as a local rate of change. For example, a weighted average can be used to integrate the multiple rates of change in different preset directions to obtain the local rate of change.
[0044] Optionally, the local rate of change can be represented using the norm of the Jacobian matrix.
[0045] In another possible implementation, the local rate of change can be determined based on the local variance or standard deviation. For example, for each pixel in the second viewpoint, a window of a specified size (e.g., 2×2, 3×3, etc.) can be selected centered on that pixel, and the variance or standard deviation can be calculated using the coordinate values at the corresponding window location in the first coordinate mapping relationship. As mentioned earlier, the first coordinate mapping relationship includes two graphs (or matrices), so the variance or standard deviation corresponding to the window can be determined for each graph (or matrix), and then the average or maximum value can be calculated based on the two variances or standard deviations to serve as the local rate of change.
[0046] Using the above method, the local rate of change at each pixel on the image plane in the second viewpoint can be determined. Then, based on the local rate of change, the region that is not visible in the first viewpoint but is visible in the second viewpoint can be identified as the occlusion region.
[0047] As mentioned earlier, the magnitude of the local rate of change characterizes the degree of deformation at a pixel location. A larger local rate of change indicates a more severe and greater degree of deformation after the pixel is mapped from the second viewpoint to the first viewpoint. Normally, when the viewpoint changes, the coordinate mapping relationship corresponding to continuously visible surfaces in the same scene should change smoothly. However, when occlusion exists, this continuity is broken, resulting in discontinuous jumps in the mapping relationship. For example, an object may be visible in the second viewpoint, but in the first viewpoint, the object may be occluded by a foreground object due to the change in viewpoint. Based on this, the local rate of change can be used to determine whether a pixel location is occluded, thereby identifying the occluded area in the first image.
[0048] Alternatively, the occlusion area can be determined in the following ways: For each pixel in the second viewpoint, determine whether the pixel is excessively deformed based on the local rate of change corresponding to the pixel. Identify the target pixel that is excessively deformed; Based on the first coordinate mapping relationship, determine the corresponding mapped pixel in the first image for the target pixel; The occlusion area is determined based on the mapped pixels.
[0049] As mentioned earlier, the local rate of change can reflect the occlusion situation under viewpoint switching. Therefore, for each pixel in the second viewpoint, it can be determined whether the pixel is excessively deformed based on the local rate of change corresponding to the pixel.
[0050] In one possible implementation, whether a pixel is excessively deformed can be determined in the following way: Obtain the preset local rate of change reference value; If the local rate of change corresponding to a pixel is greater than the local rate of change reference value, the pixel is determined to be excessively deformed. If the local rate of change corresponding to a pixel is less than or equal to the reference value of the local rate of change, it is determined that the pixel has not been excessively deformed.
[0051] Therefore, pixels with a local rate of change greater than a reference value can be selected as target pixels. Furthermore, based on the coordinates of the target pixels and the first coordinate mapping relationship, the corresponding mapped pixels in the first image can be determined, and all mapped pixels constitute the occlusion region.
[0052] In step 14, a mask image corresponding to the first image is generated based on the occlusion area.
[0053] Optionally, the mask image corresponding to the first image can be generated in the following way: Generate an initial image corresponding to the first image; The pixel values of the pixels corresponding to the occluded area in the initial image are set to the first value, and the pixel values of the remaining pixels in the initial image are set to the second value to obtain the mask image.
[0054] The initial image is a default image with the same size as the first image. Based on the initial image, the pixel values of the pixels corresponding to the occluded area can be set to a first value, and the pixel values of the remaining pixels can be set to a second value to obtain the mask image. For example, the first value can be 1, and the second value can be 0.
[0055] In step 15, the first image and the mask image are identified as a sample data pair associated with the first viewpoint.
[0056] After obtaining the mask image corresponding to the first image, the first image and the mask image can be used as a sample data pair associated with the first viewpoint, and used as training data for subsequent image filling tasks.
[0057] The above technical solution determines the spatial depth information of the first image from a first-viewpoint. Based on the spatial depth information and the relative positional relationship between the second and first viewpoints, it determines the first coordinate mapping relationship from the second viewpoint to the first viewpoint. By determining the local rate of change of the first coordinate mapping relationship, it identifies occlusion regions in the first image and generates a mask image corresponding to the first image based on these occlusion regions. The first image and the mask image are then used as sample data pairs associated with the first viewpoint. Therefore, regions potentially occluded in the first viewpoint can be determined using the second viewpoint and the coordinate mapping relationship between the two viewpoints. Mask image generation can be achieved based solely on a monocular image, improving the flexibility of sample data pair generation and making it applicable to more data generation scenarios. Furthermore, utilizing the local change characteristics of the coordinate mapping to determine occlusion regions results in more accurate mask images with more natural and smoother boundaries. Additionally, the data generation process does not rely on complex algorithms, improving processing efficiency. Based on this, the generation quality and efficiency of sample data pairs can be effectively improved, providing more and higher-quality usable data for data imputation techniques.
[0058] Figure 2 This is a block diagram of a data generation apparatus provided according to one embodiment of the present disclosure. Figure 2 As shown, the device may include: The first determining module 21 is used to determine the spatial depth information corresponding to the first image based on the first image from the first perspective. The second determining module 22 is used to determine a first coordinate mapping relationship from the second viewpoint to the first viewpoint based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint. The third determining module 23 is used to identify the occluded area in the first image by determining the local rate of change of the first coordinate mapping relationship; The generation module 24 is used to generate a mask image corresponding to the first image based on the occlusion area; The fourth determining module 25 is used to determine the first image and the mask image as a sample data pair associated with the first viewpoint.
[0059] Optionally, the second determining module 22 includes: The first determining submodule is used to determine a second coordinate mapping relationship from the first viewpoint to the second viewpoint based on the spatial depth information and the relative position relationship; The second determining submodule is used to determine the mapping coordinates of each pixel in the second viewpoint in the first image according to the second coordinate mapping relationship, so as to obtain the first coordinate mapping relationship.
[0060] Optionally, the third determining module 23 includes: The third determining submodule is used to determine the local rate of change at each pixel in the second viewpoint based on the first coordinate mapping relationship. The identification submodule is used to identify, based on the local change rate at each pixel in the second viewpoint, the region that is not visible in the first viewpoint but is visible in the second viewpoint, as the occlusion region.
[0061] Optionally, the local rate of change is the norm of the Jacobian matrix.
[0062] Optionally, the identification submodule includes: The fourth determination submodule is used to determine whether there is excessive deformation at each pixel in the second viewpoint based on the local change rate corresponding to the pixel. The fifth determination submodule is used to determine the target pixel points of excessive deformation; The sixth determining submodule is used to determine the mapped pixel point corresponding to the target pixel point in the first image based on the first coordinate mapping relationship; The seventh determining submodule is used to determine the occlusion area based on the mapped pixels.
[0063] Optionally, the fourth determining submodule includes: The acquisition submodule is used to acquire preset local rate of change reference values; The eighth determining submodule is used to determine excessive deformation at the pixel if the local rate of change corresponding to the pixel is greater than the local rate of change reference value. The ninth determining submodule is used to determine that the pixel has not been excessively deformed if the local rate of change corresponding to the pixel is less than or equal to the local rate of change reference value.
[0064] Optionally, the generation module 24 includes: A generation submodule is used to generate an initial image corresponding to the first image; The setting submodule is used to set the pixel value of the pixel point corresponding to the occluded area in the initial image to a first value, and set the pixel value of the remaining pixel point in the initial image to a second value, so as to obtain the mask image.
[0065] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0066] Based on the same inventive concept, this disclosure also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the data generation method provided in any embodiment of this disclosure.
[0067] Based on the same inventive concept, this disclosure also provides an electronic device, including: A storage device on which computer programs are stored; A processing device is configured to execute the computer program in the storage device to implement the steps of the data generation method provided in any embodiment of the present disclosure.
[0068] Based on the same inventive concept, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the data generation method provided in any embodiment of this disclosure.
[0069] The following is for reference. Figure 3 The diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0070] like Figure 3 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0071] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0072] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0073] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0074] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0075] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0076] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: Based on the first image from a first-view perspective, determine the spatial depth information corresponding to the first image; Based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint, a first coordinate mapping relationship from the second viewpoint to the first viewpoint is determined; By determining the local rate of change of the first coordinate mapping relationship, the occluded area in the first image is identified; Based on the occlusion area, a mask image corresponding to the first image is generated; The first image and the mask image are identified as a sample data pair associated with the first viewpoint.
[0077] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0078] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0079] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the modules do not necessarily limit the module itself; for example, the first determining module can also be described as "a module that determines the spatial depth information corresponding to a first image based on a first image from a first viewpoint".
[0080] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0081] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0082] According to one or more embodiments of this disclosure, a data generation method is provided, the method comprising: Based on the first image from a first-view perspective, determine the spatial depth information corresponding to the first image; Based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint, a first coordinate mapping relationship from the second viewpoint to the first viewpoint is determined; By determining the local rate of change of the first coordinate mapping relationship, the occluded area in the first image is identified; Based on the occlusion area, a mask image corresponding to the first image is generated; The first image and the mask image are identified as a sample data pair associated with the first viewpoint.
[0083] According to one or more embodiments of this disclosure, a data generation method is provided, wherein determining a first coordinate mapping relationship from the second viewpoint to the first viewpoint based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint includes: Based on the spatial depth information and the relative positional relationship, a second coordinate mapping relationship from the first viewpoint to the second viewpoint is determined; Based on the second coordinate mapping relationship, the mapping coordinates of each pixel in the second viewpoint in the first image are determined to obtain the first coordinate mapping relationship.
[0084] According to one or more embodiments of this disclosure, a data generation method is provided. The step of identifying the occluded region in the first image by determining the local rate of change of the first coordinate mapping relationship includes: Based on the first coordinate mapping relationship, determine the local rate of change at each pixel in the second viewpoint; Based on the local change rate at each pixel in the second viewpoint, the region that is invisible in the first viewpoint but visible in the second viewpoint is identified as the occlusion region.
[0085] According to one or more embodiments of this disclosure, a data generation method is provided, wherein the local rate of change is the norm of the Jacobian matrix.
[0086] According to one or more embodiments of this disclosure, a data generation method is provided, wherein identifying regions invisible in the first view but visible in the second view, based on the local change rate at each pixel in the second view, as the occlusion region, includes: For each pixel in the second viewpoint, determine whether the pixel is excessively deformed based on the local rate of change corresponding to the pixel. Identify the target pixel that is excessively deformed; Based on the first coordinate mapping relationship, determine the mapped pixel point corresponding to the target pixel point in the first image; The occlusion area is determined based on the mapped pixels.
[0087] According to one or more embodiments of this disclosure, a data generation method is provided, wherein determining whether a pixel is excessively deformed based on the local rate of change corresponding to the pixel includes: Obtain the preset local rate of change reference value; If the local rate of change corresponding to the pixel is greater than the reference value of the local rate of change, it is determined that the pixel is excessively deformed. If the local rate of change corresponding to the pixel is less than or equal to the reference value of the local rate of change, it is determined that the pixel has not been excessively deformed.
[0088] According to one or more embodiments of this disclosure, a data generation method is provided, wherein generating a mask image corresponding to a first image based on the occlusion region includes: Generate an initial image corresponding to the first image; The pixel values of the pixels corresponding to the occluded area in the initial image are set to a first value, and the pixel values of the remaining pixels in the initial image are set to a second value to obtain the mask image.
[0089] According to one or more embodiments of the present disclosure, a data generation apparatus is provided, the apparatus comprising: The first determining module is used to determine the spatial depth information corresponding to the first image based on the first image from a first perspective. The second determining module is used to determine a first coordinate mapping relationship from the second viewpoint to the first viewpoint based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint. The third determining module is used to identify the occluded area in the first image by determining the local rate of change of the first coordinate mapping relationship; The generation module is used to generate a mask image corresponding to the first image based on the occluded area; The fourth determining module is used to determine the first image and the mask image as a sample data pair associated with the first viewpoint.
[0090] According to one or more embodiments of this disclosure, a data generation apparatus is provided, wherein the second determining module includes: The first determining submodule is used to determine a second coordinate mapping relationship from the first viewpoint to the second viewpoint based on the spatial depth information and the relative position relationship; The second determining submodule is used to determine the mapping coordinates of each pixel in the second viewpoint in the first image according to the second coordinate mapping relationship, so as to obtain the first coordinate mapping relationship.
[0091] According to one or more embodiments of this disclosure, a data generation apparatus is provided, wherein the third determining module includes: The third determining submodule is used to determine the local rate of change at each pixel in the second viewpoint based on the first coordinate mapping relationship. The identification submodule is used to identify, based on the local change rate at each pixel in the second viewpoint, the region that is not visible in the first viewpoint but is visible in the second viewpoint, as the occlusion region.
[0092] According to one or more embodiments of this disclosure, a data generation apparatus is provided, wherein the local rate of change is the norm of the Jacobian matrix.
[0093] According to one or more embodiments of this disclosure, a data generation apparatus is provided, wherein the identification submodule includes: The fourth determination submodule is used to determine whether there is excessive deformation at each pixel in the second viewpoint based on the local change rate corresponding to the pixel. The fifth determination submodule is used to determine the target pixel points of excessive deformation; The sixth determining submodule is used to determine the mapped pixel point corresponding to the target pixel point in the first image based on the first coordinate mapping relationship; The seventh determining submodule is used to determine the occlusion area based on the mapped pixels.
[0094] According to one or more embodiments of this disclosure, a data generation apparatus is provided, wherein the fourth determining submodule includes: The acquisition submodule is used to acquire preset local rate of change reference values; The eighth determining submodule is used to determine excessive deformation at the pixel if the local rate of change corresponding to the pixel is greater than the local rate of change reference value. The ninth determining submodule is used to determine that the pixel has not been excessively deformed if the local rate of change corresponding to the pixel is less than or equal to the local rate of change reference value.
[0095] According to one or more embodiments of this disclosure, a data generation apparatus is provided, the generation module comprising: A generation submodule is used to generate an initial image corresponding to the first image; The setting submodule is used to set the pixel value of the pixel point corresponding to the occluded area in the initial image to a first value, and set the pixel value of the remaining pixel point in the initial image to a second value, so as to obtain the mask image.
[0096] According to one or more embodiments of the present disclosure, a computer-readable medium is provided having a computer program stored thereon that, when executed by a processing device, implements the steps of the data generation method provided in any embodiment of the present disclosure.
[0097] According to one or more embodiments of this disclosure, an electronic device is provided, comprising: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the data generation method provided in any embodiment of the present disclosure.
[0098] According to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the data generation method provided in any embodiment of the present disclosure.
[0099] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0100] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0101] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A data generation method, characterized in that, The method includes: Based on the first image from a first-view perspective, determine the spatial depth information corresponding to the first image; Based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint, a first coordinate mapping relationship from the second viewpoint to the first viewpoint is determined; By determining the local rate of change of the first coordinate mapping relationship, the occluded area in the first image is identified; Based on the occlusion area, a mask image corresponding to the first image is generated; The first image and the mask image are identified as a sample data pair associated with the first viewpoint.
2. The method according to claim 1, characterized in that, Determining the first coordinate mapping relationship from the second viewpoint to the first viewpoint based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint includes: Based on the spatial depth information and the relative positional relationship, a second coordinate mapping relationship from the first viewpoint to the second viewpoint is determined; Based on the second coordinate mapping relationship, the mapping coordinates of each pixel in the second viewpoint in the first image are determined to obtain the first coordinate mapping relationship.
3. The method according to claim 1, characterized in that, The step of identifying the occluded region in the first image by determining the local rate of change of the first coordinate mapping relationship includes: Based on the first coordinate mapping relationship, determine the local rate of change at each pixel in the second viewpoint; Based on the local change rate at each pixel in the second viewpoint, the region that is invisible in the first viewpoint but visible in the second viewpoint is identified as the occlusion region.
4. The method according to claim 3, characterized in that, The local rate of change is the norm of the Jacobian matrix.
5. The method according to claim 3, characterized in that, The step of identifying the region that is invisible in the first view but visible in the second view, based on the local change rate at each pixel in the second view, as the occlusion region includes: For each pixel in the second viewpoint, determine whether the pixel is excessively deformed based on the local rate of change corresponding to the pixel. Identify the target pixel that is excessively deformed; Based on the first coordinate mapping relationship, determine the mapped pixel point corresponding to the target pixel point in the first image; The occlusion area is determined based on the mapped pixels.
6. The method according to claim 5, characterized in that, The step of determining whether a pixel is excessively deformed based on the local rate of change corresponding to the pixel includes: Obtain the preset local rate of change reference value; If the local rate of change corresponding to the pixel is greater than the reference value of the local rate of change, it is determined that the pixel is excessively deformed. If the local rate of change corresponding to the pixel is less than or equal to the reference value of the local rate of change, it is determined that the pixel has not been excessively deformed.
7. The method according to claim 1, characterized in that, The step of generating a mask image corresponding to the first image based on the occlusion area includes: Generate an initial image corresponding to the first image; The pixel values of the pixels corresponding to the occluded area in the initial image are set to a first value, and the pixel values of the remaining pixels in the initial image are set to a second value to obtain the mask image.
8. A data generation apparatus, characterized in that, The device includes: The first determining module is used to determine the spatial depth information corresponding to the first image based on the first image from a first perspective. The second determining module is used to determine a first coordinate mapping relationship from the second viewpoint to the first viewpoint based on the spatial depth information and the relative positional relationship between the second viewpoint and the first viewpoint. The third determining module is used to identify the occluded area in the first image by determining the local rate of change of the first coordinate mapping relationship; The generation module is used to generate a mask image corresponding to the first image based on the occluded area; The fourth determining module is used to determine the first image and the mask image as a sample data pair associated with the first viewpoint.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method described in any one of claims 1-7.
10. An electronic device, characterized in that, include: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1-7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.