Image processing method, device, apparatus and storage medium
By segmenting the target object and adjusting its depth map, and combining it with the depth map of the virtual object, accurate rendering of the occlusion relationship between the virtual object and the target object is achieved, solving the problem of inaccurate occlusion relationships in existing technologies and improving the realism of the image.
Patent Information
- Application Number
- CN202210451633.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-04-26
AI Technical Summary
In existing technologies, when determining the occlusion relationship between virtual objects and target objects using standard virtual models, it is impossible to match various target objects, resulting in inaccurate occlusion relationships and affecting the realism of the image.
By segmenting the target object into defined parts, a first depth map of the virtual object and a second depth map of the standard virtual model are obtained. Based on these two maps, the initial mask map is adjusted to obtain a target mask map. Rendering is then performed based on the target mask map. Finally, the defined part map is overlaid with the virtual object map to achieve accurate addition of the virtual object.
It improves the fit between virtual objects and target objects, enhancing the realism of the images.
Smart Images

Figure CN114782659B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of image processing, and particularly relate to an image processing method, device, equipment and storage medium. BACKGROUND
[0002] It is a common application scenario in the current augmented reality to add a virtual object on a target object. In this application scenario, when adding the virtual object, the occlusion relationship between the virtual object and the target object needs to be determined, and the virtual object and the target object are rendered based on the occlusion relationship.
[0003] In the prior art, a standard virtual model is used to determine the occlusion relationship between the virtual object and the target object. However, since the standard virtual model has a fixed size and cannot match various target objects, the determined occlusion relationship is inaccurate, and the virtual object and the target object are not well fitted, which affects the authenticity of the image. SUMMARY
[0004] Embodiments of the present disclosure provide an image processing method, device, equipment and storage medium, which can add a virtual object to a target object and improve the authenticity of the virtual object.
[0005] In a first aspect, embodiments of the present disclosure provide an image processing method, comprising:
[0006] segmenting a set part of a target object to obtain an initial mask image;
[0007] obtaining a first depth map of a virtual object and a second depth map of a standard virtual model related to the target object;
[0008] adjusting the initial mask image based on the first depth map and the second depth map to obtain a target mask image;
[0009] rendering the set part based on the target mask image to obtain a set part image, and rendering the virtual object to obtain a virtual object image;
[0010] superimposing the set part image and the virtual object image to obtain a target image.
[0011] In a second aspect, embodiments of the present disclosure further provide an image processing device, comprising:
[0012] an initial mask image obtaining module configured to segment a set part of a target object to obtain an initial mask image;
[0013] a depth map obtaining module configured to obtain a first depth map of a virtual object and a second depth map of a standard virtual model related to the target object;
[0014] an object mask obtaining module, configured to adjust the initial mask based on the first depth map and the second depth map, to obtain an object mask;
[0015] a rendering module, configured to render the set part based on the object mask, to obtain a set part image; and render the virtual object, to obtain a virtual object image.
[0016] an object image obtaining module, configured to superimpose the set part image and the virtual object image, to obtain an object image.
[0017] In a third aspect, an electronic device is provided, and the electronic device includes:
[0018] one or more processing devices;
[0019] a storage device configured to store one or more programs;
[0020] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the image processing method according to the embodiments of the present disclosure.
[0021] In a fourth aspect, a computer readable medium is provided, and the computer readable medium stores a computer program, which, when executed by a processing device, implements the image processing method according to the embodiments of the present disclosure.
[0022] The embodiments of the present disclosure provide an image processing method and device, an electronic device, and a storage medium. A set part of a target object is segmented to obtain an initial mask. A first depth map of a virtual object and a second depth map of a standard virtual model related to the target object are obtained. The initial mask is adjusted based on the first depth map and the second depth map to obtain an object mask. The set part is rendered based on the object mask to obtain a set part image. The virtual object is rendered to obtain a virtual object image. The set part image and the virtual object image are superimposed to obtain an object image. The image processing method provided by the embodiments of the present disclosure can add the virtual object to the target object and improve the authenticity of the virtual object by rendering the set part based on the object mask and superimposing the set part image and the virtual object image. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 is a flowchart of an image processing method according to an embodiment of the present disclosure;
[0024] Figure 2 is an initial mask after segmentation of a face according to an embodiment of the present disclosure;
[0025] Figure 3ais a depth map of a virtual object in embodiments of the present disclosure;
[0026] Figure 3b is a depth map of a standard virtual model in embodiments of the present disclosure;
[0027] Figure 4 is an example image of a two-dimensional image in embodiments of the present disclosure;
[0028] Figure 5 is an example image of an adjusted mask in embodiments of the present disclosure;
[0029] Figure 6 is an example image of processing frame delay in embodiments of the present disclosure;
[0030] Figure 7 is an example image of adding a virtual object to a target object in embodiments of the present disclosure;
[0031] Figure 8 is a structural schematic diagram of an image processing apparatus in embodiments of the present disclosure;
[0032] Figure 9 is a structural schematic diagram of an electronic device in embodiments of the present disclosure. DETAILED DESCRIPTION
[0033] Embodiments of the present disclosure will be described in more detail by referring to the drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided so as to more completely and thoroughly understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of the present disclosure.
[0034] It is understood that each of the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0035] The term “comprising” and variations thereof as used herein are open-ended, that is, “including but not limited to”. The term “based on” is “based, at least in part, on”. The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Related definitions are given throughout the description.
[0036] It should be noted that the terms "first", "second", and the like in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0037] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0038] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0039] Figure 1 A flowchart of an image processing method provided by an embodiment of the present disclosure, the embodiment can be applied to the case of adding a virtual object in an image, the method can be executed by an image processing device, which can be composed of hardware and / or software, and can generally be integrated in a device with image processing function, which can be an electronic device such as a server, a mobile terminal or a server cluster. As shown in the figure, Figure 1 The method specifically includes the following steps:
[0040] S110, segmenting the set part of the target object to obtain an initial mask map.
[0041] The target object can be a real object identified in the current scene, or a real object set to add a virtual object, such as a human body, a plant, a vehicle, a building, etc. The set part can be a part that has a blocking relationship with the added virtual object, which can be determined according to the position of the added virtual object. For example, if the target object is a human body and the virtual object is added on the head, the set part can be the hair area, and if the virtual object is added on the neck, the set part can be the face area, etc.
[0042] Specifically, the process of segmenting the set part of the target object to obtain the initial mask map can be: detecting the set part of the image containing the target object, obtaining the confidence of each pixel point in the image belonging to the set part, taking the confidence as the pixel value of each pixel point, and thus obtaining the initial mask map. For example, Figure 2 is the initial mask map after segmenting the face in the embodiment, as shown in the figure, Figure 2 The white area represents the face area and the black area is the non-face area. The face can be segmented through the initial mask map.
[0043] S120, obtaining a first depth map of the virtual object and a second depth map of a standard virtual model related to the target object.
[0044] The virtual object can be any constructed virtual object, and the virtual object can be an irregularly shaped object, such as a virtual animal (e.g., a cat, a dog, etc.), a virtual headwear, a virtual earwear, a virtual object that can be added to a neck (e.g., a virtual necklace, a virtual neck pillow, etc.), which is not limited herein. In this embodiment, the number of virtual objects is not limited, for example, a combination of a virtual animal and a virtual neck pillow.
[0045] The standard virtual model can be a virtual model related to the target object, and can be a virtual model for depth detection instead of the target object. The standard virtual model can be a virtual model of the shape of the target object, or a virtual model associated with the shape of the target object, or a virtual model constructed based on the target object in the current frame. For example, if the target object is a human head, the standard virtual model is a virtual model of the shape of a human head, and the shape associated with the shape of the target object can be a virtual model of a cube shape or a cylinder shape. The virtual model of the shape of the target object and the virtual model associated with the shape of the target object can be pre-created. The virtual model constructed based on the target object in the current frame can be understood as constructing the standard virtual model based on the target object in real time.
[0046] The method of constructing a virtual model based on the target object in the current frame can be 3D scanning the target object in the current frame to obtain 3D data of the target object, and constructing the standard virtual model based on the 3D data. In this embodiment, the virtual model associated with the shape of the target object can reduce the amount of calculation. Constructing the standard virtual model based on the target object in real time can improve the accuracy of depth detection.
[0047] In this embodiment, the first depth map of the virtual object can be a depth map after the virtual object is added to the target object, and the second depth map of the standard virtual model can be a depth map after the standard virtual model is added to the target object.
[0048] In this embodiment, the method of obtaining the first depth map of the virtual object can be tracking the object addition site based on a set tracking algorithm to obtain position information of the object addition site, adding the virtual object to the object addition site based on the position information, obtaining depth information of the added virtual object, and obtaining the first depth map.
[0049] The object adding part is a part where the virtual object is added to the target object. The setting tracking algorithm can be any existing tracking algorithm, which is not limited herein. The position information can be represented by a transformation matrix. Specifically, in the process of tracking the object adding part by the setting tracking algorithm, a transformation matrix corresponding to the object adding part is obtained. The matrix corresponding to the virtual object is multiplied by the transformation matrix, so as to realize the operation of adding the virtual object to the object adding part. Finally, the depth information of the added virtual object is obtained by using a virtual camera, and a first depth map is obtained. For example, Figure 3a is a depth map of a virtual object in this embodiment. As shown in Figure 3a , this map is a depth map of a "virtual neck pillow". In this embodiment, adding a virtual object to a target object when obtaining a depth map can improve the accuracy of subsequent depth detection.
[0050] In this embodiment, the way to obtain the second depth map of the standard virtual model can be: tracking the setting part based on the setting tracking algorithm to obtain the position information of the setting part; adding the standard virtual model to the setting part based on the position information; obtaining the depth information of the added standard virtual model to obtain the second depth map.
[0051] The position information can be represented by a transformation matrix. The setting part can be a face. Specifically, in the process of tracking the setting part by the setting tracking algorithm, a transformation matrix corresponding to the setting part is obtained. The matrix corresponding to the standard virtual model is multiplied by the transformation matrix, so as to realize the operation of adding the standard virtual model to the setting part. Finally, the depth information of the added standard virtual model is obtained by using a virtual camera, and a second depth map is obtained. For example, Figure 3b is a depth map of a standard virtual model in this embodiment. As shown in Figure 3b , this map is a depth map of a "virtual human head".
[0052] S130, adjusting the initial mask map based on the first depth map and the second depth map to obtain a target mask map.
[0053] In this embodiment, the principle of adjusting the initial mask map based on the first depth map and the second depth map can be understood as: judging whether the corresponding pixel point in the setting part is blocked by the virtual object according to the depth value of the corresponding pixel point in the first depth map and the second depth map. If it is blocked, the pixel value of the corresponding pixel point in the initial mask map is adjusted, so that the adjusted mask map reflects the blocking relationship between the virtual object and the setting part of the target object.
[0054] Optionally, the process of adjusting the initial mask based on the first depth map and the second depth map to obtain the target mask can be: obtaining a near plane depth value and a far plane depth value of the virtual camera; performing linear transformation on the first depth map and the second depth map respectively according to the near plane depth value and the far plane depth value; and adjusting the initial mask based on the linearly transformed first depth map and the linearly transformed second depth map to obtain the target mask.
[0055] The near plane depth value and the far plane depth value can be directly obtained from the configuration information (such as the field of view angle) of the virtual camera. The linear transformation on the first depth map and the second depth map can be understood as transforming the depth values in the first depth map and the second depth map to the range of the near plane depth value and the far plane depth value. Optionally, the formula for linearly transforming the first depth map and the second depth map can be represented as: In the formula, L(d) represents the linearly transformed depth value, d represents the depth value before linear transformation, zNear represents the near plane depth value, and zFar represents the far plane depth value. In this embodiment, linearly transforming the depth values in the first depth map and the second depth map to the range of the near plane depth value and the far plane depth value can improve the accuracy of adjusting the mask.
[0056] Specifically, the way of adjusting the initial mask based on the first depth map and the second depth map to obtain the target mask can be: if the first depth value in the first depth map is greater than the second depth value in the second depth map, keeping the pixel value of the corresponding pixel point in the initial mask unchanged; and if the first depth value is less than or equal to the second depth value, adjusting the pixel value of the corresponding pixel point in the initial mask to a set value.
[0057] In the formula, the first depth value can be represented by one of the channel values (such as the R channel) in the first depth map; the second depth value can be represented by one of the channel values (such as the R channel) in the second depth map; and the pixel value of each pixel point in the initial mask can be represented by one of the channel values (such as the R channel). Since the first depth map, the second depth map and the initial mask are all grayscale images, the values of the three color channels (red, green and blue, RGB) are equal, so one of the channel values can be arbitrarily selected.
[0058] The set value can be set to 0. In this embodiment, if the first depth value in the first depth map is greater than the second depth value in the second depth map, it indicates that the virtual object is behind the set part and the set part is not occluded by the virtual object. In this case, the pixel value of the corresponding pixel in the initial mask map remains unchanged. If the first depth value is less than or equal to the second depth value, it indicates that the virtual object is in front of the set part and the virtual object occludes the set part. In this case, the pixel value of the corresponding pixel in the initial mask map is adjusted to 0. In this embodiment, when the virtual object occludes the set part, the pixel value of the corresponding pixel in the initial mask map is adjusted to 0, which can improve the speed of adjusting the mask map.
[0059] Optionally, the initial mask image can be adjusted based on the first depth image and the second depth image to obtain the target mask image. This can be done by: obtaining a two-dimensional image of the virtual object; if the first depth value in the first depth image is greater than the second depth value in the second depth image, then keeping the pixel value of the corresponding pixel in the initial mask image unchanged; if the first depth value is less than or equal to the second depth value, then adjusting the pixel value of the corresponding pixel in the initial mask image by subtracting the set channel value of the corresponding pixel in the two-dimensional image to obtain the final pixel value.
[0060] One method for obtaining a 2D image of a virtual object is to project each 3D point constituting the virtual object onto a 2D plane to obtain a 2D image corresponding to the virtual object. The channel can be the A channel in a 2D image. In this embodiment, the 2D image contains four channels: RGBA, where RGB represents three color channels, and the A channel is the alpha channel. The value of the A channel is between 0 and 1, representing the transparency of a pixel. If the A channel value is 0, the pixel is transparent; if the A channel value is greater than 0, the pixel is opaque. For example... Figure 4 This is an example diagram of a two-dimensional image in this embodiment. For example... Figure 4 As shown, this image is a two-dimensional representation of a "virtual neck pillow," as follows: Figure 4 As shown, the black area is the transparent area, meaning the A channel value is 0.
[0061] Specifically, if the first depth value in the first depth map is greater than the second depth value in the second depth map, it indicates that the virtual object is behind the designated area and the designated area is not occluded by the virtual object. In this case, the pixel value of the corresponding pixel in the initial mask map remains unchanged. If the first depth value is less than or equal to the second depth value, it indicates that the virtual object is in front of the designated area and the virtual object occludes the designated area. In this case, the pixel value of the corresponding pixel in the initial mask map is subtracted from the A channel value of the corresponding pixel in the 2D image. For example, Figure 5 This is an example diagram of the adjusted mask image in this embodiment. As can be seen from the diagram, compared to... Figure 2 The initial mask image in the image. Figure 5In the target mask image, the pixel value of the pixel point in the region blocked by the virtual object is adjusted. In this embodiment, if the virtual object blocks the setting position, the pixel value of the corresponding pixel point in the initial mask image is subtracted by the A channel value of the corresponding pixel point in the two-dimensional image, so that the smoothness of the edge can be improved, the blocking transition is smooth, and the rendered image is more realistic.
[0062] In S140, the setting position is rendered based on the target mask image to obtain a setting position image, and the virtual object is rendered to obtain a virtual object image.
[0063] In this embodiment, the target mask image can represent which pixel points of the setting position are blocked by the virtual object, so when the setting position is rendered based on the mask image, only the unblocked pixel points can be rendered, or the transparency of the blocked pixel points can be adjusted to 0.
[0064] Optionally, the manner of rendering the setting position based on the target mask image can be that image information corresponding to the target mask image is fused with image information corresponding to the original image of the setting position to obtain fused image information, and the setting position is rendered based on the fused image information.
[0065] The image information can be represented by a matrix with the same size as the image, and each element of the matrix represents the pixel value of the corresponding pixel point. The manner of fusing the image information corresponding to the target mask image with the image information corresponding to the original image of the setting position can be that the pixel values of the corresponding pixel points in the target mask image and the original image of the setting position are multiplied to obtain the fused image information. In this embodiment, since the fused image information contains the final pixel value of each pixel point, the setting position is rendered based on the final pixel value of each pixel point. In this embodiment, the target mask image is fused with the original image before rendering, so that the rendered setting position image accurately reflects the blocking relationship with the virtual object.
[0066] Optionally, the manner of rendering the setting position based on the target mask image can be that the transparency information of each pixel point of the setting position is determined according to the target mask image, and the setting position is rendered based on the transparency information.
[0067] In the target mask image, the pixel value of the pixel point represents the transparency. The pixel value of the pixel point in the target mask image is a value between 0 and 1. If the pixel value is 0, it means that the corresponding pixel point in the setting position image is transparent. If the pixel value is greater than 0, it means that the corresponding pixel point in the setting position image is non-transparent, and the transparency is determined by the specific value. For example, if the pixel value is 1, the transparency is 100%, and if the pixel value is 0.5, the transparency is 50%. In this embodiment, the setting position is rendered based on the transparency information, which can reduce the amount of calculation.
[0068] Optionally, if the virtual object includes multiple virtual objects, the manner of obtaining the first depth map of the virtual object can be: obtaining a first depth map corresponding to each of the multiple virtual objects, and obtaining multiple first depth maps; correspondingly, the manner of rendering the virtual object can be: rendering the multiple virtual objects based on the multiple first depth maps.
[0069] Specifically, the process of rendering the multiple virtual objects based on the multiple first depth maps can be: comparing the depth values in each depth map, and rendering the pixel points in the virtual object with the smallest depth value. Assuming that the virtual object includes two virtual objects, virtual object A and virtual object B, and the first depth maps corresponding to the two virtual objects are first depth map a and first depth map b, for a certain pixel point, if the first depth value a is smaller than the first depth value b, the pixel point in the virtual object A is rendered. In this embodiment, the pixel points in the virtual object with the smallest depth value are rendered, so that the rendered virtual object reflects the respective occlusion relationship, and the image is more realistic.
[0070] Optionally, if the virtual object includes multiple virtual objects, a set virtual object is determined from the multiple virtual objects; the second depth map of the standard virtual model is obtained by: obtaining the depth information of the standard virtual model and the set virtual object by using a virtual depth, and obtaining the second depth map.
[0071] The set virtual object can be selected by a user according to the actual shape of the virtual object (for example, ear ornaments). In order to ensure that the set virtual object is more consistent with the target object, the depth information of the set virtual object and the standard virtual model is placed in the same depth map during the depth detection, so that the set virtual object is better integrated with the target object. Specifically, the depth information of the standard virtual model and the set virtual object is obtained by using the same virtual camera, and the second depth map is obtained.
[0072] S150, superimposing the set part map and the virtual object map to obtain a target image.
[0073] The manner of superimposing the set part map and the virtual object map can be: superimposing the set part map on the virtual object map. In this embodiment, the set part map and the virtual object map are rendered by using different virtual cameras, and are respectively in different layers. The set part map is superimposed on the virtual object map, so that the set part map covers the virtual object map, and the target image is generated. In this embodiment, when the first depth value is smaller than or equal to the second depth value, the pixel points in the set part map in the region are transparent or not rendered, so the virtual object is not occluded. When the first depth value is greater than the second depth value, the pixel points in the set part map in the region are rendered with color, and the virtual object is occluded, so that the target image is more realistic.
[0074] Optionally, after the target mask image is obtained, the target mask image is cached. Correspondingly, the rendering of the set part based on the target mask image to obtain the set part image, and the rendering of the virtual object to obtain the virtual object image can be: for the current frame, the rendering of the set part based on the target mask image corresponding to the set forward frame to obtain the set part image, and the rendering of the virtual object corresponding to the set forward frame to obtain the virtual object image.
[0075] In this embodiment, due to the difference in processing speed of different algorithms, there is a frame delay problem. In order to ensure that the overall algorithm result meets the actual display effect, the current screen display picture is replaced with the picture N frames ago. Wherein, N can be determined according to the actual processing speed of each algorithm, for example: it can be 3.
[0076] In this embodiment, the determined target mask image is cached first. For the current frame, the target mask image corresponding to the set forward frame is obtained from the cache to render the set part. The virtual object corresponding to the set forward frame is rendered. Since the first N frames have no algorithm data, the data determined by the 0th frame can be used for rendering. Exemplarily, Figure 6 is an example diagram for processing frame delay in this embodiment. As Figure 6 shown, render0-render3 are rendered using the data in buffer0. The rendering picture after that is rendered using the data corresponding to three frames before it. In this embodiment, the problem of frame delay in the rendered picture can be avoided.
[0077] Exemplarily, Figure 7 is an example diagram for adding a virtual object to a target object in this embodiment, taking adding a virtual object to a human neck as an example. As Figure 7 shown, the camera of the terminal captures human images in real time, tracks the neck using a neck tracking algorithm to obtain the neck position, and mounts the virtual object to the neck position. The face is tracked using a face tracking algorithm to obtain the face position, and the virtual head model is mounted based on the face position. The face is segmented using a face segmentation algorithm to obtain an initial face mask image. The first depth image of the virtual object and the second depth image of the virtual head model are obtained using a depth camera. The depth values of the first depth image and the second depth image are compared, and the pixel values in the initial face mask image are adjusted based on the comparison result to obtain a target face mask image. The face image and the virtual object image are rendered on different layers, and the face image is superimposed on the virtual object image to obtain a neck occlusion effect image.
[0078] For example, one application scenario in the embodiment is to add a virtual animal (such as a kitten or a puppy) to the shoulder of a person, in which case the above technical solutions are used to determine the occlusion relationship between the virtual animal and the shoulder and the face of the person, and the rendering of the virtual animal and the shoulder and the face of the person is based on the occlusion relationship, so as to obtain a special effect of adding the virtual animal to the shoulder of the person. In another application scenario, a virtual neck pillow can be hung on the neck of a person, in which case the above technical solutions are used to determine the occlusion relationship between the virtual animal and the neck and the face of the person, and the rendering of the virtual neck pillow and the neck and the face of the person is based on the occlusion relationship, so as to obtain a special effect of hanging the virtual neck pillow on the neck of the person.
[0079] The technical solution of the embodiment of the present disclosure segments the set part of the target object to obtain an initial mask image; acquires a first depth image of a virtual object and a second depth image of a standard virtual model; adjusts the initial mask image based on the first depth image and the second depth image to obtain a target mask image; renders the set part based on the target mask image to obtain a set part image; renders the virtual object to obtain a virtual object image; and superimposes the set part image on the virtual object image to obtain a target image. The image processing method provided by the embodiment of the present disclosure renders the set part based on the target mask image and superimposes the set part image on the virtual object image, which can realize adding the virtual object to the target object and improve the authenticity of the virtual object.
[0080] Figure 8 is a structural schematic diagram of an image processing device disclosed by the embodiment of the present disclosure, as Figure 8 shown, the device comprises:
[0081] An initial mask image acquisition module 810 is configured to segment a set part of a target object to obtain an initial mask image.
[0082] A depth image acquisition module 820 is configured to acquire a first depth image of a virtual object and a second depth image of a standard virtual model related to the target object.
[0083] A target mask image acquisition module 830 is configured to adjust the initial mask image based on the first depth image and the second depth image to obtain a target mask image.
[0084] A rendering module 840 is configured to render the set part based on the target mask image to obtain a set part image, and render the virtual object to obtain a virtual object image.
[0085] A target image acquisition module 850 is configured to superimpose the set part image and the virtual object image to obtain a target image.
[0086] Optionally, the depth image acquisition module 820 is further configured to:
[0087] track the object adding part based on the set tracking algorithm to obtain position information of the object adding part; wherein the object adding part is a part where the virtual object is added to the target object;
[0088] add the virtual object to the object adding part based on the position information;
[0089] obtain depth information of the added virtual object to obtain a first depth map.
[0090] Optionally, the standard virtual model is a virtual model of a shape of the target object, or a virtual model associated with the shape of the target object, or a virtual model constructed based on the target object in the current frame.
[0091] Optionally, the target mask map obtaining module 830 is further configured to:
[0092] obtain a near plane depth value and a far plane depth value of the virtual camera;
[0093] perform linear transformation on the first depth map and the second depth map respectively according to the near plane depth value and the far plane depth value;
[0094] adjust the initial mask map based on the linearly transformed first depth map and second depth map to obtain the target mask map.
[0095] Optionally, the target mask map obtaining module 830 is further configured to:
[0096] if the first depth value in the first depth map is greater than the second depth value in the second depth map, keep the pixel value of the corresponding pixel point in the initial mask map unchanged;
[0097] if the first depth value is less than or equal to the second depth value, adjust the pixel value of the corresponding pixel point in the initial mask map to a set value.
[0098] Optionally, the target mask map obtaining module 830 is further configured to:
[0099] obtain a two-dimensional map of the virtual object;
[0100] if the first depth value in the first depth map is greater than the second depth value in the second depth map, keep the pixel value of the corresponding pixel point in the initial mask map unchanged;
[0101] if the first depth value is less than or equal to the second depth value, adjust the pixel value of the corresponding pixel point in the initial mask map by subtracting the set channel value of the corresponding pixel point in the two-dimensional map to obtain the final pixel value.
[0102] Optionally, the rendering module 840 is further configured to:
[0103] Fuse image information corresponding to the target mask map with image information corresponding to the original image of the set part to obtain fused image information;
[0104] Render the set part based on the fused image information.
[0105] Optionally, the rendering module 840 is further configured to:
[0106] Determine transparency information of each pixel point of the set part according to the target mask map, wherein a pixel value of a pixel point in the target mask map represents transparency;
[0107] Render the set part based on the transparency information.
[0108] Optionally, if the virtual object includes a plurality of virtual objects, the depth map acquisition module 820 is further configured to:
[0109] Acquire a plurality of first depth maps corresponding to the plurality of virtual objects respectively to obtain a plurality of first depth maps;
[0110] The rendering module 840 is further configured to:
[0111] Render the plurality of virtual objects based on the plurality of first depth maps.
[0112] Optionally, if the virtual object includes a plurality of virtual objects, a set virtual object is determined from the plurality of virtual objects;
[0113] The depth map acquisition module 820 is further configured to:
[0114] Acquire depth information of the standard virtual model and the set virtual object by using a virtual camera to obtain a second depth map.
[0115] Optionally, the device further comprises a cache module configured to:
[0116] Cache the target mask map;
[0117] The rendering module 840 is further configured to:
[0118] For the current frame, render the set part based on the target mask map corresponding to the set forward frame to obtain a set part image, and render the virtual object corresponding to the set forward frame to obtain a virtual object image.
[0119] The device described above can execute the method provided by all the preceding embodiments of the present disclosure, and has corresponding function modules and beneficial effects for executing the above method. Technical details not described in detail in the present embodiment can be referred to the method provided by all the preceding embodiments of the present disclosure.
[0120] The following refers to Figure 9The diagram illustrates a structural schematic of an electronic device 300 suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs, desktop computers, or various forms of servers, such as standalone servers or server clusters. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0121] like Figure 9 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a memory device 305 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0122] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0123] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing a method of word recommendation. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 309, or installed from a storage device 305, or installed from a ROM 302. When the computer program is executed by a processing device 301, it performs the functions defined above in the methods of embodiments of this disclosure.
[0124] Note that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer readable program code is contained. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF, etc., or any suitable combination thereof.
[0125] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0126] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and can be accessed via the electronic device.
[0127] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: segment a set part of a target object to obtain an initial mask image; acquire a first depth image of a virtual object and a second depth image of a standard virtual model related to the target object; adjust the initial mask image based on the first depth image and the second depth image to obtain a target mask image; render the set part based on the target mask image to obtain a set part image; render the virtual object to obtain a virtual object image; and superimpose the set part image and the virtual object image to obtain a target image.
[0128] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0129] The flow and block diagrams in the drawings show architectural, functional, and operational architectures of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code that comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0130] The units described in the embodiments of the present disclosure can be implemented by means of software, or by means of hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.
[0131] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0132] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0133] According to one or more embodiments of the embodiments of this disclosure, the embodiments of this disclosure disclose an image processing method, comprising:
[0134] Segmenting a set part of a target object to obtain an initial mask map;
[0135] Obtaining a first depth map of a virtual object and a second depth map of a standard virtual model related to the target object;
[0136] Adjusting the initial mask map based on the first depth map and the second depth map to obtain a target mask map;
[0137] Rendering the set part based on the target mask map to obtain a set part map; and rendering the virtual object to obtain a virtual object map;
[0138] Superimposing the set part map and the virtual object map to obtain a target image.
[0139] Further, obtaining a first depth map of a virtual object comprises:
[0140] Tracking an object added part based on a set tracking algorithm to obtain position information of the object added part; wherein the object added part is a part of the target object to which the virtual object is added;
[0141] adding the virtual object to the object adding position based on the position information;
[0142] obtaining a first depth map by acquiring depth information of the added virtual object.
[0143] Further, the standard virtual model is a virtual model of the target object shape, or a virtual model associated with the target object shape, or a virtual model constructed based on the target object of the current frame.
[0144] Further, adjusting the initial mask map based on the first depth map and the second depth map to obtain a target mask map, comprising:
[0145] obtaining a near plane depth value and a far plane depth value of the virtual camera;
[0146] linearly transforming the first depth map and the second depth map respectively according to the near plane depth value and the far plane depth value;
[0147] adjusting the initial mask map based on the linearly transformed first depth map and second depth map to obtain a target mask map.
[0148] Further, adjusting the initial mask map based on the first depth map and the second depth map to obtain a target mask map, comprising:
[0149] if the first depth value in the first depth map is greater than the second depth value in the second depth map, keeping the pixel value of the corresponding pixel point in the initial mask map unchanged;
[0150] if the first depth value is less than or equal to the second depth value, adjusting the pixel value of the corresponding pixel point in the initial mask map to a set value.
[0151] Further, adjusting the initial mask map based on the first depth map and the second depth map to obtain a target mask map, comprising:
[0152] obtaining a two-dimensional map of the virtual object;
[0153] if the first depth value in the first depth map is greater than the second depth value in the second depth map, keeping the pixel value of the corresponding pixel point in the initial mask map unchanged;
[0154] if the first depth value is less than or equal to the second depth value, adjusting the pixel value of the corresponding pixel point in the initial mask map by subtracting the set channel value of the corresponding pixel point in the two-dimensional map to obtain the final pixel value.
[0155] Further, rendering the set part based on the target mask map comprises:
[0156] Fusing image information corresponding to the target mask map and image information corresponding to the original image of the set part to obtain fused image information;
[0157] Rendering the set part based on the fused image information.
[0158] Further, rendering the set part based on the target mask map comprises:
[0159] Determining transparency information of each pixel point of the set part according to the target mask map; wherein the pixel value of the pixel point in the target mask map represents the transparency;
[0160] Rendering the set part based on the transparency information.
[0161] Further, if the virtual object includes multiple, obtaining a first depth map of the virtual object comprises:
[0162] Obtaining a first depth map corresponding to each of the multiple virtual objects to obtain multiple first depth maps;
[0163] Rendering the virtual object comprises:
[0164] Rendering the multiple virtual objects based on the multiple first depth maps.
[0165] Further, if the virtual object includes multiple, determining a set virtual object from the multiple virtual objects;
[0166] Obtaining a second depth map of the standard virtual model comprises:
[0167] Obtaining the depth information of the standard virtual model and the set virtual object by using a virtual camera to obtain a second depth map.
[0168] Further, after obtaining the target mask map, further comprising:
[0169] Caching the target mask map;
[0170] Rendering the set part based on the target mask map to obtain a set part image; rendering the virtual object to obtain a virtual object image, comprising:
[0171] For the current frame, rendering the set part based on the target mask map corresponding to the set forward frame to obtain a set part image; rendering the virtual object corresponding to the set forward frame to obtain a virtual object image.
[0172] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technical solutions of the present disclosure can be achieved, which are not limited herein.
[0173] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. An image processing method, characterized by, The method comprises the following steps: segmenting a set part of a target object to obtain an initial mask image, wherein the set part is a part having an occlusion relationship with an added virtual object, and the initial mask image is determined by detecting the confidence of each pixel point in an image containing the target object belonging to the set part, and taking the confidence as the pixel value of the corresponding pixel point; obtaining a first depth map of the virtual object and a second depth map of a standard virtual model related to the target object, wherein the first depth map is a depth map of the virtual object after being added to the target object, and the second depth map is a depth map of the standard virtual model after being added to the target object; adjusting the initial mask image based on the first depth map and the second depth map to obtain a target mask image; rendering the set part based on the target mask image to obtain a set part image; and rendering the virtual object to obtain a virtual object image; superimposing the set part image and the virtual object image to obtain a target image; wherein the adjusting of the initial mask image based on the first depth map and the second depth map to obtain a target mask image comprises: obtaining a near plane depth value and a far plane depth value of a virtual camera; linearly transforming the first depth map and the second depth map according to the near plane depth value and the far plane depth value; adjusting the initial mask image based on the linearly transformed first depth map and second depth map to obtain a target mask image.
2. The method of claim 1, wherein, The method comprises the following steps: tracking an object adding part based on a set tracking algorithm to obtain position information of the object adding part; wherein the object adding part is a part of the target object to which the virtual object is added; adding the virtual object to the object adding part based on the position information; obtaining the depth information of the added virtual object to obtain a first depth map.
3. The method of claim 1, wherein, The standard virtual model is a virtual model of the shape of the target object, or a virtual model associated with the shape of the target object, or a virtual model constructed based on the target object of the current frame.
4. The method of claim 1, wherein, The adjusting of the initial mask image based on the first depth map and the second depth map to obtain a target mask image comprises: if a first depth value in the first depth map is greater than a second depth value in the second depth map, keeping the pixel value of the corresponding pixel point in the initial mask image unchanged; if the first depth value is less than or equal to the second depth value, adjusting the pixel value of the corresponding pixel point in the initial mask image to a set value.
5. The method of claim 1, wherein, The adjusting of the initial mask image based on the first depth map and the second depth map to obtain a target mask image comprises: obtaining a two-dimensional image of the virtual object; if a first depth value in the first depth map is greater than a second depth value in the second depth map, keeping the pixel value of the corresponding pixel point in the initial mask image unchanged; if the first depth value is less than or equal to the second depth value, adjusting the pixel value of the corresponding pixel point in the initial mask image by subtracting a set channel value of the corresponding pixel point in the two-dimensional image to obtain a final pixel value.
6. The method of claim 1, wherein, The target mask image is used to render the set part, including: The image information corresponding to the target mask image is fused with the image information corresponding to the original image of the set part, to obtain fused image information; The set part is rendered based on the fused image information.
7. The method of claim 6, wherein, The target mask image is used to render the set part, including: The transparency information of each pixel point of the set part is determined according to the target mask image; wherein the pixel value of the pixel point in the target mask image represents the transparency; The set part is rendered based on the transparency information.
8. The method of claim 1, wherein, If the virtual object includes multiple virtual objects, a first depth map of the virtual object is obtained, including: A first depth map corresponding to each of the multiple virtual objects is obtained, to obtain multiple first depth maps; The virtual object is rendered, including: The multiple virtual objects are rendered based on the multiple first depth maps.
9. The method according to claim 1 or 8, characterized in that, If the virtual object includes multiple virtual objects, a set virtual object is determined from the multiple virtual objects; A second depth map of a standard virtual model is obtained, including: Depth information of the standard virtual model and the set virtual object is obtained by using a virtual camera, to obtain a second depth map.
10. The method of claim 1, wherein, After the target mask image is obtained, further including: The target mask image is cached; The target mask image is used to render the set part, to obtain a set part image; the virtual object is rendered, to obtain a virtual object image, including: For a current frame, the set part is rendered based on the target mask image corresponding to a set forward frame, to obtain a set part image; the virtual object corresponding to the set forward frame is rendered, to obtain a virtual object image.
11. An image processing apparatus characterized by comprising: Including: An initial mask image obtaining module is configured to segment a set part of a target object to obtain an initial mask image, wherein the set part is a part having an occlusion relationship with an added virtual object, and the initial mask image is determined by detecting a confidence degree of each pixel point in an image containing the target object belonging to the set part and taking the confidence degree as a pixel value of the corresponding pixel point; A depth map obtaining module is configured to obtain a first depth map of a virtual object and a second depth map of a standard virtual model related to the target object, wherein the first depth map is a depth map of the virtual object after being added to the target object, and the second depth map is a depth map of the standard virtual model after being added to the target object; A target mask image obtaining module is configured to adjust the initial mask image based on the first depth map and the second depth map to obtain a target mask image; A rendering module is configured to render the set part based on the target mask image to obtain a set part image, and render the virtual object to obtain a virtual object image; A target image obtaining module is configured to superimpose the set part image and the virtual object image to obtain a target image; The target mask image obtaining module is further configured to: Obtain a near-plane depth value and a far-plane depth value of a virtual camera; Linearly transform the first depth map and the second depth map according to the near-plane depth value and the far-plane depth value, respectively; Adjust the initial mask map based on the first depth map and the second depth map after linear transformation to obtain a target mask map.
12. An electronic device, comprising: The electronic device includes: one or more processing devices; a storage device for storing one or more programs; When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the image processing method as claimed in any one of claims 1-10.
13. A computer readable medium having stored thereon a computer program, characterized in that, The program is executed by the processing device to implement the image processing method as claimed in any one of claims 1-10. The program is executed by the processing device to implement the image processing method as claimed in any one of claims 1-10.
Citation Information
Patent Citations
Image processing method and device, processor, electronic equipment and storage medium
CN110889890A
Image processing method and device, electronic equipment and computer readable storage medium
CN112102340A