Image processing method and device, electronic equipment and storage medium
By acquiring and synthesizing images of target objects in virtual scenes, the problem of weak user participation in online activities has been solved, realizing the interaction between users and images and an immersive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING XINTANG SICHUANG EDUCATIONAL TECH CO LTD
- Filing Date
- 2023-02-24
- Publication Date
- 2026-07-21
AI Technical Summary
Online activities have a weaker sense of user engagement because the user's actual environment does not match the environment required for offline activities.
By acquiring the first image data, the image of the target object is obtained by cutting out the image. The virtual scene is then acquired, and the display data of the target object's image in the virtual scene is determined. This data is then combined with the virtual scene, and finally, image acquisition is performed to generate the image to be displayed.
It enhances the user's sense of participation in the virtual scene, enables the user to interact with the images, and enhances the immersive experience.
Smart Images

Figure CN116258738B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an image processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] Due to the pandemic, many activities (such as teaching, entertainment, and office work) have been moved online to ensure their continued operation.
[0003] However, when conducting online activities, user engagement is weak because the actual environment users are in does not match the requirements of offline activities. Therefore, how to improve user engagement in online activities is an urgent problem to be solved. Summary of the Invention
[0004] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this disclosure provides an image processing method, apparatus, electronic device and storage medium.
[0005] According to one aspect of this disclosure, an image processing method is provided, comprising:
[0006] Acquire first image data; the first image data includes the target object and the background;
[0007] The first image data is cut out to obtain the image of the target object;
[0008] Obtain the virtual scene;
[0009] Determine the display data of the image of the target object in the virtual scene, wherein the display data includes at least one of the position, orientation, and size of the image of the target object in the virtual scene;
[0010] Based on the displayed data, the image of the target object is composited with the virtual scene;
[0011] The synthesized virtual scene is then image-captured to obtain the image to be displayed.
[0012] According to another aspect of this disclosure, an image processing apparatus is provided, comprising:
[0013] A first acquisition module is used to acquire first image data; the first image data includes a target object and a background.
[0014] The image cutout module is used to cut out the first image data to obtain the image of the target object;
[0015] The acquisition module is used to acquire virtual scenes;
[0016] A determining module is used to determine the display data of the image of the target object in the virtual scene, wherein the display data includes at least one of the position, orientation, and size of the image of the target object in the virtual scene;
[0017] A compositing module is used to composite the image of the target object with the virtual scene based on the displayed data;
[0018] The second acquisition module is used to acquire images of the synthesized virtual scene to obtain the image to be displayed.
[0019] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0020] Processor; and
[0021] Stored program memory,
[0022] The program includes instructions that, when executed by the processor, cause the processor to perform the image processing method as described above.
[0023] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the image processing method as described above.
[0024] One or more technical solutions provided in this application embodiment simulate offline activity scenarios by constructing a virtual world, and make the virtual world contain logically reasonable images that correspond to real users, and enable users and their corresponding images to interact, which can improve the user's sense of participation. Attached Figure Description
[0025] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0026] Figure 1 A flowchart of an image processing method provided in an embodiment of this disclosure;
[0027] Figure 2 A schematic diagram of a first image data provided in an embodiment of this disclosure;
[0028] Figure 3 A schematic diagram of an image of a virtual scene provided in an embodiment of this disclosure;
[0029] Figure 4 A schematic diagram of an image to be displayed, provided for an embodiment of this disclosure;
[0030] Figure 5This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present disclosure;
[0031] Figure 6 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0032] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0033] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0034] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0035] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0036] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0037] Figure 1This is a flowchart illustrating an image processing method provided in an embodiment of this disclosure. This embodiment is applicable to image processing performed on a client-side basis. The method can be executed by an image processing device, which can be implemented in software and / or hardware. This device can be configured in an electronic device, such as a terminal, including but not limited to smartphones, PDAs, tablets, wearable devices with displays, desktop computers, laptops, all-in-one computers, and smart home devices. Alternatively, this embodiment is applicable to image processing performed on a server-side basis. The method can be executed by an image processing device, which can be implemented in software and / or hardware. This device can be configured in an electronic device, such as a server. The image processing method provided in this application can be applied to various scenarios such as online teaching, office work, and entertainment.
[0038] In this application, for ease of understanding, a teaching scenario is used as an example to illustrate the image processing method provided. See also Figure 1 The image processing method includes:
[0039] S110. Acquire first image data; the first image data includes the target object and the background.
[0040] The target object refers to the object that is ultimately intended to be presented in the image to be displayed, that is, the participants in the activity of the real scene simulated by the virtual scene. For example, the target object can be a person. For example, if we want to simulate an offline class scene and create a scene of students and teachers having class in a classroom in the virtual world, both the teacher and the students are the target objects.
[0041] In one embodiment, the target object is the user of the terminal that performs the image processing method provided in this application.
[0042] There are multiple ways to implement this step, and this application does not limit this one. For example, if the executing entity of this step is a server, optionally, the implementation method of this step includes: receiving first image data from a terminal, whereby the terminal is used to acquire images of a target object in the real world. If the executing entity of this step is a terminal, optionally, the implementation method of this step includes: acquiring images of the target object in the real world using an image acquisition device to obtain the first image data. The image acquisition device can be built into the terminal, or configured to communicate with the terminal.
[0043] Optionally, the camera function can be invoked using the Unity3D real-time engine editor to capture images of real-world target objects. The specific acquisition process can be implemented using APIs (Application Programming Interfaces) such as Application.RequestUserAuthorization, WebCamTexture, and WebCamDevice.
[0044] It should be noted that in practice, in online classes, teachers and students are scattered in different geographical locations in the real world. Multiple terminals are needed to collect images of teachers and students, resulting in multiple first image data. These multiple first image data need to be processed (such as image matting and compositing).
[0045] S120. Perform image matting on the first image data to obtain the image of the target object.
[0046] There are multiple ways to implement this step, and this application does not limit this one. In one embodiment, the background in the first image data is a solid color background; the specific implementation method of this step includes: clustering the pixels in the first image data to obtain a first clustering result and a second clustering result, wherein the color of the pixels in the first clustering result is the same as the background color, and the color of the pixels in the second clustering result is different from the background color; removing the pixels in the first clustering result from the first image data to obtain the image of the target object. This setting can reduce the difficulty of image cutout, thereby improving the efficiency of image processing.
[0047] Furthermore, the system can be configured to cluster the pixels in the first image data to obtain a first clustering result and a second clustering result, including: determining the background color; determining color judgment conditions based on the background color; setting the alpha channel of pixels in the first image data that meet the color judgment conditions to a first value; setting the alpha channel of pixels in the first image data that do not meet the color judgment conditions to a second value; clustering the pixels with the alpha channel set to the first value into one group and the pixels with the alpha channel set to the second value into another group.
[0048] The color judgment condition is used to determine whether the color of a pixel matches the background color, and whether it can be identified as a pixel that constitutes the background. If a pixel meets the color judgment condition, it means that the pixel can be regarded as a pixel that constitutes the background.
[0049] The purpose of "setting the alpha channel of pixels in the first image data that meet the color judgment criteria to a first value" is to mark these pixels. The purpose of "setting the alpha channel of pixels in the first image data that do not meet the color judgment criteria to a second value" is to mark these pixels. The first value and the second value are different. Therefore, pixels with the alpha channel set to the first value are the background pixels. Pixels with the alpha channel set to the second value are the target object pixels. Subsequently, pixels with the alpha channel set to the first value need to be removed from the first image data. Optionally, the pixels with the alpha channel set to the first value can be made transparent to obtain only pixels with the alpha channel set to the second value, thus obtaining the image of the target object.
[0050] For example, before image acquisition, the user can be prompted to set up the shooting location to form a solid color background (such as blue, red, or green). After the solid color background is set up, the target object enters the set-up shooting location, and image acquisition is performed to obtain the first image data with a solid color background.
[0051] Assuming the background color is green, when determining the color judgment criteria, we can set the G value in the RGB values to be greater than a first set threshold, the R value to be less than a second set threshold, and the B value to be less than a third set threshold. Subsequently, when determining whether each pixel in the first image data meets the color judgment criteria, we can directly compare the RGB values of each pixel with the color judgment criteria.
[0052] Alternatively, image matting can be achieved using ChromaKey Shader key code.
[0053] Considering that directly applying transparency values to the first image data based on the alpha channel values of each pixel in the first image data would result in a heavily jagged edge appearance in the second image data, the following alternative approach can be taken: First, after setting a first or second value for the alpha channel values of each pixel in the first image data, the following steps are performed: First, obtain the first image data after adjusting the alpha channel; second, normalize the first image data after adjusting the alpha channel, adjusting the maximum and minimum grayscale values to the (0, 1) range; third, determine the non-feathered regions in the first image data after adjusting the alpha channel; fourth, perform a feathering operation on the non-feathered regions in the shader to blur the edges of the target object in the first image data after adjusting the alpha channel. During the feathering operation, interpolation methods can be used to address the jagged edge issue. Furthermore, during the feathering operation, by setting the parameters appropriately, the final second image data is made to closely resemble the real situation and is visually harmonious. In the fifth step, based on the transparency channel values of each pixel in the first image data, the feathered first image data is made transparent to obtain the second image data that only includes the target object.
[0054] During the feathering operation in the shader, at least one of the following parameters can be set: KeyColor, TintColor, Cutoff, ColorFeatering, Sharpening, DespillStrength, DespillLuminanceAdd, and RenderQueue.
[0055] Here, KeyColor is the target color value, i.e., the color value of the target image being cut out. TintColor is the original color value, i.e., the color value not being cut out. Cutoff sets the transparency, with a value of 0 or 1, where 0 represents opaque and 1 represents transparent. ColorFeatering is used to feather the color, creating a hazy effect at the edges of the selected area; the larger the color feathering, the greater the haziness. Sharpening is the sharpness, with a default value of 0.5. DespillStrength is the color intensity, with a default value of 0.5. DespillLuminanceAdd represents the increase in dehazing brightness, with a default value of 0. RenderQueue represents the rendering queue, with a default value of 3000+ (transparent).
[0056] In this step, the image of the target object is a two-dimensional image.
[0057] S130, Obtain the virtual scene.
[0058] A virtual scene is a part of a virtual world used to simulate real-world activities. A virtual scene is a three-dimensional environment. It can include virtual objects, such as physical objects or people.
[0059] There are various ways to implement this step, and this application does not limit the specific implementation. For example, the implementation method of this step may include: loading the selected virtual scene in response to the selection operation of the virtual scene.
[0060] Alternatively, this step can also include: in response to a selection operation of virtual scene configuration materials, generating a virtual scene based on the selected virtual scene configuration materials. For example, after the user selects virtual scene configuration materials, the virtual scene configuration materials are imported using the Unity3D real-time engine editor to build a virtual world, thereby obtaining the virtual scene.
[0061] S140. Determine the display data of the target object's image in the virtual scene. The display data includes at least one of the target object's image's position, orientation, and size in the virtual scene.
[0062] The position of the target object's image in the virtual scene refers to information describing where the target object should be placed in the virtual scene. The orientation of the target object's image in the virtual scene refers to information describing which direction the target object should face in the virtual scene. The size of the target object's image in the virtual scene refers to information describing the size the target object should appear in the virtual scene.
[0063] There are various ways to implement this step, and this application does not limit the specific implementation. In one embodiment, the implementation method of this step includes: acquiring the pose data of the target object; and determining the display data of the target object's image in the virtual scene based on the pose data of the target object. The pose data of the target object refers to data describing the position and / or pose number of the target object in the real world.
[0064] Those skilled in the art will understand that, in practice, for a virtual scene to simulate a real offline activity, it is necessary to ensure that the composite result of the target object's image and the virtual scene is reasonable and consistent with the real situation. For example, if a virtual scene simulates a teacher giving a physical education class on a playground, with students in fixed positions and the teacher running towards them from a distance (the teacher being the target object), in the real world, the teacher's position will move closer to the students over time, making the teacher appear larger to the students. Obtaining the target object's pose data means re-determining the display data of the target object's image in the virtual scene each time this application is executed, to ensure that the final composite result is reasonable, consistent with the real situation, and reflects the user's (i.e., the target object's) intent.
[0065] Furthermore, the setup can be configured to acquire the pose data of the target object, including: acquiring pose data collected by a pose detection device worn on the target object; and obtaining the pose data of the target object based on the pose data collected by the detection device. The pose detection device includes, but is not limited to, electronic devices with pose data detection functions such as handles, head-mounted devices, smartwatches, and smart glasses. This setup utilizes existing detection equipment to obtain the pose data of the target object, reducing the difficulty of acquiring this data.
[0066] Furthermore, if the displayed data includes the position of the target object's image in the virtual scene, determining the display data of the target object's image in the virtual scene includes: obtaining the target object's position information in the real scene; obtaining the size conversion relationship between the virtual scene and the real scene: based on the target object's position information in the real scene and the size conversion relationship between the virtual scene and the real scene, determining the position information of the target object's image in the virtual scene. The purpose of this setting is to map the position in the real world to the position in the virtual world, which facilitates allowing users to control the movement of their image in the virtual environment by moving in the real world, giving users a sense of immersion.
[0067] Furthermore, based on the target object's position information in the real scene and the size conversion relationship between the virtual and real scenes, the position information of the target object's image in the virtual scene is determined. This includes: obtaining the origin of the virtual scene and the calibration point in the real scene; the correspondence between the calibration point and the origin; determining the relative positional relationship between the target object and the calibration point based on the target object's position information in the real scene; determining the relative positional relationship between the target object's image and the origin based on the relative positional relationship between the target object and the calibration point, and the size conversion relationship between the virtual and real scenes; and determining the position information of the target object's image in the virtual scene based on the relative positional relationship between the target object's image and the origin. The origin of the virtual world coordinate system is the origin of the virtual world. The origin of the real world coordinate system is the calibration point in the real world. The size conversion relationship between the real and virtual worlds is the conversion relationship between the virtual world coordinate system and the real world coordinate system. This allows for a one-to-one correspondence between points in the real world and points in the virtual world, thus ensuring a one-to-one correspondence between the target object's position in the real world and its position in the virtual scene. This improves the rationality of the interaction between the user and the corresponding image, giving the user a sense of immersion.
[0068] S150. Based on the display data, the image of the target object is composited with the virtual scene.
[0069] Optionally, this step can be implemented by placing the image of the target object in a virtual scene to ensure that the effect of the target object in the virtual scene matches the displayed data. This makes the final virtual scene, including the image of the target object, logical and reasonable, thus improving the user experience.
[0070] Furthermore, the display data includes the position, orientation, and size of the target object's image in the virtual scene. The specific implementation method for this step includes: adjusting the size of the target object's image so that its size matches the size in the display data; placing the adjusted target object's image in the virtual scene so that its position and orientation match the display data. This ensures that the effect of the target object in the virtual scene is consistent with the display data, guaranteeing that the final composite result is reasonable and consistent with reality.
[0071] S160. Acquire images of the synthesized virtual scene to obtain the images to be displayed.
[0072] Optionally, a virtual camera can be used to capture images of the synthesized virtual scene to obtain the image to be displayed.
[0073] A virtual camera is used to simulate an observer. The position of the virtual camera corresponds to the observer's position. The image captured by the virtual camera corresponds to what the observer's eyes see. The virtual camera can be placed anywhere in the virtual environment and can be positioned in any manner. Changing the virtual camera's position, rotation angle, pitch angle, and other parameters can alter the observer's field of view.
[0074] The above technical solution involves: acquiring first image data, including a target object and a background; performing image matting on the first image data to obtain the image of the target object; acquiring a virtual scene to depict a virtual world; determining the display data of the target object's image in the virtual scene, including at least one of the target object's image's position, orientation, and size in the virtual scene; compositing the target object's image with the virtual scene based on the display data; and acquiring images of the composite virtual scene to obtain the image to be displayed. Essentially, this technical solution simulates offline activity scenarios by constructing a virtual environment, creating logically consistent images within the virtual environment that correspond to real users, and allowing interaction between users and their corresponding images, thereby enhancing user engagement.
[0075] Based on the above technical solution, optionally, the virtual scene is associated with compositing conditions; S140 further includes: if the pose data of the target object meets the compositing conditions associated with the virtual scene, determining the display data of the target object's image in the virtual scene. Compositing conditions are used to limit under what conditions the image of the target object is allowed to be composited with the virtual scene. The purpose of setting compositing conditions is to ensure that the final composite result is reasonable. This application does not limit the specific content of the compositing conditions.
[0076] Optionally, in one embodiment, the virtual camera is configured to only allow changes to parameters such as position, rotation angle, and pitch angle within a specified range. In this case, optionally, the compositing conditions include the shooting angle of the target object's image. The shooting angle of the target object's image refers to the relative positional relationship between the front of the target object and the image acquisition device acquiring the first image data. For example, a shooting angle of 0° indicates that the front of the target object is facing the image acquisition device when acquiring the first image data. A shooting angle of 180° indicates that the back of the target object is facing the image acquisition device when acquiring the first image data.
[0077] For example, if a virtual scene is used to simulate a soccer goalkeeper standing at the goalpost, waiting for the ball to come, the goalkeeper should stand with his back to the goal and facing the ball. The virtual camera is limited to capturing images only from the direction directly facing the goal, ensuring that the displayed image includes a frontal view of the goalkeeper. To this end, the compositing condition associated with this virtual scene is set to a shooting angle of 0°. If the shooting angle of the target object's image is 0°, satisfying the compositing condition associated with the virtual scene, the display data for the target object's image in the virtual scene is determined, and the target object's image is subsequently composited with the virtual scene. If the shooting angle of the target object's image is 90°, not satisfying the compositing condition associated with the virtual scene, the display data for the target object's image in the virtual scene is not determined, and the target object's image is not subsequently composited with the virtual scene.
[0078] Furthermore, a virtual scene can be set up to associate one or more compositing conditions.
[0079] Furthermore, when multiple compositing conditions are associated with a virtual scene, optionally, different compositing conditions correspond to different identity restriction conditions. Before determining the display data of the target object's image in the virtual scene, the method further includes: determining the target object's identity information; based on the target object's identity information, determining the target compositing condition among the multiple compositing conditions associated with the virtual scene, wherein the target object's identity information satisfies the identity restriction conditions of the target compositing condition; and if the target object's pose data satisfies the target compositing condition, determining the display data of the target object's image in the virtual scene.
[0080] For example, the compositing conditions associated with virtual scene A include shooting angles of 0° and 180°. The identity restriction condition corresponding to a shooting angle of 0° is "teacher," and the identity restriction condition corresponding to a shooting angle of 180° is "student." This means it's allowed to composite a student's reversed image with virtual scene A, and to composite a teacher's frontal image with virtual scene A. If the target object's identity information in the image is "student," then a shooting angle of 180° is used as the target compositing condition. The system determines whether the shooting angle of the target object's image is 180°. If yes, the target compositing condition is met, and the display data for the target object's image in the virtual scene is determined. Otherwise, the display data for the target object's image in the virtual scene is not determined.
[0081] Figure 2 This is a schematic diagram of a first image data provided in an embodiment of the present disclosure. Figure 3 This is a schematic diagram of an image of a virtual scene provided in an embodiment of this disclosure. Figure 4 This is a schematic diagram of an image to be displayed, provided as an embodiment of this disclosure. In a physical education class, the teacher is demonstrating a movement. An image of the teacher demonstrating the movement is captured, and the result is as follows... Figure 2 As shown, the first image data is obtained. The first image data is then cut out to obtain the teacher's image. A virtual scene corresponding to the physical education class is then acquired. Images of this virtual scene are captured, and the results are as follows. Figure 3 As shown. The display data for the teacher's image in the virtual scene is determined. The teacher's image is placed in the virtual scene to ensure that the teacher's appearance in the virtual scene matches the display data. Layer acquisition is performed on the composited virtual scene to obtain the image to be displayed. The image to be displayed is as follows. Figure 4 As shown.
[0082] It should be noted that in some cases, virtual scenes may also include virtual objects. For the composited virtual scene, at any given time, the position and size of the virtual objects within the virtual scene are fixed, and the position and size of the target object's image within the virtual scene are also fixed. By acquiring images of the composited virtual scene, the resulting image to be displayed can reflect the front-back positional relationship between the virtual objects and the target object's image. That is, using the technical solution provided in this application, the occlusion relationship between virtual objects and the target object can be accurately presented.
[0083] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0084] Figure 5 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this disclosure. See also... Figure 5 The image processing apparatus includes:
[0085] The first acquisition module 310 is used to acquire first image data; the first image data includes a target object and a background.
[0086] The image cutout module 320 is used to cut out the first image data to obtain the image of the target object;
[0087] Module 330 is used to acquire virtual scenes;
[0088] The determining module 340 is used to determine the display data of the image of the target object in the virtual scene, wherein the display data includes at least one of the position, orientation, and size of the image of the target object in the virtual scene;
[0089] The compositing module 350 is used to composite the image of the target object with the virtual scene based on the displayed data;
[0090] The second acquisition module 360 is used to acquire images of the synthesized virtual scene to obtain images to be displayed.
[0091] Further, a module is identified for:
[0092] Obtain the pose data of the target object;
[0093] Based on the pose data of the target object, the display data of the image of the target object in the virtual scene is determined.
[0094] Further, a module is identified for:
[0095] Acquire pose data collected by a pose detection device, which is worn on the target object;
[0096] Based on the pose data collected by the detection device, the pose data of the target object is obtained.
[0097] Furthermore, the virtual scene is associated with composition conditions; the determining module is used for:
[0098] If the pose data of the target object satisfies the compositing conditions associated with the virtual scene, the display data of the image of the target object in the virtual scene is determined.
[0099] Furthermore, the virtual scene is associated with multiple synthesis conditions, and different synthesis conditions correspond to different identity restrictions. The determining module is used for:
[0100] Determine the identity information of the target object;
[0101] Based on the identity information of the target object, a target synthesis condition is determined from multiple synthesis conditions associated with the virtual scene, wherein the identity information of the target object satisfies the identity restriction condition of the target synthesis condition;
[0102] If the pose data of the target object satisfies the target synthesis conditions, the display data of the image of the target object in the virtual scene is determined.
[0103] Furthermore, the synthesis module is used for:
[0104] The image of the target object is placed in the virtual scene so that the effect of the target object in the virtual scene is consistent with the display data.
[0105] Furthermore, the display data includes the position, orientation, and size of the target object's image in the virtual scene; the compositing module is used for:
[0106] The size of the target object's image is adjusted so that the size of the target object's image after the size adjustment is consistent with the size in the display data;
[0107] The image of the target object, after being resized, is placed in the virtual scene such that the position of the image of the target object in the virtual scene is consistent with its position in the display data, and the orientation of the image of the target object in the virtual scene is consistent with its orientation in the display data.
[0108] Furthermore, the displayed data includes the position of the target object's image in the virtual scene, and the determination module is used for:
[0109] Obtain the location information of the target object in the real scene;
[0110] Obtain the size conversion relationship between the virtual scene and the real scene:
[0111] Based on the location information of the target object in the real scene and the size conversion relationship between the virtual scene and the real scene, the location information of the image of the target object in the virtual scene is determined.
[0112] Further, a module is identified for:
[0113] Obtain the origin of the virtual scene and the calibration point of the real scene; the calibration point corresponds to the origin;
[0114] Based on the location information of the target object in the real scene, the relative positional relationship between the target object and the calibration point is determined;
[0115] Based on the relative positional relationship between the target object and the calibration point, and the size conversion relationship between the virtual scene and the real scene, the relative positional relationship between the image of the target object and the origin is determined;
[0116] Based on the relative positional relationship between the image of the target object and the origin, the positional information of the image of the target object in the virtual scene is determined.
[0117] Furthermore, the image cutout module is used for:
[0118] The pixels in the first image data are clustered to obtain a first clustering result and a second clustering result. The color of the pixels in the first clustering result is the same as the color of the background, while the color of the pixels in the second clustering result is different from the color of the background.
[0119] The image of the target object is obtained by removing the pixels from the first clustering result in the first image data.
[0120] Furthermore, the image cutout module is used for:
[0121] Determine the color of the background;
[0122] Based on the background color, determine the color judgment criteria;
[0123] Set the alpha channel of pixels in the first image data that meet the color judgment condition to a first value; set the alpha channel of pixels in the first image data that do not meet the color judgment condition to a second value;
[0124] Pixels with the first alpha channel value are grouped into one group, and pixels with the second alpha channel value are grouped into another group.
[0125] The apparatus disclosed in the above embodiments can implement the process flow of the methods disclosed in the above method embodiments and has the same or corresponding beneficial effects. To avoid repetition, it will not be described again here.
[0126] Exemplary embodiments of this disclosure also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of this disclosure.
[0127] Exemplary embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to embodiments of this disclosure.
[0128] Exemplary embodiments of this disclosure also provide a computer program product, including a computer program, wherein, when executed by a processor of a computer, the computer program is used to cause the computer to perform a method according to an embodiment of this disclosure.
[0129] refer to Figure 6 The present invention describes a structural block diagram of an electronic device 800 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0130] like Figure 6As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0131] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, output unit 807, storage unit 808, and communication unit 809. Input unit 806 can be any type of device capable of inputting information to electronic device 800. Input unit 806 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 807 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 804 may include, but is not limited to, disk and optical disk. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0132] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above. For example, in some embodiments, the image processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. In some embodiments, the computing unit 801 may be configured to perform image processing by any other suitable means (e.g., by means of firmware).
[0133] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0134] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0135] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0136] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0137] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0138] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
Claims
1. An image processing method, characterized in that, include: Acquire the first image data; The first image data includes the target object and the background; The first image data is cut out to obtain the image of the target object; Obtain the virtual scene; Determine the display data of the image of the target object in the virtual scene, wherein the display data includes at least one of the position, orientation, and size of the image of the target object in the virtual scene; Based on the displayed data, the image of the target object is composited with the virtual scene; The synthesized virtual scene is subjected to image acquisition to obtain the image to be displayed; The data used to determine the display of the image of the target object in the virtual scene includes: The pose data of the target object is obtained; the virtual scene is associated with compositing conditions; If the pose data of the target object satisfies the compositing conditions associated with the virtual scene, determine the display data of the image of the target object in the virtual scene; The virtual scene is associated with multiple synthesis conditions, and different synthesis conditions correspond to different identity restriction conditions.
2. The image processing method according to claim 1, characterized in that, The step of obtaining the pose data of the target object includes: Acquire pose data collected by a pose detection device, which is worn on the target object; Based on the pose data collected by the detection device, the pose data of the target object is obtained.
3. The image processing method according to claim 1, characterized in that, Before determining the display data of the image of the target object in the virtual scene, the method further includes: Determine the identity information of the target object; Based on the identity information of the target object, a target synthesis condition is determined from multiple synthesis conditions associated with the virtual scene, wherein the identity information of the target object satisfies the identity restriction condition of the target synthesis condition; If the pose data of the target object satisfies the target synthesis conditions, the display data of the image of the target object in the virtual scene is determined.
4. The image processing method according to claim 1, characterized in that, The step of compositing the image of the target object with the virtual scene based on the displayed data includes: The image of the target object is placed in the virtual scene so that the effect of the target object in the virtual scene is consistent with the display data.
5. The image processing method according to claim 4, characterized in that, The display data includes the position, orientation, and size of the target object's image in the virtual scene; the step of compositing the target object's image with the virtual scene based on the display data includes: The size of the target object's image is adjusted so that the size of the target object's image after the size adjustment is consistent with the size in the display data; The image of the target object, after being resized, is placed in the virtual scene such that the position of the image of the target object in the virtual scene is consistent with its position in the display data, and the orientation of the image of the target object in the virtual scene is consistent with its orientation in the display data.
6. The image processing method according to claim 1, characterized in that, The display data includes the position of the target object's image in the virtual scene, and determining the display data of the target object's image in the virtual scene includes: Obtain the location information of the target object in the real scene; Obtain the size conversion relationship between the virtual scene and the real scene: Based on the location information of the target object in the real scene and the size conversion relationship between the virtual scene and the real scene, the location information of the image of the target object in the virtual scene is determined.
7. The image processing method according to claim 6, characterized in that, The step of determining the position information of the target object's image in the virtual scene based on the target object's position information in the real scene and the size conversion relationship between the virtual scene and the real scene includes: Obtain the origin of the virtual scene and the calibration point of the real scene; the calibration point corresponds to the origin; Based on the location information of the target object in the real scene, the relative positional relationship between the target object and the calibration point is determined; Based on the relative positional relationship between the target object and the calibration point, and the size conversion relationship between the virtual scene and the real scene, the relative positional relationship between the image of the target object and the origin is determined; Based on the relative positional relationship between the image of the target object and the origin, the positional information of the image of the target object in the virtual scene is determined.
8. The image processing method according to claim 1, characterized in that, The background in the first image data is a solid color; the step of cutting out the first image data to obtain the image of the target object includes: The pixels in the first image data are clustered to obtain a first clustering result and a second clustering result. The color of the pixels in the first clustering result is the same as the color of the background, while the color of the pixels in the second clustering result is different from the color of the background. The image of the target object is obtained by removing the pixels from the first clustering result in the first image data.
9. The image processing method according to claim 8, characterized in that, The step of clustering the pixels in the first image data to obtain a first clustering result and a second clustering result includes: Determine the color of the background; Based on the background color, determine the color judgment criteria; Set the alpha channel of pixels in the first image data that meet the color judgment condition to a first value; set the alpha channel of pixels in the first image data that do not meet the color judgment condition to a second value; Pixels with the first alpha channel value are grouped into one group, and pixels with the second alpha channel value are grouped into another group.
10. An image processing apparatus, characterized in that, include: The first acquisition module is used to acquire the first image data; The first image data includes the target object and the background; The image cutout module is used to cut out the first image data to obtain the image of the target object; The acquisition module is used to acquire virtual scenes; A determining module is used to determine the display data of the image of the target object in the virtual scene, wherein the display data includes at least one of the position, orientation, and size of the image of the target object in the virtual scene; A compositing module is used to composite the image of the target object with the virtual scene based on the displayed data; The second acquisition module is used to acquire images of the synthesized virtual scene to obtain images to be displayed. The data used to determine the display of the image of the target object in the virtual scene includes: The pose data of the target object is obtained; the virtual scene is associated with compositing conditions; If the pose data of the target object satisfies the compositing conditions associated with the virtual scene, determine the display data of the image of the target object in the virtual scene; The virtual scene is associated with multiple synthesis conditions, and different synthesis conditions correspond to different identity restriction conditions.
11. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-9.