Generating virtual representations using media assets
By capturing sensor data during the registration process and supplementing it with user media content, and combining role networks and visual artifact networks, the efficiency and accuracy of virtual representation generation on power-constrained devices are solved, enabling virtual representation supplementation when data is insufficient.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2026-04-10
AI Technical Summary
Existing avatar generation systems are unable to efficiently generate accurate virtual human, animal, or plant life representations on power-constrained mobile devices, and struggle to generate accurate virtual representations when data capture is incomplete.
By capturing sensor data during the registration process and supplementing the sensor data with the user's media content, a virtual representation is generated. An accurate virtual representation is generated by combining a role network and a visual artifact network.
It improves the efficiency and accuracy of generating virtual representations on power-constrained devices, and can generate more accurate virtual representations by supplementing digital assets when data is insufficient.
Smart Images

Figure CN121844359A_ABST
Abstract
Description
Background Technology
[0001] The computerized character representing a user is often called an avatar. Avatars can take many forms, including virtual humans, animals, and plant life. Existing avatar generation systems often fail to accurately represent users, require high-performance general-purpose and graphics processors, and are generally not very effective on power-constrained mobile devices such as smartphones or computing tablets.
[0002] Sometimes, difficulties arise in generating realistic-looking avatars because avatars are based on data captured at a specific time. Therefore, improvements are needed to avatar generation. Attached Figure Description
[0003] Figure 1 A flowchart is shown illustrating a technique for generating object characters and accessories according to some implementation schemes.
[0004] Figure 2 A flowchart is shown for a technique for supplementing virtual representation data with media content, according to one or more embodiments.
[0005] Figure 3 A flowchart is shown of a technique for determining geometric virtual representation data according to some implementation schemes.
[0006] Figure 4 A flowchart is shown for a technique for predicting joint position based on derived subsequent data, according to some implementation schemes.
[0007] Figure 5 An example network diagram for sharing digital assets is depicted according to one or more implementation schemes.
[0008] Figure 6 A simplified system diagram according to one or more implementation schemes is shown in block diagram form.
[0009] Figure 7 A computer system according to one or more implementation schemes is shown in block diagram form. Detailed Implementation
[0010] This disclosure relates in its entirety to techniques for enhancing registration for generating virtual representations of objects. More specifically, but not limitingly, this disclosure relates to techniques and systems for enhancing virtual representations of objects using media that capture image data of the objects.
[0011] This disclosure relates to systems, methods, and computer-readable media for generating a virtual representation of an object by capturing sensor data during the registration process and supplementing the sensor data used to generate the virtual representation with additional media content including the user. In some embodiments, the virtual representation is generated to present an accurate or realistic representation of the object. During the registration process, the user can use a personal device to capture one or more images or other sensor data pointing at the user, from which registration data can be derived. In some cases, the sensor data captured during the registration process can be enhanced by incorporating additional media items of the user, such as stored images. Additionally, the user's appearance at the time of registration may not be an accurate representation of the user's typical appearance, for example, due to temporary conditions such as scars or other skin conditions, clothing that substantially covers the face or other parts of the body, etc.
[0012] The embodiments described herein generate a virtual representation of an object by capturing sensor data of the object during the registration process. Furthermore, the device can acquire digital assets containing the object on local or remote storage. Modern consumer electronics devices enable users to accumulate large amounts of digital assets (e.g., images, videos, etc.). These digital assets can be tagged with specific individuals. For example, facial recognition can be used to detect digital assets including the user. The embodiments described herein can utilize these digital assets to obtain additional data from which an alternative but still accurate virtual representation of the user can be generated.
[0013] In some implementations, a user's digital assets may be processed to extract visual artifacts of the user. These visual artifacts may include, for example, texture components such as clothing and makeup. Additionally, visual artifacts may include geometric artifacts such as hairstyles, large jewelry, or other components with geometric elements. In some implementations, these visual artifacts may be used as input to a persona network to generate a virtual representation of the user. Additionally or alternatively, these visual artifacts may be incorporated into an accessory library from which the user can modify accessories or other components of their virtual representation in a manner that depicts a realistic version of an object.
[0014] In some implementations, digital assets can be used to supplement sensor data captured during registration when sensor data is insufficient. For example, it may be determined that the sensor data fails to meet one or more quality metrics. For instance, a portion of a user may not have been well captured by the sensor data. Instead, image data of that portion of the user can be identified in digital assets and supplemented into the persona network to generate a virtual representation.
[0015] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the disclosed concepts. As part of this specification, some of the accompanying drawings of this disclosure are shown in block diagram form to avoid obscuring the novel aspects of the disclosed embodiments. In this context, it should be understood that references to numbered drawing elements without associated identifiers (e.g., 100) refer to all instances of drawing elements with identifiers (e.g., 100a and 100b). Additionally, as part of this specification, some of the accompanying drawings of this disclosure are provided in the form of flowcharts. The boxes in any particular flowchart may be presented in a particular order. However, it should be understood that any particular flow in any flowchart is merely illustrative of one embodiment. In other embodiments, any various components depicted in the flowchart may be omitted, or components may be performed in a different order, or even simultaneously. Furthermore, other embodiments may include additional steps not depicted as part of a flowchart. The language used in this disclosure has been chosen primarily for readability and instructional purposes and may not have been chosen to describe or limit the disclosed subject matter. In this disclosure, reference to “an implementation” or “implementation” means that at least one implementation includes a particular feature, structure or characteristic described in connection with that implementation, and multiple references to “an implementation” or “implementation” should not be construed as necessarily referring to the same or different implementations.
[0016] It should be understood that in any actual implementation of development (as in any development project), numerous decisions must be made to achieve the developer's specific goals (e.g., compliance with system and business-related constraints), and these goals will vary between different implementations. It should also be understood that such development work can be complex and time-consuming, but nevertheless, it remains routine work for those of ordinary skill in the art who benefit from the image captures provided in this disclosure.
[0017] The term "digital asset" (DA) refers to data / information bundled or grouped in a manner meaningfully presented by computing devices for viewing, reading, and / or listening by humans or other computing devices / machines / electronic devices. Digital assets may include media items such as photographs, recordings, and data objects (or simply "objects"), as well as video and audio files. Image data associated with photographs, recordings, data objects, and / or video files may include information or data necessary for electronic devices to display or render images (such as photographs) and videos. Audiovisual data may include information or data necessary for electronic devices to present videos and content with visual and / or auditory components.
[0018] For the purposes of this application, the term "role" refers to a virtual representation of a subject generated to accurately reflect the subject's physical characteristics, movement, etc.
[0019] For the purposes of this application, the term "coexistence environment" refers to a shared extended reality (XR) environment among multiple devices. Components within the environment typically maintain consistent spatial relationships to preserve spatial truth.
[0020] Go to Figure 1 The diagram illustrates flowcharts for generating object roles and accessories according to some implementation schemes. For illustrative purposes, various processes are depicted and described as being performed by specific components. However, it should be understood that various actions may be performed by alternative components. Furthermore, some actions may be performed simultaneously, and some actions may be unnecessary, or additional actions may be added.
[0021] The flowchart begins with registration data 100. According to one or more embodiments, registration data 100 may include sensor data captured during the registration process, from which a virtual representation of the object captured in the registration data is derived. Registration data may include, for example, image data and depth data. Image data may be captured by one or more user-facing cameras, which capture one or more images of the object. In some embodiments, one or more images of a user performing one or more facial expressions may be captured, and these images may be captured from one or more viewpoints associated with the object. Depth data may be captured by one or more depth sensors simultaneously with the image data being captured. The depth data may indicate the relative depth of the object's surface to the viewpoint of the device capturing the sensor data.
[0022] Registration data 100 is applied to a character network 110 to generate one or more object characters 130. According to one or more embodiments, the character network 110 can use the registration data 100 to predict the user's geometric and textural characteristics, enabling the generation of accurate representations of objects in the form of one or more object characters 130. For example, the character network 110 can utilize registration data including image data and depth data to determine the geometric representation of the object, such as in the form of a 3D mesh, texture, etc. In some embodiments, other characteristics can be determined during registration to enhance the character in a virtual environment, such as a coexistence environment. For example, a skeleton, including the positions of individual joints and the relationships between them, can be determined. The determination of the skeleton can be performed by the character network 110 or another network. The skeleton can be used to drive the character during runtime so that the character moves in a manner that accurately represents the movement of an object.
[0023] According to one or more embodiments, the role network 110 further considers visual artifacts from the object digital asset 105. That is, the registration data 100 can be supplemented with additional data derived from the digital assets including the object. Digital assets (DAs) may include media content, such as image data, video data, etc. Digital assets may be in the form of 2D or 3D image data and may include metadata indicating the context associated with the image, or stored together with metadata indicating the context associated with the image, such as people or objects detected in the image or otherwise identified as being present in the image. Digital assets may be obtained from a variety of sources. For example, in different scenarios, DAs may be stored locally, stored on a server, or a combination thereof. Thus, DAs can be processed locally to identify DAs with objects. Alternatively, DAs may be requested from remote sources such as individual client devices, servers, etc. In some embodiments, the identity of the object in the registration data 100 can be determined or obtained. For example, if the registration data is performed during a registration process for a specific user account, that user account can be used to identify digital assets including the user associated with that user account. Alternatively, unique identification information of objects can be derived from registration data 100, and this unique identification information can then be used, for example, using facial recognition technology, to identify a subset of digital assets containing objects. In this example, object digital assets 105 include three images: image A 108A, image B 108B, and image C 108C. According to one or more embodiments, each of image A 108A, image B 108B, and image C 108C can be determined to include an object. As shown, each of image A 108A, image B 108B, and image C 108C may include an object captured at different times or in different contexts. Thus, the digital assets can provide additional contextual data about the visual characteristics of objects in registration data 100. According to one or more embodiments, digital assets utilized by the visual artifact network 120 can be selected based on other considerations, such as the determination of the quality of the digital assets, whether the digital assets include portions of objects that are not available in sufficient registration data 100, etc.
[0024] In some embodiments, the visual artifact network 120 may be configured to generate visual artifacts to supplement the registration data 100 for use by the character network 110 in generating the object character 130. For example, the visual artifact network 120 may generate data corresponding to portions of the object for consideration by the character network 110. In some embodiments, the visual artifacts may include portions of the object that are not well captured by the registration data 100. Thus, the visual artifacts provided by the visual artifact network 120 may be provided upon request from a preprocessing procedure for the character network 110, which determines a health score or other similar quality metric for the registration data 100 or portions thereof. The visual artifacts may determine a 2D or 3D representation of the object or a portion thereof based on one or more images including the object or a portion thereof. For example, if registration data for a particular side of a user's face is unavailable, several images of the sides of the user's face may be used to supplement the registration data. Additionally, in the example shown, different hairstyles of the object in images B 108B and C 108C of the object digital asset 105 can be extracted by the visual artifact network 120 and provided to the character network 110. In some embodiments, this additional data can be used to generate alternative characters. As shown here, the hairstyle in character A 132A matches the hairstyle depicted in registration data 100, while the hairstyle in character B 132B is generated based on the hairstyles shown in images B 108B and C 108C. For example, the hairstyle can be generated by the character network 110 by fitting image data of the hairstyle to a geometric representation of the object's head.
[0025] In some implementations, the visual artifact network 120 may additionally use object digital assets 105 to construct a wardrobe or accessory set that a user can use to complement or further personalize a character. Example accessories may include jewelry, clothing, hats or other headwear, scarves, etc. As shown here, the visual artifact network can use object detection in the object digital assets to identify accessories, clothing, or other items shown as being worn by an object. The visual artifact network 120 can be trained to detect these items and generate object accessories in a form that can be used to enhance or complement one or more object characters. In this example, the visual artifact network 120 is shown as having generated accessory A 142A as a necklace worn in image A 108A, and accessory B 142B as a necklace worn in image B 108B. Accessories can be generated by detecting objects in the digital assets, extracting items from the digital assets, and transforming items from image data to texture data or another format that can be used to complement the object character. As an example, character accessories can be generated as textures to cover a portion of the geometry of a virtual representation of an object, or as three-dimensional components to replace or enhance the geometry of a virtual representation of a person, etc. Although Figure 1Object accessory 140 is illustrated as including only items, but the implementation is not limited to this. In some implementations, other features of the user's appearance that are not items detected from object digital assets 105 may be included in object accessory 140. For example, if a user desires the hairstyle in character B 132B instead of the hairstyle in her registered image (e.g., the hairstyle in character A 132A), the hairstyle in character B may be included in object accessory 140 for the user to choose from.
[0026] In some implementations, the character network 110 may provide data about objects to the visual artifact network to facilitate the transformation of image data into textures or geometries. For example, using the geometric characteristics of an object captured by the character network 110, the visual artifact network 120 may use one or more images with specific accessories to generate a texture of an accessory from the digital asset of the object wearing the accessory. Thus, geometric information from the registration data 100 can be used to infer the geometry of the same object in the object digital asset 105 to extract and transform accessories. Consequently, a user can modify one of the object characters 130 to wear a necklace or other accessory that will be actually worn by the user, thereby enhancing the ability to provide an accurate virtual representation of the user that differs from the appearance captured during registration.
[0027] In another example, the persona network 110 can use digital assets captured after registration data to enhance the object persona 130 to maintain an accurate representation of the object. For example, if new physical properties of the object become apparent in the object's digital assets, the visual artifact network 120 can extract those physical properties and feed them into the persona network 110 to provide additional or revised object personas.
[0028] Although registration data for the object's face and upper body is shown, in some embodiments, other parts of the object, such as the object's hands or arms, can be captured during registration. For example, in XR environments, and particularly in coexistence environments, a virtual representation of a user's hand can be used to represent an object interacting with an item. The accuracy of the virtual representation of hands and arms is important to maintain spatial truth and ensure accurate representation. However, hands or arms may be obscured during registration. In some embodiments, the character network can obtain hand or arm information extracted from the object's digital assets 105 by a visual artifact network 120. The image data can then be transformed into texture and / or geometric information by the visual artifact network 120 and used to register the hands of the object character 130.
[0029] Go to Figure 2 This presents a flowchart depicting a technique for generating virtual representations of objects according to one or more implementation schemes. For illustrative purposes, [further details will be provided]. Figure 1The following steps are described in the context of [the document / document]. However, it should be understood that various actions can be performed by alternative components. Furthermore, various actions can be performed in different orders. Additionally, some actions can be performed simultaneously, and some actions may be unnecessary, or additional actions may be added.
[0030] The flowchart begins at box 205, where image data of the object is obtained. In some embodiments, the image data may be captured, for example, during a registration period in which the user uses a personal device to capture an image pointing towards the user's face, from which registration data can be derived for presenting avatar data associated with the user. In some embodiments, additional image data, such as the object's hands or arms, may be captured. The image data may be captured by one or more cameras on the user's device. For example, the user may register on a device (such as a head-mounted display) that the user will use in an XR environment. Alternatively, the object may register on a separate device (such as a mobile device, desktop computer, etc.) communicatively connected to the head-mounted display.
[0031] The flowchart continues to block 210, where depth data corresponding to the image captured at 205 can be obtained. In some embodiments, depth sensor data may be captured simultaneously by one or more depth sensors while capturing an image at block 205. The depth sensor data may indicate the relative depth of the object's surface from the viewpoint of the device capturing the image / sensor data. In some embodiments, the image data may include depth, or depth may be derived from image data. For example, multiple cameras (such as a stereo camera system) may be used to capture images of the object from which depth can be determined. Alternatively, the image data may be captured at least partially by a depth camera that provides both depth information and image information.
[0032] Optionally, at box 215, a virtual representation of the object is generated from the image and depth information. In some embodiments, the virtual representation data may include geometric information, texture information, and other information about the object, which can be used to drive the character during runtime to render an accurate representation of the object. For example, the virtual representation data may include a skeleton such as a series of joints, which can be used to drive the movement of the geometry of the virtual representation to provide an accurate representation of the object's movement. Reference will be made below. Figure 3 A sample technique for generating virtual representation data is described in more detail.
[0033] Flowchart 200 continues at box 220, where media content containing objects is accessed. That is, additional data derived from the DA including the object can supplement the image and / or depth data. Digital assets may be in the form of 2D or 3D image data and may include metadata indicating context associated with the image, or stored together with metadata indicating context associated with the image, such as people or objects detected or otherwise identified as being present in the image. DAs can be obtained from various sources. For example, in different scenarios, DAs may be stored locally, stored on a server, or a combination thereof. Therefore, DAs can be processed locally to identify DAs containing objects. Alternatively, DAs can be requested from remote sources such as individual client devices, servers, etc. In some implementations, DAs containing objects can be obtained by requesting DAs with metadata “tagged” by the object or otherwise indicating the presence of the object in the DA. Alternatively, information related to the object's identity (such as image data or features derived from image data) can be used to perform facial recognition on the DA to identify the DA containing the object.
[0034] At box 225, visual artifacts are extracted from the media content. Visual artifacts may include visual components of objects in the DA. This may include texture characteristics of the object, such as skin tone, makeup, scars, or the absence of scars. Therefore, visual artifacts may include a 2D or 3D representation of an object or a portion of an object based on one or more images that include the object or a portion of the object. Other examples of digital artifacts include components with geometric shapes, such as hairstyles, headwear, etc. Additionally, other examples include items associated with the user that can be used to build a virtual wardrobe, such as clothing, accessories, etc. In some implementations, the character network may provide data about objects to the visual artifact network to facilitate the transformation of image data into textures or geometric shapes. For example, using the geometric characteristics of the object generated by the character network at box 215, the visual artifact network may use one or more images with specific accessories or characteristics to generate the texture of an accessory from a digital asset where the object is wearing an accessory. According to one or more implementations, the DA used to generate the accessory may be selected by the user. For example, the user may select a set of DAs with specific accessories, or a set of DAs that the user wants to present in the character data.
[0035] Optionally, as shown in box 230, the visual artifact network 120 may additionally use DA to construct a wardrobe or accessory set that a user can use to complement or further personalize a character. Example accessories may include jewelry, clothing, hats or other headwear, scarves, etc. For example, the visual artifact network 120 may use object detection in the object's digital assets to identify accessories, clothing, or other items shown as frequently worn by the object. The visual artifact network 120 may be trained to detect these items and generate object accessories in a form that can be used to enhance or complement one or more object characters. As an example, character accessories may be generated as textures to cover a portion of the geometry of the object's virtual representation, or as 3D components to replace or enhance the geometry of the person's virtual representation, etc.
[0036] The flowchart ends at box 235, where virtual representation data and extracted visual artifacts are used to generate a virtual representation of an object. In some implementations, the virtual representation can be generated by using the virtual representation generated at optional step 215 and enhancing or modifying the virtual representation based on visual artifacts (such as by replacing a portion of the virtual representation with generated character accessories or enhancing the virtual representation). Additionally or alternatively, visual artifacts can be fed into character network 110 to enhance the generation of one or more characters. For example, a first character can be generated using registration data without considering additional predefined DAs, while alternative characters can be generated using visual artifacts from the DAs.
[0037] Figure 3 An example flowchart depicts a technique for generating virtual representation data during the registration process, according to one or more implementation schemes. It should be understood that... Figure 3 The example process described herein is just one of many techniques that can be deployed to generate virtual representations of objects according to one or more implementation schemes.
[0038] Flowchart 300 begins with input image data 302. Input image data 302 may be one or more images of a user or other object. In some embodiments, image 302 may be captured, for example, during a registration period in which the user uses a personal device to capture an image pointing to the user's face or other body parts, from which registration data may be derived for presenting virtual representation data associated with the object. According to some embodiments, image data 302 may be supplemented by image data obtained from a DA including the user or data derived from a DA including the user (such as visual artifacts generated from the DA by the visual artifact network 120).
[0039] In addition to image 302, depth data 304 corresponding to the image can also be obtained. That is, depth data 304 can be captured by one or more depth sensors corresponding to the object in image data 302. Additionally or alternatively, image 302 can be captured by a depth camera, and depth and image data can be captured simultaneously. Therefore, depth data 304 can indicate the relative depth of the surface of the subject from the viewpoint of the device that captured the image / sensor data. In some embodiments, the DA having the object may include 3D image data from which depth data can be determined and used to supplement depth data 304.
[0040] According to one or more embodiments, image data 302 may be applied to feature network 310 to obtain a set of features 312 of image data 302. Feature network 310 may additionally use posterior depth data 306 and front depth data 308. In some embodiments, front depth data 308 may be obtained from depth data 304 and includes depth information of the portion of an object facing a sensor, such as a depth sensor. Posterior depth 306 may be inferred or derived based on the captured depth data 304 and image data 302. In one or more embodiments, feature network 310 is configured to provide feature vectors for a given pixel in an image. A given sampled 3D space point will have X, Y, and Z coordinates. Feature vectors are selected from the features 312 of the image based on the X and Y coordinates. In some embodiments, feature network 310 may additionally use data from DA to determine features. For example, visual artifacts generated by visual artifact network 120 may be used to generate feature vectors.
[0041] In some implementations, each feature vector in the feature vector can be combined with the corresponding Z-coordinate of a given sampled 3D point to obtain a feature vector 312 for each sampled 3D point in the image. According to one or more implementations, the feature vector 312 can be applied to a classification network 314 to determine a classification value for each input vector for a specific sampled 3D point. For example, returning to example image 302, a classification value can be determined for a given sampled 3D point. In some implementations, the classification network can be trained to predict the relationship between the sampled point and the surface of an object presented in input image 302. For example, in some implementations, the classification network 314 can return values between 0 and 1, where 0.5 is considered to be on the surface, and 1 and 0 are considered to be inside and outside the 3D volume of the object outlined by that surface, respectively. Thus, a classification value is determined for each sampled 3D point on the input image. A 3D occupancy domain 316 for the user can be derived from the combination of classification values from classification network 314. For example, this set of classification values can be analyzed to recover the surface of the 3D object presented in the input image. In some implementations, the 3D occupancy domain 316 can then be used to generate a representation of the user or a portion of the user, such as an avatar representation of the user.
[0042] In addition to the 3D occupancy domain, the user's joint position can be determined from feature 312. The joint network 318 can use image data 302 as well as depth information. Depth may include front depth data 308 and back depth data, which is either derived from back depth 306 or based on the back depth determined by the classification network 314 for the 3D occupancy domain 316. The joint network can be trained to predict the user's joint position 320 in the image for its predicted 3D occupancy domain 316.
[0043] As described above, in some implementations, the sensor data used to generate a virtual representation of an object may be insufficient to generate an accurate representation of the object. Therefore, as Figure 4 As shown, some implementations involve supplementing the sensor data with content from the DA including the object when the quality level of the sensor data is insufficient to generate an accurate representation of the object. For illustrative purposes, [further details will be provided]. Figure 1 and Figure 3 The following steps are described in the context of [the document / document]. However, it should be understood that various actions can be performed by alternative components. Furthermore, various actions can be performed in different orders. Additionally, some actions can be performed simultaneously, and some actions may be unnecessary, or additional actions may be added.
[0044] Flowchart 400 begins at box 405, where image data of the object is obtained. In some embodiments, the image data may be captured, for example, during a registration period, during which the user uses a personal device to capture an image pointing towards the user's face, from which registration data can be derived for presenting avatar data associated with the user. In some embodiments, additional image data, such as the object's hands or arms, may be captured. The image data may be captured by one or more cameras on the user's device. For example, the user may register on a device (such as a head-mounted display) that the user will use in an XR environment. Alternatively, the object may register on a separate device (such as a mobile device, desktop computer, etc.) communicatively connected to the head-mounted display.
[0045] The flowchart continues to block 410, where depth data corresponding to the image captured at 405 can be obtained. In some embodiments, depth sensor data may be captured simultaneously by one or more depth sensors while capturing an image at block 405. The depth sensor data may indicate the relative depth of the object's surface from the viewpoint of the device capturing the image / sensor data. In some embodiments, the image data may include depth, or depth may be derived from image data. For example, multiple cameras (such as a stereo camera system) may be used to capture images of the object from which depth can be determined. Alternatively, the image data may be captured at least partially by a depth camera that provides both depth information and image information.
[0046] At box 415, virtual representation data is generated from image and depth information. In some embodiments, the virtual representation data may include data from sensor data from which geometrical shape information, texture information, and other information of the object can be derived, which can be used to drive the character to present an accurate representation of the object during runtime. In some embodiments, one or more preprocessing steps, such as determining feature 312, back depth 306, front depth 308, etc., may be performed before generating the user's virtual representation.
[0047] Flowchart 400 proceeds to box 420, where it is determined whether the virtual representation data of box 415 meets one or more quality parameters. For example, a portion of the user may not be well captured by the sensor data. The sensor data captured during registration can be analyzed to determine a health score or other quality parameters related to the quality of the data and the likelihood of an accurate representation of the object generated from the data. Therefore, quality parameters can be predefined and used to compare registration data with a health score or other metrics to determine whether additional media content should be used to supplement the registration data. In some implementations, quality parameters can be determined for different portions of the user / virtual representation, making it possible to identify which portions have a DA. For example, if the left side of the face is blurred in the registration data, the DA visible on the left side of the face can be selected. If it is determined at box 420 that the virtual representation data (such as registration data 100) meets one or more quality parameters, the flowchart ends at box 425, and a virtual representation of the object is generated using the virtual representation data captured during registration. For example, as described above regarding... Figure 3 The virtual representation data described herein is generated without considering additional data from the DA including the object. Additionally, in some implementations, the obtained virtual representation data may be supplemented, enhanced, or adjusted based on role accessories, as described above. Figure 2 As described.
[0048] Returning to box 420, if it is determined that the virtual representation data (such as registration data 100) fails to meet one or more quality parameters, the flowchart proceeds to box 430 and accesses the media content containing the object. Digital assets in the form of 2D or 3D image data can be accessed and may include metadata indicating the context associated with the image, or stored together with metadata indicating the context associated with the image, such as people or objects detected or otherwise identified as being present in the image. Digital assets can be obtained from a variety of sources. For example, in different scenarios, DAs may be stored locally, stored on a server, or a combination thereof. Thus, DAs can be processed locally to identify DAs containing objects. Alternatively, DAs can be requested from remote sources such as individual client devices, servers, etc. In some embodiments, DAs containing objects can be obtained by requesting DAs with metadata “tagged” with the object or otherwise indicating that the object is present in the DA. Alternatively, information related to the identity of the object (such as image data or features derived from image data) can be used to perform facial recognition on the DA to identify the DA containing the object. In other embodiments, the user can choose which DAs to consider for generating the virtual representation. Flowchart 400 proceeds to box 435, where physical properties of an object are extracted from the media content. These physical properties may be in the form of visual artifacts and may be extracted by the visual artifact network 120. Visual artifacts may include visual components of the object in the DA. This may include texture characteristics of the object, such as skin tone, makeup, scars, or the absence of scars. Therefore, visual artifacts may include a 2D or 3D representation of an object or a portion of an object based on one or more images that include the object or a portion of the object. Other examples of digital artifacts include components with geometric shapes, such as hairstyles, headdresses, etc. In some embodiments, the character network may provide data about the object to the visual artifact network to facilitate the transformation of image data into textured or geometric forms. For example, using the geometric properties of the object generated by the character network at box 415, the visual artifact network may use one or more images with specific properties to generate image or geometric data from which the character network 110 may generate a virtual representation of the object in the form of a character.
[0049] Flowchart 400 ends at box 440, where virtual representation data from box 415 and extracted physical properties from box 435 are used to generate a virtual representation of the object. For example, the extracted physical properties, or data derived from them, can be fed into the character network to enhance the generation of the object's virtual representation, as described above. Figure 3 As described.
[0050] Figure 5A system 500 with networked end-user equipment according to one embodiment is illustrated in block diagram form, from which DAs can be generated, stored, and transferred. In system 500, end-user equipment 502A-502N includes a DA interface 550 from which DAs can be generated, for example, from a local camera or other sensor, or obtained, for example, from network device 506, and the end-user equipment may include a DA library therein from which DAs can be stored. End-user equipment 504 may be a device on which registration is performed. End-user equipment 504 may additionally include a DA interface 555 from which DAs can be generated, for example, from a local camera or other sensor, or obtained, for example, from network device 506 or client devices 502A-502N. In some embodiments, end-user equipment 504 may include a registration module 565 which can register a user using locally captured sensor data, and / or image data of DAs from local storage and / or remote storage (e.g., in network device 506 or client devices 502A-502N). Additionally, the virtual artifact module 560 can be configured to generate virtual artifacts from the DA, which can be used by the registration module 565 and / or used to generate character accessories.
[0051] In system 500, DAs can be transmitted and received between end-user devices 502A-502N and end-user device 504 via network device 506 and corresponding messaging or sending applications. Network device 506 includes wired and / or wireless communication interfaces. In some example embodiments, network device 506 provides cloud storage and / or computing options for end-user devices 502A-502N and end-user device 504, including storage for personal DA libraries 512A-512N and shared DA libraries 514A-514N. Personal DA libraries 512A-512N may be provided for specific end-user devices or user profiles associated with end-user devices, and / or available via subscription. Similarly, shared DA libraries 514A-514N may be provided for specific end-user devices or user profiles associated with end-user devices, and / or available via subscription. Some of these DAs are more meaningful to the shared DA libraries, while others are less meaningful. For example, the shared DA library 514A-514N may include a library with DAs derived from client devices 502A-502N, wherein the user of the end user device 504 is tagged.
[0052] refer to Figure 6An alternative simplified network diagram 600 is presented, including end-user equipment 504, which shows additional details of end-user equipment 504. End-user equipment 504 can be used to generate virtual representations of objects and may include electronic devices such as telephones, tablet computers, personal digital assistants, portable music / video players, wearable devices, head-mounted devices, base stations, laptop computers, desktop computers, mobile devices, network devices, or any other electronic device with the ability to capture image data and generate virtual representations of users.
[0053] End-user device 504 may include one or more processors 616, such as a central processing unit (CPU). Processor 616 may include a system-on-a-chip (such as those present in mobile devices) and may include one or more dedicated graphics processing units (GPUs) or other graphics hardware. Additionally, processor 616 may include multiple processors of the same or different types. End-user device 504 may also include memory 610. Memory 610 may include one or more memories of different types that can be used in conjunction with processor 616 to perform device functions. Memory 610 may store various programming modules for execution by processor 616, including a DA interface 555, a virtual artifact module 560, a registration module 565, and various other potential applications.
[0054] The end-user device 504 may also include a storage device 612. The storage device 612 may include registration data 634, which may include data such as user-specific profile information or user-specific preferences. The registration data 634 may additionally include data for generating user-specific avatars, such as the user's 3D mesh representation, the user's joint positions, the user's skeleton, etc. The storage device 612 may also include an image storage area 636. The image storage area 636 can be used to store a series of images from which the registration data can be determined, such as the input images described above, from which three-dimensional information of objects in the images can be determined. Furthermore, the image storage area 636 may include additional DAs, for example, in the form of a DA library. The storage device 612 may also include a character storage area 638, which may store data for generating virtual representations of objects, such as geometric data, texture data, predefined characters, etc. Furthermore, the character storage area 638 may include a character accessory library in which character accessories generated by the virtual artifact module 560 are stored.
[0055] In some embodiments, the end-user equipment 504 may include other components for role registration, such as one or more cameras 618 and / or other sensors, such as one or more depth sensors 620. In one or more embodiments, each of the one or more cameras 618 may be a conventional RGB camera or a depth camera, etc. The one or more cameras 618 may capture input images of a subject for determining 3D information based on 2D images. Additionally, the cameras 618 may include a stereo or other multi-camera system.
[0056] While client device 602 is depicted as including the numerous components described above, in one or more embodiments, the various components and their functionality may be distributed in different ways across one or more additional devices (e.g., across a network). For example, in some embodiments, any combination of storage devices 612 may be deployed partially or entirely on additional devices such as network device 606. Additionally, various components of the end-user device may be configured, for example, to access DA or other data or functions from other devices (such as client device 602 and network device 606) across network 608 using network interface 622.
[0057] According to one or more embodiments, the end-user device may include a display 614 that can be used to facilitate the registration process. For example, a preview of a character, character accessories, etc., may be presented to the user on the display 614. Many different types of electronic systems enable people to sense and / or interact with various XR environments. Examples include: head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed for placement on a person's eyes (e.g., similar to contact lenses), headsets / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablet computers, and desktop / laptop computers. A head-mounted system may have an integrated opaque display and one or more speakers. Alternatively, a head-mounted system may be configured to receive an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems can have transparent or semi-transparent displays, rather than opaque displays. The transparent or semi-transparent display can have a medium through which light representing the image is directed to the viewer's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium can be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In some implementations, the transparent or semi-transparent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection techniques that project graphic images onto the viewer's retina. Projection systems can also be configured to project virtual objects into a physical environment, such as as holograms or onto a physical surface.
[0058] Additionally, in one or more embodiments, the end-user equipment 504 may consist of multiple devices in the form of an electronic system. For example, an input image may be captured from a camera on an accessory device communicatively connected to the client device 602 across network 608 or a local area network. Similarly, some or all of the computational functions described as being executed by computer code in memory 610 may be offloaded to an accessory device communicatively coupled to the client device 602, such as a network device like a server. Therefore, although certain calls and transmissions are described herein with respect to a particular system depicted, in one or more embodiments, various calls and transmissions may be routed differently based on different distributed functions. Furthermore, additional components may be used, and certain combinations of the functions of any components may be combined.
[0059] Now for reference Figure 7This document illustrates a simplified functional block diagram of an exemplary multi-functional electronic device 700 according to one embodiment. Each electronic device in the document may be a multi-functional electronic device, or may have some or all of the components described herein. The multi-functional electronic device 700 may include a processor 705, a display 710, a user interface 715, graphics hardware 720, device sensors 725 (e.g., proximity / ambient light sensors, accelerometers, and / or gyroscopes), a microphone 730, an audio codec 735, a speaker 740, communication circuitry 745, digital image capture circuitry 750 (e.g., including a camera system), a memory 760, a storage device 765, and a communication bus 770, and certain combinations thereof. The multi-functional electronic device 700 may be, for example, a mobile phone, a personal music player, a wearable device, and a tablet computer.
[0060] Processor 705 can execute necessary instructions to implement or control the operation of various functions performed by device 700. Processor 705 may, for example, drive display 710 and receive user input from user interface 715. User interface 715 allows a user to interact with device 700. For example, user interface 715 may take various forms, such as buttons, keypad, dial pad, click wheel, keyboard, display screen, or touchscreen. Processor 705 may also be, for example, a system-on-a-chip, such as those present in mobile devices, and may include a dedicated GPU. Processor 705 may be based on a Reduced Instruction Set Computer (RISC) or Complex Instruction Set Computer (CISC) architecture or any other suitable architecture, and may include one or more processing cores. Graphics hardware 720 may be dedicated computing hardware for processing graphics and / or assisting processor 705 in processing graphics information. In one embodiment, graphics hardware 720 may include a programmable GPU.
[0061] Image capture circuit 750 may include one or more lens assemblies, such as 780A and 780B. The lens assemblies may have various combinations of characteristics, such as different focal lengths. For example, lens assembly 780A may have a shorter focal length relative to the focal length of lens assembly 780B. Each lens assembly may have a separate associated sensor element 790. Alternatively, two or more lens assemblies may share a common sensor element. Image capture circuit 750 can capture still images, video images, and enhanced images, etc. The output from image capture circuit 750 may be processed at least partially by a video codec 755 and / or a processor 705 and / or graphics hardware 720, and / or a dedicated image processing unit or pipeline incorporated within circuit 745. Images thus captured may be stored in memory 760 and / or storage device 765.
[0062] Memory 760 may include one or more different types of media used by processor 705 and graphics hardware 720 to perform device functions. For example, memory 760 may include memory cache, read-only memory (ROM), and / or random access memory (RAM). Storage device 765 may store media (e.g., audio, image, and video files), computer program instructions or software, preference information, device profile information, and any other suitable data. Storage device 765 may include one or more non-transitory computer-readable storage media, including, for example, magnetic disks (fixed hard disks, floppy disks, and removable disks) and magnetic tape, optical media (such as CD-ROMs and digital video optical discs (DVDs)), and semiconductor storage devices (such as electrically programmable read-only memory (EPROM) and electrically erasable programmable read-only memory (EEPROM)). Memory 760 and storage device 765 may be used to tangibly hold computer program instructions or computer-readable code organized into one or more modules and written in any desired computer programming language. When executed by, for example, processor 705, such computer program code may implement one or more of the methods described herein.
[0063] A physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic devices. A physical environment can include physical features such as physical surfaces or physical objects. For example, a physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People can directly sense and / or interact with a physical environment through senses such as sight, touch, hearing, taste, and smell. In contrast, an XR environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic devices. For example, an XR environment can include augmented reality (AR) content, mixed reality (MR) content, and / or virtual reality (VR) content, etc. In the case of an XR system, a subset of a person's physical motion or a representation thereof is tracked, and in response, one or more properties of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. As an example, an XR system can detect head movement and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. For example, an XR system can detect movement of electronic devices (e.g., mobile phones, tablets, laptops, etc.) presenting an XR environment, and in response, adjust the graphical content and sound field presented to the user in a manner similar to how such views and sounds would change in the physical environment. In some cases (e.g., for accessibility reasons), an XR system can adjust the characteristics of the graphical content in the XR environment in response to representations of physical movement (e.g., voice commands).
[0064] It should be understood that the above description is intended to be illustrative and not restrictive. Material has been presented to enable any person skilled in the art to make and use the disclosed subject matter protected by the claims and to provide that material in the context of a particular embodiment, variations of which will be readily apparent to those skilled in the art (e.g., some of the disclosed embodiments may be used in combination with each other). Therefore, Figures 1 to 4 The specific arrangement of the steps or actions shown or Figures 5 to 7 The arrangement of elements shown should not be construed as limiting the scope of the disclosed subject matter. Therefore, the scope of the invention should be determined by referring to the appended claims and the full scope of their equivalents. In the appended claims, the terms "comprising" and "wherein" are used as common Chinese equivalents to the corresponding terms "comprising" and "characterized in".
Claims
1. A non-transitory computer-readable medium comprising computer-readable code, said computer-readable code being executable by one or more processors to: Obtain sensor data of the object at the local device; Obtain media assets including the described objects from the digital asset library; Generate a visual artifact of the object based on the media assets; as well as The local device uses the sensor data and the visual artifacts to generate a virtual representation of the object.
2. The non-transitory computer-readable medium of claim 1, wherein the computer-readable code for obtaining sensor data of the object comprises computer-readable code for performing the following operations: The geometric and textural properties of the object are determined based on the sensor data.
3. The non-transitory computer-readable medium of claim 1, wherein the computer-readable code for identifying the media asset including the object includes computer-readable code for performing the following operations: Request one or more media items from the digital asset library associated with the object.
4. The non-transitory computer-readable medium of claim 1, wherein the computer-readable code for generating a visual artifact of the object based on the media asset further comprises computer-readable code for performing the following operations: Extract image data, including object accessories, from the media asset containing the object; The image data including the object accessory is applied to a network trained to generate textures corresponding to the object accessory to obtain object accessory components; and Generate an accessory library that includes the accessory components of the object.
5. The non-transitory computer-readable medium of claim 4, wherein the object accessory component includes an object accessory texture, and wherein the computer-readable code for generating the one or more virtual representations of the object includes computer-readable code for performing the following operations: The object accessory texture is applied to at least a portion of one or more virtual representations of the object.
6. The non-transitory computer-readable medium of claim 4, wherein the one or more virtual representations generated using the sensor data and the object accessory components are provided as one or more alternative virtual representations of the object.
7. The non-transitory computer-readable medium of claim 4, wherein the computer-readable code for generating the one or more virtual representations of the object includes computer-readable code for performing the following operations: Use the sensor data to generate an initial virtual representation; and The initial virtual representation is enhanced by utilizing the object accessory components.
8. The non-transitory computer-readable medium of claim 4, wherein the object accessory component includes a three-dimensional hairstyle component.
9. The non-transitory computer-readable medium of claim 1, wherein the media asset including the object is obtained from a digital asset library based on the failure of the sensor data to meet the determination of one or more quality parameters.
10. A method, the method comprising: Obtain sensor data of the object at the local device; Obtain media assets including the described objects from the digital asset library; Generate a visual artifact of the object based on the media assets; as well as The local device uses the sensor data and the visual artifacts to generate one or more virtual representations of the object.
11. The method of claim 10, wherein obtaining sensor data of the object comprises: The geometric and textural properties of the object are determined based on the sensor data.
12. The method of claim 10, wherein identifying the media asset including the object comprises: Request one or more media items from the digital asset library associated with the object.
13. The method of claim 10, wherein generating the visual artifact of the object based on the media asset further comprises: Extract image data, including object accessories, from the media asset containing the object; The image data, including the object accessory, is applied to a network trained to generate textures corresponding to the object accessory to obtain object accessory components; as well as Generate an accessory library that includes the accessory components of the object.
14. The method of claim 13, wherein the object accessory component includes an object accessory texture, and wherein generating the one or more virtual representations of the object includes: The object accessory texture is applied to at least a portion of one or more virtual representations of the object.
15. The method of claim 13, wherein the one or more virtual representations generated using the sensor data and the object accessory components are provided as alternative one or more virtual representations of the object.
16. The method of claim 13, wherein generating the one or more virtual representations of the object comprises: The sensor data is used to generate an initial virtual representation; as well as The initial virtual representation is enhanced by utilizing the object accessory components.
17. The method of claim 13, wherein the object accessory component includes a three-dimensional hairstyle component.
18. The method of claim 10, wherein the media asset including the object is obtained from a digital asset library based on the failure of the sensor data to meet the determination of one or more quality parameters.
19. A system comprising: One or more processors; and One or more computer-readable media, the one or more computer-readable media including computer-readable code, the computer-readable code being executable by the one or more processors to: Obtain sensor data of the object at the local device; Obtain media assets including the described objects from the digital asset library; Generate a visual artifact of the object based on the media assets; as well as The local device uses the sensor data and the visual artifacts to generate one or more virtual representations of the object.
20. The system of claim 19, wherein the computer-readable code for obtaining sensor data of the object includes computer-readable code for performing the following operations: The geometric and textural properties of the object are determined based on the sensor data.
21. The system of claim 19, wherein the computer-readable code for identifying the media asset including the object includes computer-readable code for performing the following operations: Request one or more media items from the digital asset library associated with the object.
22. The system of claim 19, wherein the computer-readable code for generating a visual artifact of the object based on the media asset further comprises computer-readable code for performing the following operations: Extract image data, including object accessories, from the media asset containing the object; The image data including the object accessory is applied to a network trained to generate textures corresponding to the object accessory to obtain object accessory components; and Generate an accessory library that includes the accessory components of the object.
23. The system of claim 22, wherein the object accessory component includes an object accessory texture, and wherein the computer-readable code for generating the one or more virtual representations of the object includes computer-readable code for performing the following operations: The object accessory texture is applied to at least a portion of one or more virtual representations of the object.
24. The system of claim 22, wherein the one or more virtual representations generated using the sensor data and the object accessory components are provided as alternative one or more virtual representations of the object.
25. The system of claim 22, wherein the computer-readable code for generating the one or more virtual representations of the object includes computer-readable code for performing the following operations: Use the sensor data to generate an initial virtual representation; and The initial virtual representation is enhanced by utilizing the object accessory components.
26. The system of claim 22, wherein the object accessory component includes a three-dimensional hairstyle component.
27. The system of claim 19, wherein the media asset including the object is obtained from a digital asset library based on the failure of the sensor data to meet the determination of one or more quality parameters.