System and method for texture replacement in multimedia

CN115187686BActive Publication Date: 2026-08-07BLACK SESAME TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BLACK SESAME TECH CO LTD
Filing Date
2022-06-23
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

尽管该论文提供了在背景中的特定纹理替换,但是仍然缺乏在任何其他多媒体中进行纹理替换的适用性

Benefits of technology

[0017] Another object of the present invention is to provide a fusion module that automatically changes the hue of a new texture template in accordance with the original multimedia.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187686B_ABST
    Figure CN115187686B_ABST
Patent Text Reader

Abstract

Systems and methods for texture replacement in multimedia are disclosed. The system is an AI-based multimedia processing system for replacing original background textures of multimedia with texture templates. The system applies foreground masks to hide and protect foreground regions and multiple background textures. The system uses deep learning to segment specific textures from an image or video sequence. The system replaces the textures of the original input image with texture templates to form a processed image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to systems and methods for replacing textures in multimedia backgrounds. The system applies a foreground mask to hide and protect foreground regions and multiple background textures. More specifically, this invention relates to an AI-based multimedia processing system for replacing original textures in a multimedia background with texture templates. Background Technology

[0002] Texture replacement aims to replace specific texture patterns in multimedia such as images, animations, or videos without altering the original lighting, shadows, and occlusions. Traditionally, all methods are based on correlational classification using color constancy, Markov random fields, etc. All these methods consider the relationships between pixels but not their semantic information, leading to inaccurate segmentation results. For example, if a foreground object contains a similar color to a background texture, a color classification method might classify a portion of the foreground as part of the background texture. This results in imperfect or inaccurate multimedia as the final product.

[0003] U.S. Patent 7,309,639, belonging to National Semiconductor Corp., discloses a technique related to region of interest (ROI) selection for texture replacement. Furthermore, the patent discloses a comparison of the ROI's color characteristics with other pixels in the frame and pixels with similar color characteristics classified into the same texture group. The invention provides characteristic-based color classification, which leads to inaccurate results. This can affect the integrity of foreground objects.

[0004] Another US patent, 8503767, belonging to Microsoft Corporation, discloses techniques related to texture region segmentation applied only to images. Although the system segments unique features within an image, this invention has failed to provide its application in other multimedia applications.

[0005] Another U.S. patent, 9,503,685, belonging to International Business Machines Corp., provides a solution for replacing the background in video conferencing. Although this invention is an improvement on existing inventions, it still lacks the ability to replace a specific part of the background, instead replacing the entire background.

[0006] The research paper "Texture Replacement in Real Images," attributed to Yanghai Tsin, discloses a technique for texture replacement in real images, such as interior design, digital filmmaking, and computer graphics. Furthermore, the paper discloses a system for replacing specific texture patterns in an image while preserving lighting effects, shadows, and occlusion. Although the paper provides specific texture replacement for backgrounds, its applicability for texture replacement in any other multimedia remains lacking.

[0007] This invention seeks to provide improvements in the field of texture replacement in multimedia, and more specifically, but not exclusively, in the field of deep neural learning texture recognition. Furthermore, this invention proposes a unique texture and foreground selection based on semantics using deep learning. The selected texture is replaced while preserving foreground region exclusion, thus maintaining the integrity of the foreground when applying texture replacement.

[0008] Therefore, to overcome the shortcomings of existing technologies, there is a need to provide an AI-based image processing system. This system is applied to texture region segmentation of images or videos. Furthermore, the system uses a texture motion tracker to track the motion of selected textures and refines the region segmentation results from frame to frame. Motion tracking results in smoother segmentation. And the replaced texture will also follow the motion of the previous texture, leading to a more realistic appearance. In view of the foregoing invention, there is a need in the art for a system to overcome or mitigate the aforementioned shortcomings of the prior art.

[0009] It is now apparent that numerous methods and systems suitable for various purposes have been developed in the prior art. Furthermore, even if these inventions are applicable to their specific purposes, they are not applicable to the purposes of this invention as described above. Therefore, there is a need for an advanced texture replacement system that uses deep neural networks for recognition to identify textures in multimedia backgrounds in real time. Summary of the Invention

[0010] A texture recognition and replacement system is disclosed, which identifies multiple textures in a multimedia background. The system includes modules for recognizing textures in the background and their replacements. These modules include a segmentation module, a tracking module, a blending module, and a replacement module.

[0011] The segmentation module divides the multimedia into a foreground region and a background region with multiple textures. Furthermore, the segmentation module compares these multiple textures with predefined textures to generate several identified textures. The segmentation module further includes a portrait map unit and a texture map unit. The portrait map unit protects the foreground region. The texture map unit replaces one or more identified textures with a texture template.

[0012] The tracking module includes a first tracker unit and a second tracker unit. The first tracker unit is used to track feature matches of multiple identified textures to guide the texture template. Furthermore, the second tracker unit is used to track the motion of the foreground and background regions. Additionally, the motion of the background region guides the motion of the texture template.

[0013] The fusion module generates processed textures by adjusting the hue of multimedia texture templates. This fusion module is a Generative Adversarial Network (GAN) module. The fusion module also includes an encoder and a decoder. The encoder encodes several identified textures and texture templates to produce the processed texture, and the decoder decodes the processed texture into one or more identified textures.

[0014] Finally, the replacement module replaces one or more textures with the processed texture. Furthermore, the replacement module combines the processed texture with the foreground area to form texture-replaced multimedia.

[0015] Smartphones are now increasingly embedded with motion sensors for a variety of applications. The advantages of these sensors are being extended to texture recognition systems. These systems are trained to recognize the unique textures of background regions. Neural networks are robust to any setup of multi-modal sensors, including those lacking sensors on the device. Ultimately, the extracted feature vectors utilize information beyond still images or videos to produce accurate texture-replaced multimedia.

[0016] The primary objective of this invention is to provide deep learning for segmenting specific textures from image or video sequences, segmenting portraits or foregrounds requiring protection. A deep neural network training system is used to assign several predefined textures to unique textures within the multimedia content. Furthermore, the deep neural network utilizes probabilistic gating techniques to predict the probability of a set of predefined textures by analyzing various factors.

[0017] Another object of the present invention is to provide a fusion module that automatically changes the hue of a new texture template in accordance with the original multimedia.

[0018] Another objective of this invention is to provide a tracking module for tracking the movement of a portrait area or foreground area and simulating texture motion.

[0019] Another object of the present invention is to provide alternative background textures for multimedia using post-processed texture templates.

[0020] Other objects and aspects of the invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings, which illustrate features according to embodiments of the invention.

[0021] In order to achieve the above and related objectives, the present invention may be implemented in the form illustrated in the accompanying drawings. However, it should be noted that the drawings are merely illustrative and that the specific structures illustrated and described may be modified within the scope of the appended claims.

[0022] Although the invention has been described above with reference to various exemplary embodiments and implementations, it should be understood that the various features, aspects, and functions described in one or more individual embodiments are not limited to their applicability to the particular embodiments in which they are described, but can be applied individually or in various combinations to one or more other embodiments of the invention, whether or not such embodiments are described and whether or not such features are presented as part of the described embodiments. Therefore, the breadth and scope of the invention should not be limited by any of the exemplary embodiments described above.

[0023] In certain circumstances, the presence of broadening words and phrases such as “one or more,” “at least,” “but not limited to,” or other similar phrases should not be interpreted as implying an intention or need for a narrower situation where such a broadening phrase may not exist. Attached Figure Description

[0024] The objects and features of the present invention will become clearer from the following description and the appended claims in conjunction with the accompanying drawings. It should be understood that these drawings only illustrate exemplary embodiments of the invention and are therefore not intended to limit its scope. The invention will be described and explained with additional specificity and detail using the drawings, in which:

[0025] Figure 1 A texture replacement system according to the present invention is shown;

[0026] Figure 2A The segmentation module within the texture replacement system is shown;

[0027] Figure 2B The segmentation module according to the present invention is shown;

[0028] Figure 3A The tracking module within the texture replacement system is shown;

[0029] Figure 3B The tracking module according to the present invention is shown;

[0030] Figure 4A This illustrates the fusion module in the texture replacement system;

[0031] Figure 4B The fusion module according to the present invention is shown;

[0032] Figure 5 The replacement module in the texture replacement system is shown;

[0033] Figure 6 This demonstrates a method for replacing textures in multimedia.

[0034] Figure 7 A flowchart of texture replacement according to the present invention is shown. Detailed Implementation

[0035] Due to limitations imposed by lighting, clouds, or other uncontrollable weather factors, photographers may not achieve the desired results. Therefore, good photos or videos rely not only on the photographer's skill but also on post-production. Photographers use digital imaging software to adjust image lighting, saturation, and hue, or manually add or alter textures within the image. Just as images depend on post-production, videos also rely on texture replacement to create unique effects.

[0036] Manually labeling specific textures can be tedious, especially for video. Therefore, automating the entire texture segmentation and labeling process is appealing. The goal of texture replacement is to replace specific texture patterns without altering the original lighting, shadows, and occlusion.

[0037] Traditional methods include classification based on color constancy, Markov random fields, and so on. All of these methods consider the relationships between pixels but not their semantic information, which leads to inaccurate segmentation results. For example, if a foreground object contains a similar color to the background texture, a color classification method might classify a portion of the foreground as the background texture.

[0038] Foreground objects will be affected after texture replacement. Currently, AI techniques such as image segmentation are applied to texture replacement. Most of these methods only segment the background region, which has a low error tolerance. If the background segmentation is inaccurate, it may affect the foreground object. Furthermore, texture replacement is often based on copy and paste, which results in rough edges. In video applications, texture replacement often does not consider the relationship between frames, leading to inconsistent texture replacement results. In this disclosure, we use an AI model to segment a specific texture and use a portrait mask or foreground mask to protect the portrait or foreground.

[0039] Furthermore, we track the motion of the portrait or foreground and texture, and use this information to guide the motion of the replaced texture. Additionally, we add a fusion module to adjust the hue of the replaced texture to match the original texture. One approach to solving the texture replacement problem in related work is to leverage machine learning models to find patterns with information similar to the selected texture, and Markov random fields are used to model spatial lighting variation constraints. Visually satisfactory results can be obtained using this statistical method, but deep learning methods such as image segmentation are used to improve texture segmentation results. U-Net (encoder and decoder architecture) is often applied to provide a deep learning solution for the background removal problem. Furthermore, depth maps are used to improve the quality of the background mask.

[0040] Figure 1 A texture recognition and replacement system 100 is shown. System 100 recognizes textures in a multimedia background. The system includes several modules for recognizing textures in the background and their replacements. The modules in the system are a segmentation module 200, a tracking module 300, a fusion module 400, and a replacement module 500.

[0041] The segmentation module 200 segments the multimedia into a foreground region and a background region with multiple textures. Furthermore, the segmentation module compares the multiple textures with predefined textures to generate several identified textures. The segmentation module 200 further includes portrait map units and texture map units. The portrait map units protect the foreground region. The texture map units replace one or more identified textures with a texture template.

[0042] The tracking module 300 includes a first tracker unit and a second tracker unit. The first tracker unit is used to track feature matching of the plurality of identified textures to guide the texture template. In addition, the second tracker unit is used to track the motion of the foreground region and the background region, wherein the motion of the background region guides the motion of the texture template.

[0043] The fusion module 400 adjusts the hue of the texture template based on multimedia to generate a processed texture, wherein the fusion module is a Generative Adversarial Network (GAN) module. The fusion module 400 also includes an encoder and a decoder. The encoder encodes the plurality of identified textures and the texture template to generate a processed texture, and the decoder decodes the processed texture into one or more identified textures.

[0044] Finally, the replacement module 500 replaces the one or more textures with the processed texture. Furthermore, the replacement module 500 combines the processed texture with the foreground area to form texture-replaced multimedia.

[0045] Figure 2AA segmentation module in a texture replacement system 200A is illustrated. The segmentation module 200 segments multimedia into a foreground region and a background region having one or more textures. Further, the segmentation module compares the one or more textures with predefined textures to generate one or more identified textures. The segmentation module also includes a portrait map unit 204 and a texture map unit 202. The portrait map unit 204 protects the foreground region by covering it with a foreground mask. The texture map unit 202 replaces the one or more identified textures with a texture template.

[0046] The segmentation module 200 uses artificial intelligence and machine learning algorithms to segment the background and foreground regions. Furthermore, the segmentation module 200 uses artificial intelligence and machine learning algorithms to compare the one or more textures with predefined textures to generate the one or more identified textures.

[0047] The feature matching of the one or more identified textures is based on an optical flow algorithm, which determines the patterns of apparent motion of objects, surfaces, and edges in multimedia. The feature matching of the one or more identified textures is also based on a feature mapping algorithm, including SIFT, which determines patterns of altered scale, intensity, and rotation.

[0048] Figure 2B The architecture of segmentation module 200B is illustrated. The segmentation module includes a deep learning application for training texture maps and portrait or foreground maps. The segmentation module is applied to the input image 206, where a foreground mask is applied to hide or protect the foreground region 208 and multiple textures (210a, 210b) of the background. AI is used to predefine some textures of interest, such as sky, walls, water, etc. The user selects one or more textures to replace from these multiple textures (210a, 210b). This texture is referred to as texture A (210a). The graph of texture A (210a) is used as a guide to replace texture A (210a) with the selected texture template B.

[0049] Portrait or foreground images are used to protect the portrait or foreground area. The area being replaced should exclude the portrait or foreground image. For example... Figure 2B As shown, the proposed neural network segments the pixels of an image into foreground object regions or masks, predefined textures, and unknown textures. In the proposed system, foreground objects can be people, cats, dogs, buildings, etc. Background textures can be sky, water, trees, etc.

[0050] Figure 3AA tracking module in a texture replacement system 300a is shown. The tracking module 300 includes a first tracker unit 304 and a second tracker unit 306. The first tracker unit 304 tracks feature matches of one or more identified textures to guide the texture template. The second tracker unit 306 tracks the motion of the foreground and background regions. The motion of the background region guides the motion of the texture template.

[0051] First, the first type of tracking module is based on image feature mapping algorithms, such as optical flow algorithms and SIFT feature matching.

[0052] Alternatively, the first type of tracking module is based on image feature matching, such as Harris Corner, Speeded Up Robust Feature (SURF), Features from Accelerated Segment Test (FAST), or Oriented FAST and Rotated BRIEF (ORB).

[0053] The second type of tracking module is based on the device's motion sensors, such as gyroscope sensors and accelerometer sensors.

[0054] Ideally, after detecting points of interest, we proceed to compute a descriptor for each of them. Descriptors can be divided into two categories: Local descriptors: These are compact representations of a point's local neighborhood. Local descriptors attempt to have a similar shape and appearance only in the local neighborhood around the point, and are therefore well-suited for representing it in a matching manner. Global descriptors: Global descriptors describe the entire image. They are generally not very robust, as a change to a part of the image can cause it to fail, as this will affect the resulting descriptor.

[0055] Figure 3B The architecture of tracking module 300b is shown. Tracking module 300 is used for video texture replacement. A first type of tracking module is based on image feature mapping performed by tracker 310 on different frames of image 308, such as optical flow algorithms, SIFT feature mapping, etc. A second type of tracking module is based on the device's motion sensors, such as gyroscope sensors and accelerometers. This motion is represented as rotation, translation, and scaling.

[0056] Two types of tracking modules (312a, 312b) can be used independently or in combination. It predicts the motion of foreground objects and background textures. The motion of the foreground is used to refine the mask of the portrait or the foreground, and the motion of background texture A guides the motion of texture template B. This guidance of the texture template's motion is based on sensing the motion of the texture template via a motion sensor in the electronic device. These create correlations between nearby frames, resulting in smoother and less jittery video.

[0057] Figure 4A A fusion module in a texture replacement system 400a is shown. The fusion module 400 adjusts the hue of a texture template based on multimedia to generate a processed texture. The fusion module 400 encodes the selected texture into a feature code and uses this code as a guide to transfer the texture template to the domain of the selected texture to form the processed texture.

[0058] The fusion module is based on a Generative Adversarial Network (GAN) model. To ensure fusion, the GAN model maintains consistency in emission, color temperature, and hue. The GAN model's loss consists of three components: VAE loss, GAN loss, and cycle consistency loss. VAE loss controls the reconstruction from latent code to the input image and from the image to latent code. GAN loss controls the accuracy of the discriminator. Cycle consistency loss ensures that the image transformation from domain A to domain B can be reversed.

[0059] The fusion module 400 includes an encoder 402 and a decoder 404. The encoder 402 is used to encode one or more identified textures and texture templates to produce a processed texture, and the decoder 404 is used to decode the processed texture into one or more identified textures.

[0060] Figure 4B The architecture of the fusion module 400b is shown. The fusion module generates a consistent tone between the original input image 206 and texture template B. This fusion model can be a GAN model 408 with original texture A 210a and texture template B 406 as input. The output will be a tunable texture B. For example, the fusion model takes texture A 210a and texture template B 406 from the original image as input. The fusion module encodes texture A into a feature code and uses this feature code as a guide to pass texture template B to the domain of texture A to create output 410. The loss of GAN model 408 includes three components: VAE loss, GAN loss, and cycle consistency loss.

[0061] or

[0062] VAE loss control is applied to the reconstruction from latent factors to the input image and from the image to latent factors.

[0063] or

[0064] GAN loss control discriminator accuracy.

[0065] or

[0066] Cyclic consistency loss ensures that an image transformation from domain A to domain B can be transformed back.

[0067] Figure 5 The architecture of the replacement module 500 is shown. The replacement module 500 replaces one or more textures with processed textures. The replacement module includes a merger 502 to combine the processed textures with foreground regions to form texture-replaced multimedia 504.

[0068] Figure 6 A method for replacing textures in multimedia is illustrated. The method includes the following steps: First, once the computing device receives multimedia, one or more textures from a background region and a foreground region are segmented 602. During segmentation, the one or more textures are compared with a plurality of predefined textures to generate one or more identified textures 604. After segmentation, feature matching of the one or more identified textures is tracked to guide a texture template 606. Motion 608 of the predefined textures and foreground region is tracked, where a tracking module simulates texture motion. The hue of the texture template 610 is then adjusted to match at least one of the one or more identified textures. A texture template is retrieved for a texture selected by the user from the one or more identified textures to form a processed texture. The selected texture is then replaced with the processed texture 612, and finally, the processed texture is merged with the foreground region to form texture-replaced multimedia 614.

[0069] While various embodiments of the invention have been described above, it should be understood that they are presented by way of example only and not as limitations. Similarly, the figures may depict exemplary architectures or other configurations for use in the invention, made to aid in understanding the features and functions that may be included in the invention. The invention is not limited to the exemplary architectures or configurations shown, but various alternative architectures and configurations can be used to achieve the desired features.

[0070] Although the invention has been described above with reference to various exemplary embodiments and implementations, it should be understood that the various features, aspects, and functions described in one or more individual embodiments are not limited to their applicability to the particular embodiments in which they are described, but can be applied individually or in various combinations to one or more other embodiments of the invention, whether or not such embodiments are described and whether or not such features are presented as part of the described embodiments. Therefore, the breadth and scope of the invention should not be limited by any of the exemplary embodiments described above.

[0071] In certain circumstances, the presence of broadening words and phrases such as “one or more,” “at least,” “but not limited to,” or other similar phrases should not be interpreted as implying an intention or need for a narrower situation where such a broadening phrase may not exist.

Claims

1. A system for texture replacement in multimedia, wherein the system comprises: A segmentation module, wherein the segmentation module segments the multimedia into a foreground region and a background region having one or more textures, and wherein the segmentation module compares the one or more textures with predefined textures to generate one or more identified textures, and wherein the segmentation module includes: Portrait image unit, wherein the portrait image unit protects the foreground area; and A texture map unit, wherein the texture map unit replaces the one or more identified textures with a texture template; Tracking module, wherein the tracking module includes: A first tracker unit is configured to track feature matching of the one or more identified textures to guide the texture template; and The second tracker unit is used to track the motion of the foreground region and the background region, wherein the motion of the background region guides the motion of the texture template. A blending module, wherein the blending module adjusts the hue of the texture template based on the multimedia to generate a processed texture; and A replacement module, wherein the replacement module replaces the one or more textures with the processed texture, and wherein the replacement module combines the processed texture with the foreground region to form texture-replaced multimedia.

2. The system for texture replacement in multimedia according to claim 1, wherein, The system is equipped with electronic devices.

3. The system for texture replacement in multimedia according to claim 2, wherein, The electronic device is any one of a smartphone, tablet computer, or camera.

4. The system for texture replacement in multimedia according to claim 2, wherein, The predefined texture is stored in the memory of the electronic device.

5. The system for texture replacement in multimedia according to claim 1, wherein, The multimedia refers to any of the following: images, videos, and animations.

6. The system for texture replacement in multimedia according to claim 1, wherein, The segmentation module uses artificial intelligence and machine learning algorithms to segment the background region and the foreground region.

7. The system for texture replacement in multimedia according to claim 6, wherein, The segmentation module uses artificial intelligence and machine learning algorithms to compare the one or more textures with predefined textures to generate one or more identified textures.

8. The system for texture replacement in multimedia according to claim 1, wherein, The portrait image unit protects the foreground area by using a portrait mask.

9. The system for texture replacement in multimedia according to claim 1, wherein, The feature matching of the one or more identified textures is based on an optical flow algorithm.

10. The system for texture replacement in multimedia according to claim 9, wherein, The optical flow algorithm determines the apparent motion patterns of objects, surfaces, and edges in the multimedia.

11. The system for texture replacement in multimedia according to claim 1, wherein, The feature matching of the one or more identified textures is based on a feature mapping algorithm.

12. The system for texture replacement in multimedia according to claim 11, wherein, The feature mapping algorithm determines the patterns for changing the scale, intensity, and rotation.

13. The system for texture replacement in multimedia according to claim 2, wherein, The movement of the texture template is guided by sensing the movement of the texture template through the motion sensor of the electronic device.

14. The system for texture replacement in multimedia according to claim 13, wherein, The motion sensor is an accelerometer or a gyroscope.

15. The system for texture replacement in multimedia according to claim 13, wherein, The motion is any one of rotational motion, translational motion, and scaling motion.

16. The system for texture replacement in multimedia according to claim 1, wherein, The fusion module is based on a generative adversarial network model.

17. The system for texture replacement in multimedia according to claim 1, wherein, The fusion module includes an encoder for encoding the one or more identified textures and the texture template to generate the processed texture.

18. The system for texture replacement in multimedia according to claim 1, wherein, The fusion module includes a decoder for decoding the processed texture into one or more identified textures.

19. A method for texture replacement in multimedia, the method being used in the system of claim 1, wherein the method comprises: One or more textures from the background region are separated from the foreground region, wherein the one or more textures are compared with a plurality of predefined textures to generate one or more identified textures; Track the motion of the predefined texture and the foreground region to simulate texture motion; The hue of the texture template is adjusted to match at least one of the one or more identified textures, wherein the texture template is retrieved for a texture selected by the user from the one or more identified textures to form the processed texture; Replace the selected texture with the processed texture; as well as The processed texture is merged with the foreground region to form texture-replaced multimedia.

20. A computer-readable storage medium for use in the system of claim 1, the computer-readable storage medium having computer program logic for enabling at least one processor in a computer system to replace the texture of multimedia via a software platform, the computer program logic comprising: One or more textures from the background region are separated from the foreground region, wherein the one or more textures are compared with a plurality of predefined textures to generate one or more identified textures; Track the motion of the predefined texture and the foreground region to simulate texture motion; The hue of the texture template is adjusted to match at least one of the one or more identified textures, wherein the texture template is retrieved for a texture selected by the user from the one or more identified textures to form the processed texture; Replace the selected texture with the processed texture; as well as The processed texture is merged with the foreground region to form texture-replaced multimedia.

Citation Information

Patent Citations

  • Method of forming a metal trace with reduced RF impedance resulting from the skin effect

    US7309639B1

  • Textual attribute-based image categorization and search

    US8503767B2

  • Background replacement for videoconferencing

    US9503685B2

  • Background removal in a live video

    CN101326514A

  • Method and equipment for tracking object

    CN102385754A