Digital Space Styling with Multi-Modal AI for Décor Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Customers face challenges in visualizing how home décor items will look in their homes, with style mismatches often becoming apparent only after purchase, leading to dissatisfaction.

Innovation Solution

A multi-modal generative artificial intelligence system is used to style a digital space by segmenting images, applying target styles, and recommending complementary items based on dominant colors and deep learning models, enabling virtual visualization before purchase.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If customers purchase home décor items without virtual visualization, then purchasing process is simple and quick, but style mismatches occur and customer satisfaction decreases

Engineering Contradiction:
Improvepurchasing decision accuracyVSAvoidtime for visualization
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary action by generating virtual visualization of décor items in the customer's space before the purchase decision is made. The image generation model creates preview images showing how the item will look in the customer's environment, allowing customers to assess style compatibility in advance and make informed purchasing decisions without time loss.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If traditional image generation models are used, then system complexity is low, but visualization accuracy and style matching are insufficient

Engineering Contradiction:
Improvestyle matching accuracyVSAvoidsystem architecture
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system applies segmentation by dividing the image generation process into multiple specialized components: an image encoder to extract features from reference images, a style encoder to capture aesthetic characteristics, and an image generation model to synthesize the final visualization. This segmented architecture enables each component to specialize in specific tasks, improving style matching accuracy while managing system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses composite materials concept by combining multiple encoding models and generation models into a unified system. The image encoder, style encoder, and image generation model work together as composite components, where each model contributes specific capabilities (feature extraction, style recognition, image synthesis) to achieve superior visualization accuracy that individual models cannot achieve alone.

Inventive Principle:
Principle #40Composite materials

3Measurement precision

If detailed image analysis is performed to ensure accurate styling, then visualization quality improves, but processing time increases

Engineering Contradiction:
Improvevisualization qualityVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The system applies partial action by focusing detailed analysis only on the most critical aspects of the décor item and surrounding space. The image encoder extracts key features from reference images, and the style encoder identifies essential aesthetic characteristics, rather than analyzing every detail. This selective approach maintains high visualization quality while reducing unnecessary processing time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250245493A1Styling a digital space using multi-modal image generative artificial intelligence
Publication Date: 2025.07.31 WALMART APOLLO LLC
  • US20250245493A1 patent drawing
  • US20250245493A1 patent drawing
  • US20250245493A1 patent drawing

AI summary

A system including a processor and a non-transitory computer-readable media storing computing instructions that, when executed on the processor, cause the processor to perform certain operations: obtaining an image of a digital space; extracting a depth map and a segmentation map of the image; passing each of the depth map and the segmentation map through a respective model of two parallel image diffusion models using stable diffusion with controlled image generation; prompting a selection of a target style for the digital space; segmenting, using image segmentation, the image in a target stylized digital space; and determining, using dominant color filtering, visual images of complementary items. Other embodiments are described.