Information processing system
Patent Information
- Application Number
- CN202610238098.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-02-28
- Publication Date
- 2026-09-22
AI Technical Summary
一,很多系统仅依据气温和天气给出通用化穿衣建议,无法结合用户衣橱中实际拥有的服装进行针对性搭配;
[0013] Through the above structural configuration and processing flow, the system provided by the present invention can comprehensively utilize environmental information, personal wardrobe information, online commodity information and user emotional information within the same platform, and by virtue of the generative artificial intelligence model and emotion analysis algorithm, provide users with dynamic, accurate and highly personalized clothing matching suggestions, thereby effectively solving the problems of single recommendation, poor context adaptability and insufficient interactive experience in the prior art.
Smart Images

Figure CN122797474A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.
[0003] The main problem this invention aims to solve is that existing clothing matching suggestion systems typically rely solely on weather information or simple user preferences, lacking comprehensive utilization of multi-source information such as the user's actual clothing ownership, online product information, and the user's emotional state. This results in insufficient personalization and poor situational adaptability of the recommendations, making it difficult to provide users with optimal clothing matching solutions that match their current environment, psychological state, and style preferences in a timely and accurate manner. Specifically, the existing technology suffers from the following problems: First, many systems only provide general clothing suggestions based on temperature and weather, and cannot make targeted combinations based on the actual clothes in the user's wardrobe; Second, existing recommendation technologies based on online shopping websites are mostly limited to statistical analysis of historical purchase or browsing behavior, lacking linkage with users' current needs, the real environment, and their own clothing inventory, which can easily lead to redundant or impractical recommendations. Third, traditional recommendation systems often ignore users' current emotions and make recommendations based only on static preferences or historical data, making it difficult to respond in a timely manner to users' dynamic needs for color, style, occasion and other aspects under different emotional states. Fourth, user interactions with the system often rely on fixed forms or simple keyword input, lacking the use of generative artificial intelligence models to automatically generate personalized prompts based on user input and emotional state. This limits the sophistication of recommended content and the natural conversational experience.
[0004] Therefore, how to provide a system that can comprehensively utilize environmental information, user wardrobe information, online product information, and user emotional information, and use generative artificial intelligence models for intelligent reasoning and copywriting generation to provide users with more refined, personalized, and context-appropriate clothing matching suggestions has become a technical issue that urgently needs to be addressed in this field. Summary of the Invention
[0005] To address the aforementioned issues, this invention provides an information processing system that includes a processor, sensor devices, a database, and hardware components such as a camera and microphone. Based on this system, through specific software function configurations and information processing flows, it achieves intelligent clothing matching recommendations based on multi-source information fusion.
[0006] Specifically, the present invention is achieved through the following technical means: First, the system is equipped with sensor devices, which are controlled by a processor to acquire the user's current temperature and weather information. Based on the acquired temperature and weather information, the processor generates clothing matching schemes for the user according to preset environmental adaptation rules, thereby ensuring that the recommended results match the external environment in terms of warmth, waterproofness, and comfort.
[0007] Second, the system sets up a database accessible to the processor. The processor stores the user's clothing information in the database, including clothing category, color, material, style tags, and seasonal attributes. When making clothing combinations, the processor uses the clothing information stored in the database to call a generative artificial intelligence model to generate the optimal clothing combination scheme that matches the user's personal wardrobe. This prioritizes the use of the user's existing clothing in the recommendations, improving practicality and cost-effectiveness.
[0008] Third, the system retrieves product information from online shopping websites, and the processor filters and selects this information based on user preferences and styles. Combining user preference parameters recorded in the database with historical behavior analysis results, the processor selects clothing from online products that matches the user's style, color preferences, and budget range, generating online clothing recommendations to provide reasonable supplementary options when the user's existing wardrobe is insufficient to meet their needs.
[0009] Fourth, the system is equipped with a camera and microphone, controlled by a processor to acquire image and voice information related to the user's emotions. The processor uses an emotion analysis algorithm to analyze the acquired images and voice, identifying the user's emotional state, such as happiness, sadness, or tension. Based on the emotion analysis results, the processor adjusts the original clothing recommendation rules. For example, when the user is in a low mood, it recommends bright colors or more comfortable outfits; when the user needs more formal attire, it recommends more conservative styles, thus making the recommendations more aligned with the user's current psychological needs.
[0010] V. The system uses a generative artificial intelligence model, wherein the processor generates corresponding prompt text based on user input information and the parsed user emotional state. The prompt text is used as an input prompt or dialogue context for the generative model, guiding the model to output more personalized and context-aware clothing recommendation copies. Through filtering and optimizing the generated text, the processor finally presents the user with the optimal clothing recommendation suggestion in natural language form, realizing a linked response to the user's input and emotion.
[0011] VI. According to a preferred embodiment of the present invention, the processor is further configured to analyze the user's past purchase history and browsing history using data mining technology, and extract the user's preference characteristics in terms of color, style, style and price range. Based on the above characteristics, the processor performs secondary filtering and sorting on product information from online shopping websites, thereby further improving the matching degree and hit rate in the online recommendation part.
[0012] VII. According to another preferred embodiment of the present invention, the processor receives various forms of input information from the user through a preset interface, including text input, voice input, and interactive behavior information collected by the terminal. The processor stores this input information in a database to form a user's long-term behavior and preference profile. Meanwhile, when generating clothing recommendations, the processor uses an emotion engine to comprehensively consider the user's long-term preferences and current emotional state, realizing dual adaptation to the user's personality characteristics and immediate emotion.
[0013] Through the above structural configuration and processing flow, the system provided by the present invention can comprehensively utilize environmental information, personal wardrobe information, online commodity information and user emotional information within the same platform, and by virtue of the generative artificial intelligence model and emotion analysis algorithm, provide users with dynamic, accurate and highly personalized clothing matching suggestions, thereby effectively solving the problems of single recommendation, poor context adaptability and insufficient interactive experience in the prior art.
[0014] "System" refers to an overall technical device composed of a processor, a sensor device, a database, a camera, a microphone and related software programs, which is used for acquiring, processing and integrating user environmental information, personal clothing information, online commodity information and emotional information, so as to generate clothing matching and recommendation results.
[0015] "Processor" refers to an electronic computing unit that is used to execute program instructions, process, analyze and calculate data obtained from sources such as sensor devices, databases, cameras, microphones and online shopping websites, and generate clothing matching schemes and recommendation results based on the data, which may be a single or multiple physical processors or logic processing units.
[0016] "Sensor device" refers to a hardware component controlled by a processor for acquiring information related to the user's current environment, including at least a sensor module for acquiring environmental parameters such as temperature and weather, and providing the acquired data to the processor for further processing.
[0017] "User's current temperature and weather information" refers to the ambient temperature data and weather conditions corresponding to the user's location and current time, including but not limited to temperature values, weather types such as sunny, rainy, and snowy, humidity, wind force, and other environmental parameters that can be used to make clothing matching decisions.
[0018] "Database" refers to a processor-accessible storage system used to store and manage data related to the system of this invention, including but not limited to user-owned clothing information, user input information, user preference information, and online product information, which can be implemented using relational or non-relational data storage structures.
[0019] "Clothing information owned by the user" refers to structured data about the clothing actually owned by the user that is registered and stored in the database. This includes information such as clothing category, color, material, size, style tag, seasonal attributes, and images or additional descriptions related to the clothing.
[0020] "Online shopping website" refers to an e-commerce platform that provides services for displaying and selling clothing and other goods through the internet. This includes e-commerce systems in the form of web pages or applications, from which the system can obtain product information related to clothing for recommendations.
[0021] "Product information" refers to clothing-related data obtained from online shopping websites, including product name, category, color, size, material, style tags, price, inventory status, image links, and user reviews, which can be used for recommendation calculations.
[0022] "User preferences and styles" refers to a comprehensive description of user preferences in terms of color, style, fit, style type (such as casual, business, sports, etc.), brand, and price range, obtained through explicit user settings, analysis of historical purchase and browsing behavior, or other interactive data mining.
[0023] A "camera" refers to an image acquisition device controlled by a processor, used to collect image information related to the user, especially images of the user's face or posture, for emotion analysis and related recommendation adjustments.
[0024] A "microphone" is an audio acquisition device controlled by a processor, used to collect user voice or ambient sound signals for voice content recognition or voice emotion analysis, thereby providing data for emotion analysis and recommendation adjustments.
[0025] "User emotion-related information" refers to data collected by cameras and microphones that can be used to infer a user's emotional state, including facial expression images, voice tone, speech rate, volume changes, and other multimodal signals related to emotional characteristics.
[0026] "Emotion analysis algorithms" refer to a set of algorithms executed by a processor to analyze and identify user emotion-related information, including but not limited to facial expression recognition algorithms, voice emotion analysis algorithms, and multimodal emotion fusion judgment algorithms based on machine learning or deep learning.
[0027] The "emotion engine" refers to a functional module that integrates emotion analysis algorithms and their result application logic. This module is called by the processor and is used to comprehensively consider the user's current emotional state and adjust the recommendation strategy when generating clothing recommendations.
[0028] "Generative artificial intelligence models" refer to artificial intelligence models that are invoked and executed by a processor to automatically generate text content or intermediate semantic representations based on the contextual information of the input, including but not limited to natural language generation models, dialogue models, or multimodal generation models based on deep learning.
[0029] "Prompt text" refers to text content generated by the processor based on user input information and the user's emotional state, used as input prompts or context for generative artificial intelligence models to guide the model to output clothing recommendation results that meet specific semantic and style requirements.
[0030] "Clothing matching scheme" refers to a set of results generated based on the user's current environment information, personal clothing information and emotional information, which combines and selects tops, bottoms, coats, shoes and accessories, and can be presented in the form of structured data or natural language description.
[0031] "Clothing recommendations" refers to the suggestions related to clothing selection that the system outputs to users. These include matching schemes based on the user's own clothing and product recommendations based on product information from online shopping websites. The system can be personalized based on user preferences, style, and mood.
[0032] "Optimal clothing matching scheme" refers to a clothing matching scheme that, under given constraints, comprehensively considers factors such as temperature, weather, occasion, user preferences, style preferences, user emotions, and the user's existing clothing, and is evaluated by a generative artificial intelligence model or a predetermined algorithm, so that the overall matching degree or satisfaction index reaches the predetermined optimal level.
[0033] "Filtering" refers to the process by which the processor filters product information obtained from online shopping websites based on conditions. This process eliminates products that do not meet the conditions based on user preferences, style preferences, price range, size fit, and other constraints, thereby obtaining a candidate recommendation set.
[0034] "Data mining technology" refers to statistical analysis and machine learning methods used by processors to analyze users' past purchase and browsing history. These methods include cluster analysis, association rule mining, recommendation algorithms, feature extraction, and preference modeling, and are used to discover user preference patterns from large-scale behavioral data.
[0035] "User's purchase history and browsing history" refers to the data set recorded in the system or online shopping website about user behaviors such as viewing products, adding items to the shopping cart, and placing orders. It can be used to infer users' implicit preferences and habits in clothing selection.
[0036] An "interface" refers to a hardware or software channel used for data interaction between a user and a system. This includes graphical user interfaces, application programming interfaces (APIs), and communication protocols. Through this interface, users can input information, and the system can output recommendation results.
[0037] "User input information" refers to various types of data provided by users to the system through the interface, including text input, voice input, option selection, click behavior, and other interactive information that can be recognized and processed by the system, in order to guide the system to generate corresponding clothing matching and recommendations. Attached Figure Description
[0038] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0039] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0040] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0041] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0042] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0043] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0044] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0045] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0046] Figure 9 This represents an emotion map that maps multiple emotions.
[0047] Figure 10 This represents an emotion map that maps multiple emotions.
[0048] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0049] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0050] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0051] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0052] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.
[0053] First, let me explain the terminology used in the following instructions.
[0054] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0055] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0056] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0057] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0058] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0059] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0060] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0061] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0062] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0063] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0064] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0065] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0066] Figure 2The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0067] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0068] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0069] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0070] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0071] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0072] In existing technologies, computer-based clothing recommendation systems typically rely solely on static rules or simple recommendation algorithms, resulting in the following technical problems: First, existing systems often process weather information, user wardrobe information, and online product information separately, lacking a mechanism to structurally fuse multi-source heterogeneous data within a unified data processing flow and dynamically input it into a generative AI model. This leads to insufficient expression of model input information and inconsistent context, affecting the accuracy and stability of the generated results. Second, existing systems often only send requests to the AI model based on fixed templates, failing to dynamically construct or adjust prompts based on comprehensive changes in real-time weather conditions, user-held items, user preferences, and user emotional states. This results in insufficient utilization of key constraints during the model's inference process, making it difficult for the generated results to accurately match the current scenario and individual needs. Third, traditional product recommendation modules and natural language generation modules are disconnected. Product selection is mostly based on simple tag matching or collaborative filtering, lacking technical means to feed back semantic-level clothing proposal information output by the generative AI model into the product selection process. This prevents the formation of a closed-loop optimization process at the system level: "environment—user—model output—product selection." Furthermore, existing systems mostly handle user emotions by displaying single emotion tags, lacking the ability to structurally represent emotion-related information and inversely constrain the input and output formats of generative artificial intelligence models. Consequently, they cannot adaptively adjust generation strategies and interactive presentation methods within the computer's internal workflow.
[0073] In summary, the key technical challenge is how to build a computer system capable of: 1) obtaining location information from user terminals and linking it with external environmental information services; 2) combining user-owned item data, preference data, historical behavior data, and emotion-related information; 3) automatically generating and dynamically adjusting prompts for generative artificial intelligence models; and 4) using natural language proposal information output by the model to drive structured matching reasoning and product selection. This will improve the intelligence, response accuracy, and computational resource utilization efficiency of the entire recommendation process at the system architecture and data flow levels.
[0074] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0075] In this invention, the server includes: a processing unit for acquiring environmental information including current weather conditions from an external information providing device via a communication network based on location information from a user terminal, and generating basic proposal information related to clothing combinations; a processing unit for reading user-owned item information and user preference information from a storage device based on user identification information, and associating them with the environmental information as clothing attribute information; a processing unit for constructing prompt statements input to a generative artificial intelligence model based on the environmental information, the owned item information, and the preference information, and calling the generative artificial intelligence model to obtain clothing proposal information in natural language form; and a processing unit for inferring the use of clothing proposal information. A processing unit generates matching proposal information based on the combination patterns of items held by the user; a processing unit performs attribute extraction and filtering on product group information obtained from external services via a communication network, based on the environmental information, preference information, and clothing proposal information, thereby generating candidate product information associated with the matching proposal information; and a processing unit obtains input information from the user and emotion-related information associated with the clothing proposal information, dynamically adjusts the prompt statement content based on the emotion-related information to call the generative artificial intelligence model again to generate updated clothing proposal information, and sends the matching proposal information and the candidate product information to the user terminal for interactive presentation. This allows for the construction of a closed-loop data processing link within the server for the generative artificial intelligence model, enabling joint optimization of environmental information, multi-dimensional user profile information, and model input prompt statements at the system level. This not only improves the consistency and personalization accuracy between the model-generated results and actual usage scenarios but also reduces invalid calls and redundant calculations, improving the overall performance and resource utilization efficiency of computer-based clothing recommendation processing.
[0076] A "system" refers to a comprehensive information processing unit consisting of at least one server, at least one user terminal, and multiple functional units interconnected through a communication network, used to perform environmental information acquisition, data storage and analysis, generative artificial intelligence model invocation, and result presentation.
[0077] A "server" is an electronic information processing device equipped with a processor and storage device, capable of interacting with user terminals and external services through a communication network, and performing data processing, model invocation, and result generation.
[0078] "User terminal" refers to an electronic device operated by a user to input information, receive and display server output results, including but not limited to smartphones, tablet computers, personal computers or wearable devices.
[0079] "Communication network" refers to wired or wireless network infrastructure used to transmit data between servers, user terminals and external services, including but not limited to the Internet, mobile communication networks and local area networks.
[0080] "External information providing device" refers to an external computing device or online service platform that operates independently of this system and is used to provide environmental information or product information and other data to the server through an interface.
[0081] "Environmental information" refers to a set of data provided by an external information providing device and obtained by a server, used to characterize the external environment in which the user is located, including but not limited to current meteorological conditions, temperature, weather conditions, humidity, and wind speed.
[0082] "Current meteorological conditions" refers to atmospheric state information corresponding to a specific time and geographical location, including but not limited to parameters such as temperature, precipitation, wind force, cloud cover, and humidity.
[0083] "Location information" refers to data used to identify the geographical location of a user, including but not limited to latitude and longitude coordinates, city name, area code, or geographic identifiers generated by location services.
[0084] "Basic proposal information" refers to intermediate recommendation information generated by the server based on environmental information and preset rules or models, which serves as the basis for constructing subsequent clothing recommendations and prompts.
[0085] "User identification information" refers to the identification information used to uniquely or quasi-uniquely distinguish different users in the system, including but not limited to user ID, account name or other identifiers.
[0086] "Storage device" refers to physical or logical storage resources used to persistently or temporarily store data in servers or other devices, including but not limited to hard drives, solid-state storage, database systems, or cloud storage.
[0087] "User-owned item information" refers to a collection of data related to the clothing items owned by the user, including but not limited to item name, category, color, material, size, seasonal applicability, and functional attributes.
[0088] "User preference information" refers to a set of data that reflects users' preferences in terms of clothing style, color, fit, price range, etc. It can be explicitly set by users or inferred by the system based on historical behavior.
[0089] "Clothing attribute information" refers to structured attribute data generated by the server after associating environmental information with user-owned item information and user preference information, used to characterize the relationship between clothing items and environment and preferences.
[0090] "Conditional information" refers to the set of input data used to construct prompts and input them into generative artificial intelligence models. This data is composed of environmental information, information on items owned by the user, information on user preferences, and other data related to recommendations.
[0091] "Generative artificial intelligence models" refer to artificial intelligence models that can generate output results such as natural language text based on input conditional information and prompts through machine learning or deep learning methods.
[0092] "Prompt statements" refer to natural language or structured text that is constructed by the server based on conditional information and input into the generative artificial intelligence model to instruct the model to perform specific generative tasks and constrain the output range.
[0093] "Natural language clothing suggestion information" refers to recommendations on clothing choices and combinations expressed in natural language text, output by a generative artificial intelligence model based on prompts.
[0094] "Combination pattern" refers to the combination relationship and matching structure of multiple clothing items inferred by the server based on the user's stored item information and clothing proposal information.
[0095] "Outfit Proposal Information" refers to information generated by the server based on the combination pattern, which specifically describes the outfit matching schemes for tops, bottoms, shoes, and accessories, along with the rationale behind them.
[0096] "Product group information" refers to a set of product-related data obtained from external services through communication networks, including but not limited to product identifiers, categories, attributes, prices, image links, and purchase links.
[0097] "Candidate product information" refers to a set of product data that is suitable as part of clothing matching, obtained by the server after extracting and filtering attributes of product group information based on environmental information, user preference information, and clothing proposal information.
[0098] "Attribute extraction" refers to the process by which a server identifies and extracts feature fields or tags from environmental information, user-related information, and product group information for filtering and matching.
[0099] "Filtering" refers to the process by which the server filters and sorts product information based on preset rules, preference information, or model results to select products that meet the criteria.
[0100] "Emotion-related information" refers to data that reflects the user's current emotional state, including but not limited to the emotion category, intensity, and tendency inferred by emotion recognition algorithms from input information or interaction behavior.
[0101] "Dynamic change" refers to the server modifying the content, parameters, or structure of prompt statements in real time or near real time based on new input or status information to adapt to the current situation.
[0102] "Natural language input information" refers to language data that users input into the system in the form of text, speech transcription, or other recognizable language to express their needs or provide feedback.
[0103] "State information" refers to a set of data that comprehensively describes the current interaction state, consisting of natural language input information, emotion-related information, and other contextual data related to the current session or environment.
[0104] "Input / output device" refers to a human-computer interaction device used to receive user input and present output results to the user, including but not limited to touch screens, keyboards, mice, microphones, speakers, and displays.
[0105] The embodiments of the present invention will be described based on a typical network architecture. In the following description, the subjects are limited to "server", "terminal" and "user".
[0106] I. Overall System Composition The server is equipped with a processor, main memory, non-volatile storage devices, and a network interface. The server runs an operating system, such as a Linux-based server operating system, and web service middleware, such as HTTP server software and application server software. The server also runs a database management system, such as relational database management software, to store information about user-held items, user preferences, and historical behavior data.
[0107] A terminal is an electronic device held by a user, including a display device, input devices (touchscreen, keyboard, microphone, etc.), and a network communication module. The terminal runs a client program, which can be a mobile application or web front-end code executed by a browser. The terminal communicates bidirectionally with the server via a communication network.
[0108] Users operate this system through a terminal. Users input natural language questions, authorize the terminal to obtain location information, enter clothing and preference information, and view clothing proposals and candidate product information generated by the server.
[0109] The server accesses external information providing devices via a communication network. These external information providing devices include those providing environmental information and those providing commodity information. The environmental information providing device provides the server with environmental information such as current weather conditions, while the commodity information providing device provides the server with commodity group information.
[0110] II. Data Structure and Storage Method The server defines structured data tables for various types of data in the storage device. The server stores fields in the user's owned items information table: item identifier, user identifier, category, color, material, size, applicable season tag, waterproof attribute, warmth attribute, etc. The server stores fields in the user preference information table: style preference tag, color preference tag, price range, scenario preference (commuting, dating, sports, etc.), etc. The server stores fields in the environment information cache: location information, temperature, weather conditions, humidity, wind speed, and data timestamp. The server stores fields in the session information table: session identifier, user identifier, most recent user input, most recently generated clothing proposal, and most recently inferred emotion-related information.
[0111] The server uses standardized field naming and indexing structures for the above data structures. The server creates composite indexes on the user identifier and timestamp fields to improve query speed based on users and time. The server creates inverted indexes on the item category and feature tag fields to accelerate conditional filtering during item matching inference and product selection.
[0112] III. Generative Artificial Intelligence Models and Prompt Statement Construction In this embodiment, the server uses a generative artificial intelligence model, which is a deep neural network based on a Transformer structure. During the pre-training phase, the generative artificial intelligence model used by the server employs a massive amount of natural language corpus, models the context on a multi-head self-attention layer, and includes several encoding and decoding layers, with the number of parameters reaching billions. The server invokes an instance of this generative artificial intelligence model deployed on external computing resources via a remote interface.
[0113] During model training, the server uses an autoregressive language modeling objective, sets cross-entropy as the loss function, updates model parameters via backpropagation, and employs stochastic gradient descent or its variants, such as momentum or adaptive learning rate optimization algorithms. In the fine-tuning phase, the server performs supervised fine-tuning using labeled datasets relevant to the clothing scenario, enabling the model to better understand the clothing scenario, weather conditions, and matching rules.
[0114] The server does not retrain the model during the inference phase, but it controls the model's inference behavior by constructing different prompts. Before each invocation of the generative AI model, the server constructs a well-structured prompt based on environmental information, user-owned items, and user preferences. The server explicitly marks environmental conditions, clothing lists, style requirements, and output format requirements in the prompt, allowing the model's internal attention mechanism to focus more intently on task-related aspects and reduce interference from irrelevant noise in the generated results.
[0115] In one embodiment, the server constructs the following prompt statement: You are a professional fashion consultant.
[0116] Current weather information for the city: temperature 16.5℃, light rain, humidity 78%.
[0117] The user's question is: "It's a bit cold today and it might rain. I'm going on a date, what should I wear?" List of clothing owned by the user: 1. Black waterproof jacket (suitable for rainy days) 2. Blue jeans 3. White sneakers 4. Gray knit sweater (relatively warm) Users prefer a style that is casual yet slightly formal, and not too flashy.
[0118] Based on the above weather conditions and the user's existing clothing, please provide a complete outfit suggestion for the date.
[0119] Please explain in sections: 1) How to match the top; 2) How to match bottoms; 3) How to choose shoes and accessories; 4) The reasons for this combination in terms of rain protection, cold protection, and aesthetics.
[0120] Please answer in Simplified Chinese.
[0121] The server provides explicit task descriptions, input elements, and output format requirements to the generative AI model through these structured prompts. By controlling key phrases and segmented descriptions in the prompts, the server allows the model output to be directly mapped to structured result fields defined internally by the server, thereby reducing server-side post-processing complexity and significantly lowering the error parsing rate.
[0122] IV. Server-side data processing and calculation refinement After receiving the location information uploaded by the terminal, the server invokes the environmental information service. The server sends a request to the environmental information service interface via an HTTP client library, including latitude and longitude or city name in the request. Upon receiving the response, the server parses the returned data and writes fields such as temperature, weather conditions, and humidity into the environmental information cache. During parsing, the server uses default values or interpolation strategies for missing fields to avoid errors in downstream processing.
[0123] When generating an outfit proposal, the server retrieves all clothing items owned by the user from the user's inventory table based on the user's identifier. The server then selects items suitable for the current season using filtering criteria; for example, it performs logical checks based on season tags and the current temperature threshold, excluding items clearly unsuitable for the current climate from the candidate set. Simultaneously, the server retrieves the user's style preference tags and color preference tags from the user preference information table.
[0124] The server extracts features from environmental information and user-owned item information. For environmental information, the server generates feature vectors, such as discretizing temperature into four interval labels: "cold," "cool," "mild," and "hot," and encoding weather conditions into category labels such as "sunny," "rainy," "snowy," "cloudy," and "foggy." For user-owned item information, the server generates feature sets, such as generating multi-dimensional attribute vectors for each garment, including category, color, material, waterproof rating, and warmth level. Finally, the server generates preference vectors from user preference information to constrain matching strategies.
[0125] The server internally uses a rule engine to make preliminary inferences about the above features, and determines the categories of items that should be given priority based on preset rules. For example, when the server detects the "rain" tag, it increases the priority of coats and shoes that are truly waterproof. When the server detects a "date" scenario, it decreases the priority of overly sporty items. The server embeds this priority information into the prompts, so that the generative AI model implicitly follows this structured preference in natural language generation.
[0126] After receiving the output from the generative AI model, the server parses the natural language text. The server prompts the model to use explicit segmentation markers, such as "Top:", "Bottoms:", "Shoes:", and "Accessories:". The server uses these markers as delimiters to segment the text into multiple fields during parsing. These fields are then mapped to corresponding fields in the styling suggestion information. For sections that cannot be parsed successfully, the server employs a fallback strategy, such as using keyword matching and similarity calculations, to categorize the text segments into the most similar fields.
[0127] The server employs a multi-layered filtering algorithm during the product selection process. First, it filters out products irrelevant to the current pairing based on category and functional tags. Then, it performs a second layer of filtering based on environmental and user preference characteristics. Finally, it conducts a third layer of semantic matching based on key attributes provided by a generative artificial intelligence model (such as "waterproof," "simple," and "slightly formal"). This multi-stage filtering significantly narrows the product candidate set, reducing the computational load of subsequent sorting and thus decreasing server load and response time.
[0128] V. Adjustment of Emotion-Related Information and Dynamic Prompt Statements During user interaction with the system, users may input natural language containing emotional nuances, such as "I'm feeling a bit down today and don't want to wear anything too flashy." The terminal sends this natural language text to the server. The server extracts emotional features from the text using a sentiment analysis module. The server can use a neural network-based text classification model to determine the emotion category and intensity. The server encodes emotion categories such as "down," "excited," and "nervous" as discrete labels and maps emotion intensity to numerical levels.
[0129] The server explicitly incorporates emotion-related information when constructing prompts. For example, the server might add a description like, "The user is currently feeling down; please avoid overly bright and exaggerated color combinations and choose comfortable, low-saturation colors instead." By encoding emotion-related constraints into the model input in this way, the generative AI model can automatically adjust tone and color suggestions during generation, thus achieving emotion-based generation strategy adjustment at the computational level, rather than simply labeling emotions at the result display level.
[0130] When emotional characteristics are strong, the server can dynamically adjust the parameters of the generative AI model. For example, it can lower the temperature parameter to reduce output randomness and shorten the maximum output length to provide more concise and predictable suggestions. By combining parameter control with prompt adjustments, the server can finely manage model behavior in the computation process, thereby controlling inference resource consumption while ensuring recommendation accuracy.
[0131] VI. Technical Effects and Improvements in Computer Technology The server improves the signal-to-noise ratio of the generative AI model input through the aforementioned structured data modeling and multi-stage data processing flow. By explicitly encoding environmental information, owned items, preferences, and emotions in the prompt statements, the server enables the model to focus its internal attention allocation more on task-relevant subsequences, thereby reducing the interference of irrelevant features on the output. Compared with the traditional method of simply piecing together user questions, this prompt statement construction strategy achieves higher consistency and a lower mismatch rate under the same inference resource conditions, resulting in improved accuracy and reduced error.
[0132] By introducing multi-layered filtering logic with model output as constraints during the product selection phase, the server transforms the product recommendation process from a simple retrieval based on user history or fixed tags into a closed loop of "environmental information → model-generated proposals → product semantic matching." Internally, this algorithm reduces the computational complexity of ranking by decreasing the size of candidate products, reducing database accesses and network data transmission, thereby lowering communication load and overall response time.
[0133] By setting up a combined index on the data table structure based on user identifiers and timestamps, the server enables sub-second response times for both historical and real-time queries, improving data management and I / O access efficiency. Furthermore, by using inverted indexes and tag clustering on item attribute fields, the server effectively improves the speed of attribute matching and combination inference, thus allowing stable response performance to be maintained even in high-concurrency scenarios.
[0134] In the emotion-related processing section, the server encodes emotional features into prompt statements and links them with model-generated parameters, achieving a technical solution that automatically adjusts the generated behavior based on state information. This adjustment is not a simple manual rule, but rather a collaborative effort of structured input and parameter control, enabling the computation process to adapt to different interaction states, improving user satisfaction and interaction quality while maintaining controllable computational load.
[0135] VII. Multiple Implementation Forms and Alternative Structures In one optional implementation, the server can provide services without relying on specific environmental information, instead obtaining environmental information from a local cache or other meteorological data sources. In another implementation, the generative AI model can be deployed on a local GPU server instead of a remote service, allowing the system to run even within an intranet environment. In this configuration, the server directly processes prompts through a local model inference module, reducing external communication latency and further shortening response time.
[0136] In another variant embodiment, the server can employ different generative artificial intelligence model structures, such as using a sequence-to-sequence model with an encoder-decoder structure, or adding a task-specific adaptation layer to a pre-trained model. In this configuration, the server can adjust the model size and number of layers according to different computing power environments, enabling flexible deployment between edge servers and cloud servers.
[0137] In one embodiment, the terminal can perform some preprocessing locally. For example, before sending the user's natural language input, the terminal can first perform local speech recognition, convert the audio into text, and then send it. In another embodiment, the terminal can cache the most recent clothing proposal information locally. When network conditions are poor, the terminal can first display the most recent result to the user and wait for new results from the server in the background, thereby improving the user experience.
[0138] In one implementation, users can take images of clothing using a terminal. After the terminal sends the images to the server, the server automatically extracts the clothing category and color tags using an image recognition module and writes them into the user's inventory information table. This visual feature extraction method reduces the burden of manual data entry for users and also improves consistency in attribute labeling, making the feature basis for subsequent matching inferences more reliable.
[0139] Through the above-mentioned implementation forms, this system not only realizes the functions of clothing matching and product recommendation, but also improves the computer technology itself in terms of generative artificial intelligence model input construction, data structure design, feature processing flow and resource utilization strategy, thereby achieving comprehensive technical effects in terms of accuracy, speed, resource consumption and scalability.
[0140] use Figure 11 The processing flow is explained.
[0141] Step 1: Users input their clothing preferences and authorize location information on the terminal. Users can input clothing-related questions in natural language on the terminal's input interface, such as "It's a bit cold today and it might rain. I'm going on a date, what should I wear?", and authorize the terminal's location function to obtain the current location.
[0142] Input: The user's natural language text, and the user's authorized location information (city name or latitude and longitude).
[0143] Output: The terminal generates request data containing the question text and location information.
[0144] The terminal saves the question text to a memory variable based on the input event, reads the location information from the operating system's location interface and combines it into structured data in preparation for sending it to the server.
[0145] Step 2: The terminal sends a clothing consultation request to the server. The terminal sends a request to a predetermined interface of the server through the communication network, encapsulating data such as user question text, location information, and user identifier into a request message.
[0146] Input: Structured request data (user questions, location information, user ID, etc.) in terminal memory.
[0147] Output: The request message transmitted to the server over the network.
[0148] The terminal calls the network communication module to serialize the request into a message in a format such as JSON, sets the request path and header information, sends it to the server via the HTTPS protocol, and displays a "generating suggestions" message on the interface.
[0149] Step 3: The server parses the request and verifies the basic parameters. After receiving a request from the terminal, the server parses the request body using the web application framework, extracts fields such as user identifier, question text, and location information, and verifies the existence and format of the fields.
[0150] Input: Request messages from the terminal (including user questions, location information, user ID, etc.).
[0151] Output: The server's internal request context object (containing validated fields) or an error response.
[0152] The server checks for empty or incorrectly formatted required fields. When an error occurs, it generates an error code and returns an error message to the terminal. When the verification passes, the field is stored in the current session context for subsequent processing modules to access.
[0153] Step 4: The server obtains environmental information from external information providing devices. The server constructs a query for external environmental information services based on location information, sends a request to obtain current weather conditions, and parses the returned results into an internally unified format.
[0154] Input: Location information in the request context.
[0155] Output: Structured environmental information (temperature, weather conditions, humidity, etc.).
[0156] The server calls an HTTP client library, sends the location information as a parameter to the environmental information service interface, receives the returned meteorological data, extracts fields and converts units in the JSON response, writes the temperature, weather conditions, etc. into an environmental information object, and caches it in memory.
[0157] Step 5: The server reads the user's stored items and preference information. The server retrieves the user's registered clothing and preference data from the storage device based on the user's identifier and generates a set of attributes for subsequent processing.
[0158] Input: User ID in the request context.
[0159] Output: A set of information on items owned by the user and a set of information on user preferences.
[0160] The server uses database query statements or ORM interfaces to read the user's clothing records from the clothing information table, and reads style preferences, color preferences, price ranges, etc. from the preference information table. It then converts the query results into a list or dictionary structure and associates them with the current request in memory.
[0161] Step 6: The server performs feature extraction and pre-screening. The server performs feature encoding on environmental information, stored item information, and preference information, and performs preliminary filtering of clothing that is not suitable for the current environment based on preset rules.
[0162] Input: Environmental information object, collection of owned items information, and collection of preference information.
[0163] Output: Encoded environmental feature vectors, item feature sets, and preference feature sets, as well as a filtered subset of candidate clothing items.
[0164] The server maps temperature to interval labels, converts weather conditions into category labels, converts clothing categories, materials, waterproof properties, etc. into multi-dimensional features, and converts preference labels into vectors. Based on rules such as "prioritizing waterproof in rainy weather" and "prioritizing high warmth in low temperatures," the server removes obviously unsuitable items from the clothing set and marks the remaining items as a candidate list.
[0165] Step 7: Prompt statements for the server to construct generative artificial intelligence models The server uses environmental information, a list of candidate clothing items, user preference information, and user natural language questions to construct structured prompts to control the input to the generative artificial intelligence model.
[0166] Input: Environmental characteristics, candidate clothing information, preference characteristics, and user question text.
[0167] Output: Prompt text for generative artificial intelligence models.
[0168] The server follows a predefined template to sequentially write the weather description, user question, clothing list, and style requirements into the prompt statement, adding output format requirements and segmentation instructions, for example: "You are a professional fashion consultant."
[0169] Current weather information for the city: temperature 16.5℃, light rain, humidity 78%.
[0170] The user's question was: 'It's a bit cold today and it might rain. I'm going on a date. What should I wear?' List of clothing owned by the user: 1. Black waterproof jacket (suitable for rainy days) 2. Blue jeans 3. White sneakers 4. Gray knit sweater (relatively warm) Users prefer a style that is casual yet slightly formal, and not too flashy.
[0171] Based on the above weather conditions and the user's existing clothing, please provide a complete outfit suggestion for the date.
[0172] Please explain in sections: 1) How to match the top; 2) How to match bottoms; 3) How to choose shoes and accessories; 4) The reasons for this combination in terms of rain protection, cold protection, and aesthetics.
[0173] Please answer in Simplified Chinese. The server generates complete text using string concatenation or template rendering algorithms and stores it in variables for the model to use.
[0174] Step 8: The server calls a generative artificial intelligence model to generate clothing proposal information. The server takes the prompt as input and calls a generative artificial intelligence model through an interface to obtain clothing proposals in natural language.
[0175] Input: Prompt text, model call parameters (such as temperature, maximum generation length, etc.).
[0176] Output: Clothing proposal text in natural language format.
[0177] The server sends a request to the model service, packages the prompts and control parameters into call data, and waits for the model to return the generated results; the server extracts the generated text content from the response and stores it as clothing proposal information in the current session context.
[0178] Step 9: The server parses the clothing proposal text and structures the matching information. The server parses the generated clothing proposal text and breaks down the descriptions of tops, bottoms, shoes, and accessories into structured matching information.
[0179] Input: Natural language clothing proposal text.
[0180] Output: Structured outfit proposal information (including explanations of each garment component and the reasons behind it).
[0181] Based on the segmentation identifiers agreed upon in the prompt statement, the server performs string search and segmentation on the proposal text, extracting parts such as "tops", "bottoms", "shoes", "accessories", and "reason explanation" separately, and storing the corresponding content in a structured data structure for subsequent display and product matching.
[0182] Step 10: The server retrieves product group information and performs multi-level filtering. The server obtains information on clothing-related product groups from external product information services and performs multi-level filtering based on environmental characteristics, user preferences, and clothing proposals to form a candidate product set.
[0183] Input: Structured matching proposal information, environmental characteristics, and preference characteristics.
[0184] Output: A list of candidate product information that matches the current scenario.
[0185] The server calls the product service interface to obtain a preliminary product list by category and functional keywords; the server filters unsuitable products based on the user's style and price range preferences; the server then matches and scores the product attribute fields based on key attributes extracted from the clothing proposal text, such as "waterproof" and "slightly formal", retains the products with higher scores as candidates, and associates these products with the corresponding matching parts.
[0186] Step 11: The server dynamically adjusts the prompts based on emotion-related information and generates correction proposals (if applicable). If a user expresses emotions during the interaction, the server will add emotion-related information to subsequent prompts to obtain a more appropriate clothing suggestion that matches the user's state.
[0187] Input: Emotional features extracted from the user's natural language input, and original clothing proposal information.
[0188] Output: New prompts containing emotional constraints, as well as new clothing proposal text and structured information.
[0189] The server uses a sentiment analysis algorithm to identify the category and intensity of emotions, such as "low mood, moderate intensity," and adds text such as "The user is currently in a low mood, please avoid overly bright and exaggerated combinations" to the new prompt. The server then inputs the new prompt into the generative artificial intelligence model to obtain the updated proposal text, and repeats the parsing process to generate new structured matching information to replace or supplement the original proposal.
[0190] Step 12: The server integrates matching suggestions and candidate product information and generates a response. The server integrates the finalized pairing proposal information with the corresponding candidate product list to generate response data for terminal display.
[0191] Input: Structured matching proposal information and a list of candidate product information.
[0192] Output: A response object containing text suggestions and product recommendations.
[0193] The server converts the pairing information into a field layout suitable for front-end display, binds each pairing component to several product candidates, attaches the overall suggested description text to the response, and serializes the fields to form a response message with a unified format.
[0194] Step 13: The terminal receives the server's response and displays the result on the interface. After receiving the data returned by the server, the terminal parses the response and presents the clothing suggestion text and candidate products to the user in a graphic format.
[0195] Input: Response message from the server (including pairing proposal information and candidate product information).
[0196] Output: A visual interface for clothing suggestions and product recommendations displayed on the terminal device.
[0197] The terminal parses the response object, displays suggestions for tops, bottoms, shoes, etc. in text form, displays each candidate product in card or list form with pictures and prices, and opens the product link through a browser or embedded web view when the user clicks on a product.
[0198] Step 14: Users can provide feedback or update clothing information based on the results. After reviewing the suggestions, users can input feedback or update their clothing list via the terminal, such as adding newly purchased items or deleting items that are no longer used.
[0199] Input: Feedback text entered by the user on the terminal, and data on adding or editing clothing items.
[0200] Output: The update request sent by the terminal to the server, and the updated user-owned item information and preference information on the server side.
[0201] The terminal encapsulates the user's feedback and new clothing attributes into request data and sends it to the server; the server parses the request and writes it into the database, updating the clothing information table or preference information table, so that the subsequently generated prompts and matching suggestions can be calculated based on the latest data.
[0202] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0203] Existing computer-based clothing recommendation technologies typically employ fixed rules or simple recommendation algorithms, generating recommendation lists based solely on basic user attributes or historical purchase records. On one hand, traditional systems have limitations in acquiring environmental context, failing to automatically and precisely utilize real-time weather information corresponding to the user's current location, thus hindering timely adjustments to clothing recommendation strategies. On the other hand, traditional systems often rely on predefined templates when generating recommendation content, lacking comprehensive semantic understanding and natural language generation capabilities for multi-source heterogeneous data (weather data, user wardrobe data, store inventory data, etc.), resulting in recommendations lacking personalization and interpretability, and leading to a poor user experience.
[0204] Furthermore, existing systems also have shortcomings in terms of computer resource utilization and data processing workflows: location parsing, weather query, recommendation calculation, and interface display are usually completed by several independent services, lacking a unified processing pipeline oriented towards "prompt statements + generative artificial intelligence models," and unable to uniformly model multi-source data in a structured manner before inputting it into generative artificial intelligence models; at the same time, there is insufficient support for the linkage between real-time store environment data triggered by users in physical store scenarios through code information (such as QR codes) and online recommendation services, making it difficult to apply online generative artificial intelligence reasoning capabilities to offline in-store real-time shopping guidance scenarios.
[0205] Furthermore, in existing technologies, when generative AI models are applied to recommendation scenarios, they often simply take the user's natural language questions as input without specifically designing and optimizing the composition of the "prompt statements," dynamic update strategies, and the way they are associated with user historical behavior data, meteorological data, and store inventory data at the system level. This results in insufficient model input information and discontinuous context, thereby limiting the quality of model output and the overall intelligence level of the system.
[0206] Therefore, an improved computer implementation is needed. By building a unified processing chain on the server side for data acquisition, structured modeling, prompt generation, and generative artificial intelligence model invocation, meteorological information, multi-dimensional user feature information, and physical store environment information can be effectively integrated. This will improve the clothing recommendation system at both the algorithm and system architecture levels, thereby enhancing the real-time performance, personalization, and interpretability of recommendation results, as well as improving the efficiency and scalability of computer resource utilization.
[0207] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0208] In this invention, the server includes a functional module for actively invoking an external information providing device to obtain meteorological data based on location information from a user terminal via a communication network, and converting the meteorological data into structured environmental data containing temperature and meteorological status information; a functional module for uniformly reading user attribute information, information on the user's owned clothing, and user historical behavior information from a storage device, and associating them with the structured environmental data to generate user state data representing the user's current situation; a functional module for automatically constructing prompt statements containing environmental factors, user preferences, and clothing resource constraints based on the user state data, and sending the prompt statements as input to a generative artificial intelligence model to obtain semantic proposal information related to clothing combinations; a functional module for matching and filtering the semantic proposal information with user's owned information and product information obtained from the external providing device to generate clothing proposal information that distinguishes between "readily usable owned clothing combinations" and "candidate products available for purchase"; and a functional module for associating physical store identifiers and in-store inventory information with the clothing proposal information based on code recognition information from the user terminal to generate customized clothing proposals adapted to the store environment and feeding them back to the user terminal. This allows for the formation of a unified data processing and inference chain on the server side for generative artificial intelligence models, enabling multi-source heterogeneous data to be structured and contextualized before being input into the model, thus improving the relevance and effectiveness of model inference. At the same time, by collaboratively processing online weather services, multi-dimensional user data, and offline store environment data in the same computer system, real-time and personalized clothing recommendations for end users in different scenarios can be achieved, thereby improving the data processing flow and enhancing the scalability and intelligence of the system at the computer technology level.
[0209] "Information processing device" refers to a computing device with computing and storage capabilities, used to execute program code and process input data, including but not limited to servers, computer terminals, virtual machines, or cloud computing nodes.
[0210] "User terminal" refers to an electronic device carried or operated by a user for data interaction with an information processing device, including but not limited to smartphones, tablets, personal computers, or wearable devices.
[0211] "Location information" refers to spatial information used to represent the current geographical location of a user terminal, including but not limited to longitude, latitude, altitude, and parameters related to positioning accuracy.
[0212] "Communication network" refers to network infrastructure used to transmit data between information processing devices, user terminals and external information providing devices, including but not limited to the Internet, mobile communication networks, local area networks or any combination thereof.
[0213] "External information providing device" refers to a computing system or service platform that operates outside of the information processing device and is used to provide data services to other systems through an interface, such as a meteorological information service device or a commodity information service device.
[0214] "Meteorological information" refers to data that represents the environmental state of a geographical location at a specific time, including but not limited to temperature, perceived temperature, humidity, precipitation, descriptions of weather phenomena, wind speed, and wind direction.
[0215] "Structured data" refers to the data format obtained by organizing and representing raw data according to a predetermined data format or data model, so that each data item is stored and processed with explicit fields or labels.
[0216] "Temperature information" refers to data extracted from meteorological information that represents ambient temperature or perceived temperature, including numerical values, units, and optional temperature range labels.
[0217] "Meteorological status information" refers to data used to describe the categories and characteristics of weather conditions, including but not limited to sunny, cloudy, rainy, snowy, partly cloudy, and corresponding text descriptions or classification labels.
[0218] "Storage device" refers to a hardware or logical component used to store program code and data for a long period of time or temporarily, including but not limited to disk storage, solid-state storage, memory devices, and database systems built on them.
[0219] "Information management area" refers to the storage space or data structure in a storage device that is pre-divided or logically defined for storing user-related data, meteorological data, commodity data and their relationships.
[0220] "User attribute information" refers to data used to describe a user's basic characteristics, including but not limited to attributes such as gender, age group, height, body type, style preference, color preference, and place of residence.
[0221] "Item information" refers to data representing a list of clothing or related items currently owned by the user and their attributes, including but not limited to clothing type, size, color, material, applicable season, and frequency of use.
[0222] "User status information" refers to comprehensive data obtained by processing environmental information, user attribute information, and property information, which is used to represent the user's situation and available resource status at a specific point in time.
[0223] "Generative artificial intelligence model" refers to an artificial intelligence model that is based on machine learning or deep learning technology and learns from a large amount of training data, and can automatically generate text, images or other forms of output based on input content. In this invention, it is mainly used to generate text information for clothing combination suggestions.
[0224] "Prompt statements" refer to text content constructed by an information processing device based on user status information and system context, which is provided as input to a generative artificial intelligence model to guide the model in generating outputs related to the target task.
[0225] "Proposal information related to clothing combinations" refers to text or structured data that includes clothing matching schemes, reasons for recommendation, and related explanations, output by a generative artificial intelligence model based on prompts.
[0226] "Clothing elements" refer to specific clothing items or their category attributes involved in clothing combinations or proposal information, including but not limited to information units such as outerwear, tops, bottoms, footwear, and accessories.
[0227] "User preference information" refers to data obtained from the analysis of users' historical behavior data or user input, reflecting users' preferences in terms of style, color, price range, brand style, etc.
[0228] "Product information" refers to data related to purchasable products provided by external information providing devices, including but not limited to product name, category, specifications, price, inventory status, image address, and purchase link.
[0229] "Candidate clothing information" refers to a set of clothing or related products that have been filtered from product information and meet the requirements of generative artificial intelligence model proposals and user preferences.
[0230] "Code information" refers to machine-readable tags set in a physical environment to identify a specific object or scene, including but not limited to QR codes, barcodes, or graphic codes containing identification data.
[0231] "Identification information" refers to identification data obtained by the user terminal by reading code information, which can be used to uniquely identify a store, region, or service entry point on the server side.
[0232] "Store information" refers to data related to physical sales locations, including but not limited to store logo, geographical location, business category, business hours, and the types of services that can be provided.
[0233] "Inventory information" refers to data indicating the current availability and quantity of goods for sale in a specific store, including but not limited to product identification, inventory quantity, size distribution, and display location.
[0234] "Clothing proposal information" refers to clothing recommendation results displayed to users after comprehensive processing of the output of generative artificial intelligence models, user data, and product data by information processing devices. It includes matching schemes using already owned clothing and recommendations for purchasing candidate products.
[0235] "Dialogue-style display" refers to a display method on the user terminal interface that presents clothing proposal information one by one in a manner similar to a chat dialogue or question-and-answer message bubble.
[0236] "List display" refers to a display method where multiple clothing proposals or product recommendations are presented in a list of items or cards on the user terminal interface.
[0237] "Historical purchase record information" refers to data accumulated during system operation that is related to a user's past product purchase behavior, including but not limited to purchase time, purchased products, purchase quantity, and payment amount.
[0238] "Browsing history information" refers to data related to a user's browsing behavior of products, content, or pages in a terminal or system interface, including but not limited to browsing time, dwell time, click behavior, and favorite behavior.
[0239] "Data mining processing" refers to the process of using statistical analysis methods or machine learning methods to perform pattern recognition, clustering, or predictive analysis on historical behavioral data in order to extract user preferences, trend characteristics, or potential correlations.
[0240] "Style preference information" refers to data features extracted from users' historical behavior through data mining, used to represent users' long-term preferences in clothing style, matching habits, etc.
[0241] "Interface device" refers to hardware or software components used to enable data input, output and interaction between user terminals and information processing devices, or between different functional modules, including but not limited to graphical user interfaces, application programming interfaces or network interfaces.
[0242] "Natural language input information" refers to requests, questions, or instructions that users input into the system in the form of human natural language (such as text or text converted from speech recognition).
[0243] "Emotional state" refers to the state that reflects a user's psychological or emotional tendency, inferred from user input, behavior, or other relevant information, such as pleasure, depression, tension, or relaxation.
[0244] In a preferred embodiment of the present invention, the server, terminal, and user each assume different functional roles, and the three collaborate through a communication network to complete clothing recommendation processing based on meteorological information and user status information. At the hardware level, the server can be a computing device with a multi-core central processing unit, main memory, and non-volatile memory, such as a rack-mounted server or cloud computing node running a general-purpose operating system. At the hardware level, the terminal can be a mobile computing device equipped with a display device, touch input device, GPS module, and camera module, such as a smartphone or tablet. At the software level, the server can run operating system-based applications, including web server software, application server software, database management system software, and generative artificial intelligence model calling client libraries. At the software level, the terminal can run native applications or browser applications, utilizing the positioning interface, camera interface, and network communication interface provided by the operating system.
[0245] In one exemplary implementation, the server uses a Unix-like operating system, such as Linux, and runs a web application framework implemented in an interpreted or compiled language, such as a Python-based web framework or a scripting language-based web framework. The server uses a relational database management system or a document-oriented database management system as data management software on the storage device to store user attribute information, item information, historical behavior data, and store inventory data. The server uses an HTTP client library to communicate with external information providers, such as using common network request libraries to call weather information service interfaces and product information service interfaces. The server uses a generative artificial intelligence model calling library to communicate with external artificial intelligence inference services; in a preferred example, the server calls a text generation model based on a multi-layered self-attention mechanism, which is a large-scale language model.
[0246] In one embodiment, the terminal uses a mobile operating system, such as a mobile device operating system, and obtains location information through system-provided location services (such as a fusion service based on satellite positioning and base station / Wi-Fi assisted positioning). The terminal captures image data through a camera and parses code information using a local decoding library. The terminal communicates encrypted with the server via a secure version of Hypertext Transfer Protocol (HTTP) to transmit user input, location information, and code recognition information, and receives clothing proposal information generated by the server. The user inputs natural language questions through the terminal's touchscreen, allowing the terminal to access their location information and scan code information using the camera in physical stores, thereby initiating recommendation requests in different contexts.
[0247] In terms of data structure design, to improve data management and subsequent computation efficiency, the server divides the data related to clothing recommendations into various logical tables or sets. The server maintains the following tables in its storage: a basic user information table (containing fields such as user ID, gender, age group, and residential region); a user preference table (containing fields such as color preference, style preference, price preference, and sensitivity to cold); a user inventory table (containing fields such as the category, material, color, size, and applicable season of clothing owned by the user); a historical behavior table (containing browsing and purchase records, including timestamps, product IDs, dwell time, clicks, and purchase quantities); an environmental information table (caching meteorological data corresponding to different latitudes, longitudes, and times to reduce the number of calls to external meteorological services during frequent queries within a short period); and a store information table and a store inventory table (describing the store's ID, location, available products, and inventory status).
[0248] In terms of data processing, the server performs structured transformation on raw meteorological data from external meteorological information services. The server parses fields such as temperature, perceived temperature, precipitation type, precipitation intensity, humidity, wind speed, and weather description from the response messages returned by the external meteorological services, and converts them into unified structured data objects. For example, the server normalizes raw numerical fields to uniform units, maps weather description fields to predefined category labels (such as "sunny," "cloudy," "light rain," "heavy snow," etc.), and generates several derived labels (such as "cold," "muggy," "need wind protection," "need rain protection," etc.) based on temperature ranges and wind speed thresholds. Through this structuring and labeling process, the server transforms raw meteorological data, which would otherwise be difficult to directly use as input for generative artificial intelligence models, into environmental features that are easy to describe and combine. This significantly reduces the complexity of string concatenation in the subsequent prompt generation stage, improving the overall processing efficiency of the system.
[0249] In user behavior analysis, the server uses data mining algorithms to perform statistical analysis and pattern mining on browsing and purchase records in the historical behavior table. The server can use matrix factorization-based collaborative filtering algorithms, style clustering methods based on clustering algorithms, or preference prediction models based on supervised learning models. In one preferred implementation, the server uses a gradient boosting tree model or a shallow neural network to model user historical behavior, outputting several high-level preference features, such as user preference scores for different style tags (business, casual, sports), different color categories, and different warmth levels. The server stores these preference score vectors in a user preference table and writes them into a natural language description as part of the feature input when generating prompts. Through this preference modeling that is updated offline or online in advance, the server does not need to repeatedly perform complex calculations from the original behavior data each time a prompt is generated, thereby reducing the real-time computing load and improving the recommendation response speed.
[0250] In one embodiment, the server uses a large language model based on a multi-layer Transformer architecture for selecting and invoking generative AI models. This model contains multiple encoder-decoder layers or stacked decoder layers, each layer containing a multi-head self-attention sublayer and a feedforward neural network sublayer. During the training phase, the server can pre-train and fine-tune the base language model using large-scale general text corpora and domain-specific corpora related to clothing, styling, and weather. During fine-tuning, the server uses a cross-entropy loss function as the objective function and employs adaptive optimization algorithms (such as adaptive moment estimation) to update the model weights. The server can also use data augmentation techniques, such as synonym replacement, syntactic transformation, and context rewriting, to expand the training samples and enhance the consistency and robustness of the generated recommendation text under different expressions.
[0251] During the inference phase, the server uses a generative AI model to call an interface, taking the constructed prompt as input. It sets parameters for maximum generated length, sampling temperature, and probability truncation to control the length, diversity, and stability of the output text. In one example, the server constructs the prompt into multiple segments of structured natural language, clearly distinguishing between sections such as "current weather," "user information," "clothing owned by the user," and "task requirements." For example, the server can generate the following prompt: You are a professional fashion consultant.
[0252] Current weather - City: Tokyo - Weather: Light rain - Temperature: 12℃, feels like 9℃ - Humidity: 80% - Wind speed: 5 m / s - Weather tag: Cold, waterproof required User Information - Gender: Male - Age range: 25-34 years old - Style preference: Casual with a touch of business style - Color preference: Black, navy blue, gray - Characteristics: Not too afraid of the cold, but doesn't like getting wet in the rain. Clothing owned by the user - Outerwear: Dark blue waterproof windbreaker, black lightweight down jacket - Tops: White T-shirt, grey crew neck sweatshirt, blue denim shirt - Bottoms: Dark denim jeans, black slim-fit casual pants - Shoes: White sneakers, black ankle boots (waterproof) User Issues What should I wear today? Task 1. Based on the current weather and user information, create an outfit suitable for commuting and daily activities in the city from the user's existing clothing.
[0253] 2. Clearly state which item you will choose for the coat, top, bottoms, and shoes, and give a brief reason.
[0254] 3. Please remind us separately if we need to prepare additional rain gear (such as an umbrella).
[0255] Please list your suggestions in Simplified Chinese, in bullet points.
[0256] By explicitly embedding structured tags and task descriptions into the prompt statements, the server departs from the traditional approach of simply piecing together user questions. This allows the generative AI model to comprehensively consider weather characteristics, user preferences, and wardrobe resources within a unified context. This specific prompt statement construction rule represents a non-conventional improvement to the internal computer processing flow in this invention, resulting in clothing proposal information output by the model that is significantly superior in relevance and consistency to the traditional approach that uses only user question text as input.
[0257] In the clothing element matching and product selection stages, the server further transforms the output of the generative artificial intelligence model into actionable recommendation results through explicit data structures and algorithmic constraints. Upon receiving the natural language text output by the model, the server uses keyword extraction algorithms or lightweight sequence labeling models to identify clothing categories, colors, and functional descriptions (such as "waterproof jacket," "dark trousers," and "ankle boots"), mapping these elements to a predefined clothing category dictionary. In this process, the server does not rely entirely on manual rules for matching; instead, it uses techniques such as statistical word frequency analysis, contextual co-occurrence relationships, and word vector similarity to map natural language descriptions to an internally unified clothing element encoding, thereby reducing matching errors caused by differences in expression.
[0258] After obtaining the clothing element codes, the server first matches them against the user's inventory table, filtering out candidate clothing combinations that the user owns and that meet functional requirements and weather conditions. Then, the server filters purchasable items in the product information database based on the same codes, color preferences, and price ranges, generating a purchase candidate list. In this two-stage matching process, the server uses a ranking algorithm, employing weather matching degree, preference matching degree, and inventory availability as ranking features to improve the accuracy and usability of the recommendation results. Because the clothing elements are already coded and pre-indexed, this retrieval and ranking process can be completed within logarithmic complexity, thereby improving the system's response speed under conditions of a large-scale product database.
[0259] In a physical store setting, the server processes code recognition information from the terminal, combining online recommendations with the offline store environment. The terminal captures images via a camera and uses a 2D barcode decoding library to parse the code content, obtaining the store identifier or area identifier. After receiving the recognition information, the server searches for the corresponding record in the store information table and extracts the list of currently available products from the store inventory table. When constructing the prompt statement, the server adds a "brief list of available products" and "store style tags" to a part of the prompt statement, causing the generative AI model to automatically prioritize in-store purchasable products when generating proposals. For example, the server can add the following to the prompt statement: Items available for sale in the store 1. Camel wool coat (mid-length) 2. Black down jacket (short style) 3. Gray turtleneck sweater 4. Black wool wide-leg pants 5. Platform leather boots (black) 6. Cashmere scarf (light gray) In this way, the server incorporates offline inventory constraints into the prompt statements. The generative AI model, through its internal attention mechanism, will be more inclined to output combinations that match the aforementioned list, thus technically achieving automatic alignment between online inference results and offline supply capacity. This approach of explicitly injecting constraint sets through prompt statements, compared to manual selection of solutions by humans as an intermediary, reduces multiple human-computer interactions and manual editing operations. It allows the computer system to complete complex conditional and constraint inference in a single model call, improving overall computational efficiency.
[0260] In terms of computer technology improvements, the server also reduces the communication load and processing burden caused by repeated calls to generative artificial intelligence models through strategic design of the prompt generation module. In one implementation, the server shares some environmental segments for users making multiple requests within the same geographical area within a short period, dynamically rewriting only the user-personalized parts and task descriptions. The server can also pre-generate template prompts for common weather patterns and user type combinations, inserting only a small amount of incremental information during real-time invocation, thus significantly reducing the amount of string manipulation and data transmission required for each prompt construction. Through this layered prompt strategy, the server can support more concurrent access from terminals per unit time, thereby achieving performance improvements at the system level.
[0261] In terms of implementation, the terminal converts real-world location and code information into digital signals that the server can process by calling the system's positioning interface, camera interface, and network communication interface. Because the terminal uses local image decoding and positioning preprocessing functions, the server does not need to process the raw image data and raw signals, thus reducing data upload volume and server computational load. Furthermore, when displaying clothing proposal information, the terminal can render the structured results returned by the server (including text suggestions, a list of owned clothing, and a list of potential purchase items) as dialogue messages and product cards, respectively. The terminal locally controls the display order, pagination, and lazy loading strategies to avoid loading large amounts of images or long text at once, further reducing network bandwidth and rendering latency.
[0262] During use, users ask questions to the terminal using natural language, and trigger different scenarios through authorized location and code scanning, enabling the system to collect multi-source state information. Users do not need to master complex operating procedures; they only need to make simple requests such as "What should I wear today?" or "What should I wear on a rainy day?" to receive personalized clothing suggestions based on the current weather, personal preferences, and the store environment. Internally, the system performs complex calculations through a unified data structure, specific prompt statement construction strategies, and multi-stage matching algorithms, thus extending the human decision-making process rather than simply automating it.
[0263] In summary, this invention achieves the following technical effects by designing a unified data processing link for generative artificial intelligence models on the server side, a specific prompt generation strategy, and a clothing element encoding and matching mechanism: First, structured and labeled environmental and user state data reduce model input noise, improving the relevance and stability of recommendation results; second, pre-constructing user preference feature vectors and reusable prompt fragments reduces real-time computation and communication overhead, improving system response speed under high concurrency conditions; third, clothing element encoding and a two-stage matching algorithm maintain high retrieval efficiency and matching accuracy under large-scale product database conditions; and fourth, by injecting physical store inventory information into prompts and participating in model inference, it achieves a technical coupling between online generative recommendations and offline inventory constraints. These improvements are concentrated in the internal data structure design, algorithm flow, and model calling method of the computer, enabling this invention to not only provide clothing recommendation functionality at the application level but also substantially improve the processing efficiency, data management, and intelligent reasoning capabilities of the computer technology itself.
[0264] use Figure 12 The processing flow is explained.
[0265] Step 1: The terminal receives user requests and collects location information. After the user opens the application, the terminal displays an input interface. The user enters a natural language question (e.g., "What clothes should I wear today?") and clicks send. The terminal's input consists of the user's text and the current application state. The terminal calls the location service interface provided by the operating system to drive the GPS and network positioning modules to obtain the current latitude, longitude, and accuracy information. The terminal filters and denoises the raw location data sampled multiple times, selecting the most recent and most accurate record as the location information. The terminal assembles the user identifier, the user's input text, and the processed location information into a request data structure. The terminal's output is structured request data containing the user's question and location information.
[0266] Step 2: The terminal sends a request to the server. The terminal uses a network communication module to establish a connection with the server via a secure transmission protocol. The terminal's input is the structured request data generated in step 1. The terminal calls an HTTP client library to encode this data into a request message and sends it to the server's pre-configured interface address. During transmission, the terminal serializes the request data and performs necessary encryption to ensure data integrity and confidentiality during network transmission. The terminal's output is the request message that has been sent to the server over the network.
[0267] Step 3: The server parses the request and verifies the basic data. The server receives request messages from the terminal at the application layer. The server's input is structured request data sent by the terminal. The server uses a web framework to parse the HTTP message, extracting the user identifier, user question text, and location information. The server performs format validation on the location information field, checking if the latitude and longitude are within the valid range; it also performs a validity check on the user identifier and checks if the question text is empty. The server generates an error response when validation fails and writes the valid request data to the access log or request queue list when validation succeeds. The server's output is the validated user request record and the internal data objects required for subsequent processing.
[0268] Step 4: The server acquires meteorological data based on location information and performs structured processing. The server takes the location information output in step 3 as input and sends a query request to an external meteorological information service via an HTTP client library. The server constructs request parameters based on latitude and longitude, requesting current weather conditions and temperature information. The server receives raw meteorological data packets returned by the external meteorological service and parses out fields such as temperature, perceived temperature, weather phenomena, precipitation, humidity, and wind speed. The server performs unit standardization, numerical formatting, and outlier filtering on these fields, and generates labels such as "cold," "hot," and "needs rain protection" based on temperature ranges and weather types. The server organizes these numerical fields and labels into a structured environmental data object. The server's output is environmental data containing standardized temperature information, meteorological status information, and derived labels.
[0269] Step 5: The server obtains user attributes, owned items, and historical behavior data. The server reads records corresponding to user identifiers from the database in the storage device. The server's input is the user identifier. The server queries the user's basic information table for gender, age group, and place of residence, the user preference table for style preferences, color preferences, and price preferences, the inventory table for the user's registered clothing items and their attributes, and the historical behavior table for browsing and purchase records. The server performs statistical aggregation on the historical behavior data, calculates the preference weights for each style tag and color category, and forms a preference feature vector. The server combines the basic attributes, preference features, and inventory list into basic user status data. The server's output is a structured set of user attribute information, inventory information, and preference features.
[0270] Step 6: The server generates user status information. The server takes the environmental data generated in step 4 and the user basic data obtained in step 5 as input. Through a data integration module, the server organizes environmental tags (such as "cold" or "light rain"), user attributes (such as gender and age group), preference characteristics (such as "leisure-oriented" or "neutral color-oriented"), and the list of owned items into a single user status information object. The server fills in inconsistencies or missing fields and handles default values; for example, it uses statistical means or general preferences when some preferences are missing. Through this data integration, the server maps multi-source heterogeneous data to a unified key-value structure, facilitating subsequent insertion into the prompt statements by paragraph. The server's output is user status information containing environmental, user characteristics, and resource constraints.
[0271] Step 7: Prompt statements for the server to construct generative artificial intelligence models The server takes user status information and user-input questions as input and initiates the prompt generation module. Following a predefined template, the server generates natural language paragraphs from elements such as "current weather," "user information," "user's owned clothing," and "task requirements." The server inserts corresponding text descriptions based on environmental tags, such as converting temperature values to "Temperature: 12℃, Feeling temperature: 9℃, Weather: Light rain"; describing style preferences as "Leaning towards casual, with a slight business feel"; and listing available clothing items based on the inventory list. The server appends the user's original question and clear task instructions to the end of the prompt, such as requiring structured suggestions and brief reasons. The server combines this information into a complete prompt text using string concatenation and template rendering algorithms. The server's output is a prompt text for generative artificial intelligence models.
[0272] Step 8: The server invokes a generative artificial intelligence model to generate clothing proposal information. The server takes the prompt generated in step 7 as input and initiates an inference request through the generative artificial intelligence model call interface. In the call request, the server specifies the model type as a text generation model based on a multi-layer self-attention mechanism and sets parameters such as maximum generation length, sampling temperature, and probability truncation. The server submits the prompt as input text to the model server. Internally, the model uses pre-trained word embedding layers, self-attention layers, and feedforward networks to encode the prompt and predicts the output word by word based on language modeling objectives, generating complete clothing matching suggestion text. The server parses the generated text from the model response, removing redundant blank lines and irrelevant prefixes. The server's output is a clothing proposal text containing specific matching schemes and explanations.
[0273] Step 9: The server extracts clothing elements from the proposal text and encodes them. The server takes clothing proposal text output by a generative artificial intelligence model as input and identifies the clothing types, colors, and functional descriptions mentioned in the text through keyword extraction algorithms or lightweight sequence labeling models. Using a predefined clothing category dictionary and thesaurus, the server maps different expressions (such as "waterproof jacket," "raincoat," and "rainproof windbreaker") to internal clothing element codes. For each clothing item appearing in the proposal, the server generates a structured entry with a category code, function code, and color code. Through this element extraction and encoding process, the server converts free text into standardized data that can be used for database matching. The server's output is a list of clothing element codes.
[0274] Step 10: The server matches clothing elements from the stored item information. The server takes a list of clothing element codes and the user's owned items as input and executes a matching algorithm. For each clothing element code, the server retrieves clothing records from the owned items table that match in category, function, and color attributes. When multiple candidates exist, the server sorts them based on usage frequency, applicable season, and user preference score, selecting the best or several best options. The server reassembles the selected owned items into a complete outfit, such as an outerwear + top + bottoms + shoes, and records elements that cannot be satisfied by owned items. The server's output is a candidate list of outfits based primarily on the user's owned clothing and a list of missing elements.
[0275] Step 11: The server filters and selects potential clothing items from the product information. The server takes a list of clothing element codes, a list of missing elements, and a product information database as input and performs a query operation on the product information table. Based on the category, function, and color attributes in the element codes, the server quickly locates matching product entries using an index. When there are many candidate products, the server further filters and sorts them based on price range, brand tags, and inventory status. For each missing element, the server selects several of the most matching products as purchase candidates, adding fields such as product name, price, image link, and purchase link to the results. The server's output is a set of purchase candidate clothing information for each missing element.
[0276] Step 12: The server links physical store information with clothing proposal information. In the in-store scenario, the server takes code recognition information from the terminal as input, queries the store information table for the geographical location and operating style corresponding to the store identifier, and retrieves the list of currently available products for sale in the store from the store inventory table. The server performs intersection matching between the products in the store inventory and the purchase candidates filtered in step 11, prioritizing the retention of products that are in stock in the store, and distinguishing and marking products that are not in store inventory but can be purchased online. The server integrates this store-related product information with clothing proposal text and inventory matching information to generate clothing proposal information corresponding to the store environment, marked as "products available for try-on in this store" and "products available for purchase online only". The server output is structured clothing proposal data with store constraints and inventory labels.
[0277] Step 13: The server generates the final response data for display on the terminal. The server takes clothing proposal text, a list of owned items, a list of potential purchase items, and store association information as input, and constructs a unified response data structure. The server places the natural language suggestion text in the "Description Section," organizes the owned items into a "List of Owned Clothing Combinations" in sequence, and arranges the potential purchase items according to recommendation priority into a "Recommended Purchase List," attaching product image addresses and links. The server compresses and simplifies the response data, removing intermediate calculation fields unnecessary for terminal display to reduce network transmission burden. The server's output is compact response data suitable for direct rendering on the terminal.
[0278] Step 14: The terminal receives the server's response and parses the data. The terminal takes the response message returned by the server as input and parses the response header and body using an HTTP client. The terminal uses a JSON parsing library to convert the response body into an internal object structure, extracting fields such as suggested text, owned clothing combinations, and recommended purchase items. The terminal performs preliminary validation of product image URLs and applies length limits to text fields to prevent abnormal data from affecting the interface layout. The terminal's output is a local data object suitable for interface rendering.
[0279] Step 15: The terminal displays clothing proposal information in the user interface. The terminal takes the local data object parsed in step 14 as input and calls the interface rendering engine to draw content on the screen. In the chat area, the terminal displays the suggested text output by the generative AI model in a conversational format, showing different suggested items in separate lines or segments. Below, the terminal displays a list of "Owned Clothing Combinations," showing the clothing name and key attributes for each item. In the product recommendation area, the terminal displays purchase candidate product cards, including images, names, prices, and a purchase button. Responding to user touch input, the terminal opens an embedded browser or an external browser to access the product purchase link when the user clicks on a product card. The terminal's output is a visual interface presented on the display screen and user-interactive controls.
[0280] Step 16: Users select and provide feedback based on the proposals. Users input clothing proposals displayed on the terminal, read styling suggestions on the interface, and confirm whether to adopt a recommended combination of clothing they already own or click to enter the product purchase page. Users can provide feedback to the system by selecting "satisfied," "unsatisfied," or providing a brief rating for a particular outfit. In performing these actions, users essentially transform their subjective judgments of the recommendation results into behavioral data for the system to learn and optimize later. User output consists of new interaction records and feedback tags.
[0281] Step 17: The server records user feedback and updates the preference model. The server takes user feedback data uploaded from terminals as input, classifying and counting the feedback results. It aggregates positive and negative user feedback on various style tags, colors, and clothing categories into a user preference table, fine-tuning the corresponding preference weights. The server can periodically retrain or incrementally update the preference prediction model using accumulated feedback data, making the model parameters more closely reflect users' actual preferences. Through this iterative process, the server gradually improves the accuracy of preference descriptions in subsequent prompts and the ranking quality during product selection. The server's output is updated preference features and behavioral statistics, providing a more accurate data foundation for subsequent recommendation loops.
[0282] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0283] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0284] In the field of computer-based clothing styling suggestions, existing technologies typically rely on rule engines or simple recommendation algorithms to provide static, patterned clothing suggestions based on temperature, weather, or pre-defined styling rules. This approach suffers from several problems: First, the system struggles to accurately analyze the fine-grained styling needs expressed by users in natural language, failing to extract semantic elements such as target clothing, usage scenarios, and styling ranges from complex queries, resulting in generated suggestions that do not align with the user's true intentions. Second, existing systems often treat the generative model merely as a one-way "black-box text generator," failing to construct structured prompts tailored to the specific user and incorporating clothing database information, scenario information, and preference information before generating the model. This fails to fully utilize the generative capabilities of generative AI models, leading to inconsistent and unexecutable results that do not match the user's existing clothing data. Third, existing technologies often lack structured post-processing capabilities for the generated results, making it difficult to integrate the natural language output of the generative model with specific data in the database. The lack of precise mapping between clothing data (identifiers, images, attributes) prevents the terminal from intuitively displaying matching schemes in a text-image linkage format, reducing human-computer interaction efficiency. Fourth, traditional recommendation systems utilize user preferences and emotional states in a rather crude way, making it difficult to embed preference and emotional information as constraints or optimization goals during the prompt construction stage. This results in the system's inability to dynamically adjust matching strategies based on the user's historical behavior and current emotions. Fifth, existing systems lack a complete end-to-server collaborative technical process, including an integrated processing path for collecting multimodal input from the terminal, performing semantic parsing and data retrieval on the server side, constructing prompt statements, calling generative artificial intelligence models, and performing structured post-processing on the response results before returning them to the terminal for display. Consequently, the system suffers from deficiencies in overall response efficiency, data consistency, and scalability.
[0285] Therefore, it is necessary to provide an improved computer implementation method and system architecture. By introducing a prompt statement construction mechanism for generative artificial intelligence models, a semantic retrieval and matching mechanism based on clothing databases, a dynamic adjustment mechanism combining user preferences and emotional information, and a structured result generation mechanism for terminal display on the server side, the quality of clothing matching suggestions can be improved while the overall optimization of computer data processing flow and human-computer interaction process can be achieved.
[0286] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0287] In this invention, the server includes a data acquisition and storage unit for acquiring image data and attribute data of clothing owned by the user through an information input device and an image acquisition device, structuring the data, and storing it in a clothing information dataset; a natural language understanding unit for parsing natural language questions input by the user through a terminal device and extracting target clothing information, matching scenario information, and matching requirement type information; a data retrieval and organization unit for retrieving clothing record data corresponding to the above parsing results from the clothing information dataset based on the user identifier and generating a data set for use by a generative artificial intelligence model; and a data retrieval and organization unit for combining the clothing record data with the natural language parsing results and user preferences. A prompt statement generation unit combines information and emotional information to generate prompt statements containing system role information, user clothing list information, matching scenario information, and matching task instructions according to a predetermined text template. A model invocation unit provides the prompt statements to a generative artificial intelligence model via a communication device and receives natural language responses containing at least one clothing matching scheme. A result structuring and output unit further processes the natural language responses, matches the clothing name information in the generated result with clothing record data in the clothing information dataset, generates structured matching result data containing clothing identification information, image reference information, and matching reasons, and sends it to the terminal device. This creates a data processing pipeline optimized for generative artificial intelligence models on the server side: automatically extracting structured semantic information from user natural language questions and the clothing database; constructing high-quality prompt statements through the prompt statement generation unit and inputting them into the generative artificial intelligence model; and then precisely binding the model output with specific clothing data through the result structuring and output unit. This achieves clothing matching suggestions with both semantic understanding and data consistency, thereby improving the processing efficiency and result quality of computers in multimodal data processing, natural language interaction, and intelligent recommendation.
[0288] "Processing device" refers to an electronic device with computing power and the ability to execute program instructions. It may include a central processing unit, a graphics processing unit, a memory and related chipsets, and is used to process input data and control the hardware or combination of hardware and software of the various functional units of the system.
[0289] "Storage device" refers to a computer-readable storage medium used to store program data and user-related data, which may include semiconductor memory, magnetic storage medium, optical storage medium or a combination thereof, for persistently storing clothing information datasets, user information datasets and preference models, etc.
[0290] "Terminal device" refers to an electronic device operated by a user and interacting with a server, which may include mobile communication devices, portable computing devices, desktop computing devices or other devices with human-computer interaction interfaces, used to input user information and display matching suggestions.
[0291] "Information input device" refers to an input device used to obtain text, options or control instructions from a user, and may include a touch screen, keyboard, mouse, voice input interface or a combination thereof, for collecting text information and operation instructions from the user.
[0292] "Image acquisition device" refers to a device used to acquire still or moving image data, which may include a camera, scanning device or sensor with imaging function, for acquiring images of clothing or users.
[0293] "Sound acquisition device" refers to a device used to acquire audio signals, which may include a microphone array or a single microphone, for acquiring user voice data or ambient sound data.
[0294] "Sensing device" refers to sensing hardware used to detect and collect environmental parameters or status information. It may include temperature sensors, humidity sensors, position sensors, acceleration sensors, or weather information acquisition modules provided by communication networks to acquire environmental information such as temperature and weather.
[0295] "Communication device" refers to a communication module used for sending and receiving data between a server and terminal devices and external services. It may include wired network interfaces, wireless network interfaces, and software modules that implement communication protocols, and is used to transmit request and response data over a network.
[0296] "Clothing information dataset" refers to a structured collection of data related to clothing owned by a user, stored in a storage device. Each record includes at least clothing type information, color information, style information, occasion information, season information, image reference information, and identification information.
[0297] "Clothing record data" refers to structured data units that represent the attributes of one or more garments. These records are extracted from clothing information datasets and are used to describe the basic characteristics, usage scenarios, and association with image data of a particular garment.
[0298] "User identification information" refers to identification data used to uniquely or distinguish a specific user in the system, which may include account identifiers, device identifiers, or other markers that can be used to distinguish different users.
[0299] "Natural language query data" refers to text data entered by users in natural language to request clothing matching suggestions, including question sentences, explanatory sentences, or compound descriptive sentences.
[0300] "Word segmentation" refers to the automatic segmentation of natural language query data, which divides a continuous text stream into words or basic language units for subsequent keyword extraction and intent recognition.
[0301] "Keyword extraction processing" refers to the process of identifying and selecting semantically representative or important words or phrases from natural language query data to determine core elements such as color, clothing type, and scene.
[0302] "Intent recognition processing" refers to the process of analyzing natural language question data through natural language processing algorithms to infer the user's target needs, task type, and constraints, in order to determine the matching target and scope.
[0303] "Target clothing information" refers to clothing-related information obtained from natural language query data that is clearly used as a matching benchmark or key object, including color, type, or descriptive characteristics.
[0304] "Outfit context information" refers to semantic information determined from natural language query data or other inputs that indicates the purpose and environment of clothing outfits, including activity type, formality, season, etc.
[0305] "Category information for matching needs" refers to semantic information describing which clothing categories the user wants the system to recommend, such as tops, bottoms, shoes, or accessories.
[0306] "Prompt statements" refer to text data constructed for generative artificial intelligence models, which includes system role settings, user clothing lists, matching scenarios, and task descriptions. These are used as input to the generative artificial intelligence models to guide them in generating the target output.
[0307] "Generative AI models" refer to AI models trained using machine learning and deep learning techniques that can automatically generate natural language text or structured output based on input text. They typically employ neural network structures and are used to generate clothing matching suggestions.
[0308] "Natural language response information" refers to the response content in natural language form output by the generative artificial intelligence model based on the prompt statement, which includes one or more clothing matching schemes and descriptions generated for the matching task.
[0309] "Structured matching result data" refers to matching result data generated by the server after parsing and matching natural language response information, represented by a predetermined data structure, including clothing identification information, image reference information, and matching reason information in each matching scheme.
[0310] "Image reference information" refers to the identification information that can be used to access or associate clothing image data, including image file path, Uniform Resource Locator (URL) or internal image identifier, used to display the corresponding clothing image on the terminal.
[0311] "User preference information" refers to data that reflects a user's preferences in terms of color, style, clothing category, brand, etc., obtained based on the user's historical behavior, evaluation information, or explicit settings, and is used to adjust preferences during the recommendation and generation process.
[0312] "Online product information data" refers to structured information data about products provided by external services or online platforms, including product types, colors, prices, brands, images, and descriptions, which can be used to generate or expand clothing matching suggestions.
[0313] "Emotion analysis algorithm" refers to an algorithm or model that analyzes and infers a user's current emotional state based on image data, sound data, or text data, and outputs emotional information representing the category or intensity of emotion.
[0314] "Emotional information" refers to semantic data obtained by emotion analysis algorithms that represents the user's current psychological state. It can include categories such as pleasure, calmness, and tension, as well as their intensity, and is used to adjust matching strategies or prompt statements.
[0315] A “Natural Language Understanding Unit” refers to software or a combination of software and hardware that runs on the server side and is used to segment, extract keywords, and identify intent from user natural language questions to obtain a structured semantic representation.
[0316] The "Data Retrieval and Organization Unit" refers to a functional module used to retrieve relevant records from the clothing information dataset based on the semantic information obtained from parsing, and to organize and filter the retrieval results according to dimensions such as type, color, and purpose, so as to provide a basis for generating prompts and calling models.
[0317] The "prompt statement generation unit" refers to a functional module used to generate prompt statements suitable for use as input to a generative artificial intelligence model, based on clothing record data, natural language parsing results, user preference information, and emotional information, according to a predetermined text template.
[0318] The "model invocation unit" refers to a functional module used to provide prompts to generative artificial intelligence models via communication devices, receive natural language responses, and set generation parameters as needed.
[0319] The “Result Structuring and Output Unit” refers to the functional module used to parse the natural language response information output by the generative artificial intelligence model, match it with the clothing information dataset, generate structured matching result data, and send the result data to the terminal device.
[0320] In an embodiment of this invention, the system includes a server, a terminal, and an input / output interface operated by a user. The server runs on a computer device with computing and network communication capabilities, such as a server equipped with a general-purpose central processing unit and a graphics processing unit, and the operating system can be a Unix-like server operating system. The server utilizes database management software (e.g., a relational database management system), application server software (e.g., a general-purpose web server and backend framework), and a deep learning inference framework (e.g., a framework based on tensor computation) to work together to realize functions such as data storage management, natural language processing, generative artificial intelligence model invocation, and result structured processing. The terminal is a smart terminal device, which can be a mobile device equipped with a camera and a display screen or a general-purpose computing terminal, running a terminal operating system and a graphical user interface, used to perform user input collection and result display.
[0321] During the initialization phase, the server creates multiple data tables and index structures in the storage device. The server constructs logical tables in the database, including at least a clothing information dataset, a user information dataset, a preference information dataset, and a log dataset. Each record in the clothing information dataset contains fields such as clothing identifier, user identifier, category (e.g., classification codes for "tops," "bottoms," "footwear"), color (using normalized color coding), style (casual, formal, etc.), occasion (commuting, dating, etc.), season, image reference, and timestamp. The server improves subsequent retrieval efficiency by indexing the user identifier field and the composite key (user identifier, category, color), thus enabling low-latency responses to data requests from generative artificial intelligence models even with large-scale clothing data.
[0322] When a user registers clothing on the terminal, the user takes a picture of the clothing using the terminal's image acquisition device. The terminal then uses the camera interface provided by the operating system to acquire the raw image data. The terminal performs preliminary image processing locally, such as scaling the high-resolution image to medium resolution using an image processing library to reduce network transmission burden. The terminal also provides a form interface where the user enters information such as clothing color, type, brand, applicable season, applicable occasion, and material. The terminal standardizes this text information, for example, mapping Chinese color names to internal color codes and category names to preset enumerated categories, thereby generating a set of structured attributes. The terminal then transmits the structured attributes and image data to the server.
[0323] After receiving data from the terminal, the server calls the backend logic program to convert the clothing attribute data into database insertion records. The server groups and stores clothing records according to the user identifier field and saves the corresponding clothing image files in a file storage system (such as object-oriented file storage or a distributed file system). Only the image reference path or Uniform Resource Identifier is stored in the clothing record data. Optionally, the server uses an image processing library (such as general-purpose image processing software) to perform operations such as color histogram extraction and contour feature extraction on the uploaded clothing images, appending the image feature vectors to the clothing record data for use as auxiliary features when generating prompts or performing similar clothing searches. This structured storage and feature preprocessing allows the server to quickly access the required clothing information without re-decoding the original image, improving data processing efficiency.
[0324] When users need styling suggestions, they can input their styling questions through the natural language input interface on the terminal. Users can input the following text: “I want to wear that red midi dress to my friend’s wedding. Can you recommend a top and shoes for me?” The terminal sends the text as a natural language question to the server, along with a user identifier. Upon receiving the natural language question, the server invokes a natural language understanding unit (NLU) in memory. The NLU can be implemented based on a Chinese word segmentation library, a part-of-speech tagging module, and a self-trained or pre-trained intent classification model. The server performs word segmentation on the question text, dividing the sentence into word sequences, such as "I / want / to / use / that / red / mid-length skirt / to / attend / a friend's / wedding / help / me / recommend / a top / and / shoes". The server identifies the color phrase "red", the clothing phrases "mid-length skirt", "top", and "shoes" using a combination of rules and statistical models, and identifies "wedding" as the matching scenario using a scene dictionary and classification model. The server also determines "red mid-length skirt" as the target clothing and "top and shoes" as the clothing categories to be recommended using dependency parsing or sequence labeling methods. This semantic structure is represented internally by the server as key-value pairs or a tree structure for subsequent data retrieval and prompt generation unit consumption.
[0325] The server invokes the data retrieval and processing unit based on the extracted target clothing information. The server executes queries in the clothing information dataset according to user identifiers and attribute conditions, such as "type = skirt and color = red," to find records of red skirts owned by the user. When multiple records are retrieved, the server can sort them by weights such as recent usage time and user historical selection frequency to select the most suitable clothing record as the benchmark. The server simultaneously searches for other clothing categories belonging to the user, such as retrieving records of "type: top or coat" as a candidate set of tops, and records of "type: footwear" as a candidate set of shoes. The server can use pre-calculated image features or style tags to filter the candidate set, prioritizing record groups that match the target clothing in the style and occasion fields. This reduces irrelevant input to the generative AI model, thereby reducing model input length and inference burden, and improving overall response speed.
[0326] After obtaining the target outfit and the list of candidate outfits, the server invokes the prompt statement generation unit. The prompt statement generation unit first organizes the system role information, user objective, outfit list, matching scenario, and task requirements into natural language text based on a preset text template. For example, the server generates the following prompt statement: "You are a professional fashion stylist."
[0327] The user's goal: to wear a red midi dress to a formal occasion, a friend's wedding.
[0328] Users want you to recommend suitable tops and shoes from their existing clothing collection.
[0329] List of clothing owned by the user (retrieved from the database): 1. Skirt: Red midi skirt, brand A, suitable for weddings and dates.
[0330] 2. Top: White silk shirt, brand B, formal style.
[0331] 3. Top: Beige suit jacket, brand C, business formal style.
[0332] 4. Shoes: Black high heels, brand D, formal style.
[0333] 5. Shoes: Nude high heels, brand E, suitable for a wedding.
[0334] Task: Using only the clothing items listed above, recommend at least two outfits suitable for attending a friend's wedding. Each outfit includes a top and shoes, and you should explain the reasoning behind each combination. Please answer in Simplified Chinese. When constructing prompts, the server precisely limits candidate clothing items to the user's existing clothing set and constrains the output space of the generative AI model through explicit task descriptions. This ensures that the model essentially performs combinatorial search and language expression on a finite, structured subset of data, rather than generating results arbitrarily in an open space. Through this prompt structure, the server transforms the unstructured natural language task into a form of "clothing record set + scene constraints + output format constraints," which facilitates the generation of executable and easily parsed results. This approach differs from traditional rule engines; it utilizes the sequence modeling capabilities of generative AI models to search for reasonable solutions in the combinatorial space, while effectively narrowing the search scope through the design of the prompts. This ensures diversity and creativity while improving consistency between the generated results and the database data.
[0335] The server then invokes the model invocation unit to send the aforementioned prompt to the generative AI model via a communication device. The generative AI model can be deployed on a local inference server or an external cloud inference service. The model architecture can employ a multi-layer Transformer neural network architecture, consisting of word embedding layers, positional encoding, several self-attention layers, and feedforward neural network layers. During the training phase, the model is pre-trained using large-scale text and dialogue corpora. It utilizes a self-supervised learning objective (e.g., masked language modeling or autoregressive language modeling), employs cross-entropy as the loss function, and uses stochastic gradient descent optimization algorithms or their improved learning algorithms to iteratively update the weights. The model can be further fine-tuned using clothing matching-related corpora to enhance its understanding of clothing terminology and scene semantics.
[0336] During the inference phase, the generative AI model receives prompts and first converts each Chinese character or word into an embedding vector. Then, it calculates the association weights between different words using a multi-head self-attention mechanism, thereby establishing semantic relationships between entities such as "red," "mid-length skirt," "wedding," "white silk shirt," and "high heels." The model generates output text through progressive decoding. During the decoding process, temperature parameters and top-k or top-p sampling strategies can be set to control the diversity and stability of the results. The server can adjust these parameters according to the clothing matching scenario, such as lowering the temperature to reduce unstable output, thereby improving the consistency and reliability of the recommendations.
[0337] The natural language response information output by the generative AI model is returned to the server. The server calls the result structuring and output unit to parse the response text. The server can segment the continuous text into multiple matching scheme blocks using predefined format keywords (e.g., "matching scheme one", "top:", "shoes:", "reason:") or through regular expressions and semantic segmentation algorithms. The server analyzes the clothing name phrases in each scheme block and performs fuzzy matching or tag matching operations based on name, color, and type in the clothing information dataset. For example, by calculating name string similarity and type consistency, the server selects the most likely corresponding clothing record. The server adds the clothing identifier field and image reference field of these matched clothing records to the structured matching result and retains the reason text output by the model. This post-processing process establishes a one-to-one correspondence between the generative text and the structured database records, enabling the terminal to accurately match text suggestions with specific images during display, thereby achieving the unification of "semantics-data-image" in human-computer interaction.
[0338] After receiving the structured outfit matching results data from the server, the terminal uses its image loading module to retrieve the corresponding clothing images from the server or content delivery network based on the image reference information. The terminal uses a graphical interface component to display each outfit in a combination of text and images: "Target Clothing + Matching Top + Matching Shoes." Users can view details of different outfits by swiping or clicking. Users can also provide feedback on each outfit on the terminal, such as clicking "Like," "Dislike," or "Change Top." The terminal encodes these interactions as preference feedback data and sends it back to the server.
[0339] After receiving preference feedback, the server updates the preference information dataset. The server can use counting or incremental updates to assign weights to user preferences such as color, clothing type, and style. The server can further train or update a simple preference model, such as a logistic regression model or a small recommendation model based on matrix factorization, to score clothing records based on the user's historical selections, generating a preference vector. When constructing subsequent prompts, the server incorporates this preference information into the prompts in natural language, for example, by adding the following text: "Additional information: Based on user history, users prefer bright colors and formal style tops, and do not like all-black outfits." In this way, when the generative AI model sees explicit preference descriptions in the prompts, it implicitly adjusts the strength of relevant features during internal attention calculations, making it more likely to output combinations that match the user's preferences. This approach of guiding model generation through "prompts + preference semantics," compared to the traditional method of only filtering rules on the output, incorporates preference constraints early in the model's decoding process. This reduces the number of invalid solutions generated, lowers the burden of subsequent filtering, and improves overall computational efficiency.
[0340] The server further reduces communication load and computation time through multi-level caching and data compression in its system design. For example, the server can build a cache structure for a user's frequently used clothing list. When a user repeatedly requests similar outfit combinations within a short period, the server can directly read the already organized clothing list text from the cache when generating the suggestion message, without having to access the database again and reorganize the list. The server can also compress image files into multi-resolution versions, prioritizing the delivery of lower-resolution images when network conditions are poor, to ensure that end-to-end latency is within an acceptable level.
[0341] The technical advantage of this invention lies in the fact that the server does not simply pass on user questions to the generative artificial intelligence model, but instead introduces a specific data structure and processing flow: first, the user's clothing data is structured and indexed; then, the natural language questions are semantically parsed and mapped to database search conditions; and finally, a prompt generation unit reconstructs the input space of the generative artificial intelligence model at the text level, transforming the open-ended question into a question of "generating combinations within a specific clothing set." Through this unconventional pre-processing and post-processing flow, the server significantly improves the controllability, consistency, and computational efficiency of the generative artificial intelligence model in clothing recommendation tasks without changing the basic neural network structure, achieving a substantial improvement over traditional computer recommendation technology.
[0342] In an alternative implementation, the server can employ a locally deployed generative AI model. In this case, the server uses a deep learning inference framework to load pre-trained model weights onto a local machine equipped with a graphics processing unit. The server can quantize or prune the model to reduce memory usage and inference time, making it particularly suitable for scenarios handling a large number of user requests simultaneously. The server can also continuously learn or fine-tune the model, for example, by using user feedback to build small incremental training sets and making minor updates to the model's high-level parameters, thereby gradually improving the fit of the matching suggestions.
[0343] In another implementation, the server can add a rule-based security filtering module to filter and truncate inappropriate content or output unrelated to clothing matching in the text before sending the prompt statement to the generative artificial intelligence model and before sending the model's response information to the terminal, so as to ensure the stability and controllability of the system in an open generative environment.
[0344] Through the above-described embodiments, the server, terminal, and user work closely together in terms of data structure, processing flow, and model invocation method. This makes the present invention not merely a simple automated reproduction of human collaboration, but rather the construction of an optimized data flow path between the generative artificial intelligence model and the structured database. Through prompt statement design, preference and emotion embedding, structured result matching, and caching mechanisms, it achieves multiple computer technology effects such as improved processing accuracy, faster response speed, reduced communication load, and optimized data management.
[0345] use Figure 13 The processing flow is explained.
[0346] Step 1: Users register clothing information via terminal. Users open the clothing registration interface on the terminal, take pictures of the clothing using the terminal's camera, and fill in attributes such as color, type, style, applicable season, and applicable occasion in the input boxes.
[0347] Inputs include: user operation instructions, raw data of captured clothing images, and various clothing attribute text entered by the user on the interface.
[0348] The terminal compresses and scales the resolution of the image data, standardizes the attribute text into internal encoding (such as color encoding and category enumeration), and encapsulates the image data and structured attributes into request data.
[0349] Output: Upload request data containing binary data of the clothing image and structured attribute fields.
[0350] Step 2: The terminal sends a clothing registration request to the server. The terminal sends the request data generated in step 1 to the server via a network protocol, based on the server's preset interface address.
[0351] Input: Request data containing user identifier, clothing image data, and clothing attribute data.
[0352] The terminal converts the requested data into a network transmission format (such as an HTTP request body) and initiates encrypted transmission through the communication module.
[0353] Output: A clothing registration request message sent to the server over the network.
[0354] Step 3: The server parses and stores the clothing data. The server receives a registration request from the terminal, parses the request body, and separates the clothing image data and attribute fields.
[0355] Input: Clothing registration request message.
[0356] The server performs data validation (checking required fields and data format), maps clothing attributes to database fields, and generates insert statements; at the same time, it writes image data to file storage or object storage, generates image reference paths, and writes them to the clothing record.
[0357] Output: One or more clothing records stored in the clothing information dataset, along with the corresponding image files and image reference information.
[0358] Step 4: Users input natural language collocation questions via the terminal. Users enter a natural language question on the device's outfit matching query interface, such as "I want to wear that red midi dress to my friend's wedding, please recommend a top and shoes for me," and then click the submit button.
[0359] Input: The user's natural language question text and submission instructions.
[0360] The terminal combines the text with the user identifier, encapsulates it into a structured request body, and prepares to send it to the server.
[0361] Output: A query request containing the user identifier and natural language question data.
[0362] Step 5: The terminal sends a matching query request to the server. The terminal uses the communication module to send the query request to the interface address specified by the server.
[0363] Input: Matching query request body.
[0364] The terminal encodes text and metadata into network packets and sends network requests.
[0365] Output: The matching query request message transmitted to the server over the network.
[0366] Step 6: The server parses natural language questions. The server receives a matching query request and reads the user identifier and the natural language query text.
[0367] Input: Matching query request message.
[0368] The server calls the natural language processing module to perform word segmentation, part-of-speech tagging, named entity recognition, and intent classification on the query text, extracting color words (such as "red"), clothing words (such as "mid-length skirt", "top", "shoes"), scene words (such as "wedding"), etc., and generating a structured semantic representation.
[0369] Output: Semantic parsing results data containing target clothing information, matching scenario information, and matching requirement type information.
[0370] Step 7: The server retrieves the user's clothing database. The server retrieves records that match the target criteria from the clothing information dataset based on the user identifier and semantic parsing results.
[0371] Input: User identifier and semantic parsing results containing target clothing conditions and demand categories.
[0372] The server generates database query conditions (e.g., type = skirt and color = red), retrieves the target clothing, and retrieves all records that can be used as candidate tops and shoes; the server can sort the results by time, occasion matching degree, and style similarity.
[0373] Output: The target clothing record set (e.g., a red midi skirt), the candidate top record set, and the candidate footwear record set.
[0374] Step 8: The server organizes the clothing text list. The server converts the clothing records obtained in step 7 into a list of readable text descriptions to be inserted into the prompt statement.
[0375] Input: The target clothing record set and the candidate clothing record set.
[0376] The server extracts fields such as type, color, brief description, occasion, and style from each clothing record, and generates multi-line text descriptions in numerical order, such as "1. Skirt: Red midi skirt, suitable for weddings and dates."
[0377] Output: A list of clothing description text used for the prompt statement.
[0378] Step 9: The server combines preference and emotion information The server reads the user's historical preference records from the preference information dataset and can obtain the current sentiment information from the sentiment analysis module (if available).
[0379] Inputs: User ID, user preference information data, and (optional) sentiment information data.
[0380] The server generates a brief text description based on statistically obtained preferences (such as preference for bright colors or formal styles), and organizes this information into "additional constraints" to influence the output of subsequent generative artificial intelligence models.
[0381] Output: Natural language description text representing user preferences and sentiments.
[0382] Step 10: The server generates a prompt statement. The server calls the prompt statement generation unit to combine system roles, user goals, clothing lists, matching scenarios, and preference information into a complete prompt statement.
[0383] Input: Clothing list description text, target clothing information, matching scenario information, matching requirement type information, and preference and emotion description text.
[0384] The server performs string concatenation and placeholder replacement according to a predefined template to form structured natural language, which includes parts such as "You are a professional fashion stylist", "The user's list of clothing", and "Task: Use only the clothing in the above list".
[0385] Output: The complete text of the prompt statement that served as input to the generative artificial intelligence model.
[0386] Step 11: The server invokes a generative artificial intelligence model to generate a response. The server sends prompts to the generative artificial intelligence model via a communication device and sets decoding parameters (such as temperature and maximum generation length).
[0387] Input: Prompt text and model call parameters.
[0388] The generative artificial intelligence model internally encodes prompts into vector representations based on a Transformer structure, and performs inference through multi-layer self-attention and feedforward networks, using an autoregressive approach to gradually generate natural language response text; the server receives the text results returned by the model.
[0389] Output: Natural language response information containing one or more matching schemes and corresponding reasons.
[0390] Step 12: The server parses the model response and structures the results. The server parses the natural language responses output by the model to identify each outfit combination and the names and reasons for the clothing in each combination.
[0391] Input: Natural language response text.
[0392] The server uses key tags (such as "Outfit Option 1", "Top:", "Shoes:") or regular expression parsing to split the text into multiple scheme units, extract the clothing name field and the reason field, and generate intermediate structured data.
[0393] Output: Intermediate structured outfit data containing multiple outfit combination entries (each containing an outfit name and a reason).
[0394] Step 13: The server maps the clothing names in the response to database records. The server uses intermediate structured matching data to match the clothing names in each scheme with the centralized records of clothing information data.
[0395] Input: Intermediate structured matching data and clothing information dataset.
[0396] For each garment name, the server performs string similarity calculations and attribute filtering (such as matching by type and color), selects the most matching garment record, and reads its garment identifier and image reference path; the server writes these fields into the matching scheme entry.
[0397] Output: Final structured matching result data containing clothing identification information, image reference information, and reasoning text.
[0398] Step 14: The server sends the structured matching results to the terminal. The server encapsulates the structured matching results data into a response message and sends it to the terminal via a communication device.
[0399] Input: Final structured matching result data.
[0400] The server may compress or trim the results (such as limiting the number of schemes) before generating a network response message.
[0401] Output: The matching result response message transmitted to the terminal via the network.
[0402] Step 15: The terminal receives and renders the matching results. The terminal receives the server's response and parses the structured matching result data.
[0403] Input: Matching result response message.
[0404] The terminal loads the corresponding clothing image based on the image reference information and uses a graphical interface component to display each outfit in the form of a card. The card contains an image of the target clothing, an image of the matching top, an image of the matching shoes, and a text explanation. The terminal renders these elements on the screen according to a preset layout.
[0405] Output: A visual matching scheme interface presented to the user on the terminal display screen.
[0406] Step 16: Users view and provide feedback on outfit suggestions Users can browse various outfit combinations on the terminal. Users can select a particular combination or rate the combination by clicking on controls such as "like", "dislike" or "change to another combination".
[0407] Input: The interface displayed on the terminal and the user's touch or click operations.
[0408] The terminal converts the user's selection behavior into structured feedback data (including scheme identifier, operation type, and timestamp) and prepares to upload it to the server.
[0409] Output: Feedback request data containing user feedback information.
[0410] Step 17: The terminal sends user feedback to the server. The terminal sends the feedback request to the server via the network.
[0411] Input: Feedback request data.
[0412] The terminal encapsulates data into a network packet and sends a request.
[0413] Output: User feedback messages arriving at the server.
[0414] Step 18: Server updates user preference information The server parses the user feedback message and extracts the scheme identifier, the selected clothing identifier, and the operation type.
[0415] Input: User feedback messages and current preference information dataset.
[0416] The server adds or removes preference weights for corresponding clothing records and related features (color, type, style), and writes the new weight values or statistical counts back to the preference information dataset; the server may optionally update the preference model parameters.
[0417] Output: Updated user preference information data and (optionally) new preference model state, providing more refined personalized constraints for subsequent prompt generation.
[0418] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0419] Existing clothing recommendation systems mostly use simple rule matching or a single recommendation algorithm, typically filtering based solely on static user preferences or historical purchase records, which presents the following technical problems: (1) At the information acquisition level, servers often process multi-source heterogeneous data such as weather information, user location, user emotional state, self-owned wardrobe data, and online product data in a scattered manner. The lack of a unified data structure and processing flow makes it impossible for the recommendation engine to comprehensively consider environmental factors and user state in a single calculation, thereby limiting the relevance and real-time performance of the recommendation results.
[0420] (2) At the model interaction level, although existing technologies can call generative artificial intelligence models, they usually only use the user's simple text input as prompts. They do not uniformly arrange and constrain the multi-dimensional structured data such as weather information, emotional state, user profile, own clothing and candidate products. This makes it difficult for generative artificial intelligence models to understand the relationship between various information elements, and the output results are unstable and difficult to control, making it difficult to form predictable and reusable recommendation logic.
[0421] (3) At the feedback learning level, although traditional recommendation systems can record user clicks or purchase behavior, they rarely model the relationship between “emotional state × clothing attributes (color, style, etc.) × adoption result” as a calculable acceptance index. The server cannot dynamically adjust the weight of clothing attributes under different emotions in the time dimension, resulting in insufficient adaptive ability of the system to changes in user state.
[0422] (4) At the system architecture level, there is a lack of a complete technical solution for computer implementation that organically integrates multi-source data acquisition, sentiment analysis, user profile generation, automatic construction of generative artificial intelligence model prompts, structured analysis of generated results, and dynamic updates of recommendation logic based on feedback data into a unified server-side processing flow. This results in low efficiency of computing resource utilization, redundant interface calls, and difficulties in system expansion and maintenance.
[0423] Therefore, a new computer implementation method and system are needed to enable servers to: Unified collection and structured multi-source data at the acquisition level; At the reasoning level, it efficiently drives generative artificial intelligence models by automatically constructing constrained prompts. At the learning level, the recommendation logic is dynamically updated by combining user emotional and behavioral feedback; This not only enhances the user experience but also substantially improves the computer technology performance of the recommendation engine in areas such as multi-source data fusion, model controllability, and adaptive learning.
[0424] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0425] In this invention, the server includes a processing module for acquiring user location information and corresponding meteorological information from an external information providing device, and determining suitable clothing combinations for the user accordingly; a processing module for collecting and structurally storing user's own clothing attributes and image information through an information terminal; a processing module for reading from an information storage device and generating user preference parameters and user profiles through statistical processing; a processing module for acquiring user voice, images, or text, invoking an emotion analysis processing device to infer emotional state, and storing it in association with time and context; a processing module for acquiring clothing product information from online information sources and forming a candidate product set based on preference parameters and meteorological information; and a processing module for automatically generating prompt statements for use by a generative artificial intelligence model based on meteorological information, user profile, own clothing information, candidate product set, and emotional state, parsing and structuring the model output, combining local and online clothing data to create specific combinations, outputting the results to the user terminal in a virtual or augmented reality manner, receiving preference and purchase feedback, and then dynamically updating the statistical processing and generative artificial intelligence model recommendation logic based on the feedback data. This enables integrated structured processing of multi-source heterogeneous data on the server side, controllable prompting and driving of generative artificial intelligence models, and adaptive recommendation updates based on emotion and behavioral feedback. From a computer technology perspective, this improves the accuracy, stability, and real-time performance of recommendation decisions, reduces redundant interface calls and invalid calculations, and optimizes the overall system resource utilization efficiency.
[0426] "User location information" refers to spatial data obtained by the positioning function of an information terminal to represent the user's current geographical location, including but not limited to longitude, latitude, and city or region identifiers derived therefrom.
[0427] "Meteorological information" refers to a dataset returned by an external information providing device that represents the environmental state of a specific geographical location, including but not limited to temperature, weather conditions, precipitation, wind speed, humidity, and their timestamps.
[0428] "External information providing device" refers to a device or service system that operates outside this system and provides environmental or commodity data to the server through a communication network, including but not limited to meteorological information service platforms and online commodity information platforms.
[0429] "Information terminal" refers to an electronic device operated by a user and interacting with a server, including but not limited to smartphones, tablet computing devices, wearable devices, desktop computing devices, and the applications running on them.
[0430] "Clothing attribute information" refers to structured data used to describe the characteristics of clothing items, including but not limited to fields such as clothing category, color, size, material, style tag, brand, applicable season, and price range.
[0431] "Image information" refers to visual data related to clothing or the user, acquired by an imaging device and stored in digital form, including but not limited to still image data and image data used to generate virtual three-dimensional images.
[0432] "Information storage device" refers to storage resources used to persistently store the data required for the operation of this system, including but not limited to database systems, file storage systems, and cache storage components that work together with them.
[0433] "Historical purchase behavior information" refers to data records related to product purchase events completed by users in the past, including but not limited to purchase time, product identifier, price, quantity, and corresponding clothing attribute information.
[0434] "Browsing behavior information" refers to data records related to a user's viewing behavior of goods or content in online information sources or application interfaces, including but not limited to the identifier of the viewed object, browsing time, dwell time, and type of interaction.
[0435] "Statistical processing device" refers to software or hardware resources that perform data aggregation, counting, weighting, normalization and other operations in a server to generate statistical results, and are used to calculate user preference parameters and other statistical characteristics based on historical data.
[0436] "Preference parameters" refer to numerical or symbolic data calculated by statistical processing devices based on historical purchase and browsing behavior information, used to quantify the degree of user preference in terms of color, category, style, price range, etc.
[0437] "User profile" refers to a set of information stored in the form of a data structure that comprehensively reflects user preference parameters, usage habits, historical behavior and other relevant characteristics, and is used to characterize user features in recommendation processing and prompt generation.
[0438] "Voice information" refers to acoustic data containing user voice content acquired through audio acquisition devices, used to identify text content and infer user emotional state.
[0439] "Image information (for sentiment analysis)" refers to image data used to identify a user's facial expressions or body postures in order to infer the user's emotional state and its intensity.
[0440] “Text information” refers to text data input by the user or transcribed by a speech recognition component, used to represent the user’s request content, subjective description, or feedback information.
[0441] "Emotional analysis processing device" refers to a processing component used to analyze voice information, image information, or text information to infer the user's emotional state and emotional intensity, including rule-based algorithm modules and machine learning or deep learning-based models.
[0442] "Emotional state" refers to an identifier that represents a user's psychological state or emotional type at a specific moment, inferred by an emotion analysis and processing device, including but not limited to joy, calmness, frustration, tension, excitement, and combinations thereof.
[0443] "Emotional intensity" refers to a numerical or level indicator used to quantify the degree of emotional state, representing the strength of a user's current emotions.
[0444] "Online information source" refers to a service platform accessible through a communication network that provides information or related content about clothing products, including but not limited to e-commerce platforms, product catalog services, and product information interfaces.
[0445] The "preliminary candidate set" refers to the first-stage candidate clothing set selected from clothing products obtained from online information sources based on user preference parameters and meteorological information, which is used for subsequent clothing recommendation processing.
[0446] A "candidate group" refers to a set of candidate clothing data that is formed from a preliminary candidate set or further processed on the basis of it, and can be referenced by generative artificial intelligence models or rule logic in the clothing recommendation process.
[0447] "Generative AI models" refer to AI models that can automatically generate natural language text or other data outputs based on input prompts, including but not limited to generative models based on machine learning or deep learning.
[0448] "Prompt statements" refer to texts automatically generated by the server based on meteorological information, user profiles, personal clothing information, candidate groups, and emotional states. These texts describe the input conditions and constraints in natural language and are used as input for generative artificial intelligence models.
[0449] "Natural language output" refers to the clothing matching scheme, explanation of reasons, and description of usage scenarios generated by the generative artificial intelligence model in the form of natural language text after receiving prompts.
[0450] "Structured results" refer to data structures that can be processed by programs, formed by parsing natural language output and mapping elements such as clothing categories, colors, styles, and scenes to predefined data fields.
[0451] "Specific clothing item combination" refers to a specific matching scheme consisting of multiple clothing items, determined based on structured results and with reference to the user's own clothing information and online product information, including the identification and attributes of each clothing item.
[0452] "Virtual 3D image information" refers to 3D visual data used to present the effect of clothing wearing in virtual or augmented reality space, including but not limited to 3D model data, posture data, and rendering parameters.
[0453] "Preference information" refers to subjective feedback data from users regarding the recommendation results, including but not limited to information such as liking, disliking, rating, saving, and whether to use.
[0454] "Purchase Result Information" refers to data records related to a user's actual purchase behavior for recommended or other products, including but not limited to whether a purchase was made, the quantity purchased, the purchase time, and the corresponding product identifier.
[0455] "Learning data" refers to a set of training data consisting of the correspondence between preference information, purchase result information, and the preference parameters and emotional state at that time, which is used to update the recommendation logic of statistical processing devices or generative artificial intelligence models.
[0456] The "acceptance index" is a quantitative evaluation value calculated based on learning data to represent the degree to which users accept relevant clothing recommendations under a specific emotional state and a specific combination of color or style attributes.
[0457] "Recommendation logic" refers to the set of decision rules and calculation processes used by the server when performing clothing recommendations, including screening rules based on statistical processing, generation strategies based on generative artificial intelligence models, and combinations of the two.
[0458] The embodiments of this invention will be described focusing on the collaborative processing of the server, terminal, and user. The following embodiments are merely examples, and those skilled in the art can make various modifications and variations without departing from the spirit of this invention.
[0459] I. Overall System Composition The server is implemented using a computing device with a multi-core central processing unit and a large-capacity random access memory. The server runs an operating system, network communication middleware, and an application server framework. In one embodiment, the server uses a backend framework based on an interpreted language (e.g., a general-purpose scripting language and its web framework) and a relational database management system (e.g., a general-purpose relational database or a lightweight file-based database) as its information storage device. In another embodiment, the server further deploys a deep learning framework (e.g., a neural network framework supporting tensor computation) to implement generative artificial intelligence models and neural networks related to sentiment analysis.
[0460] The terminal is implemented using a mobile computing device equipped with a display, touch input, camera, microphone, and positioning module, such as a smartphone, tablet, or head-mounted display. In one embodiment, the terminal runs an application built on a cross-platform mobile application framework, which interacts with the server via a secure communication protocol.
[0461] Users perform operations through the terminal, inputting location information for authorization, clothing data, text or voice requests, preference feedback, etc. The terminal sends this data to the server, and the server performs various data processing and calculations based on the data and data returned by the external information providing device.
[0462] II. Data Structure and Storage Format The server defines various data tables or equivalent data structures for this invention in the database.
[0463] In one embodiment, the server stores user location information as a record row containing user identifier, longitude, latitude, city name, and timestamp. After receiving data returned from an external meteorological service, the server writes fields such as temperature, weather code, humidity, wind speed, and weather description into a meteorological information table and associates it with the user location record using a foreign key field.
[0464] In one embodiment, the server stores clothing attribute information in a "user clothing table," which includes at least: clothing identifier, user identifier, category field (tops, trousers, skirts, coats, etc.), color field (which can be discrete color labels or multi-dimensional color vectors), style field (casual, business, sports, etc.), size field, material field, applicable season field, and image address field. In some implementations, the server extracts image features into vectors through an image processing module and caches them in auxiliary fields for subsequent fast matching.
[0465] In one embodiment, the server records user history using a "purchase behavior table" and a "browsing behavior table." Each record includes attributes such as user ID, product ID, behavior type (purchase or browsing), timestamp, price, category, and color. The server reads these tables using a statistical processing device to calculate preference parameters and user profiles.
[0466] In one embodiment, the server uses an "emotion log table" to record emotional states. This table includes a user identifier, emotion tags (such as joy, sad, neutral, etc.), emotion intensity value, source type (text, voice, image), context identifier (such as the corresponding matching request ID), and timestamp.
[0467] In one embodiment, the server uses a "candidate product table" to record products obtained from online information sources. Fields include product identifier, source platform identifier, category, color, style, price range, inventory status, evaluation indicators, and image address.
[0468] III. Server Data Processing and Calculation Methods 1. The Relationship Between Meteorological Information and Clothing Conditions After receiving the location information sent by the terminal, the server calls the application programming interface (API) of the external meteorological information provider to receive a structured response representing meteorological information for the current period or a predetermined time period. The server uses a parsing module to convert the external response into an internally unified format, including temperature (degrees Celsius), weather type enumeration, precipitation status, wind speed, etc.
[0469] In one embodiment, the server associates temperature ranges with clothing thickness indices, for example, pre-labeling each garment with a "warmth index." By comparing the temperature with this index range, the server can quickly filter out clothing unsuitable for the current temperature, thereby reducing the subsequent recommendation search space and lowering computational complexity. This pre-screening process is an algorithmic filtering operation based on numerical thresholds, preventing generative AI models from wasting inference resources on a large number of obviously unsuitable clothing combinations, thus improving overall processing speed.
[0470] 2. User Preference Parameters and User Profile Construction In one embodiment, the server uses a statistical processing module to read historical purchase and browsing behavior information. The server generates preference vectors through aggregation operations (e.g., grouping and counting by category, color, and style fields, and normalization). In one specific implementation, the server maps the color field to a fixed-length vector space, counts the frequency of user purchases and browsing time for each color, calculates the weights, and normalizes them into a "color preference vector"; similarly, categories, styles, etc., are mapped to corresponding "category preference vectors" and "style preference vectors".
[0471] In one embodiment, the server concatenates multiple-dimensional preference vectors into a high-dimensional user feature vector, denoted as a user profile, and stores it in a user profile table. All subsequent recommendation calculations and suggestion generation use this high-dimensional vector as one of the input features. Through this structured vector format, the server can directly input the user profile into the neural network model, thereby improving the differentiability and trainability of the recommendation calculation.
[0472] 3. Sentiment Analysis and Construction of Sentiment Features In one embodiment, the server invokes a natural language sentiment analysis service for user text information, an facial expression recognition service for user image information, and an acoustic sentiment recognition service for user voice information. The server merges the sentiment scores returned from each channel, for example, by using a weighted average or a confidence-weighted voting method, to obtain a unified sentiment label and scalar intensity.
[0473] In one embodiment, the server encodes the sentiment label as a one-hot vector, uses the sentiment intensity as a scaling factor, and performs linear scaling on the vector to generate a sentiment feature vector. When generating prompts and inputting them into a generative AI model, the server can either convert this feature into a natural language description (e.g., "the user is feeling down"), or, in an implementation using an end-to-end neural network, use this feature vector as an additional input layer to the model.
[0474] 4. Candidate Product Set and Multi-level Filtering In one embodiment, the server first filters a preliminary candidate set from products obtained from online information sources based on temperature, weather, and user profile. During this stage, the server can perform multi-level filtering: The server first filters out clothing with obviously mismatched materials based on the temperature range (for example, filtering out clothing with a high thermal insulation index at high temperatures). The server then calculates the similarity between the candidate product colors and the user's preferences based on the user's color preference vector (e.g., using cosine similarity), and only retains products with similarity exceeding a threshold. The server finally filters based on price range and style tags.
[0475] The multi-level screening process reduces the size of the candidate group, making the output space of the subsequent generative artificial intelligence model more concentrated, reducing model output noise, and improving overall recommendation accuracy.
[0476] IV. Generative Artificial Intelligence Models and Prompt Statement Construction 1. Structure of Generative Artificial Intelligence Models In one embodiment, the server uses a generative artificial intelligence model based on a converter architecture. This model includes a multi-layer self-attention encoder and a multi-layer self-attention decoder. A word embedding layer maps words in the prompt to vector representations, a position encoder incorporates sequence information, and the output is an autoregressive language generation head. In another embodiment, the server may employ a large language model with a unidirectional decoder architecture, omitting the encoder module.
[0477] During the training phase, the server uses a corpus containing extensive clothing descriptions, matching suggestions, and usage scenarios, and optimizes its parameters using the cross-entropy loss function. When updating the recommendation logic, the server continuously trains or fine-tunes using newly collected learning data. The server can introduce sentiment tags or weather tags as conditional labels during training, enabling the model to learn to output sentences with different styles under different conditions.
[0478] In one embodiment, the server sets generation parameters, such as temperature coefficients and top-k or top-p truncation strategies, to balance diversity and stability. The server strictly controls the generation length and constraints to prevent model output from containing content unrelated to the prompt constraints.
[0479] 2. Automatic construction of prompt statements In one embodiment, the server constructs prompt statements using a template engine. The server extracts key fields from weather information, user profiles, personal clothing information, candidate groups, and emotional states, and populates them into a preset natural language template.
[0480] In one embodiment, the server uses the following example prompt statement: "The current weather in the city is cloudy with a temperature of around 18 degrees Celsius. The user is feeling a bit down today and wants to dress casually and modestly. The user's wardrobe includes: a black trench coat, a gray knit sweater, dark blue jeans, and white sneakers. The user prefers a casual and minimalist style. Based on this information, please recommend an outfit suitable for daily commuting, using a gentle tone, and briefly explain the reasoning behind pairing each item." In another embodiment, the server uses a prompt statement for scenarios involving the purchase of new goods: "The user is considering buying a red casual jacket. The user already has a white shirt, black trousers, and blue jeans in their wardrobe. The current weather is sunny with a temperature of approximately 20 degrees Celsius. Based on this information, please recommend several outfits that include the red jacket, indicating which combinations are best for commuting and which are suitable for weekend outings, and using concise language to encourage the user to try new color combinations." By explicitly listing temperature, weather, mood, personal clothing list, and user preferences in the prompts, the server enables generative AI models to comprehensively consider multi-source information within a unified context. Compared to inputting only short text, these structured prompts significantly improve the controllability and relevance of the generated results and reduce output bias.
[0481] V. Result Parsing and Structured Matching After receiving the natural language text output by the generative artificial intelligence model, the server uses a natural language parsing module to perform syntactic and semantic analysis on the text. In one embodiment, the server uses a combination of keywords and rules to identify color words, clothing category words, and style modifiers from the output and maps them to standard attribute values defined in the database.
[0482] In one embodiment, the server converts the parsed results into a structured result object containing multiple clothing entries, each labeled with category, color, style, and whether it recommends purchasing new items. The server then matches these entries against records in the user's clothing table and the candidate product table. The server uses a similarity metric during matching; for example, if the output mentions "dark blue coat" while the user's wardrobe record shows "navy blue jacket," the server can determine their equivalence by comparing the similarity between the color vector and the category label.
[0483] This conversion from natural language to structured data allows the recommendation results to be seamlessly integrated with subsequent interface display, 3D rendering and other specific technical aspects, thus avoiding the problem that traditional "text recommendation" cannot directly drive graphics and interactive modules, and improving the overall automation and consistency of the system.
[0484] VI. Technical Implementation of Terminal-Side Visualization and Device Control In one embodiment, after receiving structured matching information and image or 3D model information returned by the server, the terminal loads it into the graphics rendering module. In a normal 2D interface, the terminal displays clothing combinations in a list and thumbnail format; in virtual reality or augmented reality mode, the 3D clothing model is overlaid onto the real-time image captured by the user.
[0485] In AR mode, the terminal uses a camera to capture images of the user's body, calls a local human pose estimation algorithm or a server-side pose recognition service, and performs geometric transformations and scaling on the clothing model according to the user's pose to align the virtual clothing with the user's body contours. Through this type of image processing and 3D rendering technology, this invention not only provides abstract suggestions but also directly presents the technical processing results visually in the real environment. This represents concrete control over display devices and sensors, avoiding falling into the realm of purely abstract algorithms.
[0486] VII. Feedback Learning and Recommendation Logic Adaptive Update In one embodiment, the server obtains user preference information and purchase result information for each recommendation, and appends this information as learning data to the learning data table. The server then periodically performs offline statistical and model update tasks. The server first calculates an acceptance index based on the learning data for all combinations of "emotional state × color attribute × style attribute". For example, the server counts the adoption rate and purchase rate of bright-colored casual style outfits when the emotional state is joyful, and compares it with other combinations to obtain numerical preferences.
[0487] The server then adjusts the candidate group filtering weights and prompt generation strategy based on the acceptability index. For example, when the mood is joyful, the server increases the ranking weight of bright-colored clothing in the candidate group and tends to describe bright-colored clothing more in the prompts; when the mood is depressed, the server appropriately reduces the weight of high-saturation colors and increases the proportion of neutral and low-saturation colors.
[0488] When needed, the server further fine-tunes the generative AI model using a deep learning framework. During training, sentiment labels and acceptance metrics are introduced as auxiliary supervision signals, making the model more inclined to output matching styles that have historically proven effective by acceptance metrics. Through this closed-loop adjustment, the recommendation engine of this invention automatically adapts to changes in user sentiment patterns and preferences over time, achieving higher long-term accuracy and stability than simple rules or traditional recommendation algorithms.
[0489] VIII. Technical Effects and Improvements in Computer Technology Through the combined processing of structured data modeling, multi-level filtering, prompt statement construction, result parsing, and feedback learning, the server achieves the following effects at the computer technology level: 1. The server performs temperature and preference-driven candidate filtering before prompting statements, which significantly reduces the combination space that generative artificial intelligence models need to consider, reduces model inference time and computing resource consumption, and achieves improved processing speed and reduced server load.
[0490] 2. By uniformly encoding weather information, emotional state, user profile, owned clothing, and online products into explicit conditions in the prompt statements, the server makes the input processed by the generative artificial intelligence model more structured, reduces output bias, and thus outperforms conventional solutions that rely solely on brief user input in terms of recommendation accuracy and stability.
[0491] 3. The server parses the generated output into structured results and matches them with database entities to form a data stream that can directly drive image rendering and AR display. This avoids secondary human understanding and manual operation, and realizes an automated technical process from natural language recommendation to terminal device control (display and rendering).
[0492] 4. The server constructs an acceptance index for sentiment and clothing attributes based on learning data, and uses this index to dynamically adjust the weights of filtering conditions and prompts, enabling the recommendation logic to adaptively update over time. This non-linear adjustment method based on multi-dimensional feedback differs from traditional static rules, achieving a comprehensive improvement in computational efficiency and accuracy under conditions of multi-source data and complex preferences.
[0493] IX. Optional Implementation Forms and Variations In one implementation, the server can use a single generative AI model to complete all prompt processing and recommendation text generation. In another implementation, the server can split the task into two models: a smaller, structured generative model specifically responsible for constructing intermediate templates, and a larger model specifically for language polishing and human-like expression. This reduces the number of times the large model is called while maintaining quality, further reducing communication and computational overhead.
[0494] In one embodiment, the server can directly use user profiles and sentiment features as vectors as additional inputs to a generative artificial intelligence model for joint training in a multimodal manner; in another embodiment, the server can use natural language to describe these features only in prompt statements. Both methods can achieve the technical effects of the present invention.
[0495] In one implementation, the terminal can completely delegate image rendering to the server. In another implementation, it can receive only the structured matching scheme and clothing image address from the server and complete the image combination or 3D rendering locally to adapt to different network bandwidth and terminal computing power conditions.
[0496] In one implementation, users can input their needs through text dialogue; in another, they can be directly recognized through voice and facial expressions; and in a further implementation, users can control the AR interface through gestures and body language. The processing flow of the server and terminal can be adjusted accordingly, but the core ideas of this invention regarding multi-source data fusion, prompt statements driving generative artificial intelligence models, and feedback learning to improve recommendation logic remain unchanged.
[0497] use Figure 14 The processing flow is explained.
[0498] Step 1: Users launch the application using a terminal and enter their requirements.
[0499] Input: Natural language requirements (text or voice) entered by the user on the terminal interface, scene selection (such as commuting, dating, exercising), and authorization status for whether to turn on the camera / microphone / location.
[0500] The specific actions of the terminal include: loading the application interface, initializing the camera, microphone, and GPS modules, receiving text entered by the user in the input box (such as "What should I wear today?"), or recording the user's voice and transcribing it into text locally, and packaging the scene options and text requirements into request data.
[0501] Output: The terminal generates a request data packet containing user identifier, text requirements, scene markers, and authorization status markers, and sends it to the server over the network.
[0502] Step 2: The terminal obtains location information and sends it to the server.
[0503] Input: Location service interface provided by the operating system, user authorization result.
[0504] The terminal's specific actions include: calling the system's location API to obtain longitude, latitude, and, optionally, the city name; if the user has not authorized location services, using the default location or prompting the user to provide additional city information. The terminal adds the location information to the aforementioned request data packet.
[0505] Output: The terminal generates a location information field containing user ID, longitude, latitude, and city name (if any), and sends it to the server along with the user's request.
[0506] Step 3: The server obtains and parses the meteorological information.
[0507] Input: Location information (longitude, latitude, city name) and time information from the terminal.
[0508] The server's specific actions include: constructing an HTTP request to an external meteorological information service, sending latitude, longitude, and timestamp as parameters; receiving JSON-formatted meteorological data (including temperature, weather description, humidity, wind speed, etc.); using a parsing module to extract fields such as main.temp and weather.main, converting the temperature from Kelvin to Celsius, and mapping the weather description to an internal enumeration (such as sunny, rainy, cloudy).
[0509] Data processing / calculation: The server performs unit conversion and interval judgment on numerical fields, divides the temperature into preset temperature ranges (such as "cold, moderate, hot"), and determines whether precipitation conditions exist based on the weather type.
[0510] Output: The server generates structured meteorological information objects (including temperature range, weather type, whether precipitation occurs, etc.), writes them into the meteorological record table, and caches them as input parameters for subsequent processing.
[0511] Step 4: The server retrieves and organizes the user's own clothing information.
[0512] Input: User ID, clothing records associated with the user in the database.
[0513] The server's specific actions include: executing a query statement in the user clothing table to read all clothing entries based on the user ID; extracting attributes such as category, color, style, warmth index, and applicable season for each record; and optionally, reading image feature vectors if they exist.
[0514] Data processing / calculation: The server performs preliminary screening of the clothing collection based on current weather information, such as filtering out heavy coats or overly thin clothing that are not suitable for the current temperature range, and categorizing the remaining clothing to form a "list of usable clothing".
[0515] Output: The server generates a list of available clothing (clothing lists grouped by category) and a corresponding attribute matrix, providing input for the construction and combination of subsequent prompt statements.
[0516] Step 5: The server acquires and analyzes users' historical behavior data to generate preference parameters.
[0517] Input: User ID, and corresponding historical purchase behavior records and browsing behavior records.
[0518] The server's specific actions include: executing aggregate queries (grouping and counting purchase and browsing times by category, color, style, and price range); constructing DataFrames using statistical processing modules or data analysis libraries; and calculating the frequency and weight of each attribute category.
[0519] Data processing / calculation: The server normalizes the statistical results, converts different attribute dimensions into vector forms, such as color preference vector, style preference vector, and category preference vector, and synthesizes high-dimensional user feature vectors (user profiles) by concatenation or weighting.
[0520] Output: The server generates a user profile object containing multiple dimensions of preference parameters and stores it in the user profile table. This object is also used as input for subsequent filtering and prompt generation.
[0521] Step 6: The terminal collects user emotion-related data and sends it to the server.
[0522] Input: User facial expression image, voice data, or text description (e.g., "I'm not in a good mood today").
[0523] The specific actions of the terminal include: activating the camera to capture facial images, or recording a voice message, or directly obtaining text input, based on user authorization; performing local or cloud-based recognition on the voice message to obtain text; and packaging the image data and text data into a sentiment analysis request after marking their source types and sending it to the server.
[0524] Output: The terminal generates a sentiment data request containing user ID, image or text content, and source type fields, and sends it to the server.
[0525] Step 7: The server performs sentiment analysis and generates sentiment feature vectors.
[0526] Input: Text content, facial images, or speech-to-text from the terminal.
[0527] The specific actions of the server include: The server calls the natural language sentiment analysis component on the text content to obtain sentiment labels (such as happy, sad, neutral) and sentiment scores; The server calls the facial expression recognition component to obtain facial expression labels and confidence scores from the facial images. The server merges the multi-source sentiment results according to preset weights to obtain a unified sentiment label and intensity value.
[0528] Data processing / calculation: The server encodes the sentiment tags into one-hot vectors, and then multiplies them by the sentiment intensity to obtain the sentiment feature vector; at the same time, the sentiment tags, along with the current time and scene markers, are written into the sentiment log table.
[0529] Output: The server generates a sentiment feature vector and a sentiment description text (such as "The user is feeling down today"), and provides these two results to the subsequent prompt statement generation module.
[0530] Step 8: The server retrieves online product information and generates candidate groups.
[0531] Input: User profile (preference parameters), current weather information, scene tags.
[0532] The server's specific actions include: selecting the product category to be queried based on the temperature range and scenario (e.g., spring / autumn jackets, commuter shirts); calling the online information source interface to obtain a list of products in that category; and parsing the returned data to extract product category, color, price, style tags, etc.
[0533] Data processing / calculation: The server calculates the similarity (such as cosine similarity) between each product and the user's preference vector, filtering out products below the threshold; it also filters out materials or thicknesses unsuitable for the weather by combining meteorological constraints; finally, a small preliminary candidate set is obtained, which can be further sorted by rating or sales volume to become a candidate group.
[0534] Output: The server generates a candidate group data structure (containing several products and their attributes) and stores it in memory or cache, providing input for the construction of prompt statements and matching the generated results.
[0535] Step 9: The server automatically generates the prompt message.
[0536] Inputs: meteorological information object, user profile object, available clothing list, candidate group, sentiment feature vector and sentiment description text, and user's original demand text.
[0537] The server's specific actions include: extracting key fields from the above input, providing natural language descriptions of temperature, weather, mood, preferred style, existing clothing summary, and whether new products need to be recommended; and concatenating these descriptions into a structured natural language segment according to a predefined template.
[0538] Data processing / calculation: The server executes string formatting and conditional branching logic. For example, if the emotional state is "low", the condition constraint "hope to dress casually and not too flashy" is added to the prompt statement; if there are new products in the candidate group, the prompt statement is added "consider recommending some new products".
[0539] Output: The server generates a prompt text containing the full context, for example: "The current weather in the city is cloudy with a temperature of around 18 degrees Celsius. The user is feeling a bit down today and wants to dress casually and modestly. The user's wardrobe includes: a black trench coat, a gray knit sweater, dark blue jeans, and white sneakers. The user prefers a casual and minimalist style. Based on this information, please recommend an outfit suitable for daily commuting, using a gentle tone, and briefly explain the reasoning behind pairing each item." The prompt statement is sent to the generative artificial intelligence model interface.
[0540] Step 10: The server invokes a generative artificial intelligence model to generate natural language recommendation results.
[0541] Input: Prompt text, model generation parameters (maximum length, temperature, top-k or top-p, etc.).
[0542] The server's specific actions include: encoding the prompt statement into the model input format, calling the generative artificial intelligence model deployed locally or remotely, setting the output length and sampling strategy, and receiving the natural language text returned by the model.
[0543] Data processing / calculation: The generative AI model performs multi-layer self-attention calculations internally, generating autoregressive words or sub-words based on prompts and trained parameters until a termination condition is met; the server does not change the internal weights of the model, but only controls the generation behavior according to the set parameters.
[0544] Output: The server receives a natural language output text containing specific clothing matching suggestions, reasons for color matching, and usage scenarios.
[0545] Step 11: The server parses the natural language output and transforms it into structured results.
[0546] Input: Natural language text output by a generative artificial intelligence model.
[0547] The server's specific actions include: using keyword recognition and rule matching to segment the text into sentences, identifying clothing category words (such as "jacket", "shirt", "trousers"), color words (such as "black", "red", "dark blue"), and modifiers (such as "casual" and "formal"); and mapping these words to an internal standard attribute table.
[0548] Data processing / calculation: The server constructs a structured record for each recommended item, with fields including category, color, style intent, whether it is recommended to purchase new items, etc., and groups multiple items in the same outfit into a "specific clothing item combination".
[0549] Output: The server receives a list of one or more structured outfit combinations, each consisting of multiple clothing items and their attributes, providing input for subsequent matching with database entities.
[0550] Step 12: The server will match structured outfits with the user's clothing and online products.
[0551] Input: a list of structured outfit combinations, user-owned clothing information, and candidate group product information.
[0552] The server's specific actions include: for each structured entry, searching for the closest clothing item in its own clothing table based on category and color; if the entry is marked "new purchase", then searching for matching products in the candidate group.
[0553] Data processing / calculation: The server uses attribute similarity calculations (such as color vector distance, category exact match, style tag overlap) to determine the best matching item; when multiple candidates meet the conditions, the server can sort them according to user preference parameters or price, etc., and select the priority item.
[0554] Output: The server generates a final list of "Specific Clothing Item Combinations", which includes the identifier, source (owned or online), display name, image URL, and optional purchase link for each garment.
[0555] Step 13: The server returns the recommendation results to the terminal.
[0556] Input: The final list of matching schemes, the generated natural language description text, and the corresponding image or 3D model address.
[0557] The server's specific actions include: packaging structured data and natural language descriptions into a response message, containing one or more matching schemes; and, to support virtual try-on or AR display, the server can include a clothing 3D model index or rendering parameters.
[0558] Data processing / computation: The server dynamically selects the data return format based on the terminal's capabilities (such as whether it supports AR), reducing unnecessary data transmission and thus lowering the communication load.
[0559] Output: The server sends the complete recommendation results (text + structured matching + media resource address) to the terminal via the network.
[0560] Step 14: The terminal displays the recommendation results and performs visualization rendering.
[0561] Input: A data packet of recommended results from the server, including pairing options, text descriptions, and images or 3D model addresses.
[0562] The specific actions of the terminal include: displaying natural language descriptions in the chat or results interface; loading its own clothing images and online product images, and arranging them according to the combination relationship specified by the server; and, in AR-supported mode, using the camera to capture the user's image and overlaying the 3D clothing model onto the user's body outline according to the posture and size parameters provided by the server.
[0563] Data processing / calculation: The terminal performs image scaling, cropping, geometric transformation of 3D models, and lighting rendering locally, improving rendering efficiency and image quality.
[0564] Output: The terminal presents one or more visual outfit combinations to the user and provides interactive controls (such as "Like", "Dislike", "Buy", and "Save Today's Outfit" buttons).
[0565] Step 15: Users can provide feedback on recommendations or make purchases.
[0566] Input: Control buttons on the terminal interface, user clicks or touch operations.
[0567] Users' specific actions include: browsing various outfit combinations, clicking "like" or "dislike," selecting an outfit for today, or clicking the "buy" link of an online product to be redirected to the e-commerce page to complete the order.
[0568] Output: The terminal packages the user's preference feedback (like / dislike, save or not) and purchase results (whether the user clicked and completed the payment) into feedback data and sends it to the server.
[0569] Step 16: The server records feedback and updates the recommendation logic.
[0570] Inputs: User feedback data (preference information, purchase result information), current recommendation scheme identifier, emotional state at the time, user profile, and matching attributes.
[0571] The server's specific actions include: writing each recommendation result and user feedback into a learning data table, associating sentiment tags, color attributes, style attributes with "adopted / not adopted" tags; reading accumulated learning data in periodic tasks, and calculating acceptance metrics under different sentiment states and attribute combinations.
[0572] Data processing / calculation: The server performs statistical analysis on the three dimensions of "emotional state × color × style" to calculate the adoption rate and purchase conversion rate of each combination; based on this, the server updates the preference parameters, filtering weights, and conditional weights in the prompt statement templates, such as increasing the recommendation weight of bright colors when the emotion is joy.
[0573] Output: The server generates an updated preference parameter weight table and recommendation strategy configuration, which are then referenced in subsequent recommendation processing to achieve dynamic adaptive optimization of the recommendation logic.
[0574] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0575] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0576] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0577] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0578] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0579] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0580] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0581] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0582] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0583] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0584] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0585] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0586] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0587] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0588] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0589] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0590] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0591] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0592] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0593] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0594] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0595] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0596] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0597] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0598] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0599] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0600] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0601] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0602] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0603] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0604] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0605] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0606] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0607] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0608] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0609] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0610] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0611] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0612] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0613] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0614] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0615] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0616] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0617] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0618] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0619] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0620] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0621] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0622] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0623] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0624] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0625] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0626] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0627] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0628] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0629] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0630] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0631] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0632] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0633] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0634] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0635] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0636] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0637] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0638] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0639] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0640] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0641] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0642] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0643] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0644] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0645] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0646] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0647] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0648] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0649] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0650] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0651] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0652] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0653] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0654] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0655] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0656] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0657] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0658] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0659] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0660] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0661] In addition, the following notes are provided in response to the above explanation.
[0662] Example 1 (Note 1) An information processing system, characterized in that it comprises: A unit for acquiring environmental information including current weather conditions from an external information providing device via a communication network based on location information from a user terminal, and generating basic proposal information related to a combination of clothing items based on the environmental information; A unit for reading user-owned item information and user preference information from a storage device based on user identification information, and extracting the owned item information and the preference information as clothing attribute information corresponding to the environmental information; A unit for generating prompt statements for inputting into a generative artificial intelligence model based on conditional information including the environmental information, the information on the owned items, and the preference information, and inputting the prompt statements into the generative artificial intelligence model to obtain clothing proposal information in natural language form; A unit for estimating the combination patterns of clothing items owned by the user based on the clothing proposal information, and generating matching proposal information containing the combination patterns; A unit for extracting and filtering attributes of product group information obtained from external services through a communication network based on the environmental information, the preference information, and the clothing proposal information, and outputting the filtered candidate product information in association with the matching proposal information. A unit for acquiring input information from the user and emotion-related information associated with the clothing proposal information, and dynamically changing the content of the prompts input to the generative artificial intelligence model based on the emotion-related information to regenerate the clothing proposal information; A unit for sending the pairing proposal information and the candidate product information to the user terminal and presenting them in an interactive manner on the display device of the user terminal.
[0663] (Note 2) The information processing system according to Appendix 1 is characterized in that, The system further includes a unit for obtaining the user's historical purchase record information and browsing record information from the storage device, extracting style preference information and price range preference information as the preference information through data parsing and processing, and performing filtering processing on the candidate product information in Note 1 based on the style preference information and the price range preference information.
[0664] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system further includes: a unit for acquiring natural language input information from the user through an input / output device, storing the natural language input information in a storage device, and controlling the constraints and output format of prompt statements input to the generative artificial intelligence model based on state information containing the natural language input information and the emotion-related information, thereby generating clothing proposal information and matching proposal information reflecting the state information.
[0665] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A processing means for running on an information processing device and acquiring location information from a user terminal, the processing means being configured to acquire meteorological information from an external information providing device via a communication network based on the location information, and to convert the meteorological information into structured data including temperature information and meteorological state information. A means for obtaining temperature information, meteorological status information, user attribute information, and information on the clothing owned by a user from the information management area in a storage device, and for integrating the temperature information, meteorological status information, user attribute information, and information on the clothing owned by the user to generate user status information. Means for generating prompt statements based on the user state information and inputting them into a generative artificial intelligence model, sending input information containing the prompt statements to the generative artificial intelligence model, and obtaining proposal information related to clothing combinations from the generative artificial intelligence model; This is a means of matching the clothing elements contained in the proposal information with the possession information to extract candidate combinations of clothing already owned by the user, and filtering the product information obtained from an external providing device based on the clothing elements and user preference information to extract clothing information as purchase candidates. A means for obtaining identification information from a user terminal that indicates that code information set in the store has been read, and associating the store information and inventory information corresponding to the identification information with the proposal information to generate clothing proposal information corresponding to the store environment; A means for sending the clothing proposal information to a user terminal and displaying the clothing proposal information on the user terminal in the form of a dialog or a list.
[0666] (Note 2) The information processing system according to Appendix 1 is characterized in that, The processing method is configured to perform data mining processing based on statistical analysis or machine learning methods on the user's historical purchase records and browsing records accumulated in the storage device to extract user preference information and style preference information, and reflect the style preference information in the generation of the prompt statement and the filtering of the product information.
[0667] (Note 3) The information processing system according to Appendix 1 is characterized in that, The processing method is configured to acquire natural language input information from the user through an interface device, store the input information in the storage device, and dynamically change the content of the prompt statement input to the generative artificial intelligence model based on the input information and the user state information, thereby generating clothing proposal information that reflects the user's subjective needs and emotional state.
[0668] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring environmental information using a processing unit and generating clothing matching suggestions; A device for using a processing device to obtain the user's current temperature and weather information through a sensing device, and generating clothing matching suggestions based on the obtained environmental information; An apparatus for acquiring image data and attribute data of clothing owned by a user through an information input device and an image acquisition device using a processing device, structuring the acquired clothing data into clothing record data containing category information, color information, style information, occasion information and season information, and storing the clothing record data into a clothing information dataset in a storage device. An apparatus for retrieving corresponding clothing record data from the clothing information dataset using a processing device based on the user's identification information, organizing the retrieval results by category, color, and purpose, and generating a data set as the input basis for a generative artificial intelligence model. An apparatus for using a processing device to parse natural language question data input by a user through a terminal device, and to perform word segmentation, keyword extraction and intent recognition processing on the natural language question data in order to extract target clothing information, matching scenario information and matching requirement type information. An apparatus for combining the parsing results of the clothing record data and the natural language question data using a processing device, generating a prompt statement containing system role information, user clothing list information, matching scenario information and matching task description according to a predetermined text template, and using the prompt statement as input data for a generative artificial intelligence model. An apparatus for providing the prompt statement to an external or internal generative artificial intelligence model via a communication device using a processing device, enabling the generative artificial intelligence model to generate natural language response information containing at least one set of clothing matching schemes based on the prompt statement, and receiving the natural language response information; An apparatus for post-processing the natural language response information using a processing device, matching the clothing name information therein with the clothing record data recorded in the clothing information dataset, and generating structured matching result data containing clothing identification information, image reference information and matching reason information in each matching scheme; An apparatus for transmitting the structured matching result data to a terminal device via a communication device using a processing device, and for enabling the terminal device to present clothing matching schemes to the user in the form of images and text based on the structured matching result data; An apparatus for receiving user evaluation information or operation history information on clothing matching schemes from a terminal device using a processing device, storing the evaluation information as user preference information, and appending the user preference information in text form to the prompt statement when generating subsequent prompt statements so that the output of the generative artificial intelligence model reflects the user preference. An apparatus for using a processing device to filter and sort externally provided product information based on online product information data and the user preference information, thereby generating clothing or product recommendation information that is consistent with the clothing owned by the user and the output results of the generative artificial intelligence model; An apparatus for using a processing device to acquire user state data based on an image acquisition device and a sound acquisition device, and for analyzing the state data through an emotion analysis algorithm to obtain emotion information, and for including the emotion information as an additional condition in a prompt statement or a matching generation rule to adjust clothing matching suggestion information. An apparatus for generating clothing matching suggestions suitable for a user's state and preferences based on user input information, emotional information, and the prompt statements using a generative artificial intelligence model and a processing device.
[0669] (Note 2) The information processing system according to Appendix 1 is characterized in that, The system further includes: a device for analyzing the user's historical purchase information, historical browsing information, and historical outfit selection information recorded in the storage device using a processing device, performing statistical processing and pattern extraction processing on the color features, clothing type features, style features, and brand features that appear therein, in order to generate a user preference model, and reflecting the user preference model as constraints or tendencies when filtering online product information data and generating prompt statements, thereby selecting or inducing the generation of clothing matching schemes or product recommendation information that better match the user's preferences and style.
[0670] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system further includes: a device for receiving text input information, image input information, and operation instruction information from a terminal device via an interactive interface using a processing device; associating the input information with the corresponding user identifier and storing it in a user information dataset in a storage device; simultaneously using an emotion estimation processing device to estimate the user's emotional information based on camera image data and sound data; and taking into account the emotional information along with user preference information and environmental information when generating prompt statements and calling a generative artificial intelligence model to generate clothing matching suggestions, so as to output clothing matching suggestions adapted to the user's current emotional state.
[0671] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring a user's location information, acquiring meteorological information corresponding to the location information from an external information providing device, and determining suitable clothing combinations for the user based on the meteorological information; A device for acquiring attribute information and image information of clothing owned by a user through an information terminal operated by the user, storing the clothing information in an information storage device, and structuring the clothing information into a data form that can be referenced in subsequent recommendation processing; This device is used to obtain users' historical purchase behavior information and browsing behavior information from information storage devices, use statistical processing devices to summarize and process attributes such as category, color, price range, and style, generate preference parameters that represent users' preference tendencies, and maintain the preference parameters as user profiles. A device for acquiring a user’s voice information, image information, or text information, using an external or internal emotion analysis processing device to infer the emotion state and emotion intensity from the information, and recording the emotion state in an information storage device in association with time information and context information. An apparatus for obtaining attribute and image information of clothing products from online information sources, selecting the clothing products as a preliminary candidate set based on the preference parameters and the meteorological information, and maintaining the preliminary candidate set as a candidate group for clothing recommendation processing; An apparatus for automatically generating prompt statements for inputting into a generative artificial intelligence model based on the meteorological information, the user profile, the clothing information owned by the user, the candidate group, and the emotional state, and providing the prompt statements to the generative artificial intelligence model to generate clothing matching schemes in natural language form; An apparatus for parsing the natural language output obtained from the generative artificial intelligence model, structuring it by mapping it to attributes such as color, category, and style, and comparing the resulting structured result with the clothing information owned by the user and the product information obtained from the online information source, thereby identifying specific clothing item combinations. A device for combining the specific clothing items and the natural language output, along with image information or virtual 3D image information, to send to a user terminal, so that the clothing status is visually presented on the user terminal in a virtual or augmented reality space, and generating information for receiving preference information or purchase operations from the user. An apparatus for acquiring user-inputted preference information, adopted pairing information, and purchase result information, recording them as learning data along with the correspondence between the preference parameters and the emotional state in an information storage device, and updating the recommendation logic of the statistical processing device or the generative artificial intelligence model based on the learning data.
[0672] (Note 2) According to the information processing system described in Appendix 1, the device for generating the prompt statement is configured to, when generating the prompt statement, explicitly record the attribute information of the clothing owned by the user, the weather information, the emotional state, and the user profile as different information elements in the prompt statement using natural language, and include in the prompt statement constraints for instructing the generative artificial intelligence model to generate matching guidelines, color matching reasons, and usage scenario descriptions based on the relationships between the information elements.
[0673] (Note 3) According to the information processing system described in Appendix 1, the system comprises: a device for calculating the acceptance index for each combination of emotional state and color attribute or style attribute using user-input preference information and purchase result information, and dynamically updating the screening conditions of the candidate group and the weight of the recommendation priority in the prompt statement using the acceptance index, so that the distribution of clothing attributes recommended based on the user's emotional state changes automatically over time.
Claims
1. An information processing system, characterized in that, include: processor; A sensor device is configured to be controlled by the processor to acquire the user's current temperature and weather information and to provide the temperature and weather information to the processor; The database is configured to be accessed by the processor to store information about clothing owned by the user and user input information; The camera and microphone are configured to be controlled by the processor to acquire information related to the user's emotions; The processor is configured as follows: Based on the user's current temperature and weather information obtained by the sensor device, a clothing matching scheme is generated for the user; The user's clothing information is stored in the database, and based on the clothing information stored in the database, a generative artificial intelligence model is used to generate the optimal clothing matching scheme for the user. The system obtains product information from online shopping websites and filters the product information based on the user's preferences and style to generate clothing recommendations suitable for the user. Based on the user's emotion-related information obtained from the camera and microphone, the emotion analysis algorithm is used to analyze the user's emotions, and the clothing recommendations are adjusted based on the analysis results. Using a generative artificial intelligence model, prompt text is generated based on the user's input information and the user's emotions, and optimal clothing recommendations are generated for the user based on the prompt text.
2. The information processing system according to claim 1, characterized in that, The processor is configured to use data mining techniques to analyze a user's past purchase and browsing history, and based on the analyzed user preferences and styles, select clothing items from product information provided by online shopping websites for recommending to the user.
3. The information processing system according to claim 1, characterized in that, The processor is configured to receive input information from the user through an interface, store the input information in the database, and generate clothing recommendations for the user by using an emotion engine that takes into account the user's emotions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A