Information processing system
Patent Information
- Application Number
- CN202610254357.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-03
- Publication Date
- 2026-09-22
AI Technical Summary
然而,现有技术中存在以下问题:第一,用户购买历史数据和行为模式数据虽被大量积累,但多以规则引擎或简单统计方式进行分析,难以及时、深入地挖掘潜在需求,导致销售推荐不够精准,个性化程度不足;第二,虚拟角色的应对内容通常由预设脚本或固定流程生成,无法根据实际应对结果进行高层次的语义分析和策略优化,导致销售促进策略难以动态调整和持续改进;第三,来自用户的反馈信息大多仅被用于简单满意度统计,未能通过智能解析转化为可操作的服务改进方案,导致服务质量优化依赖人工经验,效率低且主观性强
通过将模型结构和训练方式与提示语句生成机制结合,服务器不仅仅是简单调用已有模型,而是在内部建立了面向特定场景的高效调用路径和约束机制。这种设计在保证生成灵活性的同时提升了输出的业务相关性和语义一致性,通过减少无效对话轮次降低网络通信负荷,并在服务器内部减少重复计算,实现整体计算效率的提升。
Smart Images

Figure CN122798451A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.
[0003] In the process of providing goods or services to end users, operators typically need to develop personalized sales promotion strategies based on users' purchase history and behavioral patterns, and respond through virtual characters or virtual customer service representatives during user interactions. However, existing technologies suffer from the following problems: First, although a large amount of user purchase history and behavioral pattern data is accumulated, it is mostly analyzed using rule engines or simple statistical methods, making it difficult to timely and deeply uncover potential needs, resulting in inaccurate sales recommendations and insufficient personalization. Second, the responses of virtual characters are usually generated by preset scripts or fixed processes, making it impossible to conduct high-level semantic analysis and strategy optimization based on actual response results, thus hindering the dynamic adjustment and continuous improvement of sales promotion strategies. Third, user feedback is mostly used only for simple satisfaction statistics, failing to be transformed into actionable service improvement plans through intelligent analysis, resulting in service quality optimization relying on human experience, which is inefficient and highly subjective. Therefore, how to utilize generative artificial intelligence models to perform unified intelligent analysis and linkage optimization of user purchase history and behavioral patterns, virtual character response results, and user feedback information, thereby achieving efficient and dynamic improvement of sales promotion strategies and service quality, has become a technical challenge that this invention urgently needs to address. Summary of the Invention
[0004] To address the aforementioned technical challenges, this invention proposes an information processing system. This system includes a processor configured to work collaboratively with a generative artificial intelligence (AI) model to intelligently analyze user-related data and interaction results, outputting results for business optimization. Specifically, the processor achieves the solution to the aforementioned challenges through the following means: First, the processor inputs a prompt indicating the collection and analysis of user purchase history and behavioral pattern data into the generative AI model. The generative AI model then performs high-dimensional semantic analysis and pattern mining on the data based on the prompt, thereby obtaining information reflecting user preferences, price sensitivity, potential needs, and other characteristics, providing a data foundation for subsequently developing personalized sales promotion strategies. Secondly, the processor inputs a prompt indicating that the responses to the virtual character should be analyzed and that sales promotion strategies should be optimized based on the analysis results into the generative AI model. This allows the generative AI model to comprehensively analyze the dialogue content, recommendation order, and user responses of the virtual character during the reception process, outputting analysis results to identify the quality of the response scripts, the effectiveness of the recommendation strategies, and user reaction patterns. The processor then automatically or semi-automatically optimizes the sales promotion strategies based on these analysis results to achieve continuous improvement of the reception process and recommendation logic. Furthermore, the processor also inputs a prompt indicating that feedback from users should be collected and analyzed, and that the analysis results should be used for service improvement into the generative AI model. This allows the generative AI model to perform semantic understanding and cluster analysis on user subjective evaluations, free text opinions, and rating data, outputting results to identify service deficiencies, improvement needs, and priorities. The processor then adjusts the service content, service process, and virtual character response strategies based on these results. Through the above-mentioned technical means, this invention can organically combine user behavior data analysis, virtual character response result analysis, and user feedback analysis, and achieve intelligent and dynamic optimization of sales promotion strategies and service quality with the help of generative artificial intelligence models, thereby effectively solving the problems of insufficient personalization, difficulty in adaptive optimization of strategies, and low efficiency of user feedback utilization in existing technologies.
[0005] "System" refers to an overall device including at least one processor and storage devices, communication interfaces and / or other hardware or software components connected to the processor, which is a collection of technologies used to perform the data collection, prompt information generation, interaction with generative artificial intelligence models and result application processes described in this invention.
[0006] A "processor" is a hardware or virtual computing unit that can execute computer program instructions, process input data and output results. It can be a single central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), programmable logic device, or a processing module composed of multiple of the above units.
[0007] "Generative AI models" refer to AI models that can automatically generate text, code, vector representations, or other data outputs based on input prompts. They are usually trained using deep learning methods and include, but are not limited to, large language models, dialogue generation models, and generative models used to infer user characteristics or generate strategy suggestions.
[0008] "Prompt information" refers to text or structured data constructed by a processor and input into a generative artificial intelligence model to instruct the generative artificial intelligence model to perform specific parsing, reasoning, or generation tasks. The prompt information includes at least the data type to be processed, the processing target, or the processing constraints.
[0009] "User's purchase history data" refers to records related to a user's past purchasing behavior, including but not limited to the types, quantities, prices, purchase times, purchase channels, and related transaction attributes of the goods or services purchased, which are used to reflect the user's historical consumption behavior characteristics.
[0010] "Behavioral pattern data" refers to data used to characterize the behavioral features of users when using goods or services, browsing content, or participating in marketing activities, including but not limited to access frequency, dwell time, click path, frequency of function use, and response to recommended content.
[0011] "Virtual character" refers to a virtual image or virtual customer service entity that is presented on a terminal device in the form of images, audio, video or text and is used to interact with users. The virtual character can generate response content by a preset script or by a generative artificial intelligence model.
[0012] "Virtual character's response results" refers to the response content and related interactive behavior data generated by the virtual character during the interaction between the virtual character and the user, including but not limited to dialogue text, voice output content, recommended product or service lists, display order, and user responses and subsequent behavior records.
[0013] "Sales promotion strategy" refers to a set of strategies developed to increase the sales volume, conversion rate, or average order value of goods or services. These strategies include, but are not limited to, which goods or services to recommend, the order of recommendation, the segmentation of target user groups, the setting of preferential schemes, and the persuasive language and reception process used in the conversation.
[0014] "User feedback information" refers to evaluation data provided by users, whether actively or passively, after experiencing the product, service, or virtual character reception process. This includes, but is not limited to, ratings, questionnaire responses, free text opinions, complaint content, suggestions, and other forms of subjective evaluation information.
[0015] "Service improvement" refers to the process of adjusting, optimizing, or redesigning the products or services provided, service processes, customer service mechanisms, and / or virtual role response strategies based on the analysis results of user feedback information and related data, in order to improve user satisfaction, user experience, and business performance. Attached Figure Description
[0016] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0017] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0018] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0019] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0020] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0021] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0022] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0023] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0024] Figure 9 This represents an emotion map that maps multiple emotions.
[0025] Figure 10 This represents an emotion map that maps multiple emotions.
[0026] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0027] Figure 12This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0028] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0029] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0030] Hereinafter, an example of an implementation of the system to which the technology of this disclosure relates will be described with reference to the accompanying drawings.
[0031] First, let me explain the terminology used in the following instructions.
[0032] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0033] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0034] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0035] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0036] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0037] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0038] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0039] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0040] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0041] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0042] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0043] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0044] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0045] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0046] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0047] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0048] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0049] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0050] In intelligent customer service and sales scenarios targeting end users, traditional systems typically employ pre-written fixed scripts or recommendation logic based on simple rules. This presents the following technical problems: First, while the server can record users' purchase history and behavior logs, it lacks a systematic mechanism for unified preprocessing, statistical analysis, and feature extraction of multi-source behavioral data. This makes it difficult to model user attributes in detail, and the server-generated response content is insufficiently adaptable to individual user differences. Second, even with the introduction of generative AI models, existing technologies often use these models as single-use text generation tools, failing to establish a closed-loop computational process within the server: "behavioral features → prompt statement construction → generation result feedback → strategy iteration." This hinders the dynamic optimization of prompt statements and virtual representation configurations. Third, the image, tone, and dialogue of virtual representations (virtual characters) are often statically set during deployment. The server does not use user "role change requests," dialogue history, sales performance information, etc., as computable inputs for unified modeling and analysis. Consequently, the system cannot automatically adjust virtual representation attributes and dialogue strategies based on customer service performance, thus limiting the technical potential of generative AI models in improving conversion rates and user satisfaction.
[0051] Furthermore, in existing systems, when interacting with generative AI models, the generation of prompts is typically based on human experience, lacking a programmatic generation and update mechanism based on historical interaction data and sales results. The server cannot treat the prompts themselves as optimizable computational objects, hindering continuous improvement of their structure and content at the computational level, thus making it difficult to enhance the overall system response quality and stability. In summary, a complete computational process is needed on the server side, encompassing data acquisition, feature extraction, automatic prompt construction, dynamic configuration of virtual representations, and closed-loop strategy optimization. This would improve the controllability, adaptability, and long-term performance of generative AI models in intelligent reception systems, thereby achieving an overall improvement in computer data processing capabilities and human-computer interaction quality.
[0052] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0053] In this invention, the server includes a user attribute generation unit, a prompt statement generation unit, a reception control unit, a virtual representation optimization unit, a sales strategy generation unit, and a service improvement unit. This allows for unified preprocessing and feature extraction of user behavior information on the server side, automatically constructing generative artificial intelligence model prompt statements corresponding to user attributes and reception conditions. Based on virtual representation change requests, dialogue history information, and sales performance information from the terminal, the virtual representation attributes and prompt statement composition policy are iteratively parsed and updated. This forms a closed-loop computational process within the computer, covering data acquisition, feature modeling, prompt statement generation, model invocation, and strategy iterative optimization, thereby improving the controllability, adaptability, and overall processing performance of the generative artificial intelligence model in the intelligent reception system.
[0054] "Arithmetic device" refers to an electronic computing device that executes program instructions, processes and calculates input data, and outputs the processing results. It may include one or more processors, memory, and a bus structure connected to them.
[0055] "Information storage device" refers to a storage component used to store user behavior information, historical information and configuration data in a read-write manner, and may include semiconductor memory, magnetic storage medium or optical storage medium, etc.
[0056] "User behavior information" refers to a set of data used to represent a user's activity records in a specific service environment, including but not limited to purchase history information, browsing behavior information, dwell time information, and visit frequency information.
[0057] "User attribute information" refers to abstract information derived from user behavior information, used to characterize user preferences, interests, consumption tendencies, and other characteristics, and used to segment or personalize users in the system.
[0058] "Feature information" refers to a structured data set that is obtained by preprocessing, statistically analyzing and extracting features from raw user behavior information, and is used to represent user attribute information during the calculation process.
[0059] "General-purpose program" refers to a software program that runs on a computing device and is used for general information processing operations such as data preprocessing, statistical analysis, and feature extraction, and is not limited to a specific application scenario.
[0060] "Reception conditions information" refers to the constraints or settings related to the current reception scenario, including parameters such as time period, type of promotional activity, target product category, and target user group.
[0061] "Generative artificial intelligence models" refer to artificial intelligence models trained through machine learning or deep learning methods that can automatically generate corresponding output content based on input text or other forms of data.
[0062] "Prompt statements" refer to text or structured instructions input into a generative artificial intelligence model, used to indicate information such as the topic, style, length, and constraints of the content to be generated.
[0063] "Reception information" refers to text information, recommendation information, or dialogue content generated by a generative artificial intelligence model based on prompts and used to display reception information to users on terminal devices.
[0064] "Virtual representation" refers to a virtual entity displayed on a terminal device in the form of images, animations, or 3D models, used for interactive interaction with users, including virtual characters, virtual assistants, or other graphical agents.
[0065] "Display control information" refers to the set of parameters and instructions used to control the display of virtual representations on the visual output device of the terminal device, including information such as appearance style, action sequence, facial expression changes, and layout position.
[0066] "Voice playback information" refers to control data used to output sound on the auditory output device of a terminal device, including the text content to be read aloud, speech synthesis parameters, or identification information of pre-generated audio data.
[0067] "Terminal device" refers to an electronic device that communicates with a server and provides users with image display, voice playback and input collection functions, including but not limited to smart terminals, interactive information terminals or self-service terminals.
[0068] "Virtual representation change request information" refers to the request information entered by the user through the terminal device and sent to the server, indicating that the user wishes to change or adjust the current virtual representation's appearance, style, or response method.
[0069] "Virtual representation attributes" refers to the set of attributes used to define the appearance, voice, personality, language style, and other characteristics of a virtual representation.
[0070] "Virtual representation attribute candidates" refers to a variety of different virtual representation attributes that are pre-set or selectable, and can be selected or switched as needed during the reception process.
[0071] "Reception performance information" refers to statistical or recorded information related to the reception process and results of the virtual avatar, including dialogue history, user dwell time, recommendation clicks, and sales conversion.
[0072] "Reception effectiveness" refers to the effectiveness indicators obtained by evaluating the reception performance of the virtual entity based on reception performance information, which may include metrics such as user engagement, satisfaction, and conversion rate.
[0073] "Dialogue history information" refers to the collection of input and output content recorded in chronological order during the interaction between the user and the virtual representation, including user questions, system answers, and related contextual information.
[0074] "Sales performance information" refers to records related to the sales results of goods or services, including order records, transaction amounts, types of goods sold, and non-sold items.
[0075] "Analysis object information" refers to the data set that serves as the target of generative artificial intelligence models or other analysis processing, including dialogue history information, sales performance information, evaluation information, etc.
[0076] "Sales promotion policy information" refers to strategic information that is output by generative artificial intelligence models or other analysis processes based on the information of the parsed object, and is used to guide the adjustment of sales strategies, recommendation strategies or sales scripts.
[0077] "Generation conditions" refer to the control parameters and constraints used when generating prompts or reception information, including the types of feature information used, the style of the speech, length limits, and recommended targets.
[0078] "Evaluation information" refers to feedback information given by users after receiving or using the service, including ratings, comments, satisfaction selections, or other forms of subjective evaluation data.
[0079] "Reception control conditions" refers to a set of parameters used to control the reception process and the behavior of virtual avatars, including response timing, dialogue turn limits, role switching rules, and recommendation timing.
[0080] This invention will describe a system that operates collaboratively among a server, a terminal, and a user. The server, through a generative artificial intelligence model and prompts, performs feature modeling of user behavior information, automatic construction of prompts, dynamic configuration of virtual representations, and closed-loop optimization of strategies, thereby improving the data processing structure and inference flow within the computer. The following description, combining hardware, software, data structures, and algorithmic processing, will explain each component and its technical effects.
[0081] I. Overall System Composition A server comprises a computing unit and an information storage unit. The computing unit can be a computer system with a multi-core central processing unit and a graphics processing unit, running a general-purpose operating system, such as a Unix-like operating system. The information storage unit can include relational databases, log databases, and file storage systems, used to store user behavior information, feature information, historical dialogues, and model configurations, etc.
[0082] A terminal is a user-facing interactive device that can be a smart terminal device with a display screen, speaker, microphone, and touch input interface, running terminal applications. Terminal applications can use a graphics engine (such as a 3D graphics engine or a 2D / 3D rendering library) to render virtual representations, and a voice interface to play text-to-speech content.
[0083] Users are the actual users who interact with the terminal. They send requests to the server through the terminal's graphical interface, touch operation, or voice input, and receive the reception content output by the virtual representation.
[0084] II. Server User Attribute Generation Unit The server processes and extracts features from raw behavioral information using user attribute generation units to construct structured features suitable for input to generative artificial intelligence models.
[0085] 1. Server utilization of hardware and software environment The server stores user behavior data in relational tables within the information storage device. The table structure may include fields such as user identifier, timestamp, category, and amount. The server uses a relational database management system to execute queries, generating result sets in memory through database driver components.
[0086] The server runs a data analysis program on the computing device. This data analysis program can be implemented based on a scripting language environment and a data processing library. The server uses the data frame data structure in the data processing library to merge multiple database result tables into a unified behavioral dataset, and uses a numerical computing library to perform vectorized statistical operations.
[0087] 2. Server processing of specific data structures The server represents user behavior information as a sequence of records, where each record includes fields such as user ID, event type (browse, click, purchase), product category code, timestamp, and amount. The server performs the following data processing on these records: The server normalizes the timestamp field to generate the feature of "time interval between the most recent behavior and the current time"; the server counts the cumulative number of purchases and the cumulative amount for each product category and converts them into relative proportions to form a "category preference vector"; the server counts the access frequency and average dwell time in a given time window (e.g., the past 30 days) to generate "activity" related features.
[0088] The server combines the above results into a feature vector, where each dimension corresponds to a predefined feature, such as category preference percentage, price range distribution, access frequency, and time interval. The server caches this feature vector as user attribute information in memory and can write it to a feature storage table for subsequent access.
[0089] 3. Technical Effects and Causal Relationships By uniformly performing the aforementioned preprocessing and feature extraction within the server, the server transforms the originally lengthy row records into low-dimensional, computable feature vectors. This data structure transformation makes the input to subsequent generative AI models more stable and compact, reduces irrelevant noise features, and improves the consistency between the model-generated content and the user's actual preferences, thereby improving recommendation accuracy and response quality. Simultaneously, the server reduces the number of repeated scans of database records through vectorization, improving overall processing speed and efficiency in terms of both computing resources and storage access.
[0090] III. Server Prompt Statement Generation Unit The server uses a prompt generation unit to convert user characteristic information and reception condition information into natural language instructions suitable for generative artificial intelligence models, thereby controlling the output content and style of the model.
[0091] 1. Server prompt statement templates and parameter filling mechanism The server maintains various prompt templates in the information storage device. Each template corresponds to a different reception scenario, such as first-time reception, product recommendation, comparison explanation, etc. Each template contains several replaceable placeholders for inserting user attributes, product information, and strategy parameters.
[0092] Based on the current user's feature vector and reception conditions, the server selects a suitable template from the template library and maps the feature information into readable text. For example, the largest component of the "category preference vector" is transformed into the description "prefers sporting goods", and the price sensitivity feature is transformed into "acceptable mid-to-high price range".
[0093] The server inserts these descriptions into the template to generate the final prompt statement by concatenating strings or replacing placeholders.
[0094] For example, the server can generate the following prompt: "Based on the following user attributes, please generate a personalized greeting script for the store's virtual receptionist and recommend 3 products:" - User age group: around 30 years old - User preferred product categories: sports equipment (especially running shoes) Price sensitivity: Moderate, can accept mid-to-high price range - Recent purchases: I bought professional running shoes and sports socks last month. Requirements: The tone should be friendly and natural, the word count should be under 200 words, and three specific reasons for recommending the product should be given at the end. 2. The technical significance of the prompt statement structure The server systematically maps feature vectors to textual prompts, enabling the generative AI model to input refined user profiles and explicit generation constraints, rather than directly feeding raw line data into the model. This structured design reduces the burden of adapting to data structures and noise within the model, improves the controllability and consistency of the generated results, and helps maintain interface stability across different model versions or hardware environments.
[0095] IV. Server-side generative artificial intelligence model invocation and internal structure During the generative AI model invocation phase, the server employs a neural network architecture with multi-layered encoders and decoders, such as a sequence-to-sequence structure based on a self-attention mechanism. The server sends prompts to the model service deployed on the graphics processing unit via a network interface and receives text output.
[0096] 1. Model internal structure and training method The generative AI model used by the server can contain multiple stacked attention layers, each including a self-attention sub-layer and a feedforward network sub-layer. Model parameters include a word vector matrix, a multi-head attention weight matrix, and the weights and biases of the feedforward layers. The model is trained offline by optimizing an objective function, which can be a cross-entropy loss function. During the training phase, the server updates the weight parameters using a backpropagation algorithm.
[0097] The server uses a large-scale text corpus during training and can be fine-tuned by incorporating task-related corpora. The server can also introduce data expansion methods during training, such as randomly masking words or shuffling part of the sentence structure, to enhance the model's robustness.
[0098] 2. Control parameters when the server calls the model When constructing a request, the server specifies the maximum generation length, temperature parameters, and penalty parameters to control output diversity and repetition. The server encodes the prompt statement into a sequence of sub-words and sends it to the model service. The model service performs forward propagation computation on the graphics processing unit to generate a probability distribution sequence and generates the final text according to the sampling strategy.
[0099] The server performs regularization on the model output, including sensitive content filtering, length truncation, and syntax correction. These post-processing steps are implemented through rules or lightweight models to technically constrain the generated content and avoid unexpected output.
[0100] 3. Technical Effects By combining the model structure and training method with the prompt generation mechanism, the server does not simply call existing models, but internally establishes efficient calling paths and constraint mechanisms tailored to specific scenarios. This design improves the business relevance and semantic consistency of the output while ensuring generation flexibility. It also reduces network communication load by decreasing invalid dialogue rounds and reduces redundant calculations within the server, thereby improving overall computational efficiency.
[0101] V. Server reception control unit and terminal rendering control The server, through the reception control unit, converts the generated reception information into virtual display control information and voice playback information, and sends it to the terminal to achieve specific control over the real equipment.
[0102] 1. The server's structured encapsulation of reception information. The server segments the reception text output by the generative artificial intelligence model, marks the tone, emotional intensity and rhythm of different paragraphs, and determines the paragraph type, such as welcome, recommendation and explanation, through internal rules or lightweight classification models.
[0103] The server assigns predefined animation markers and facial expressions to each type of paragraph, such as "excited," "calm," and "explain," and packages these markers along with the corresponding text into a sequence of display control instructions. The server also adds speech synthesis parameters such as timbre, speech rate, and intonation to the text to control the text-to-speech engine.
[0104] 2. Specific control of display and voice functions by the terminal. After receiving control commands from the server, the terminal loads the corresponding virtual representation model on the graphics processing unit and triggers facial expressions, postures, and lip-sync animations according to the control commands. The terminal converts the text into an audio stream through local or cloud-based text-to-speech services and plays it synchronously according to the time sequence of the control commands.
[0105] Because the server pre-segments and tags the text, the terminal can directly execute animation controls without complex reasoning, thus reducing the computational burden on the terminal and saving terminal resources. The server's unified management of control logic also facilitates maintaining performance consistency across multiple terminals.
[0106] 3. Technical Effects By pre-generating display control information and voice parameters on the server side, complex scene logic can be centralized for computation on the server side, reducing the computational and storage burden on the terminal. Furthermore, precise time and state control improves the coherence and responsiveness of the virtual representation's output. This division of labor structure achieves a reasonable allocation of computational load and overall performance optimization.
[0107] VI. Server Virtual Representation Optimization Unit The server, through its virtual representation optimization unit, calculates the virtual representation attribute with the best expected reception effect based on the virtual representation change request information fed back by the terminal, various virtual representation attribute candidates, and historical reception performance information.
[0108] 1. Server attribute selection algorithm The server maintains a candidate list of virtual representation attributes in the information storage device. Each candidate contains multiple dimensions, such as gender, age, clothing style, speaking speed, and tone of voice. The server also maintains a mapping table between attributes and performance metrics, recording conversion rates, user dwell time, satisfaction feedback, and other values for different attribute combinations in historical sessions.
[0109] After receiving a change request, the server selects historical records with attributes similar to the current user from the mapping table, and uses statistical learning methods, such as a feature weight-based scoring function or a simple regression model, to calculate the expected effect score for the candidate attribute combination. The server then selects the attribute combination with the highest score as the new virtual representation configuration.
[0110] 2. Control process for regenerating prompt statements Once the server selects a new virtual representation attribute, it writes that attribute information into the new prompt statement, for example: "The user has proactively requested to change their virtual character. The user likes sporting goods and their previous character was a 'lively female character.' Please design a new opening line and first round of recommendation script for a 'calm, rational, and professional male character.'"
[0111] Require: - Language proficient but not stiff - A brief overview of user needs for athletic shoes - Recommend two running shoes with different target markets, and explain their key differences. In this way, the server enables the generative AI model to explicitly consider the new virtual persona when generating text, thereby maintaining consistency with the virtual persona's performance in terms of dialogue style and recommendation focus.
[0112] 3. Technical Effects Through the aforementioned attribute selection and optimization mechanism, the server no longer uses fixed roles but dynamically selects virtual representation configurations based on user feedback and historical performance. This data-driven attribute optimization approach automatically converges towards higher conversion rates and a better user experience, reducing the workload of manually tuning role configurations. At the computer system level, this is equivalent to implementing a strategy search and update mechanism based on historical performance within the server. The resulting virtual representation configuration is the result of algorithmic optimization, rather than static parameter settings, achieving adaptive evolution of behavioral strategies.
[0113] VII. Server Sales Strategy Generation Unit and Service Improvement Unit The server, through the sales strategy generation unit and service improvement unit, comprehensively analyzes dialogue history information, sales performance information, and user evaluation information, and updates prompts and reception control conditions.
[0114] 1. Server-side structuring and parsing of log data The server records the dialogue rounds, prompt message types, virtual representation attributes, and transaction process of each session as unified log entries. Each entry includes a timestamp, user attribute summary, prompt message identifier, generated content summary, user response, and sales result marker. The server periodically summarizes the logs into a statistical table and calculates the conversion rate of each prompt message template under different user groups and virtual representation attribute combinations.
[0115] The server inputs these statistical tables and key samples into the generative artificial intelligence model via prompts, requesting analysis and strategy suggestions in natural language, for example: "Below are the performance data of three customer service script templates in the sales of sports shoes in the past month (conversion rate, average order value, user dwell time, etc. are given in the data table).
[0116] Please analyze which type of sales pitch is more effective, and generate three new sales pitch templates based on the analysis results. Objective: To improve the conversion rate of high-end running shoes while maintaining a positive user experience. Based on the strategy text output by the model, the server updates the prompt template library and some feature weights. For example, the server can adjust the presentation order of certain features in the prompt, or increase or decrease the length of the price description.
[0117] 2. Utilization of User Review Information The server associates user feedback (ratings, textual feedback, etc.) submitted on the terminal with corresponding dialogue and virtual entity attributes, using it as additional training or fine-tuning data, or as constraints for subsequent strategy analysis. In this way, the server not only relies on transaction results but also considers subjective satisfaction, forming a multi-objective strategy optimization.
[0118] 3. Technical Effects By periodically updating its strategy using log and evaluation data, the server develops an internal self-learning mechanism, continuously optimizing the structure of prompts and the configuration of virtual representations over time. Unlike manual adjustments based solely on experience, the server's data-driven strategy optimization, characterized by non-linearity and high dimensionality, results in improved recommendation accuracy, fewer dialogue rounds, shorter response times, and increased resource utilization.
[0119] VIII. Alternative Implementation Methods and Extended Forms Servers can be deployed in different hardware environments, and the same architecture can be scaled to multiple physical machines or virtual machines to form a cluster mode to handle large-scale concurrent requests. Servers can migrate parts of generative artificial intelligence models to lightweight models deployed on the terminal, enabling collaboration between local inference and remote server inference to further reduce network latency and bandwidth consumption.
[0120] Terminals can adopt different rendering methods based on their hardware capabilities. For example, high-performance terminals use real-time 3D rendering, while low-performance terminals use pre-rendered sequences and simple animations. The server can detect the terminal's capabilities and dynamically adjust the granularity of the transmitted display control information, thereby reducing the overall system load while ensuring a good user experience.
[0121] The server can also employ different algorithms during feature extraction. For example, it can generate user group labels based on clustering algorithms, or use dimensionality reduction algorithms to compress high-dimensional behavioral features into a low-dimensional space to further improve the efficiency of model input. The server can treat these different methods as configurable modules, allowing for different combinations to be selected through parameters.
[0122] Through the various implementation forms and alternative methods described above, the system of the present invention constructs a complete technical link within the server, from raw behavioral data to feature vectors, from feature vectors to prompt statements, from prompt statements to generative artificial intelligence model output, and then to virtual representation control and strategy optimization. This achieves optimization of the computational structure and data flow for human-computer interaction scenarios, thus outperforming traditional systems that rely solely on manual scripts or single model calls in terms of accuracy, speed, and resource utilization.
[0123] use Figure 11 The processing flow is explained.
[0124] Step 1: The server receives user identification information and obtains behavioral data.
[0125] Input: User identifier from the terminal (e.g., user ID, member number, device identifier, or scan result).
[0126] Output: A set of raw behavioral data corresponding to the user (including purchase records, browsing records, dwell time records, visit frequency records, etc.).
[0127] Based on the request message sent by the terminal, the server parses the user identifier field in the communication module. The server then calls the database access module to read records related to the user from multiple behavioral data tables in the information storage device using database query statements. The server performs preliminary integration of the query results, sorting the records from different tables by time and merging them into a set of raw behavioral data, providing basic data for subsequent feature extraction.
[0128] Step 2: The server preprocesses and cleans the raw behavioral data.
[0129] Input: The set of raw behavioral data output from step 1.
[0130] Output: A cleaned set of behavioral data that has been denoised, completed, and formatted in a uniform way.
[0131] The server uses a data processing program to inspect each record of the raw behavioral data, deleting abnormal records with empty timestamps, negative amounts, or missing user identifiers. For missing but inferable fields (such as missing dwell time), the server fills in the missing data based on the average or median of similar records. The server uniformly converts time fields of different formats to standard timestamps and standardizes amount fields to the same currency unit while retaining a fixed number of decimal places. Based on a predefined field type table, the server assigns a uniform data type (numeric, enumerated, or text) to each field, generating a cleaned dataset with a consistent format.
[0132] Step 3: The server calculates statistical features from the cleaned behavioral data and generates user feature vectors.
[0133] Input: The set of cleaned behavioral data output from step 2.
[0134] Output: Feature vectors and feature label sets used to represent user attribute information.
[0135] The server performs statistical calculations on the cleaned behavioral data: It aggregates data by product category, calculates the number of purchases and purchase amount for each category, and normalizes these to a category preference ratio; it calculates the difference between the most recent purchase time and the current time to obtain the "recent purchase interval" feature; and it statistically analyzes the number of visits and average dwell time over multiple past time windows (e.g., 7 days, 30 days, 90 days) to calculate the "visit frequency" and "activity level" features. The server assembles these values into a one-dimensional numerical vector in a predefined order, simultaneously generating corresponding semantic labels (e.g., "preferred category: sporting goods," "price sensitivity: medium," etc.) as raw materials for constructing subsequent prompts.
[0136] Step 4: The server generates prompts based on user feature vectors and reception conditions.
[0137] Input: The feature vector and feature label set output from step 3, as well as the current reception conditions information (such as the product category being promoted, time period, terminal type, etc.).
[0138] Output: One or more prompt texts for use by generative artificial intelligence models.
[0139] The server first selects a suitable template from the prompt template library based on the reception scenario (e.g., "First-time Reception Template," "Recommendation Template," "Comparison Explanation Template"). Based on feature tags, the server maps information such as "age group," "preferred product categories," "price sensitivity," and "recent purchase history" into natural language descriptions and fills them into placeholders in the template. Simultaneously, the server inserts target descriptions and constraints based on the reception conditions, such as word limits, tone style, and the number of recommendations. The server ultimately generates a prompt statement like the following: "Based on the following user attributes, please generate a personalized greeting script for the store's virtual receptionist and recommend 3 products:" - User age group: around 30 years old - User preferred product categories: sports equipment (especially running shoes) Price sensitivity: Moderate, can accept mid-to-high price range - Recent purchases: I bought professional running shoes and sports socks last month. Requirements: The tone should be friendly and natural, the word count should be under 200 words, and three specific reasons for recommending the product should be given at the end. Step 5: The server sends the prompt to the generative artificial intelligence model and obtains the reception information.
[0140] Input: The prompt text output from step 4, and optional contextual information (summary of past conversations, current list of recommended products, etc.).
[0141] Output: The reception text information and recommended content structure (such as a list of recommended product descriptions) output by the generative artificial intelligence model.
[0142] The server encodes the prompts into a sequence of sub-words via the model interface module and sends them, along with control parameters (maximum generation length, temperature value, repetition penalty coefficient, etc.), to the generative AI model service deployed on the graphics processing unit. The server waits for the model to complete forward inference and receives the returned text results. The server parses the returned text, splitting it into several segments such as greetings, requirement guidance, and product recommendations, while extracting structured information such as recommended product names and feature descriptions. The server can truncate the results and filter for sensitive content, ultimately obtaining a structured set of reception information.
[0143] Step 6: The server converts the reception information into virtual presentation control information and voice playback information and sends it to the terminal.
[0144] Input: The set of reception information output from step 5 (segmented text, recommendation list, tone and intent tags, etc.).
[0145] Output: A set of display control commands and voice playback parameters sent to the terminal.
[0146] Based on the semantic type of different paragraphs (e.g., "welcome," "recommend," "explain"), the server assigns preset facial expressions, actions, and gesture markers to the virtual avatar. The server generates a sequence of display control instructions, each containing trigger time, action type, facial expression state, and text paragraph association information. The server also specifies audio playback parameters for each text segment, such as timbre identifier, speech rate, volume, and pitch curve, and encapsulates the text and parameters into audio playback information. Finally, the server packages the display control instructions and audio playback information into a data packet and sends it to the corresponding terminal over the network.
[0147] Step 7: The terminal receives display control commands and voice playback information and renders the virtual representation.
[0148] Input: The display control commands and voice playback information output from step 6.
[0149] Output: Virtual visuals on the terminal display screen and audio output through the speaker.
[0150] The terminal parses the data packets sent by the server, reading the virtual representation model identifier, action commands, and text content. The terminal loads the corresponding 2D or 3D model resources into its graphics processing module and drives the model to perform facial expressions, body movements, and lip-sync animations based on the timeline and action commands. The terminal submits the text content to a local or cloud-based text-to-speech engine to generate audio data, controlling the reading speed and pitch according to the voice parameters provided by the server. The terminal outputs the rendered image to the display screen and the synthesized audio to the speaker, achieving a user-facing visual and auditory experience.
[0151] Step 8: Users can interact with the virtual representation through the terminal and can send a request to "change the virtual representation".
[0152] Input: The virtual representation screen displayed on the terminal and the voice output (from step 7).
[0153] Output: User interaction information, such as click events, voice commands, and especially virtual representation change requests.
[0154] When users are watching or listening to the virtual avatar's presentation, they can tap the "Change Role" button on the touchscreen or speak commands such as "Change Role" or "Change Appearance." The terminal captures the touch event or records the voice and invokes a speech recognition service to convert the speech into text. Based on the recognized keywords, the terminal determines the user's intent. The terminal then constructs a virtual avatar change request containing the user's identifier, current virtual avatar attributes, and a description of the change intent, and sends it to the server over the network.
[0155] Step 9: The server optimizes the virtual representation attributes based on change requests, attribute candidates, and historical reception records.
[0156] Input: The virtual representation change request information output in step 8, the candidate list of virtual representation attributes stored in the information storage device, and historical reception performance information (conversion rate, stay time, satisfaction, etc.).
[0157] Output: New virtual representation attribute configuration and a new round of reception information adapted to those attributes.
[0158] The server parses the user identifier and current role information from the change request, and filters entries with similar attributes to the user from historical service records. The server then statistically analyzes the performance metrics of different virtual representation attribute combinations within these entries, such as calculating the average conversion rate and average satisfaction score for each combination. Based on a preset scoring function, the server weights and synthesizes multiple metrics into a comprehensive score, sorts the candidate attributes, and selects the attribute combination with the highest score as the new virtual representation configuration. Subsequently, the server inserts this attribute configuration along with the current user characteristics into a new prompt statement template, such as: "The user has proactively requested to change their virtual character. The user likes sporting goods and their previous character was a 'lively female character.' Please design a new opening line and first round of recommendation script for a 'calm, rational, and professional male character.'"
[0159] Require: - Language proficient but not stiff - A brief overview of user needs for athletic shoes - Recommend two running shoes with different target markets, and explain their key differences. The server calls the generative AI model again to obtain the reception text that matches the new attribute, and generates corresponding display and voice control information, ready to send it to the terminal for role switching.
[0160] Step 10: When a user asks a question about a product, the terminal collects the question and sends the conversation history to the server.
[0161] Input: The virtual representation currently displayed on the terminal and the recommended content (from step 7 or step 9), as well as the user's voice or touch input.
[0162] Output: A request message containing the user's question text, the current list of recommended products, and a summary of the conversation history.
[0163] Users select a recommended product on the terminal interface or ask a question via voice, such as "Which of these two running shoes is more suitable for daily commuting and weekend runs?" The terminal records the voice using its microphone and calls the voice recognition service to obtain the question text, or directly reads the text entered by the user on the screen. Simultaneously, the terminal reads the basic information of the currently recommended product and a summary of the current conversation turn from its local cache, packages this information together with the user's question text into a request message, and sends it to the server.
[0164] Step 11: The server uses a generative artificial intelligence model to generate responses based on dialogue history and recommended products.
[0165] Input: The user question text output from step 10, the information on currently recommended products, and a summary of the dialogue history.
[0166] Output: The response text and any updated recommendations.
[0167] The server combines the user's question, the differences between the recommended products, and the user's attributes into a new prompt statement, for example: User question: 'Which of these two running shoes is more suitable for daily commuting and weekend runs?' Product A: Lightweight cushioned running shoes, suitable for short-distance daily running.
[0168] Product B: Stable support running shoes, suitable for medium to long distance training.
[0169] User attributes: Frequent runner, walks a lot during daily commute, and has purchased professional running shoes in the past year.
[0170] Current virtual character design: a professional and calm male coach.
[0171] Please explain the differences between the two running shoes to users in simple and easy-to-understand language, and provide clear purchase advice based on user needs. The word count should be within 150 words. The server sends the prompt to the generative artificial intelligence model service, receives the response text generated by the model, verifies and simplifies the response according to its internal strategy, and then outputs it to the reception control unit to generate new display control information and voice playback information, which are then sent to the terminal for display.
[0172] Step 12: Users make purchasing or abandonment decisions based on the virtual representation's answers, and the terminal transmits the decisions to the server.
[0173] Input: The answer content and recommended product information displayed on the terminal (from step 11).
[0174] Output: An operation message indicating whether the user intends to purchase or abandon their intention.
[0175] After referring to the explanation from the virtual representation, users can click buttons such as "Buy," "Add to Cart," or "Don't Buy Now" on the terminal, or express their intentions via voice, such as "I want to buy these B-type running shoes." The terminal organizes the selected product identifier, quantity, intention type, and necessary payment method preferences into structured data and sends it to the server.
[0176] Step 13: The server processes purchase requests and records sales performance information.
[0177] Input: User purchase intent and product information output from step 12.
[0178] Output: generated order records, payment processing results, and updated sales performance information.
[0179] After receiving a purchase request, the server creates a new order entry in the order management module, writing information such as user ID, product ID, quantity, price, and timestamp into the database. The server then calls an external payment interface to generate a payment link or QR code address and returns this information to the terminal. After payment is completed, the server updates the order status to "paid" based on the payment result callback and writes data such as transaction amount, product category, virtual representation attributes, and prompt statement type into the sales performance table for subsequent statistics and strategy optimization.
[0180] Step 14: The server periodically summarizes the dialogue history and sales performance to generate new strategies and prompt message templates.
[0181] Input: Conversation history, sales performance, and user reviews stored in the information storage device.
[0182] Output: Updated prompt statement template, virtual representation attribute strategy, and generation condition parameters.
[0183] The server reads recent session and transaction records from the log database via a scheduled task, calculates the conversion rate, average order value, user dwell time, and satisfaction distribution for each combination of prompt template and virtual representation; the server then organizes these statistical results into a textual summary and constructs strategy analysis prompts, such as: "Below are the performance data of three customer service script templates in the sales of sports shoes in the past month (conversion rate, average order value, user dwell time, etc. are given in the data table).
[0184] Please analyze which type of sales pitch is more effective, and generate three new sales pitch templates based on the analysis results. Objective: To improve the conversion rate of high-end running shoes while maintaining a positive user experience. The server inputs the prompt statement into the generative artificial intelligence model, receives the new script templates and strategy suggestions output by the model, and then updates the prompt statement template library and some feature mapping rules, so that the prompt statements and virtual representation configurations generated in subsequent steps 4 and 9 are more in line with the latest data feedback, thereby realizing the adaptive evolution of system strategies and overall performance improvement.
[0185] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0186] In both physical stores and online service scenarios, technologies have been proposed that utilize virtual characters to provide users with product recommendations and sales promotion information. However, existing technologies typically employ fixed scripts or simple rule engines to generate customer service scripts. These methods struggle to perform granular analysis based on user behavior and transaction information in a timely manner, and are even less capable of continuously and automatically optimizing customer service effectiveness and sales performance in large-scale data environments.
[0187] Furthermore, existing systems often use generative AI models only as general text generation tools, lacking structured design and dynamic adjustment mechanisms for input prompts. They fail to organically combine multi-dimensional information such as user attribute characteristics, virtual character appearance and voice settings, and historical reception effects, resulting in generation results that are difficult to balance personalization, consistent character experience, and quantifiable sales performance.
[0188] Furthermore, traditional virtual shopping guide systems typically lack an integrated closed loop that connects the user's terminal display history, location-related information, and operation history, to the selection of prompt templates and the updating of metrics. This prevents the automatic iteration and optimization of prompts based on customer service quality and sales performance indicators. Consequently, computer resources cannot be efficiently utilized for the focused use of high-value prompt templates, and the computational power of generative AI models fails to specifically improve overall recommendation effectiveness and user satisfaction.
[0189] Therefore, it is necessary to provide a system that can automatically acquire and analyze user behavior and transaction data in information processing devices, utilize generative artificial intelligence models in conjunction with adjustable prompt statement structures, combine the appearance and voice information of virtual display elements, and adaptively optimize the prompt statements based on reception quality indicators and sales performance indicators, thereby improving the computer's processing flow, data utilization efficiency, and generation quality control capabilities in intelligent reception and sales promotion tasks.
[0190] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0191] In this invention, the server includes components for acquiring user behavior information and transaction information, summarizing the behavior information and transaction information using data processing and data analysis programs, and generating feature quantities representing attribute information and preference information for each user; components for generating prompt statements in natural language based on the feature quantities and product or marketing information; and components for inputting the prompt statements into a generative artificial intelligence model to generate response text containing product recommendation text or sales promotion information for a single user; and components for acquiring setting information related to the appearance and voice information of virtual display elements from the user and reflecting the setting information in the prompt statements, so that the response... The system includes components that generate text styles or expressions corresponding to the appearance and voice information of the virtual display elements; components that send the response text to a mobile information terminal via a communication component; components that display the response text to the user on the mobile information terminal in the form of a dialogue of the virtual display elements or through voice output; and components that, after the response text is displayed, retrieve the user's operation information and transaction information from the mobile information terminal, calculate reception quality indicators and sales performance indicators using statistical processing programs, and update the content or structure of subsequent prompts input into the generative artificial intelligence model based on the reception quality indicators and sales performance indicators. This allows for the construction of a closed-loop processing chain on the server side, encompassing multi-source user data collection, feature generation, prompt construction, generative artificial intelligence model invocation, terminal presentation, indicator feedback, and adaptive template selection. This improves the data processing and model invocation efficiency of the computer in personalized recommendation and virtual reception tasks, enabling the generated results to be automatically optimized based on quantitative indicators while maintaining role consistency and user experience, thereby improving the overall system's intelligence and sales promotion performance.
[0192] "Information processing device" refers to an electronic device with a processor, memory and communication interface, used to perform computational processing operations such as data acquisition, data processing, model calling and result output.
[0193] "User" refers to a service recipient who interacts with the system through a mobile information terminal, and whose behavioral and transaction information is acquired by the system and used to generate personalized recommendations.
[0194] "Behavioral information" refers to the data recorded by users when using mobile information terminals or in service environments, which is related to browsing, location, clicks, dwell time, input, and other operations.
[0195] "Transaction information" refers to data such as order records, payment records, product identification, price, and quantity related to goods or services purchased or attempted to be purchased by a user.
[0196] "Data processing program" refers to an application program or script that runs in an information processing device and is used to read, clean, transform, and summarize user data.
[0197] "Data analysis program" refers to an application program that runs in an information processing device and is used to perform statistical analysis, feature extraction and pattern recognition on user data, including but not limited to spreadsheet data processing software.
[0198] "Table data processing software" refers to data processing software that can perform filtering, aggregation, grouping, and statistical operations on data stored in tables or tabular structures.
[0199] "Attribute information" refers to general information derived from user behavior and transaction information, used to represent user categories, preferences, spending power, and other characteristics.
[0200] "Preference information" refers to descriptive information obtained by analyzing user behavior and transaction information, which indicates the user's preferences in terms of category, brand, price range, style, etc.
[0201] "Features" refer to numerical or symbolic data extracted from raw data and structured to represent user attribute and preference information, which are used to drive subsequent analysis and generation processes.
[0202] "Product information" refers to descriptive data related to the product or service to be recommended, including but not limited to category, function, specifications, price, and applicable scenarios.
[0203] "Marketing information" refers to promotional data such as discounts, events, bundled sales, and points rewards designed to promote the sale of goods or services.
[0204] "Natural language" refers to the written or spoken language that users use in their daily lives, including but not limited to Chinese, used to represent prompts and responses in readable text form.
[0205] "Prompt statements" refer to natural language text generated by an information processing device and input into a generative artificial intelligence model to instruct the model to perform a specific generative task.
[0206] "Prompt statement template" refers to a predefined text structure used to construct prompt statements, which includes variable parts that can be filled with feature quantities, product information, or setting information.
[0207] "Generative artificial intelligence models" refer to machine learning models that can automatically generate natural language text or other forms of content based on input prompts.
[0208] "Response text" refers to the natural language text output by generative artificial intelligence models based on prompts, including product recommendation text or sales promotion information.
[0209] "Product recommendation text" refers to natural language content in the response text used to recommend specific products or services to users and explain the reasons or advantages of the recommendation.
[0210] "Sales promotion information" refers to natural language content in response text used to explain discounts, promotions, or other favorable conditions to stimulate user purchasing behavior.
[0211] "Virtual display elements" refer to virtual characters or virtual interface elements presented on mobile information terminals in the form of images, animations, or three-dimensional images.
[0212] "Appearance information" refers to data used to define the visual characteristics of virtual display elements, including but not limited to gender style, clothing style, color style, and facial expression style.
[0213] “Voice information” refers to parameter data related to the voice output of virtual display elements, including but not limited to timbre, speech rate, tone, and gender perception.
[0214] "Settings information" refers to the personalized selections or configuration data made by users regarding the appearance and voice information of virtual display elements.
[0215] "Style" refers to the overall expression of the response text in terms of language style, tone, formality, etc.
[0216] "Expression style" refers to the specific manifestations of the response text in terms of sentence structure, word choice, and friendliness.
[0217] "Communication means" refers to wired or wireless communication interfaces and protocols used to transmit data between information processing devices and mobile information terminals.
[0218] "Mobile information terminal" refers to a portable electronic device with processing and communication capabilities, used to run applications and interact with servers.
[0219] "Dialogue-style display" refers to the display method of presenting interactive text with users on mobile information terminals in the form of message bubbles, dialog boxes, or similar interface elements.
[0220] “Voice output” refers to the process of converting text content into audible sound and playing it through the speaker of a mobile information terminal or a connected audio device.
[0221] "Operation information" refers to the interactive behavior data generated by the user on the mobile information terminal after responding to the text prompt, such as clicking, swiping, inputting, and selecting.
[0222] "Reception quality index" refers to a quantitative or qualitative evaluation value calculated based on user operation information and feedback, used to assess the effectiveness of virtual reception.
[0223] "Sales performance indicators" refer to quantitative metrics calculated based on transaction information and used to evaluate sales results and marketing effectiveness.
[0224] "Statistical processing program" refers to a program in an information processing device used to perform statistical operations such as aggregation, correlation analysis, and regression analysis to calculate reception quality indicators and sales performance indicators.
[0225] "Storage means" refers to storage devices or storage media used to store information such as user data, feature quantities, prompt statement templates, and indicator relationships.
[0226] "Correspondence" refers to a relational data structure in storage that represents the relationship between reception quality indicators or sales performance indicators and specific prompt statements or prompt statement templates.
[0227] "Pre-set benchmark" refers to the threshold or evaluation condition used to judge whether the reception quality indicator or sales performance indicator has reached the expected level.
[0228] In various embodiments of the present invention, the server, terminal, and user collaboratively constitute the whole that realizes the system functions. The server acts as an information processing device, the terminal as a mobile information terminal, and the user as the subject interacting with the system. The embodiments of the present invention will be described below from the aspects of hardware architecture, software modules, data structure, the composition and learning method of generative artificial intelligence model, and technical effects.
[0229] I. Overall System Composition The server employs a general-purpose computer hardware architecture with a multi-core central processing unit (CPU) and optional graphics processing unit (GPU), runs a server operating system such as a Unix-like operating system, and is equipped with main memory and non-volatile memory. The server includes communication interfaces for data communication with terminals via wired or wireless networks.
[0230] The server includes the following software features: The server is an application implemented using a programming language, used to implement business logic control; The server uses a network framework to implement the Hypertext Transfer Protocol interface; The server uses a data processing library to clean, aggregate, and extract features from tabular data; The server uses a database management system to store user behavior information, transaction information, characteristic data, prompt statement templates, as well as reception quality indicators and sales performance indicators.
[0231] The server also includes an interface component for invoking generative AI models, which interacts with external generative AI services via an application programming interface (API).
[0232] A terminal is a mobile information terminal equipped with a processor, memory, display unit, audio output unit, and wireless communication module, such as a smartphone, tablet computing device, or wearable device. The terminal runs a mobile operating system and applications for communicating with a server. The terminal presents virtual display elements on the display unit, outputs voice on the audio output unit, and locally records the user's display history, location-related information, and operation history.
[0233] Users interact with virtual display elements through the terminal's user interface, initiating actions such as browsing, clicking, inputting, evaluating, and setting avatars on the terminal. In real-world scenarios, users can also try out or purchase products based on terminal prompts, and the resulting transaction information is acquired by the server.
[0234] II. Software Modules and Data Structures The server can be logically divided into multiple functional modules at the application layer, including: data acquisition module, feature generation module, prompt statement generation module, generative artificial intelligence invocation module, virtual display element setting and processing module, result distribution module, and effect evaluation and optimization module.
[0235] In the data acquisition module, the server stores the behavioral and transaction information received from the terminal according to a unified data structure. Behavioral information can be represented in a tabular structure, where each record includes fields such as user identifier, event type, object identifier, timestamp, duration, location information, and terminal identifier. Transaction information uses a similar tabular structure, containing fields such as order identifier, user identifier, product identifier, unit price, quantity, total amount, and payment time. The server uses a data processing library to read and write these tables.
[0236] In the feature generation module, the server uses data analysis programs to aggregate and perform calculations on behavioral and transaction information. The server groups the data by user identifier and calculates statistical characteristics for each user. For example, the server calculates a user's browsing frequency, purchase frequency, browsing-to-purchase conversion rate, average order amount, preferred product categories, preferred price ranges, and preferred brands within a predetermined time window. The server can also construct feature quantities reflecting trend changes based on time series analysis of the differences between recent and long-term user behavior.
[0237] The server organizes these statistical features into structured feature vectors, with each feature vector corresponding to a user. Feature dimensions include behavior frequency, monetary distribution, category preferences, and time characteristics. The server's feature generation module can also perform preprocessing operations such as discretization, normalization, or embedding mapping to transform the raw multidimensional numerical and categorical data into a representation suitable for subsequent generative artificial intelligence models. The server also generates human-readable attribute description text, such as "This user has recently mainly purchased athletic shoes, with an average price in the middle range, and frequently visits the sports products page in the evening," which can be directly embedded as prompts.
[0238] The server maintains a user attribute information table for each user in the database, which records the user's feature values, the time of the most recent feature generation, and related statistical indicators. The server can quickly retrieve the feature values of a specified user through indexes, thereby supporting the generation of real-time or near real-time prompt statements.
[0239] III. Generative Artificial Intelligence Models and Prompt Statement Construction The server employs a generative AI model in its generative AI invocation module. In one implementation, this model can be a neural network model based on a multi-layer transformer architecture, including an encoder layer, a decoder layer, and a self-attention mechanism. The model learns language structure during the pre-training phase using a large corpus, and then undergoes supervised learning using scenario-specific dialogue data and reception script data during the fine-tuning phase. The model training process includes defining a cross-entropy loss function, updating model weights using a stochastic gradient descent-like optimization algorithm, and using regularization techniques to control model complexity.
[0240] During training, the server constructs training samples using paired data between user behavior and recommended phrases. User attribute descriptions, product descriptions, and marketing conditions are used as input, while historical phrases with high satisfaction rates are used as the target output. Model parameters are adjusted by minimizing the loss function. Similarly, the server can introduce a reinforcement learning phase, using service quality indicators and sales performance indicators as reward signals to optimize the combination of prompts and generated results, making the model more inclined to produce outputs that lead to high performance metrics during inference.
[0241] The server maintains multiple prompt templates in its prompt generation module. These prompt templates are natural language text structures, including placeholders that can be filled with feature values, product information, and setting information. For example, the server can use the following prompt example: "You are a sales associate in a physical store, and you need to provide personalized recommendations based on the user's historical purchasing behavior."
[0242] The user profile is as follows: This user has recently been primarily buying athletic shoes, with an average price of approximately 8,800 yen, and frequently chooses a particular sportswear brand.
[0243] There is a new product now: Lightweight running shoes, suitable for daily running and commuting, with good cushioning and breathability.
[0244] Please generate a greeting message in Chinese, no more than 80 characters, with a natural and friendly tone, suitable for online conversations with users. The server can also generate various prompts based on different objectives, for example: "The user's past purchase history is in athletic shoes. Please generate a customer service script of no more than 60 words to recommend a newly launched athletic shoe to the user in a friendly and natural tone, suitable for face-to-face communication in a physical store." "You are a footwear sales assistant."
[0245] Users often buy lightweight running shoes with good cushioning and prefer neutral color schemes.
[0246] Please recommend a newly arrived running shoe to this user, describing its highlights in no more than 80 Chinese characters, and providing a brief invitation to try it on. When constructing specific prompts, the server first populates the user profile section based on user characteristics, then fills the product description section with descriptions from the product information table. Simultaneously, it adds stylistic constraints based on the virtual display element settings, such as, "The current virtual sales assistant has a rather composed image; please use a slightly formal expression with a moderate speaking speed." The server then generates complete prompts, which are provided as input to the generative artificial intelligence model.
[0247] The server invokes the generative AI model service via a communication interface, transmitting the prompt and related parameters to the model execution environment. The generative AI model, on accelerated hardware, performs word segmentation, embedding, and attention calculations on the input, generating a probability distribution for each position and outputting the response text in sequence. The server extracts the complete response text from the returned results and splits it into product recommendations and sales promotion information according to simple rules.
[0248] IV. Virtual Display Element Setting and Terminal Presentation In the virtual display element setting processing module, the server processes the setting information uploaded from the terminal. The terminal sends setting information, including user identifier, avatar style identifier, and voice style identifier, to the server. The server stores this setting information in the user setting table, so that each user has a corresponding virtual display element configuration.
[0249] When generating prompts, the server converts this setting information into language style constraints, such as "use a lively and friendly tone" and "the response should be concise yet polite," and embeds these constraints into the prompts in text form. In this way, the generative AI model is subject to these constraints during the decoding phase, resulting in outputting response text consistent with the virtual display elements.
[0250] After receiving the response text from the server, the terminal locally parses it to obtain the text content, avatar identifier, and voice identifier. The terminal then invokes the graphical user interface component in its display unit to render pre-stored or downloaded virtual character images, animations, or 3D models onto the screen and display the response text in the form of speech bubbles. In its audio output unit, the terminal selects a local or remote text-to-speech engine based on the voice identifier to convert the response text into an audio signal and play it. The terminal thus enables virtual display elements to communicate with the user in a conversational manner.
[0251] V. Effect Evaluation and Optimization of Prompt Statements In the performance evaluation and optimization module, the server uses operation and transaction information received from the terminal to calculate service quality indicators and sales performance indicators. Through a statistical processing program, the server aggregates the performance of each prompt statement and its corresponding response text across multiple interactions, calculating metrics such as click-through rate, add-to-cart rate, order conversion rate, average order amount, user dwell time, and user ratings. The server establishes a correspondence between these metrics and the identifiers of the prompt statements or prompt statement templates, and stores them in the database.
[0252] When generating subsequent prompts, the server selects appropriate prompt templates based on reception quality and sales performance indicators. The server searches its storage for templates that meet predetermined benchmarks and prioritizes these templates. The server can also adjust the predetermined benchmark values or template weights in real time based on changes in the indicators, gradually shifting the system towards higher-performing template combinations during operation. Thus, the server achieves adaptive optimization of the prompt structure at the algorithmic level, rather than simply replacing content.
[0253] Because the server organizes behavioral and transaction information using specific data structures, encodes user profiles in the form of feature quantities, and finely controls the input to the generative AI model through automatic selection and updating of prompt templates, it can reduce invalid calls, improve the matching degree between model output and actual user needs, and reduce the probability of repeatedly generating unsuitable prompts. This control method improves resource utilization efficiency within the computer, including reducing unnecessary network requests, reducing the number of model inferences, shortening response time, and improving overall generation accuracy through targeted learning and optimization.
[0254] VI. Technical Effects and Causal Relationships The technical effects achieved by the server through the aforementioned structured processing and control mechanisms are not only reflected in personalized recommendations at the business level, but also in improvements to computer technology itself. Specifically: The server uses tabular data processing software and data analysis programs to transform multi-source behavioral and transaction information into compact and highly generalized feature quantities. This feature processing reduces redundant information input to generative artificial intelligence models, thereby reducing data transmission volume and model processing burden, improving computational efficiency and reducing communication load.
[0255] By maintaining prompt statement templates and their corresponding relationships, the server establishes a mapping between reception quality indicators and sales performance indicators and the prompt statement structure. An algorithm selects templates that meet predetermined benchmarks from the template set, automatically converging the subsequent generation process towards efficient templates. This mechanism forms a feedback-based adaptive control loop within the computer, enabling fine-grained constraints on the generation process, improving the stability and accuracy of the generated results, and reducing errors caused by fluctuations in script quality.
[0256] The server parameterizes the appearance and voice information of virtual display elements and embeds these parameters into prompts, prompting the generative AI model to follow specific styles and expressions when outputting text. This achieves a coordinated match between the model's output and the front-end presentation. This consistency control from visual and auditory settings to text style is achieved through a computer internal technology solution via specific data streams and parameter injection, which differs from the method of manually editing scripts.
[0257] By combining reception quality metrics and sales performance metrics during the training phase, the server can use these metrics as rewards or weights to weight the error function, thereby reinforcing the learning of high-performance outputs during gravity updates. This training strategy causes the model's internal parameter distribution to cluster towards regions corresponding to higher business performance, thus achieving higher recommendation accuracy and conversion rates at runtime.
[0258] In summary, this invention, through the design of data and control flows between the server, terminal, and user, cleverly combines the use of generative artificial intelligence models with feature generation, prompt template management, and indicator feedback control. This not only achieves the functional goal of personalized reception but also improves computer technology itself in terms of processing efficiency, generation accuracy, and resource utilization. The server can thus autonomously and dynamically optimize the generation process in a non-simplistic, rule-based manner within a complex and ever-changing user behavior environment, achieving efficient and stable technical results.
[0259] use Figure 12 The processing flow is explained.
[0260] Step 1: The terminal acquires and sends the user's raw data.
[0261] The terminal takes the user's browsing history, click history, location-related information, order information, etc., as input data and temporarily stores them locally in the form of an event list. The terminal adds metadata such as user identifier, timestamp, and terminal identifier to each event and serializes the event list into structured data. Based on a communication protocol, the terminal sends this structured data to the server via the network, outputting a request message containing behavioral and transaction information.
[0262] Step 2: The server receives and stores the raw data.
[0263] The server takes the request message from the terminal as input and parses the behavioral and transaction information fields through the communication interface. The server performs field and format validation on the parsed results, discarding or marking incomplete or abnormal records. The server writes the validated data into database tables, storing it in the behavioral record table and the transaction record table respectively, outputting a normalized raw dataset for subsequent analysis.
[0264] Step 3: The server generates user characteristic values.
[0265] The server reads records associated with a specified user identifier from the behavior log table and transaction log table as the input dataset. The server uses a data processing program to perform aggregation operations on this dataset, including counting pageviews and purchases within a time window, calculating the average order amount, and identifying major product categories and brands. The server normalizes, discretizes, or encodes the statistical results, converting them into multi-dimensional feature vectors and corresponding textual attribute descriptions, outputting user feature quantities and user profile text.
[0266] Step 4: The server selects and fills in the prompt statement template.
[0267] The server takes user characteristics, user profile text, and candidate product or marketing information as input. It selects an appropriate template from a stored set of prompt templates based on preset rules or historical performance metrics. The server then fills in the user profile text, product description, and marketing conditions in the placeholder positions within the template, and sets constraints such as additional text style and tone based on the user's virtual avatar and voice style, outputting a complete natural language prompt.
[0268] Step 5: The server invokes a generative artificial intelligence model to generate response text.
[0269] The server takes the prompt from step 4 as input, along with control parameters such as model identifier, maximum output length, and temperature parameters, and sends it to the generative AI model's runtime environment via the model interface. Internally, the generative AI model performs word segmentation, embedding mapping, self-attention calculation, and decoding on the prompt. It then calculates the probability of each word using a multi-layered neural network, outputting natural language text word by word. The server receives the text returned by the model, parses it into a complete response text, and outputs a response text containing product recommendations and sales promotion information.
[0270] Step 6: The server organizes the response text and generates the data to be sent.
[0271] The server takes the response text as input and uses rule-based or pattern-matching methods to identify different parts such as product recommendations and promotional messages. The server combines the identification results with user identifiers, product identifiers, session identifiers, and current virtual display element configurations to form structured response data. The server specifies the avatar style identifier and voice style identifier to be used in this data, and outputs a response data packet that can be directly sent to the terminal.
[0272] Step 7: The terminal receives and presents the response content.
[0273] The terminal takes the response data packet returned by the server as input and parses out fields such as the response text, avatar style identifier, and voice style identifier. Based on the avatar style identifier, the terminal loads the corresponding virtual character image or animation, draws virtual display elements in the display unit, and displays the response text as a speech bubble or similar format. Based on the voice style identifier, the terminal calls a local or remote text-to-speech component to convert the response text into a speech signal, which is then played through the audio output unit, outputting both user-facing image display and voice output.
[0274] Step 8: Users interact with the terminal.
[0275] Users read or listen to the dialogue displayed on the terminal as input. Based on the prompts, they perform interactive actions on the terminal interface, such as clicking "View Details," "Add to Cart," "Not Interested," and entering questions or reviews. User actions directly change the state of the terminal interface, such as navigating to a product details page, updating the shopping cart, or popping up a review window, resulting in new user action events and potential new transaction requests.
[0276] Step 9: The terminal records and reports interaction logs and setting information.
[0277] The terminal takes the user's operation events in step 8, as well as changes to the user's settings such as avatar style and voice style for virtual display elements, as input. The terminal adds a timestamp, current session identifier, and interface context information to each operation, compiling this data into an interaction log and setting information record. The terminal sends the interaction log and the latest setting information to the server via the communication interface, outputting updated user operation information and virtual display element setting information messages.
[0278] Step 10: Server updates user settings and template effect records.
[0279] The server takes the user's operation and settings information uploaded by the terminal as input. First, it updates the user's avatar style and voice style identifiers in the settings storage table. Then, the server associates the user's operation information with previously issued prompts and response texts, calculating service quality indicators such as click-through rate, conversion rate, and dwell time for each prompt, and updating the corresponding sales performance indicators from the transaction information. The server establishes or updates a mapping between these indicators and prompt template identifiers, outputting a new user setting status and a template performance record with performance metrics.
[0280] Step 11: The server adjusts its prompt message strategy based on performance metrics.
[0281] The server takes the template performance records generated in step 10 as input and reads the reception quality indicators and sales performance indicators corresponding to each prompt statement template. Based on a predetermined benchmark or optimization strategy, the server scores or ranks the templates, selecting those with better performance than the benchmark and reducing or discontinuing the use of poorly performing templates. The server updates the template selection rules or weight table and stores the new weight configuration in storage. The server then outputs an optimized set of prompt statement templates and their weight configurations, which are used by the subsequent prompt statement generation module, thereby achieving adaptive optimization of the generative artificial intelligence model input in subsequent processing cycles.
[0282] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0283] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0284] Existing technologies that utilize generative artificial intelligence models for human-computer dialogue and product or service recommendations suffer from the following technical problems: First, servers typically use generative AI models only as single-turn response engines, lacking technical solutions for structured collection and correlation analysis of user attribute information, behavioral information, and multi-turn dialogue logs. This results in the inability to automatically identify user characteristics at the computer system level, and the inability to dynamically adjust the model's input content based on user characteristics, leading to a low degree of matching between the response content and the user's actual needs.
[0285] Second, when existing systems invoke generative artificial intelligence models, most rely solely on fixed prompts or simple rules, lacking a closed-loop optimization mechanism based on historical sales data, user behavior results, and response effectiveness. The server side is unable to quantitatively calculate and compare the sales performance indicators and response quality indicators of different prompts, thus failing to form a self-iterative strategy optimization process within the computer.
[0286] Third, existing marketing support systems often separate data statistical analysis from dialogue generation, making it impossible to finely correlate structured dialogue information with user behavior results in chronological order within the same server. The server also cannot automatically generate, filter, and update prompts applicable to different user groups and different product or service categories based on this time-series data, making it difficult for the context construction of generative artificial intelligence models to be optimized for specific scenarios.
[0287] Fourth, in terms of system performance and resource utilization, existing technologies often only focus on single inference performance and lack optimized design starting from the complete processing chain of "data acquisition - model invocation - effect evaluation - prompt statement update". The server cannot automatically adjust the model invocation parameters, prompt statement usage conditions and display strategies on the dialogue device side according to the evaluation and analysis results, thus limiting the overall balance between response speed, resource consumption and marketing effect of the entire platform.
[0288] The technical problem to be solved by this invention is to provide an information processing system that enables a server to: automatically acquire and structure-record dialogue information and user behavior information within the computer; perform evaluation and analysis processing based on this data; identify user characteristics and response quality; and then automatically generate, select, and update prompt statements on the server side. The optimized prompt statements are then combined with dialogue information and input into a generative artificial intelligence model, thereby forming an integrated technical solution for dialogue generation and strategy control that can self-iterate within the computer system and optimize for sales strategies. This improves the computer technical performance of the generative artificial intelligence model in terms of personalized responses, sales conversion, and resource utilization.
[0289] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0290] In this invention, the server includes a processing unit for acquiring attribute and behavioral information of the user and generating instruction information for a generative artificial intelligence model to parse the attribute and behavioral information to identify the characteristics of the user; a processing unit for acquiring dialogue information from a dialogue device and structurally recording the dialogue information and the response information output by the generative artificial intelligence model, and generating instruction information for performing evaluation and analysis processes related to sales activities based on the structured recording information; and a processing unit for generating prompt statements input to the generative artificial intelligence model based on the results obtained from the evaluation and analysis processes, combining the prompt statements with the dialogue information to control the generative artificial intelligence model to perform response information generation processing, and further comparing multiple prompt statements based on sales performance indicators and response quality indicators to optimize sales strategies and update the usage conditions of the prompt statements. This allows for a closed-loop computational process within the server, from user characteristic identification and structured acquisition of dialogue and behavioral data to automatic generation and selection of prompt statements and repeated updates to sales strategies. This enables the input construction and invocation strategies of the generative artificial intelligence model to automatically adjust with evaluation results, thereby improving the overall technical performance of the computer system in terms of personalized dialogue generation, marketing conversion efficiency, and computational resource utilization.
[0291] "System" refers to an entirety consisting of at least one processing device, a storage device, and a communication device for communicating with dialogue devices and generative artificial intelligence models, wherein the devices work together through hardware and software to perform functions of information acquisition, processing, storage, and output.
[0292] "Users" refers to the target entities that interact with the system, including but not limited to individual users, groups of users, or customer objects identified by identification information, whose attribute and behavioral information is acquired by the system and used for analysis and processing.
[0293] "Attribute information" refers to data used to describe the static or relatively stable characteristics of the user object, including but not limited to age group, region, interest preference category, consumption level category, and other classification or numerical information that can be used to characterize the user object's characteristics.
[0294] "Behavioral information" refers to data that reflects the user's interaction with the interactive device or related services within a certain time range, including but not limited to access records, click records, browsing duration, adding to favorites, adding to cart, placing orders, cancellations, and other operations in chronological order and result information.
[0295] "Dialogue device" refers to a terminal device used for human-computer dialogue interaction with users, including but not limited to portable terminals, fixed terminals, or other devices with input / output interfaces that can communicate with servers, used to collect dialogue information and present response information to users.
[0296] "Dialogue information" refers to natural language or structured text data exchanged between a dialogue device and a user, including textual representations of questions, requests, and feedback input by the user, as well as the corresponding content output by the system or generative artificial intelligence model.
[0297] "Generative artificial intelligence models" refer to artificial intelligence models built based on machine learning or deep learning methods that can automatically generate output text based on input text or other data. They are used to perform natural language understanding and natural language generation, including but not limited to dialogue models based on neural networks.
[0298] "Response information" refers to the output generated by a generative artificial intelligence model after receiving input information (including prompts and dialogue information), which is used as the text or equivalent representation of the response to a question or request from the user.
[0299] "Structured records" refers to the process and results of organizing and storing acquired dialogue and response information according to a pre-defined data structure and field format, so that each piece of information can be retrieved and analyzed in dimensions such as time, user identifier, session identifier, and content type.
[0300] "Instruction information" refers to control data generated by the server and input into the generative artificial intelligence model. It is text or equivalent in form used to instruct the generative artificial intelligence model to perform parsing, evaluation or generation of specific data.
[0301] "Prompt statements" refer to text inputs constructed to guide generative artificial intelligence models to generate response information according to predetermined roles, tones, goals, and content scope. These include, but are not limited to, natural language content used to describe the model's role, task objectives, output format constraints, and scene-related descriptions.
[0302] "Evaluation processing" refers to the process of quantitatively or qualitatively evaluating the response information output by generative artificial intelligence models in terms of sales results, user response, and dialogue continuity, based on structured recorded information and the behavioral results of the users.
[0303] "Analysis and processing" refers to the process of using statistical analysis, pattern recognition, or machine learning methods to mine attribute information, behavioral information, dialogue information, and evaluation results in order to identify user characteristics, response patterns, and their correlation with sales results.
[0304] "Sales activities" refers to the abstract collection of overall business activities surrounding the promotion and transaction of goods or services, including but not limited to information introduction, recommendation, promotional guidance, and related service descriptions.
[0305] "Sales results" refers to the outcome data generated by sales activities, including but not limited to whether a transaction was completed, the number of orders, the transaction amount, the conversion rate, or quantitative indicators that are directly or indirectly related to sales performance.
[0306] "Sales performance metrics" refer to one or more numerical values or statistics used to quantitatively represent sales performance, including but not limited to conversion rate, click-through rate, average order value, and repurchase rate.
[0307] "Response quality metrics" refer to quantitative or qualitative indicators used to measure the quality of response information, including but not limited to dialogue rounds, user dwell time, frequency of subsequent interactions, number of positive user feedbacks, or a comprehensive score calculated from these data.
[0308] "Sales strategy" refers to a set of rules, parameters, or strategy configurations used to control recommended content, the way prompts are used, and the dialogue flow, based on sales performance indicators, response quality indicators, and the characteristics of the target audience.
[0309] "Sales strategy optimization" refers to the process of automatically or semi-automatically adjusting sales strategies based on the results of evaluation and analysis, so that the system can achieve better results under a predetermined objective function (such as sales performance indicators or response quality indicators).
[0310] "Conditions of use" refers to the set of constraints and rules associated with the applicable scenario of the prompt statement, including but not limited to the applicable user group, the applicable product or service category, the applicable time period, or the type of business activity.
[0311] "Characteristic groups" refer to a collection of users clustered or grouped based on attribute and behavioral information. Users within each group are similar in terms of preferences, behavioral patterns, or other characteristic dimensions, and are used for configuring differentiated prompts and controlling responses.
[0312] In a preferred embodiment of the present invention, the system includes a server, multiple terminals, and a communication network for connecting the server and the terminals. The server typically comprises a computing device with a multi-core central processing unit and a graphics processing unit, such as a multi-core processor, a graphics accelerator, main memory, and non-volatile memory. The server runs an operating system, such as a Unix-based server operating system, and deploys a network service framework, such as an HTTP-based application server. The server is also connected to a data storage device via the network, which may include relational data storage, key-value data storage, and object data storage. The terminals can be portable information processing devices, such as smartphones or tablets, or desktop information processing devices. The terminals run client applications for collecting user input and presenting responses generated by the server. Users interact with the system through the terminals.
[0313] The server stores various software modules and data structures in its storage device. The software modules include: a generative artificial intelligence model invocation module, a prompt statement management module, a dialogue record management module, a user characteristic analysis module, an evaluation and analysis module, a sales strategy optimization module, and a management interface module. The data structures include: a user attribute record table, a behavior record table, a dialogue log table, a prompt statement template table, a sales performance statistics table, and a model configuration table.
[0314] The server loads a generative AI model in the generative AI model invocation module. In one embodiment, this model is a neural network based on a transformer architecture, comprising an embedding layer, multi-layer self-attention encoding units, multi-layer self-attention decoding units, and an output probability calculation unit. In the embedding layer, the server maps the input sequence to a vector representation. In the encoding unit, it uses a multi-head self-attention mechanism to model the global relevance of the input sequence. Residual connections and layer normalization are used to improve training stability and inference numerical stability. Subsequently, in the decoding unit, the conditional probability distribution of the next label is calculated based on the generated labels and encoded representations, and the output text is obtained through a beam search or random sampling strategy.
[0315] The server maintains a prompt statement template table in the prompt statement management module. For each prompt statement, the prompt statement template table records: unique identifier, applicable user group, applicable product or service category, applicable business scenario, tone style, and version information. The server stores these fields in a structured format in storage, such as defining table columns using a relational data structure, to facilitate fast retrieval and filtering. When selecting a prompt statement, the server performs query operations on the prompt statement template table using indexes and conditional filtering, thereby obtaining a set of candidate prompt statements with lower computational overhead.
[0316] The server performs structured recording of dialogue information received from the terminal in the dialogue record management module. For each round of dialogue, the server generates a session identifier, round number, and timestamp, and categorizes fields such as user-input text, response information output by the generative artificial intelligence model, used prompt statements, and terminal identifiers into the dialogue log table. During recording, the server uniformly encodes text fields into a unified character set, specifies data types for numeric fields (e.g., integer or floating-point), and uses a unified time base for time fields. Through this structured recording, the server can efficiently associate different behaviors and dialogue content using aggregation and join operations in subsequent processing.
[0317] In the user characteristic analysis module, the server utilizes user attribute and behavior record tables, based on a data analysis software library, to identify user characteristics. The server uses the data analysis library to aggregate user behavior information, calculating statistics such as visit frequency, dwell time, click frequency, purchase frequency, and average purchase amount. These statistics are then vectorized into feature vectors of uniform length. Based on this, the server calls a clustering algorithm module, such as a centroid-based clustering algorithm, to cluster the feature vectors to identify several user characteristic groups. Through this computational method, the server automatically forms the classification of "characteristic groups" internally through numerical calculations, without relying on manual rules, thereby improving the granularity and scalability of user segmentation.
[0318] In the evaluation and analysis module, the server correlates dialogue information, response information, and behavioral results to calculate sales performance metrics and response quality metrics. The server connects the dialogue log table and the behavior record table using session identifiers and timestamps to determine whether actions such as clicking on product details, adding items to the cart, or generating an order occurred within a certain period after a round of responses. The server accumulates these behavioral results for each type of prompt, calculating sales performance metrics such as click-through rate, add-to-cart rate, and order conversion rate, while simultaneously calculating response quality metrics such as average session rounds, user dwell time, and repeat visit rate. Through these calculations, the server establishes a correspondence between prompts and specific metrics, thus attaching a set of comparable numerical indicators to each prompt in the internal data structure.
[0319] In the sales strategy optimization module, the server uses metrics output from the evaluation and analysis module to select and update prompt statements. The server can use sorting operations to filter out the top-performing prompt statements for specific user groups and product categories, marking these as recommended versions. Simultaneously, the server can also generate new prompt statements using a generative artificial intelligence model. To do this, the server inputs meta-prompt statements and historical high-performing phrase samples into the generative artificial intelligence model. For example, the server can input the following text-based prompt statement into the generative artificial intelligence model: "Based on the following high-performing sales dialogue samples, summarize a template of prompts suitable for promoting high-end smartphones. Please output 3 prompts in different styles, each no more than 80 characters, highlighting the high-end experience, camera capabilities, and gaming performance." After receiving multiple candidate prompt statements from the model, the server filters them using rules such as regular expression matching, length limit checks, and sensitive word filtering. The filtered prompt statements are then stored in a prompt statement template table and set to a test-ready state. Subsequently, during actual operation, the server assigns users to different prompt statement versions, performing multi-version comparisons to determine the optimal prompt statement. This selection process, based on metric feedback and multi-version comparison, essentially constitutes an iterative optimization algorithm within the server, gradually bringing the prompt statement and model context construction towards optimality.
[0320] When invoking a generative AI model, the server doesn't simply accept the user's question as input; it also combines prompts, user profile information, and summaries of recent dialogue rounds to form a complete input sequence. During input construction, the server concatenates role descriptions, task objective constraints, and the user's question into a single text according to a predefined format. For example, the server might construct the following text: "You are a virtual sales consultant. Please communicate with the user in Simplified Chinese. Based on the user's question, concisely explain the core selling points of the new smartphone, including processor performance, battery life, camera features, and price range. Avoid providing information unrelated to the product. User question: Is this phone suitable for gaming?" The server then transforms the text into a labeled sequence input to the generative AI model through word segmentation and encoding. This allows the model to generate response information under more precise contextual constraints during inference. Since the prompts are automatically optimized based on historical data evaluation results, this input construction method significantly improves the computational consistency between the model output and specific user needs and sales targets.
[0321] In this invention, the terminal is primarily responsible for collecting user input and displaying content returned by the server. When running the client application, the terminal packages the user's input natural language text and related metadata into structured data and sends it to the server via network protocols. After receiving the server's response, the terminal presents it as a speech bubble or plays audio, optionally combined with virtual character images or animations. The terminal can also record specific user actions on the local interface, such as clicking a product item or triggering a function button, and report these actions to the server. In this way, the terminal converts user behavior into structured behavioral information, providing it to the server for subsequent analysis.
[0322] Users interact with the system through terminals, asking questions about products or services or providing feedback on the responses of virtual avatars. After collecting this feedback, the server can use it as additional signals for the evaluation and analysis modules, such as incorporating explicit positive or negative reviews into the calculation of response quality indicators. In this way, the system can adjust prompts and response strategies based on user subjective satisfaction, in addition to sales results.
[0323] The technical advantages of this invention are not only reflected in the automation of business processes, but more importantly in the improved computer performance brought about by the internal data structure and algorithm design of the server. The server records dialogue and behavioral information in a structured manner, and uses vectorized features and clustering operations in the evaluation and analysis modules, enabling user characteristic identification to be completed in a high-dimensional feature space. This is more accurate and flexible than traditional user classification based on fixed rules. Regarding the selection and updating of prompt statements, the server achieves incremental strategy updates under limited computing resources through statistical index comparison and multi-version trials, avoiding the resource consumption caused by large-scale model retraining. Since the evaluation results are recorded in the prompt statement template table, the server only needs to perform a simple query and sorting before each input construction to select the optimal prompt statement. Therefore, it does not significantly increase latency during the inference stage. Instead, in long-term operation, it reduces invalid dialogue rounds through higher conversion rates, thereby reducing network communication and computational overhead overall.
[0324] Furthermore, during the training phase of the generative AI model, the server can employ a combination of supervised learning and reinforcement learning. In an offline environment, the server uses existing dialogue data and sales results as training samples, employing negative log-likelihood as the basic loss function and incorporating reward signals based on sales performance metrics to adjust the model weights. The server updates the weights in the parameter space using stochastic gradient descent and its variants, enabling the model to not only fit the linguistic structure of the dialogue data but also tend to generate response patterns that promote high sales results at the parameter level. Thus, by introducing sales performance metrics during the training phase, the system can reflect its preference for sales objectives in the internal weight representation during inference, improving overall marketing effectiveness.
[0325] Unlike simple "automation of manual operations," this invention achieves independent technical improvements within the server by designing specific data structures, feature construction methods, and strategy optimization loops. The server no longer merely generates fixed dialogue scripts for humans; instead, through the coordinated optimization of generative artificial intelligence models and prompts, it introduces dynamic adjustments and closed-loop feedback mechanisms at both the model input construction and strategy decision-making levels. This mechanism changes the way traditional dialogue systems handle fixed contexts and static rules, enabling the same hardware resources to achieve higher matching accuracy and conversion effects in fewer dialogue rounds and shorter processing times, thus bringing technical improvements in processing speed, accuracy, and resource utilization.
[0326] In other implementations, the server can employ different forms of feature extraction and clustering methods. For example, instead of being limited to centroid clustering, the server can use density-based clustering in the user characteristic analysis module to automatically identify unevenly distributed user groups. The server can also introduce time decay weights in the evaluation and analysis module, giving more weight to recent conversational behaviors in metric calculations, allowing the system to adapt more quickly to changes in user interests. These variations all fall within the scope of the present invention.
[0327] In another implementation, the server can divide the prompt statements into a multi-layered structure, including a global role description layer, a scene constraint layer, and a language style layer. When constructing the final prompt statements, the server can select different combinations from each layer, thereby flexibly adjusting the tone and content emphasis without changing the overall model. Furthermore, the server can adjust the prompt statements based on the terminal type; for example, it can prioritize generating shorter responses on small-screen terminals and allow for more detailed explanations on large-screen displays. This method of dynamically adjusting the length and level of detail of the model output based on terminal characteristics helps reduce invalid data transmission and lower network communication load.
[0328] In summary, this invention, through collaboration between servers, terminals, and users, utilizes generative artificial intelligence models and prompt management mechanisms to achieve a complete technical process within a computer, including structured processing of dialogue and behavioral information, user characteristic identification, response effectiveness evaluation, prompt generation and selection, and sales strategy optimization. This improves computer technology itself in terms of processing speed, response accuracy, data management, and resource utilization.
[0329] use Figure 13 The processing flow is explained.
[0330] Step 1: The user initiates a conversation and enters a request on the terminal. Users can enter natural language questions or requests through the application interface running on the terminal, such as "Is this phone suitable for playing games?".
[0331] Input: The text content entered by the user on the terminal and the basic operation information of the user on the terminal (such as the product item clicked).
[0332] The terminal temporarily stores the text and related metadata (user identifier, session identifier, current product identifier, timestamp) in local memory, performs basic formatting on the text (removing leading and trailing spaces, and using unified encoding), and then prepares to send it to the server.
[0333] Output: A structured request data object generated internally by the terminal, which contains at least user text, user identifier, session identifier, and timestamp, for subsequent network transmission.
[0334] Step 2: The terminal sends structured request data to the server. The terminal calls the network interface provided by the operating system and uses a secure transmission protocol to encapsulate structured request data into network packets.
[0335] Input: The structured request data object generated in step 1.
[0336] The terminal serializes the data object into a text format (such as a JSON string), sets the target server address and interface path, and sends the serialization result as the request body to the server over the network; at the same time, the terminal displays status prompts such as "generating answer" locally.
[0337] Output: Request messages transmitted over the network to the server, and request sending status records on the terminal (used for retry and timeout handling).
[0338] Step 3: The server receives the request and parses the basic information. The server receives request messages from the terminal on the network interface and uses the server-side network framework to parse the messages.
[0339] Input: A network request message sent by the terminal (containing serialized user request data).
[0340] The server first verifies the message format and integrity, then deserializes the text-based request body into an internal data structure (such as key-value pairs or objects), extracting fields such as user text, user identifier, session identifier, timestamp, and product identifier. The server also generates a unique request identifier for this request, used for log tracking.
[0341] Output: The request structure used internally by the server for subsequent processing, containing parsed fields and request identifiers.
[0342] Step 4: The server preprocesses the user-input text. The server uses a text processing module to normalize the natural language text input by the user in order to adapt it to the input format of the generative artificial intelligence model.
[0343] Input: The user text field in the request structure from step 3.
[0344] The server uses regular expressions to remove extra whitespace and special characters, performs language tagging on the text (e.g., recognizing it as Chinese), and calls a word segmentation tool to split the text into words or subwords; the server can also truncate excessively long text according to preset rules. The server stores the preprocessed text along with the original text into a temporary data structure for subsequent recording.
[0345] Output: Preprocessed user text sequence (such as a segmented string or a list of tags), and an updated request structure (with the preprocessed results appended).
[0346] Step 5: The server queries user characteristics and historical behavior data. The server reads user-related records from the data storage device based on the user identifier in order to construct personalized prompts.
[0347] Input: User ID and session ID from step 3.
[0348] The server performs queries on the user attribute record table and behavior record table to retrieve the user's historical visit count, historical purchase records, frequently browsed product categories, and previously assigned user characteristic group tags. If no record exists, the server marks the user as a new user and creates an initial record. The server then transforms the retrieved data into a user characteristic object in a standardized format.
[0349] Output: A user characteristic object containing user history behavior statistics and characteristic tags, appended to the processing context of the current request.
[0350] Step 6: The server selects or generates an applicable prompt statement template. The server selects an appropriate prompt statement from the prompt statement template table based on the user's characteristic object and the current product or service category.
[0351] Input: User characteristic object generated in step 5, and current product or service category information.
[0352] The server sets filtering conditions in the prompt statement template table based on user characteristics, product category, and scenario type, performs query and sorting operations, and retrieves several candidate prompt statements. If no direct match is found, the server degenerates into selecting a general prompt statement. The server combines one or more selected prompt statements into the base prompt information used in the current conversation.
[0353] Output: The basic prompt text corresponding to the current request, and its template identifier, are stored in the request processing context.
[0354] Step 7: The input content for the server to construct generative artificial intelligence models The server combines basic prompts, user characteristic information, summaries of recent dialogues, and the current user question into a complete model input.
[0355] Input: the preprocessed text from step 4, the basic prompt text from step 6, and the history of conversations related to the current session.
[0356] The server reads the questions and responses from the most recent rounds of the session from the dialogue log table, generates a dialogue summary string, and concatenates the following information in a predefined format: role description, task requirements, user characteristic prompts, scenario description, historical dialogue summaries, and the current user's question. For example, the server generates the following input text: "You are a virtual sales consultant. Please communicate with the user in Simplified Chinese. Based on the user's question, concisely explain the core selling points of the new smartphone, including processor performance, battery life, camera features, and price range. Avoid providing information unrelated to the product. User question: Is this phone suitable for gaming?" The server uses this text as input to a generative artificial intelligence model, preparing it for further encoding.
[0357] Output: The constructed model input text string, and the corresponding context structure record.
[0358] Step 8: The server encodes the input text for the model and sends it to the generative artificial intelligence model. The server uses a model-associated encoder to convert the input text into a sequence of tokens and vector representations that the model can process, and then performs forward operations.
[0359] Input: The model input text from step 7.
[0360] The server first calls a word segmenter or tokenizer to decompose the text into a sequence of token IDs, and then maps the token ID sequence into an embedding vector sequence. The server then calls the encoder and decoder modules of the transformer network on the graphics processing unit to perform multi-head self-attention, matrix multiplication, and nonlinear activation operations, calculating the output probability distribution. During the decoding phase, the server gradually generates the output token sequence according to the set beamwidth or sampling parameters until a termination token is generated or the length limit is reached.
[0361] Output: A sequence of labeled results generated by the model, along with the corresponding probability information, which is then converted into response content in natural language text format.
[0362] Step 9: The server performs post-processing and security filtering on the model's output text. The server performs formatting and content filtering on the raw text output by the generative artificial intelligence model to meet the requirements of the application scenario.
[0363] Input: The response text and probability information obtained in step 8.
[0364] The server removes redundant whitespace and unnecessary tags, and checks whether the text length meets the terminal display requirements. It uses a keyword filter list to check and remove unsuitable sensitive words or expressions that do not conform to business policies, replacing them with alternative expressions when necessary. The server can also adjust the tone based on user characteristics, such as providing more detailed explanations for new users and more concise descriptions for experienced users.
[0365] Output: The final response text that meets the display and policy requirements, as well as the response object that records the prompt statements and filter tags used.
[0366] Step 10: The server returns a response object to the terminal and records the dialogue log. The server packages the response object into a response message and sends it to the terminal through the network interface, and records the current round of dialogue in the data storage device.
[0367] Input: The final response text generated in step 9 and the user information and session information in the request structure.
[0368] The server writes a record to the dialogue log table, including the session identifier, round number, user question text, final response text, the prompt statement identifier used, and a timestamp. Subsequently, the server serializes the response text and necessary metadata into structured response data and returns it to the terminal over the network.
[0369] Output: The response message sent to the terminal, and a new dialogue log entry added on the server side.
[0370] Step 11: The terminal receives and displays the response text, and simultaneously collects the user's subsequent actions. The terminal receives the response message returned by the server from the network interface and displays the virtual character's response on the interface.
[0371] Input: The response message returned by the server in step 10.
[0372] The terminal parses the response data, extracts the response text, and displays it on the screen as a speech bubble. Optionally, it can call the local speech synthesis module to read the text aloud. Simultaneously, the terminal begins listening for user actions following the response, such as clicking on product details, scrolling through the page, or adding items to the cart, and assigns an identifier associated with the current response to each action.
[0373] Output: The response interface presented to the user, and a series of behavioral event objects generated locally on the terminal, waiting to be reported to the server.
[0374] Step 12: The terminal reports user behavior events, and the server records and stores them in a structured format. After a user performs an action or at the end of a session, the terminal sends the action events to the server in batches; the server parses them and records them in the action log table.
[0375] Input: A set of behavioral events generated locally on the terminal (including behavior type, occurrence time, associated response identifier, and user identifier).
[0376] The terminal packages these behavioral events into structured data and sends it to the server over the network. Upon receiving the data, the server parses each event, adds a server timestamp, and writes it to the behavior record table. The fields include user identifier, session identifier, behavior type, associated response identifier, and time information.
[0377] Output: Structured behavioral data records stored in the server behavior log table for subsequent analysis.
[0378] Step 13: The server performs user characteristic analysis when offline or during off-peak hours. The server periodically reads the conversation log table and behavior record table to update user profiles based on user characteristics.
[0379] Input: Accumulated dialogue logs and behavior records, and basic information from the user attribute record table.
[0380] The server uses a data analysis module to perform statistical calculations on each user, determining features such as visit frequency, average number of conversation rounds, and purchase conversion rate, and constructing a feature vector set. Subsequently, the server executes a clustering algorithm on this feature vector set, such as centroid-based clustering, to divide users into different characteristic groups. The server writes the clustering results back to the user attribute record table, updating the user characteristic labels.
[0381] Output: The updated set of user characteristic tags and the corresponding cluster center information, providing a basis for subsequent real-time selection of prompt statements.
[0382] Step 14: The server calculates metrics for prompt statements based on historical data and optimizes its strategies. The server performs statistical analysis and comparison on the performance of each prompt statement under different user characteristics and different product categories.
[0383] Inputs: prompt statements and response records in the dialogue log table, behavioral results associated with these responses in the behavior log table, and sales performance data.
[0384] The server uses join queries to link the prompt message identifier with the behavior results and order information, calculating sales performance metrics such as click-through rate, conversion rate, and average order amount, as well as response quality metrics such as session rounds and user dwell time. The server compares the differences in metrics between different prompt message versions, calculates a comprehensive score for each prompt message according to preset rules (such as a weighted scoring formula), and updates the scoring field and status field (such as recommended, pending testing, disabled) in the prompt message template table.
[0385] Output: A table of prompt templates with the latest ratings and statuses, used to select the best prompt in the next round of live sessions.
[0386] Step 15: The server generates new candidate prompts using a generative artificial intelligence model. When the server needs to expand its prompt statement library, it uses existing high-scoring prompts as examples and automatically generates new templates through a generative artificial intelligence model.
[0387] Input: A sample set of high-scoring prompt statements and the meta-prompt text defining the target to be generated.
[0388] The server concatenates the meta-prompt statement and multiple example statements into the model input, for example: "Based on the following high-performing sales dialogue samples, summarize a template of prompts suitable for promoting high-end smartphones. Please output 3 prompts in different styles, each no more than 80 characters, highlighting the high-end experience, camera capabilities, and gaming performance." The server encodes the input text into a labeled sequence and feeds it into a generative artificial intelligence model to obtain several candidate prompt texts. Subsequently, the server checks the length, sensitive words, and format using filtering rules, retaining only candidates that meet the specifications and writing them into a prompt template table, marking them as pending testing.
[0389] Output: A new set of prompt template records, containing text content and initial state, for use in subsequent sessions for A / B testing.
[0390] Step 16: The server applies optimized prompts in real-time sessions and continues to iterate. Each time the server processes a user request, it uses the updated prompt template table and user characteristic tags to select the optimal or test prompt, thus forming a continuously iterative closed loop.
[0391] Input: User characteristic tags for the current session, product or service category, latest rating and status information from the prompt statement template table.
[0392] The server filters a candidate set of prompt statements based on user characteristics and product categories, prioritizing those with high ratings and a "recommended" status. During strategy experimentation, it also selects prompt statements in a "to be tested" status according to a preset ratio. The server uses the selected prompt statements to construct the model input in step 7, generating new personalized responses. As new dialogue and behavioral data continuously flow in, the server continuously updates user characteristics and prompt statement ratings in steps 13 and 14, achieving automatic iterative optimization of prompt statements and response strategies.
[0393] Output: The result of dynamically adjusted prompts for different users and scenarios, and the adaptive strategy control process formed within the server.
[0394] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0395] Traditional user service systems, marketing systems, and content recommendation systems, while incorporating emotion recognition and recommendation algorithms, still suffer from the following limitations at the computer technology level: First, servers typically only perform rule-based analysis on structured business data, lacking the ability to abstract multi-source heterogeneous data (such as user behavior logs, facial images, voice signals, and feedback text) into a unified intermediate representation for inference processing, resulting in complex model call processes and poor scalability. Second, existing systems often directly input raw business parameters into fixed models for inference, lacking the ability to dynamically describe and analyze tasks and optimization goals using the strong expressive power of generative artificial intelligence models for natural language instructions (prompt statements), thus failing to adaptively generate multiple response logics, marketing strategies, and service improvement schemes within the same technical framework. Third, the control logic, response script generation logic, and strategy optimization logic of virtual representations (virtual characters) are mostly independent modules, maintained in a decentralized manner by manual rules or offline scripts. Servers struggle to automatically update the appearance settings, voice style, and response strategies of virtual representations based on real-time evaluation metrics during runtime, making it difficult for the system to continuously improve overall service quality and resource utilization efficiency.
[0396] Furthermore, when providing the aforementioned functions to the outside world in software form, traditional solutions often deploy the emotion recognition module, recommendation module, and business logic module as separate services. This is not conducive to using standardized software packaging and virtualization technology to encapsulate the entire computational process of "multimodal feature extraction - prompt statement construction - generative artificial intelligence model invocation - policy evaluation and update" into a technical solution that can be quickly deployed in different business environments. This limits the advantages of such systems in terms of computing resource reuse, cross-platform migration, and automatic scaling.
[0397] This invention solves the above-mentioned problems by providing a unified feature generation mechanism for multi-source user data on the server side, a response and strategy generation mechanism based on prompt statements-driven generative artificial intelligence models, and a computer implementation scheme that uses evaluation indicators and user feedback closed-loop to update virtual representations and response logic and provides them externally through virtualization packaging. This improves the system's ability to process multimodal data, the flexibility of model calling logic, and the adaptive optimization capabilities of services and strategies at the computer technology level.
[0398] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0399] In this invention, the server includes a device for generating a unified feature quantity for a specific user's emotional state and preference tendencies based on user attribute information, behavioral history information, and facial expression and voice information acquired by an image acquisition device and a voice acquisition device, and converting the feature quantity into a prompt statement in natural language form and inputting it into a generative artificial intelligence model to instruct the generation of response content corresponding to the user's attributes and emotions; and a device for controlling the vocal content and facial expressions of a virtual avatar on a display device using the response content output by the generative artificial intelligence model, thereby executing an individual response from the virtual avatar to the user, and collecting response result information related to the individual response. The device includes an apparatus for calculating evaluation metrics for comparing and assessing the effectiveness of multiple response modes based on response and sales results information, and converting the evaluation metrics and user feedback information into prompt statements in natural language form for inputting into a generative artificial intelligence model to instruct the generation of optimized sales promotion and response schemes, and updating the appearance settings, voice style, and response logic of the virtual representation accordingly; it also includes an apparatus for integrating multiple functional units, such as input / output processing of prompt statements, storage and processing of user-related information and response results information, and output processing of optimized sales promotion and response schemes, into a distributable package through virtualization technology. This allows for an integrated computing process within the server, encompassing "multimodal feature extraction—prompt statement construction—generative AI model invocation—result evaluation and strategy update—virtualization packaging." This enables the computer system to efficiently process heterogeneous data from images, voice, behavioral logs, and feedback text. It flexibly utilizes the natural language reasoning capabilities of generative AI models to dynamically generate various response content and marketing strategies. Furthermore, it achieves automatic iterative optimization of response logic and virtual representation configuration through evaluation metrics and user feedback. Simultaneously, it facilitates the deployment of this integrated process as a virtualized package to different information delivery environments, thereby improving the system's scalability, maintainability, and overall processing performance at the computer technology level.
[0400] "Information processing device" refers to an electronic device that includes at least one processor and memory, used to execute programs to acquire, process, store, and control the output of input data, and may be a server, terminal device, or other computing device.
[0401] A processor is a hardware circuit unit in an information processing device that executes a set of instructions to perform data processing, logical operations, and flow control. It may include a central processing unit, a graphics processing unit, or other arithmetic units.
[0402] "User-related information" refers to a set of data related to users that characterizes user features and behavioral patterns, including but not limited to user attribute information, behavioral history information, purchase records, and service usage records.
[0403] "User attribute information" refers to structured information used to represent the basic characteristics of users, including but not limited to age group, gender category, regional category, preference category, and price sensitivity category.
[0404] "Behavioral history information" refers to historical behavioral data generated based on users' operation records in physical or virtual environments, including but not limited to browsing records, dwell time, click operations, purchase behavior, and interaction logs.
[0405] "Image acquisition device" refers to a hardware device used to acquire image data of users or the environment, including but not limited to cameras, image sensors, and terminal components with imaging functions.
[0406] "Voice acquisition device" refers to a hardware device used to acquire user voice signals, including but not limited to microphones, microphone arrays, and terminal components with voice input functions.
[0407] "Facial expression information" refers to image data or its analysis results that are acquired by an image acquisition device and processed to represent the state of a user's facial expressions, including but not limited to emotion-related features such as smiling, confusion, anger, and calmness.
[0408] "Speech information" refers to the original speech signal and its analysis results acquired by the speech acquisition device, including but not limited to speech waveforms, speech-to-text transcripts, and acoustic features such as tone, volume, and speech rate.
[0409] "Emotional state" refers to the category and intensity of a user's current or recent emotions, inferred from facial expressions, voice information, and / or other behavioral cues, including but not limited to positive, neutral, negative, and their subcategories.
[0410] "Preference tendency" refers to the user's interest direction and selection tendency in terms of goods, services or content inferred based on user-related information and historical behavior data, including but not limited to category preference, style preference and price preference.
[0411] "Features" refer to quantitative or qualitative feature representations extracted from user-related information, facial expression information, and voice information and used for computation and processing, including numerical vectors, category labels, or embedding representations.
[0412] "Natural language prompts" refer to text instructions written in natural language and constructed by a processor. These instructions describe the input data, analysis objectives, and output requirements to a generative artificial intelligence model, guiding the model to generate corresponding results.
[0413] "Generative artificial intelligence models" refer to artificial intelligence models that can automatically generate text, speech or other data outputs based on input prompts and related data, including but not limited to language generation models and multimodal generation models.
[0414] "Response content" refers to the output information generated by the generative artificial intelligence model based on the prompts, used for interaction with the user, including but not limited to dialogue text, explanatory statements, recommendations, and guiding phrases.
[0415] "Display device" refers to an output device used to visually present images, text or virtual objects, including but not limited to displays, projection devices and head-mounted displays.
[0416] "Virtual representation" refers to a virtual image or virtual character presented on a display device for interaction with the user, including its appearance, expressions, actions, and associated voice output.
[0417] "Individual response" refers to the personalized response behavior generated and output by the virtual representation for a single user or a single interactive session, including voice output, facial expression changes, and action feedback.
[0418] "Response result information" refers to the result data associated with the virtual entity performing an individual response process, including but not limited to user response, interaction duration, click behavior, and dialogue content records.
[0419] "Sales results information" refers to outcome data related to the effectiveness of individual responses or a series of response behaviors in a sales activity, including but not limited to whether a transaction was completed, order amount, conversion rate, and uppurchase.
[0420] "Evaluation metrics" refer to measures used to quantitatively or qualitatively assess the effectiveness of multiple response patterns or strategies, including but not limited to conversion rate, average revenue, satisfaction score, and retention rate.
[0421] "User feedback information" refers to data provided directly or indirectly by users that reflects their feelings and opinions, including but not limited to ratings, comment texts, questionnaire responses, and behavioral feedback.
[0422] "Sales promotion plan" refers to a strategy or plan generated based on evaluation indicators and user feedback information to improve sales performance or promotion efficiency, including but not limited to discount packages, referral rules, and activity processes.
[0423] "Response plan" refers to a set of response strategies designed for the user interaction process, including but not limited to script templates, dialogue flow and response rules.
[0424] "Optimization plan" refers to the adjustment suggestions and new plans output by the generative artificial intelligence model based on evaluation indicators and user feedback information, which are used to improve sales promotion plans and response plans.
[0425] "Appearance settings" refers to configuration parameters used to define the visual presentation characteristics of a virtual representation, including but not limited to appearance style, clothing style, color scheme, and facial expression baseline.
[0426] "Voice style" refers to the comprehensive characteristics set in terms of timbre, speech rate, tone, and emotional expression when a virtual entity outputs speech.
[0427] "Response logic" refers to the set of rules or decision-making process that controls how a virtual representation selects and generates response content under different input conditions, including policy trees, dialogue state machines, and policy models.
[0428] "External information provider" refers to a business entity that provides information or content services to end users in its business environment outside of this system by using the program packages provided by this system.
[0429] A "package" is a collection of software units consisting of multiple functional units that can be deployed and executed in a target runtime environment, including executable programs, configuration files, and interface definitions.
[0430] "Emotion-specific models" refer to computational models used to infer a user's emotional state based on input facial expressions, voice information, or other behavioral data, including classification models based on machine learning or deep learning.
[0431] "Virtualization technology" refers to the technology of encapsulating computing resources or functional units into virtual resources that can be deployed and managed independently through software abstraction, including but not limited to virtual machine technology and container technology.
[0432] "Functional unit" refers to a logical module or software component that undertakes a specific processing task in the system, including data acquisition module, storage module, model interface module, and strategy output module, etc.
[0433] In the embodiments of this invention, the server, terminal, and user each undertake different technical functions, and through collaborative work, achieve personalized responses and strategy optimization based on multi-source data and generative artificial intelligence models. The various embodiments of this invention can be implemented on different hardware and software platforms; for example, the server side can use a computer with a multi-core central processing unit and a graphics processing unit, and the terminal side can use a mobile terminal or a fixed terminal equipped with a camera and microphone.
[0434] In one embodiment, the server includes at least one computer hardware unit equipped with a graphics processing unit (GPU). The server uses an operating system, database management software, and numerical computation libraries to execute programs. In another specific example, the server employs a combination of a general-purpose processor and a GPU, using the GPU to accelerate forward inference and training updates of deep neural networks. At the software level, the server can use a database management system to store user-related information and response results. The server uses a data analysis library (e.g., a data framework library) for structured data processing and a deep learning framework (e.g., a tensor computation framework) to implement sentiment-specific models and feature extraction networks.
[0435] In one embodiment, the terminal includes an electronic device equipped with a front-facing camera, microphone, display panel, and local processor. The terminal runs an operating system and a graphics rendering engine on software, and can run audio / video capture components, facial expression recognition components, and voice recognition components. The terminal can perform preliminary processing of the captured image and voice signals locally to reduce the amount of data sent to the server, thereby reducing the communication load.
[0436] In this invention, users interact with the system through a terminal. Users can engage in natural language conversations using the terminal's camera and microphone. Users can select the appearance and voice style of the virtual avatar via a touchscreen or physical buttons. Users can also input feedback information on the terminal, which is then used by the server for evaluation and optimization.
[0437] In one embodiment, the server generates a program for executing the present invention. The server defines multiple functional modules within the program, including a data acquisition interface module, a feature generation module, a prompt statement generation module, a generative artificial intelligence model interface module, an evaluation index calculation module, a strategy optimization module, and a virtualization packaging module. The server specifies the data structures, algorithm flows, and inter-module relationships within these modules to form an executable overall system.
[0438] The server employs a multimodal feature extraction network for feature generation. In one specific implementation, the server inputs facial expression images from the terminal into a convolutional neural network. This network contains multiple convolutional layers, pooling layers, and fully connected layers, from which the server extracts a fixed-length facial expression feature vector. Simultaneously, the server inputs speech signals acquired by a speech acquisition device into a speech feature extraction network. This network uses one-dimensional convolution or short-time Fourier transform to generate acoustic feature maps, which are then encoded into speech feature vectors through an attention mechanism or recurrent units. The server concatenates these visual and speech features with embedded representations of user attribute information and behavioral history information to form a unified high-dimensional feature set.
[0439] The server employs a combination of fixed format and variable parts in generating prompts. Based on the numerical values representing user attributes, emotional state, and preferences in high-dimensional features, the server uses rule templates to generate natural language descriptions. In one example, the server constructs the following prompt: "Current user attributes: male, age 30-39; historical preferences: sports equipment and wearable devices; current emotional state: slightly confused, currently looking at smartwatches. Based on the above information, please generate a sales consultant's script of no more than 120 words, requiring you to first inquire about the usage scenario, then explain the two core functions, and maintain a professional and friendly tone." In another example, the server constructs the following prompt statement to optimize the response strategy: "The following are the statistical results of three response script schemes in the most recent 1000 interactions: Script A, conversion rate 15%, average revenue 200; Script B, conversion rate 20%, average revenue 180; Script C, conversion rate 12%, average revenue 250. Please analyze the suitability of each scheme for users of different age groups and price sensitivity, provide adjustment suggestions, and generate a new combined response scheme." The server uses the aforementioned prompts as text input in the generative artificial intelligence model interface module. In one implementation, the server employs a language generation model based on a self-attention structure, which includes multi-layered encoding and decoding structures. The server sends prompts and parameters (such as maximum generation length, sampling temperature, etc.) to the model via a remote interface, and receives the response content generated by the model. After receiving the output, the server performs post-processing, including illegal word filtering, format normalization, and style correction, to generate response content suitable for direct use in virtual representations.
[0440] The server uses structured response and sales results information for calculating evaluation metrics. In the database, the server assigns a response identifier, virtual representation settings, input feature summaries, output text from the generative AI model, user behavior results (clicks, dwell time, purchases, etc.), and user feedback information for each individual response record. The server performs aggregation operations on this data through a data analysis library, calculating evaluation metrics such as conversion rate, average revenue, average satisfaction rating, and negative feedback ratio. The server then uses these evaluation metrics to construct prompts with specific numerical values, which are then input again into the generative AI model to obtain optimization solutions.
[0441] In a specific instance, the server constructs the following prompt statement: "Based on the response records of the past month, the conversion rate of Plan X is 18%, and the negative feedback rate is 5%; the conversion rate of Plan Y is 21%, and the negative feedback rate is 12%. Please provide suggestions for reducing the negative feedback rate while ensuring that the conversion rate is not lower than 20%, and provide two rewritten response scripts." In the strategy optimization module, the server parses the output of the generative AI model into structured strategy data. For example, it converts key points in the text, such as "prioritize asking about needs after self-introduction" and "avoid using a strong, intimidating tone," into rule parameters. The server then updates the label weights, word choice probabilities, and decision tree branch thresholds in the response logic, thereby automatically adopting the optimized strategy in subsequent interactions.
[0442] The terminal receives response content and auxiliary control parameters from the server for virtual representation control. The terminal uses a graphics rendering engine to drive the virtual representation's skeleton and facial mesh, adjusting the virtual representation's expressions and postures based on emotion tags and motion commands returned by the server. The terminal converts the response content into audio signals using a text-to-speech engine and plays them synchronously with the virtual representation's lip-sync animation, achieving natural speech and motion output. The terminal simultaneously displays the text response content and related recommendations in the user interface, allowing users to select via touch or cursor operation.
[0443] While using the system, users can adjust the appearance and voice style of the virtual avatar through the terminal interface. For example, users can choose a "formal style," "lively style," or "composed style," and set their preferred speaking speed and tone. The terminal sends these settings to the server, which adds restrictive descriptions to subsequent prompts, such as "Please answer with a slower speaking speed and composed tone," thus matching the output of the generative AI model with the user's preferences.
[0444] In terms of virtualization packaging, the server encapsulates each functional module into a deployable unit. Using virtualization technology, the server divides the data storage module, feature generation module, prompt statement generation module, generative artificial intelligence model interface module, and policy optimization module into multiple containers or virtual instances. During packaging, the server provides a unified interface specification for external information providers. These providers only need to deploy the package in their business environment and provide user-related information and terminal data to utilize the complete processing flow of this invention. This configuration allows the system to be easily migrated to different hardware environments and business scenarios, while simultaneously achieving elastic scaling and centralized management of computing resources.
[0445] In terms of technical effectiveness, the server significantly improves the efficiency of data management and inference within the computer by combining unified multimodal feature quantities with natural language prompts. Instead of building separate rule engines for each business scenario, the server encodes analysis objectives and constraints as prompts, enabling generative AI models to perform multiple tasks under a single interface. This design simplifies the program structure, reduces the need to write and maintain large amounts of rule code for different scenarios, and thus improves the system's maintainability and scalability.
[0446] By utilizing structures such as convolutional neural networks and attention mechanisms during feature generation, the server can compress and retain key information in a lower-dimensional feature space, reducing the amount of data that needs to be transmitted and stored during the model invocation phase, thereby lowering communication load and storage costs. Through evaluation metric-driven closed-loop optimization, the server automates the response logic update process, enabling the system to continuously improve the relevance and interactivity of response content during operation. From a computer technology perspective, this automatic parameter update mechanism based on evaluation metrics and feedback reduces the frequency of manual intervention and script updates, improving the overall system's ability to maintain performance consistency and stability in large-scale user environments.
[0447] The server can employ different emotion-specific model structures in various implementation variations. For example, the server can use a bidirectional recurrent neural network or a transformer encoder model to classify emotions in speech-to-text transcription, and it can use a network based on convolutional and residual structures to extract features from facial expression images. The server can train these models through supervised learning, using cross-entropy loss or mean squared error loss as the error function, and updating the model weights with gradients based on the target label during training. The server can also improve the robustness of the model in complex environments through data augmentation methods, such as rotating, cropping, and changing the lighting of images, and adding noise and time stretching to speech. These technical details help improve the accuracy of emotion state inference, enabling generative AI models to fully utilize more reliable emotional information in subsequent response generation.
[0448] Terminals can also employ various alternative methods for data preprocessing. They can run lightweight facial expression recognition models locally to output discrete emotion labels, reducing the frequency of transmitting raw images and thus lowering bandwidth consumption. Alternatively, terminals can send updated data to the server only when the user's emotion changes significantly, while reducing the reporting frequency when the emotion is stable. This further reduces network communication load and improves system response performance in high-concurrency scenarios.
[0449] Users can also utilize the system of this invention in other application scenarios. For example, in a communication service environment, a user initiates a package consultation request through a terminal interface. The terminal sends the user's historical usage records and the current consultation text to the server. The server extracts features and generates prompts from this data, and requests a suitable package suggestion and explanation from a generative artificial intelligence model. After receiving the results, the server displays package details and explanatory text through the terminal, allowing the user to quickly understand the options and provide feedback and ratings. Based on a large number of similar conversations, the server calculates evaluation indicators for different package explanation strategies and guides the generative artificial intelligence model to produce better explanations through prompts, thereby improving the explanation effect and user comprehension.
[0450] The system of this invention is not limited to sales or customer service scenarios; the server can also apply the same technical framework to various environments such as content recommendation, educational tutoring, and remote medical assistance. In these environments, the server drives a generative artificial intelligence model to output scenario-related response content and strategy suggestions through the same set of multimodal feature generation and prompt statement generation mechanisms. Therefore, this system provides a unified, flexible, and scalable architecture at the computer technology level, enabling efficient data processing, intelligent response, and strategy optimization capabilities across multiple business domains using the same technical foundation.
[0451] use Figure 14 The processing flow is explained.
[0452] Step 1: The terminal acquires the user's multimodal raw data.
[0453] Input: User's facial image, voice signal, and basic interaction events (such as clicks and browsing stops).
[0454] The terminal activates its camera to capture multiple frames of the user's facial images, encoding the images into a compressed format. Simultaneously, the terminal activates its microphone to capture audio, buffering the audio as an audio data stream. The terminal records the user's click locations on the interface, the page IDs viewed, and the duration of each click, generating timestamps. The terminal then resizes the images, standardizes the audio sampling rate, and packages the data into a message object containing image frames, audio segments, and interaction events.
[0455] Output: A multimodal raw data packet containing image data, audio data, and interaction event logs.
[0456] Step 2: The terminal performs local feature preprocessing and compression.
[0457] Input: Multimodal raw data packets.
[0458] The terminal uses a local lightweight convolutional network to calculate facial key points and basic expression categories from image data, generating expression labels and low-dimensional feature vectors. The terminal uses a local speech feature extraction algorithm (such as Mel frequency cepstral coefficient calculation) to convert audio into acoustic feature vectors, and simultaneously calls the speech recognition module to transcribe speech into text. The terminal combines expression features, acoustic features, transcribed text, and interaction event logs, and discards redundant frames to compress the data volume.
[0459] Output: A preprocessed data package containing facial feature vectors, acoustic feature vectors, speech-to-text transcription, and summaries of interactive events.
[0460] Step 3: The terminal sends preprocessed data to the server and waits for a response.
[0461] Input: Preprocessed data packet.
[0462] The terminal sends preprocessed data packets to the server through a secure communication channel (such as an HTTP interface based on an encrypted protocol). The terminal attaches a device identifier and a user anonymity identifier to the message header. The terminal maintains a request ID locally for association when it receives a response from the server.
[0463] Output: The data request sent to the server and the request ID recorded locally.
[0464] Step 4: The server receives preprocessed data and associates it with user history information.
[0465] Input: Preprocessed data packet from the terminal, user anonymous identifier.
[0466] The server parses the request message, reading facial features, acoustic features, transcribed text, and interaction event summaries. The server uses the user's anonymous identifier to perform a query in the database, reading the user's attribute information (age group, gender category, etc.) and behavioral history information (purchase records, content browsing records, etc.). The server merges the current features with the historical information to construct a unified data structure.
[0467] Output: A unified feature dataset containing current multimodal features and historical user-related information.
[0468] Step 5: The server generates uniform feature quantities through deep networks.
[0469] Input: Uniform feature dataset.
[0470] The server inputs facial expression features into a convolutional neural network, and inputs acoustic features and transcribed text into a sequence encoding network or transformer encoder. User attributes and behavioral history are encoded into embedding vectors. The server connects and linearly transforms the above vectors in the feature fusion layer, and obtains high-dimensional unified feature quantities through non-linear activation functions. In this process, the server uses weight matrices and bias vectors to perform matrix multiplication and addition operations to achieve weighted fusion of information from various modalities.
[0471] Output: A high-dimensional unified feature quantity describing the user's current emotional state and preference tendencies.
[0472] Step 6: The server estimates the sentiment state and preference labels.
[0473] Input: High-dimensional uniform features.
[0474] The server inputs uniform feature values into a sentiment-specific model (e.g., a classification network with a fully connected output layer), calculates the probability distribution of each sentiment category, and selects the category with the highest probability as the sentiment state label. The server also performs threshold judgment on output nodes related to product category, content category, or service type to generate one or more preference labels. The server saves the sentiment state and preference labels for use in generating subsequent prompts.
[0475] Output: sentiment state labels, preference label sets, and corresponding confidence scores.
[0476] Step 7: The server constructs natural language prompts.
[0477] Inputs: Uniform features, sentiment state labels, preference labels, user attribute information, and behavioral history information.
[0478] The server converts numerical features and categorical labels into readable text based on a predefined template and inserts them into placeholder positions within the template. The server explicitly includes the current user's attributes, historical preferences, current sentiment state, and the target response in the prompt statement. For example, the server generates the following prompt statement: "Current user attributes: male, age 30-39; historical preferences: sports equipment and wearable devices; current emotional state: slightly confused, currently looking at smartwatches. Based on the above information, please generate a sales consultant's script of no more than 120 words, requiring you to first inquire about the usage scenario, then explain the two core functions, and maintain a professional and friendly tone." The server performs length checks and sensitive word filtering on the fields in the prompt statements to ensure they are suitable for input generative artificial intelligence models.
[0479] Output: Natural language prompts for generative artificial intelligence models.
[0480] Step 8: The server invokes a generative artificial intelligence model to generate response content.
[0481] Input: Natural language prompts.
[0482] The server sends prompts and control parameters (maximum generation length, sampling strategy, etc.) to the generative AI model through the model interface; the server performs forward inference on the model side, encodes the prompts using a multi-layer self-attention structure, and then generates the response text word by word; after receiving the complete response text, the server performs post-processing operations such as sentence segmentation, removal of redundant words, and standardization of wording; the server restricts the output to prevent specific violations when necessary.
[0483] Output: Structured response text and associated control information (such as suggested tone and emotion tags).
[0484] Step 9: The server generates virtual representation control commands and sends them to the terminal.
[0485] Input: Response text, control information, and emotional state tags.
[0486] The server maps the response text and emotion tags to specific lip-sync rhythm, facial expression intensity, and gesture instructions, generating a set of control instructions that describe the virtual avatar's vocal content, facial expressions, and behavioral cues. The server encapsulates the response text and control instructions into a response message, sends it back to the requesting terminal over the network, and includes the original request ID for the terminal to match.
[0487] Output: A response message sent to the terminal containing the response text and virtual representation control instructions.
[0488] Step 10: The terminal driver virtual representation outputs a response.
[0489] Input: The response text returned by the server and the virtual representation control instructions.
[0490] The terminal uses a text-to-speech engine to convert the response text into audio data. Based on control commands, the terminal drives the facial and body skeletons of the virtual representation in the graphics rendering engine, setting expression weights, posture parameters, and lip-sync animation. The terminal synchronously plays the generated speech, aligning it with the lip-sync timeline, so that the virtual representation responds to the user with corresponding expressions and voices on the display device. The terminal displays the response text and related options in parallel in the user interface.
[0491] Output: The voice and actions of the virtual representation presented on the display device, as well as the response text and interactive controls on the interface.
[0492] Step 11: Users can further interact with the virtual representation and provide feedback.
[0493] Input: The virtual representation's voice output, action performance, and interface prompts.
[0494] Users can ask questions or express opinions to the virtual avatar via voice. Users can select recommended items or adjust the appearance settings of the virtual avatar via the touch screen (such as selecting "calm style" or "lively style"). Users can also rate the response and fill in text feedback on the terminal.
[0495] Outputs: New voice input, interactive events, and rating and text feedback data.
[0496] Step 12: The terminal collects feedback data and reports it to the server.
[0497] Inputs: user ratings, feedback text, interaction events, request IDs.
[0498] The terminal packages the user-entered rating value, text feedback, and subsequent click behavior records together with the original request ID to form a feedback data object; the terminal sends the feedback data object to the server through the network interface so that the server can update the evaluation indicators and strategies.
[0499] Output: The feedback data request sent to the server.
[0500] Step 13: The server updates the response and sales results information.
[0501] Input: Feedback data object, historical response records.
[0502] The server retrieves the corresponding response record from the database based on the request ID, and writes the user rating, feedback text, and subsequent purchase information into the record. The server updates the response result tags (e.g., satisfied, dissatisfied) and sales result fields (transaction amount, conversion status) for each record. The server archives these records according to time range or batch for subsequent analysis.
[0503] Output: An updated record set containing complete response and sales results information.
[0504] Step 14: The server calculates evaluation metrics and constructs optimization prompts.
[0505] Input: Update recordset.
[0506] The server uses a data analysis library to group the record set by script scheme, virtual representation configuration, and user attributes, calculating evaluation metrics such as conversion rate, average revenue, average satisfaction rate, and negative feedback ratio for each group. The server combines these metrics with example record summaries to form a text description, inserts it into an optimization template, and generates new prompts, such as: "Based on the response records of the past month, the conversion rate of Plan X is 18%, and the negative feedback rate is 5%; the conversion rate of Plan Y is 21%, and the negative feedback rate is 12%. Please provide suggestions for reducing the negative feedback rate while ensuring that the conversion rate is not lower than 20%, and provide two rewritten response scripts." Output: Natural language prompts for policy optimization and corresponding evaluation metrics.
[0507] Step 15: The server invokes a generative artificial intelligence model to generate an optimization solution.
[0508] Input: A prompt statement used for strategy optimization.
[0509] The server again sends the optimized prompts to the generative AI model via the interface, requesting a comparison of different response schemes and the generation of improvement suggestions. The server extracts key strategies from the model's inference results, such as adjusting the response order, changing the tone, and rules for selecting recommended content. The server parses these strategies into a set of executable rules and parameter configurations.
[0510] Output: A structured optimization solution, including new response scripts, response logic rules, and virtual representation configuration suggestions.
[0511] Step 16: The server updates the response logic and virtual representation configuration, and then packages the program.
[0512] Input: Structured optimization scheme, existing response logic, and virtual representation configuration.
[0513] The server modifies the decision tree, rule weights, and dialogue templates in the response logic according to the optimization scheme, and updates the default appearance settings and voice style parameters of the virtual representation. The server packages the updated modules and other functional modules into a distributable program package using virtualization technology, which can be deployed and used by external information providers in their business environments. During the packaging process, the server generates interface specification documents to define the data format and calling protocol between the terminal and the server.
[0514] Output: Updated response logic and virtual representation configuration, and deployable package.
[0515] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0516] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0517] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0518] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0519] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0520] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0521] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0522] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0523] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0524] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0525] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0526] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0527] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0528] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0529] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0530] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0531] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0532] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0533] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0534] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0535] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0536] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0537] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0538] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0539] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0540] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0541] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0542] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0543] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0544] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0545] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0546] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0547] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0548] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0549] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0550] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0551] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0552] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0553] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0554] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0555] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0556] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0557] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0558] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0559] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0560] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0561] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0562] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0563] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0564] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0565] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0566] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0567] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0568] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0569] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0570] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0571] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0572] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0573] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0574] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0575] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0576] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0577] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0578] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0579] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0580] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0581] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0582] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0583] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0584] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0585] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0586] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0587] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0588] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0589] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0590] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0591] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0592] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0593] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0594] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0595] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0596] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0597] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0598] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0599] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0600] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0601] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0602] In addition, the following notes are provided in response to the above explanation.
[0603] Example 1 (Note 1) An information processing system, characterized in that it comprises: The user attribute generation unit, executed by the computing device, is used to obtain user behavior information from the information storage device based on user identification information, and to preprocess and summarize the behavior information through a general information processing program, thereby generating feature information to represent user attribute information. The prompt statement generation unit is used to construct prompt statements for inputting into the generative artificial intelligence model based on the feature information and reception condition information, instruct the generative artificial intelligence model to generate individual response content, and obtain the output result of the generative artificial intelligence model as reception information; The reception control unit is used to send virtual representation display control information and voice playback information containing the reception information to the terminal device, so that the terminal device can perform reception using the virtual representation through the visual output device and the auditory output device. The virtual representation optimization unit is used to analyze multiple virtual representation attribute candidates and past reception performance information based on the virtual representation change request information issued by the user obtained from the terminal device, determine the virtual representation attribute that is presumed to have a high reception effect, and generate new reception information corresponding to the attribute using the prompt statement input to the generative artificial intelligence model. The sales strategy generation unit is used to collect dialogue history information and sales performance information obtained from the terminal device as parsing object information, generate prompt statements to instruct the parsing object information and input them into the generative artificial intelligence model, thereby obtaining sales promotion policy information, and updating the generation conditions of the prompt statements and reception information based on the sales promotion policy information. The service improvement unit is used to generate prompt statements that are input to the generative artificial intelligence model from the evaluation information obtained from the user as the parsing object information, and to change the response content and reception control conditions of the virtual representation based on the parsing results of the generative artificial intelligence model.
[0604] (Note 2) According to the information processing system described in Appendix 1, the computing device is configured to perform statistical processing and feature extraction processing on the behavioral information, including purchase history information, browsing behavior information, dwell time information and visit frequency information, through the general program to generate the feature information, and to include the generated feature information in the prompt statement and input it into the generative artificial intelligence model.
[0605] (Note 3) According to the information processing system described in Appendix 1, the computing device is configured to store the dialogue history information and sales performance information as historical information corresponding to the types of prompt statements, virtual representation attributes, recommended product groups, and transaction results, and to automatically update the virtual representation attributes and the composition policy of the prompt statements by including the historical information in the prompt statements input to the generative artificial intelligence model for parsing.
[0606] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: This is a control method used to acquire user behavior information and transaction information in an information processing device, and to summarize the behavior information and transaction information using data processing and data analysis programs to generate feature quantities representing attribute information and preference information for each user. Control means for generating prompt statements in natural language based on the aforementioned feature quantities and product or marketing information, and inputting the prompt statements into a generative artificial intelligence model to generate response text containing product recommendation text or sales promotion information for a single user; Control means for obtaining setting information related to the appearance and voice information of virtual display elements from the user, reflecting the setting information in the prompt statement, and generating the response text in a style or expression that corresponds to the appearance and voice information of the virtual display elements; Control means for sending the response text to a mobile information terminal via communication means, and displaying the response text to the user on the mobile information terminal in the form of a dialogue of the virtual display element or by voice output; This is a control method used to obtain user operation information and transaction information from the mobile information terminal after the response text prompt, calculate reception quality indicators and sales performance indicators using statistical processing programs, and update the content or structure of subsequent prompt statements input into the generative artificial intelligence model based on the reception quality indicators and sales performance indicators.
[0607] (Note 2) The information processing system according to Appendix 1 is characterized in that, When acquiring the behavior information and the transaction information, the information processing device obtains display history information, location-related information, and operation history information from the mobile information terminal, and uses tabular data processing software as the data analysis program to comprehensively process the display history information, the location-related information, and the operation history information to generate the feature quantity, and embeds the user attribute information containing the feature quantity into the prompt statement.
[0608] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device stores the correspondence between the reception quality indicators and the sales performance indicators and the prompt statements in a storage device, selects a prompt statement template from multiple types of prompt statement templates that makes the reception quality indicators or the sales performance indicators meet a predetermined benchmark, and reflects the feature quantity and the setting information into the prompt statement template to automatically generate the prompt statement.
[0609] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring attribute and behavioral information of an object, generating instruction information for a generative artificial intelligence model to parse the attribute and behavioral information to identify the characteristics of the object, and inputting the instruction information into the generative artificial intelligence model; A device for acquiring dialogue information from a dialogue device, structurally recording the dialogue information and the response information output by the generative artificial intelligence model, generating instruction information based on the structured recording information to enable the generative artificial intelligence model to perform evaluation and analysis processing related to sales activities, and inputting the instruction information into the generative artificial intelligence model. An apparatus for generating prompt statements for input to the generative artificial intelligence model based on the results obtained from the evaluation process and the analysis process, generating input information that combines the prompt statements with the dialogue information, and enabling the generative artificial intelligence model to perform response information generation processing based on the input information; An apparatus for selecting the prompt statement according to the characteristics of the user and the category of goods or services based on the results obtained from the evaluation process and the analysis process, comparing the sales results of multiple prompt statements to perform verification processing, and optimizing the sales strategy based on the comparison results; A means for updating the usage conditions of the prompt statements for the generative artificial intelligence model according to the optimized sales strategy, and applying the updated prompt statements to the dialogue processing in the dialogue device.
[0610] (Note 2) The information processing system according to Appendix 1 is characterized in that, The device for acquiring dialogue information is configured to record dialogue information acquired from the dialogue device in chronological order with the behavioral information of the user, and to divide the user into multiple characteristic groups based on the records, generate different prompt statements for each characteristic group and input them into the generative artificial intelligence model, thereby enabling the generative artificial intelligence model to perform individualized response processing for the user.
[0611] (Note 3) The information processing system according to Appendix 1 is characterized in that, The evaluation process and the analysis process are configured to associate the response information with the behavioral results of the user, calculate sales performance indicators and response quality indicators, select or generate prompt statements from multiple candidate prompt statements based on the sales performance indicators and the response quality indicators, perform dialogue processing using the selected or generated prompt statements, and repeatedly update the sales strategy through the dialogue processing.
[0612] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: An apparatus for generating, by a processor in an information processing device, a feature quantity for a specific user’s emotional state and preference tendency based on user-related information including user attribute information and behavioral history information, as well as facial expression information and voice information acquired by an image acquisition device and a voice acquisition device, and converting the feature quantity into a prompt statement in natural language form and inputting it into a generative artificial intelligence model to instruct the generative artificial intelligence model to generate response content corresponding to the user’s attributes and emotions. A device for controlling the vocal content and facial expressions of a virtual entity on a display device using the response content output by the generative artificial intelligence model, thereby executing individual responses from the virtual entity to the user, and collecting response result information and sales result information related to the individual responses; An apparatus for calculating evaluation indicators for comparing and evaluating the effectiveness of multiple response modes based on the response result information and the sales result information, and inputting the calculated evaluation indicators and user feedback information as prompts in natural language form into the generative artificial intelligence model to instruct the generative artificial intelligence model to generate optimized solutions for sales promotion plans and response plans; A device for updating the appearance settings, voice style, and response logic of the virtual representation based on the optimization scheme, and configuring the updated response logic into a package that can be provided to an external information provider.
[0613] (Note 2) The information processing system according to Appendix 1 is characterized in that, The processor is configured to input user facial expression information and voice information acquired by the image acquisition device and the voice acquisition device into an emotion-specific model to estimate the emotional state, and combine the estimated emotional state with the user-related information to generate the prompt statement and input it into the generative artificial intelligence model, thereby generating multiple candidate response content corresponding to user attributes and emotions, and selecting the response content based on the evaluation index for the virtual avatar's response.
[0614] (Note 3) The information processing system according to Appendix 1 is characterized in that, The processor is configured to enable the program to execute in the business environment of an external information provider. It integrates multiple functional units, including processing input and output of prompt statements to the generative artificial intelligence model, processing of storing user-related information and response result information, and processing of outputting optimized solutions for the sales promotion plan and the response plan, through virtualization technology, thereby forming the multiple functional units into a distributable program package.
Claims
1. An information processing system, characterized in that, include: processor; The processor is configured as follows: The prompt message indicating the collection and analysis of the user's purchase history data and behavioral pattern data is input into the generative artificial intelligence model so that the generative artificial intelligence model can analyze the data; The prompt information used to instruct the response results of the virtual character to be analyzed and the sales promotion strategy to be optimized based on the analysis results is input into the generative artificial intelligence model, so that the generative artificial intelligence model analyzes the response results of the virtual character and outputs the results for optimizing the sales promotion strategy. The prompt message indicating that feedback information from users should be collected and parsed and the parsing results should be used for service improvement is input into the generative artificial intelligence model, so that the generative artificial intelligence model can parse the feedback information and output the results for service improvement.
2. The information processing system according to claim 1, characterized in that, The processor is configured to generate prompts indicating the collection and analysis of information such as the user's purchase history and behavioral patterns, and to input the prompts into the generative artificial intelligence model.
3. The information processing system according to claim 1, characterized in that, The processor is configured to generate prompts indicating the collection and analysis of response results information for virtual characters, input the prompts into the generative artificial intelligence model for analysis, and optimize sales promotion strategies based on the analysis results output by the generative artificial intelligence model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A