Information processing system
Patent Information
- Application Number
- CN202610250912.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-03
- Publication Date
- 2026-09-22
AI Technical Summary
然而,现有的日记记录方式多依赖用户主动编辑和整理,存在以下问题:其一,用户往往缺乏持续记录的动力,难以长时间保持书写日记的习惯;其二,即便用户输入了一定量的文本信息,传统系统也仅作简单存储,无法对用户的情绪状态进行深入分析,难以生成具有情感洞察价值的记录内容;其三,现有对话式或问答式系统通常根据预设规则或固定脚本生成提问,难以及时根据用户实时情绪状态调整问题内容与难度,导致交互体验不足,无法充分引导用户表达内心想法
服务器通过将图像信息和位置信息等物理世界感知数据与用户文本输入数据在同一数据结构中进行关联,使抽象的语言生成过程直接与实际环境场景绑定。服务器利用这种多模态融合处理对记录媒体进行写入,提高了记录内容对现实情境的还原度,在例如实体场所反馈收集、环境体验日志、出行记录等应用中具有实际技术用途。
Smart Images

Figure CN122797463A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot speech in response to the user's speech.
[0003] With the development of social networks, mobile devices, and generative artificial intelligence technologies, users generate a large amount of textual information related to their personal emotions, experiences, and thoughts in their daily lives. However, existing diary recording methods mostly rely on users actively editing and organizing, which has the following problems: First, users often lack the motivation to keep recording and find it difficult to maintain the habit of writing a diary for a long time; second, even if users input a certain amount of text information, traditional systems only store it simply and cannot conduct in-depth analysis of the user's emotional state, making it difficult to generate records with emotional insight value; third, existing conversational or question-and-answer systems usually generate questions based on preset rules or fixed scripts, making it difficult to adjust the content and difficulty of questions in a timely manner according to the user's real-time emotional state, resulting in insufficient interactive experience and failing to fully guide users to express their inner thoughts.
[0004] Therefore, it is necessary to provide a system that can automatically parse user input information, assess user emotional state, and dynamically generate subsequent questions based on the emotional analysis results, so as to improve the personalization and targeting of question-and-answer interaction, enhance users' willingness to express themselves, and thus facilitate the generation of diary content that is more in line with users' true emotions and experiences. Summary of the Invention
[0005] To address the aforementioned issues, this invention provides an information processing system comprising a processor configured to: provide an interface for receiving user information, thereby acquiring natural language input from the user during a dialogue; parse the received information using natural language processing techniques, performing word segmentation, syntactic analysis, semantic understanding, and sentiment analysis on the input text to assess the user's current emotional state; and, based on the sentiment analysis results, generate prompt text to instruct a generative artificial intelligence model to generate the next question, enabling the generative artificial intelligence model to dynamically generate questions more suitable for the current context based on the user's actual emotional state and expressed content, thereby guiding the user to further express their feelings and thoughts.
[0006] In a preferred embodiment, the processor is further configured to: input the prompt text into the generative artificial intelligence model, and have the generative artificial intelligence model output the next question to be presented to the user, enabling the system to achieve continuous, multi-round conversational interaction, and to associate each round of questioning with the emotional state and answer content of the user in the previous round, thereby improving the coherence and personalization of the question-and-answer process.
[0007] In another preferred embodiment, the processor is further configured to: record the user's answers to each question into a database, forming structured interactive record data; and based on the answer content and the emotion analysis results, comprehensively analyze the user's emotional changes and event context over a period of time, automatically generating corresponding diary text. This allows the user to obtain a complete diary record with emotional descriptions and context without having to manually organize fragmented answer content. Therefore, this invention effectively improves the automation level of diary generation and the richness of emotional expression, thereby solving problems such as low user participation, insufficient emotion analysis, and difficulty in generating high-quality diary content in existing technologies.
[0008] A "system" refers to an overall device or combination of devices consisting of one or more hardware components and software modules running on them, used to perform functions such as information reception, processing, storage, and result output. It may include servers, terminal devices, and databases that are connected to them in communication.
[0009] A "processor" is a hardware unit that can execute program instructions, perform calculations and logical judgments on input data, and thus realize functions such as information parsing, emotion assessment, prompt text generation, and interaction with generative artificial intelligence models. It can be a single central processing unit (CPU), graphics processing unit (GPU), neural network processing unit (NPU), or any combination thereof.
[0010] An "interface" refers to a collection of hardware and software components controlled and provided by a processor for receiving input information from or outputting information to the user. It can include graphical user interfaces, command-line interfaces, application programming interfaces (APIs), voice input interfaces, touch input interfaces, etc.
[0011] "User information" refers to natural language text, speech-to-text text, or other content that can be converted into text form and is input into the system by the user through the interface, which relates to the user's status, events, thoughts, feelings, etc.
[0012] "Natural Language Processing" refers to a class of algorithms and models used for the automated analysis and understanding of natural language text, including but not limited to word segmentation, part-of-speech tagging, syntactic analysis, semantic analysis, sentiment analysis, entity recognition, referential resolution, and text vectorization.
[0013] "Emotional state" refers to the psychological and emotional tendencies of a user at a specific point in time or during a specific conversation, as determined by the system based on user information through natural language processing and sentiment analysis. Examples of these tendencies include pleasure, sadness, tension, relaxation, anger, anxiety, or neutrality.
[0014] "Emotion analysis" refers to the process of using natural language processing technology to analyze user information, identify and quantify the user's emotional state, including the determination of emotion category, emotion intensity, and emotion change trend.
[0015] "Generative AI models" refer to AI models that can automatically generate text content or other data forms based on input prompts, including but not limited to deep learning-based language models, dialogue generation models, and text generation models.
[0016] "Prompt text" refers to text instructions or guiding content that are automatically constructed by the processor based on the results of sentiment analysis and the current dialogue context and input into the generative artificial intelligence model. It is used to constrain and guide the generative artificial intelligence model to output the next question that meets the expected requirements.
[0017] "Question" refers to a natural language sentence or set of sentences generated by a generative artificial intelligence model based on prompt text and presented to the user through an interface to guide the user to input further information or express their emotions, thoughts, or experiences.
[0018] "User's response" refers to the natural language text content that users input into the system through the interface in response to the questions presented by the system, reflecting the user's understanding, description, feedback or emotional expression on the question.
[0019] A "database" refers to a storage system used to store and manage data sets related to system operation. It can be a relational database or a non-relational database, used to persistently save information such as user answers, sentiment analysis results, prompt text, generated questions, and automatically generated diaries.
[0020] "Diary" refers to natural language text automatically generated by the system based on the user's answers in multiple rounds of question-and-answer sessions and the corresponding sentiment analysis results. It is used to record the user's events, feelings, and emotional changes within a specific time period. Attached Figure Description
[0021] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0022] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0023] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0024] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0025] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0026] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0027] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0028] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0029] Figure 9 This represents an emotion map that maps multiple emotions.
[0030] Figure 10 This represents an emotion map that maps multiple emotions.
[0031] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0032] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0033] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0034] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0035] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.
[0036] First, let me explain the terminology used in the following instructions.
[0037] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0038] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0039] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0040] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0041] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0042] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0043] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0044] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0045] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0046] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0047] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0048] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0049] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0050] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0051] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0052] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0053] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0054] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0055] With the increasing popularity of conversational diary applications, existing technologies typically employ a pre-defined set of fixed questions or a simple rule engine to present diary questions to users in sequence. However, such systems suffer from the following technical problems: First, after acquiring user input, the processing device only performs keyword matching or simple sentiment scoring, lacking systematic modeling of the user's historical response set. This results in an inability to accurately extract the user's emotional state and areas of interest from a time-series perspective, leading to a low correlation between the generated questions and the user's current psychological state and interests, and limited interaction quality. Second, when calling generative AI models, the processing device often directly uses the user's previous input as model input, lacking structured organization of multi-round historical data and fine control over prompts. The generated questions are difficult to adjust in terms of tone, detail, and topic depth, making it difficult for dialogues to continue and deepen, and lacking personalization among different users. Third, the processing device typically stores diary content in simple text form, without establishing a clear data association structure between response content, emotional state, areas of interest, and subsequent questions. This makes it difficult to efficiently calculate and retrieve emotional changes and topic evolution, and also makes it difficult to generate useful summary diaries or statistical information for users. This is insufficient in terms of computational resource utilization and data processing flow.
[0056] In other words, existing conversational diary systems lack a comprehensive technical solution optimized for multi-turn interactions in terms of internal data flow design, collaboration between natural language processing modules and generative AI models, and structured storage and reuse of diary data. This prevents them from fully leveraging the capabilities of generative AI models and natural language processing, thus limiting improvements in interaction quality, personalization, and computational efficiency. Therefore, it is necessary to provide a new system architecture and processing flow. By improving server-side processing logic, efficient analysis and storage of multi-turn user responses can be achieved. Furthermore, a controllable method of constructing prompts can drive generative AI models, thereby improving dialogue generation quality and diary data processing capabilities at the computer technology level.
[0057] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0058] In this invention, the server includes a control information input / output device, a device for presenting users with query information for writing diaries in a conversational format and receiving response information from users, a device for parsing a set of response information using natural language processing technology based on past response information stored in a storage device, extracting the user's emotional state and areas of interest from it, and generating parsed information summarizing the extraction results, a device for constructing prompt statements for inputting into a generative artificial intelligence model based on the parsed information and the currently input response information, and instructing the generative artificial intelligence model to generate the next query information to be presented to the user, a device for obtaining the next query information output from the generative artificial intelligence model and controlling the information input / output device to present the next query information to the user, and a device for registering the correspondence between the response information and the emotional state and areas of interest extracted through natural language processing technology and the next query information as diary data in a time-sequential manner in a storage device, and a device for generating historical information that can refer to the dialogue history and emotional changes of each user based on the diary data registered in the storage device. This allows for the formation of a closed-loop data processing chain within the server, encompassing multi-round response collection, sentiment and topic extraction, structured storage, prompt statement construction, and the invocation of generative AI models. This enables the generative AI model, after receiving parsed and summarized information and control information, to generate the next query in a manner more aligned with the user's current emotional state and area of interest. This, in turn, improves the relevance and controllability of dialogue generation at the computer technology level, and enhances the conversational diary system's ability to analyze historical data and its resource utilization efficiency.
[0059] "System" refers to an information processing whole that includes at least one processing device, storage device, and information input / output device, and processes user-input information and generates output results by executing a program.
[0060] "Information input / output device" refers to a human-computer interaction component used to present query information to a user and receive user response information, including display components, input components, and optional voice interaction components.
[0061] "User" refers to the entity that interacts with the system through information input / output devices, provides response information, and receives inquiry information generated by the system.
[0062] "Response information" refers to the natural language content that users input into the system in response to the system's queries, using text, voice, or other input methods.
[0063] "Storage device" means a computer-readable storage medium used to store response information, diary data, emotional states, areas of interest, and historical information in a retrieveable form.
[0064] "Past response information" refers to the set of response information that has been entered by the user and stored by the system before the current processing, and is used as historical data for analysis.
[0065] Natural Language Processing (NLP) technology refers to algorithms and models based on computer programs that perform word segmentation, syntactic analysis, semantic understanding, sentiment recognition, topic extraction, and other processing on natural language text.
[0066] "Emotional state" refers to the category or tendency of a user's psychological emotions extracted from the response information through natural language processing technology, including but not limited to emotional labels such as happy, sad, angry, and relaxed.
[0067] "Areas of interest" refers to the main topics that users are concerned about, identified from the response information through natural language processing technology, including but not limited to topics such as work, interpersonal relationships, family, and leisure.
[0068] "Analyzed information" refers to intermediate data that summarizes and structures the emotional state and areas of interest by applying natural language processing techniques to past and current response information.
[0069] "Generative artificial intelligence model" refers to a parameterized model trained through machine learning methods that can automatically generate natural language text based on input prompts.
[0070] "Prompt statements" refer to the text content that serves as input to a generative artificial intelligence model. They provide the model with contextual information and generative constraints to guide the model in generating the next query information.
[0071] "Inquiry information" refers to natural language questions or prompts used to guide users in diary-style expression. These questions are generated by a generative artificial intelligence model based on the prompts and presented to the user.
[0072] "Next query information" refers to the query information generated by a generative artificial intelligence model based on the current response information and parsed information, which will be presented to the user in subsequent rounds.
[0073] "Control information" refers to parameters or tags attached to the parsed information when constructing prompt statements. These parameters specify the tone, level of detail, and topic depth of the inquiry, thereby controlling the output style of the generative artificial intelligence model.
[0074] "Diary data" refers to a structured record that organizes response information, corresponding emotional states, areas of interest, and related inquiries in chronological order, and is used to represent a user's conversational diary content.
[0075] "Time-sequentially associated registration" refers to attaching time information or sequence numbers to each record when writing diary data to a storage device, and establishing a traceable sequential relationship according to the order of time.
[0076] "Dialogue history" refers to a collection of interaction records consisting of multiple rounds of query and response information arranged in chronological order, used to reflect the continuous interaction process between the user and the system.
[0077] "Emotional change" refers to the change in the type or intensity of emotion over time, based on the emotional states recorded in multiple diary entries within a predetermined period.
[0078] "Historical information" refers to summary or overview data generated based on diary data for users or systems to refer to, including dialogue history, emotional changes, and related statistical results.
[0079] "Summary diary" refers to a concise text automatically generated from multiple diary entries using natural language processing technology, used to summarize a user's main experiences, emotional state, and areas of interest within a certain period of time.
[0080] "Statistical information" refers to the quantity distribution, trend curve, or proportion data calculated based on diary data and tags such as emotional state and areas of interest, used to reflect the emotional and thematic characteristics of users within a specific time period.
[0081] In a preferred embodiment of the invention, the server includes a processor, main memory, persistent storage, and a network interface. The server runs a general-purpose operating system, such as a Unix-like operating system, and builds an application server environment on top of it. The server communicates with a terminal via the network interface; the terminal can be a mobile communication terminal with a display screen and input devices, or other information processing devices.
[0082] The server loads a program in main memory to implement the functions of the system of this invention. During program initialization, the server prepares a relational data management system, such as general-purpose relational database management software (e.g., a MySQL-type or PostgreSQL-type database management system), in permanent storage and establishes a data table structure including user tables, session tables, diary tables, and sentiment tag tables. The server loads a natural language processing library in main memory, such as a text encoding model and a sentiment classification model based on a deep learning framework, and loads a generative artificial intelligence model. This model can be a neural network model with a multi-layer encoder-decoder architecture or an autoregressive architecture, such as a Transformer-like language model with a multi-layer self-attention structure.
[0083] When implementing natural language processing techniques, the server uses a text vectorization model built with a deep learning framework (such as a general tensor computation framework) to segment, tokenize, and embed the response information. The server maps each word to a fixed-dimensional real-valued vector and computes sentence-level vector representations using a multi-layer feedforward network or a multi-head self-attention network. During sentiment classification, the server appends several fully connected layers and non-linear activation functions to the sentence-level vectors and uses weights pre-learned during training to calculate the probability distribution of the input text across several predetermined sentiment categories. By comparing the output probabilities of each category, the server labels the response information as sentiment states such as "happy," "sad," or "relaxed" according to the maximum probability criterion or a set threshold. When extracting domains of interest, the server maps the response information to a set of domain labels such as "work," "interpersonal relationships," "family," and "leisure" based on the similarity between word vectors and topic vectors, or based on a trained multi-label classifier.
[0084] When implementing the generative AI model, the server uses a neural network structure with multiple layers of self-attention networks. During the training phase, the server uses a large amount of dialogue and diary text as training corpus. The server initializes the model parameters and updates them using backpropagation and gradient descent optimization methods (such as the Adam optimization algorithm). During training, the server sets a cross-entropy loss function, compares the model's output word distribution with the target word sequence, and adjusts the weight matrices of each layer based on the gradient of the loss function. The server can use data augmentation techniques during training, such as randomly masking some words or shuffling the order of synonyms, to improve the model's robustness to different expressions.
[0085] During the inference phase, the server encodes the parsed information and the current response information into a prompt statement. The server combines the sentiment state, domain of interest, summaries of recent dialogues, and control information regarding output style into a text-based prompt statement, which is then input into the input sequence of the generative AI model. Internally, the server uses a self-attention mechanism to allow the model to assign different attention weights to the vector representations of different historical responses based on the sentiment state and domain of interest identified in the parsed information when generating the next query. This weight-based calculation method allows the model to give greater influence to historical information that is more relevant to the current sentiment and topic when calculating the probability distribution of the next lexical term, thereby improving the relevance and consistency of the generated results.
[0086] When constructing the prompt statement, the server explicitly incorporates control information to constrain the tone, length, and topic depth of the query information output by the generative AI model. For example, the server embeds the following text into the prompt statement: "You are a gentle diary assistant."
[0087] Recent user sentiment analysis results: Overall, users are happy, but slightly tired.
[0088] My recent focus: work and gatherings with friends.
[0089] Please generate an open-ended question of no more than 20 characters in an encouraging tone. In another specific example, the server inputs the following prompt statement to the generative artificial intelligence model: "Summary of the user's diary for the last three days:" 1. "I worked overtime with my colleagues today. I'm a little tired, but we completed an important task." (Emotion: Slightly positive; Theme: Work) 2. 'I spent the weekend relaxing at home, watching a few movies, and felt very relaxed.' (Emotion: Positive, Theme: Leisure) 3. 'Talking to my family on the phone, I feel a little homesick.' (Emotion: Slight longing; Theme: Family) Based on the above information, please generate a question suitable for today for the user.
[0090] Require: (1) Use Simplified Chinese; (2) The question length should not exceed 30 characters; (3) Questions should encourage users to describe what they feel most strongly about at that moment. In another example, the server constructs a prompt statement based on the current response information, for example: The user just replied: "I had a great time reuniting with my college classmates today. We had dinner and chatted together." Sentiment analysis results: happy, moved.
[0091] Topic: Friends, interpersonal relationships.
[0092] Based on this answer, please generate a question for the next round of conversation that can help the user reflect more deeply on their feelings.
[0093] The question should be written in Simplified Chinese, be 15-25 characters long, and contain only the question itself, without any explanatory text. By embedding control information in textual prompts, the server transforms the reasoning process of generative AI models from a simple automated reproduction of human questions into a process where, within the model, emotional features, topic features, and control parameters are jointly encoded and processed using a parameterized structure. During the generation phase, the output content and style are dynamically adjusted based on different user states. The server thus forms a programmable dialogue generation control layer within the computer, fundamentally different from script systems with fixed rules in terms of granularity of model invocation, controllability, and adaptability.
[0094] In this embodiment, the terminal includes a display unit, an input unit, local storage, and a network communication module. The terminal runs a general-purpose mobile operating system and can execute applications. The terminal establishes an encrypted communication channel with the server through the network communication module. Upon receiving an inquiry from the server, the terminal displays the text content on the display unit in the form of a speech bubble or similar format. If necessary, the terminal converts the inquiry into audio playback by invoking a local speech synthesis engine. When receiving user input, the terminal receives text input via a soft keyboard or obtains and converts the user's voice input into text by invoking a speech recognition software development kit. The terminal performs basic text cleaning locally, such as removing leading and trailing whitespace and merging consecutive spaces. Then, it packages the response information along with the user identifier, timestamp, and session identifier, and sends it to the server via the network module.
[0095] Users generate diary entries through multiple rounds of responses on the terminal interface. When users use the system multiple times on different dates, they do not need to manually organize historical records; the server automatically organizes the diary data chronologically through the database management system. While assigning a unique identifier to each response and associating it with a timestamp and session identifier, the server stores the sentiment state and interest area tags for each record in structured fields. When the server needs to generate historical information or a summary diary entry, it quickly retrieves records within a specified time period using the database's index structure and constructs a time-series data structure in main memory for subsequent analysis and visualization output.
[0096] When generating a summary diary, the server merges multiple diary entries within a selected time period. The server uses natural language processing techniques to perform clustering or topic modeling on the multiple texts, grouping sentences with similar themes into several topic clusters. The server selects representative sentences from each topic cluster, or generates short summary texts using text summarization algorithms (such as extractive or generative summarization models based on attention mechanisms). The server then preserves the main trends in emotional state and areas of focus within the summary text, such as "This week's overall mood was relaxed, with occasional moments of tension related to work progress and family communication." This summary text is then sent by the server to the terminal, which displays it as a "Weekly Mood Summary" or "Monthly Diary Overview," etc.
[0097] Through the aforementioned structured processing flow, the server achieves several improvements at the computer technology level. First, by establishing a clear correspondence between user response information and emotional state, area of interest, and information for the next generated inquiry, and storing this information chronologically in a relational database, the server can directly utilize indexes and structured queries in subsequent retrieval and analysis, significantly reducing data scanning and thus improving query speed and reducing storage access overhead. Second, by introducing control information into prompts, the server transforms the dialogue generation process from unconstrained natural language generation to a generation process controlled by multi-dimensional parameters, making the output of the generative AI model more stable and predictable. This reduces the generation of irrelevant or repetitive questions, improving overall interaction quality and information density. Third, by pre-aggregating and summarizing historical responses, the server compresses high-dimensional text data from multi-turn dialogues into low-dimensional parsed information for input into the generative AI model, reducing the model's input length and thus lowering inference computation and communication load. In practical deployments, this significantly reduces response latency and resource consumption.
[0098] When implementing the above functions, the server does not simply mechanically copy human questions to the user. Instead, it utilizes deep neural networks, attention mechanisms, and control information construction techniques to autonomously combine query information based on emotional and topical features calculated internally by the machine and language distribution rules learned by the model parameters. This approach differs fundamentally from manually drafting questions or the automated execution of simple rule engines: on the one hand, the server models high-dimensional semantic relationships through the parameter space of the neural network model, enabling the generation of diverse question templates that are difficult for humans to exhaust in advance; on the other hand, the server explicitly incorporates emotional and topical control into the prompts, causing the model to follow a non-traditional combination rule based on feature weighting and control labels during generation, rather than simple keyword matching or fixed scripts. This approach directly improves the structure of text generation algorithms within computers.
[0099] In other implementations, the server can employ different generative AI model architectures. For example, the server can use an autoregressive language model with a decoder-only architecture, or a sequence-to-sequence model with an encoder-decoder architecture. The server can encode sentiment states and areas of interest as additional feature vectors, which are then input into the model's hidden layers along with word vectors to achieve finer-grained emotion and topic modulation. The server can also use multi-task learning to simultaneously optimize dialogue generation and sentiment classification tasks during the training phase, improving the model's sensitivity to emotional details by sharing underlying representations. In other variations, the server can employ knowledge distillation to transfer the behavior of a large model to a smaller model, thereby reducing inference computation while maintaining generation quality and further improving the system's operating efficiency in resource-constrained environments.
[0100] In another implementation, the server can combine lightweight model inference from the user device to reduce server load. The server can deploy a small sentiment classification model on the terminal to perform preliminary sentiment assessment on some response information, sending only necessary features or summaries to the server, thereby reducing the amount of raw text transmitted at the communication level and lowering network bandwidth usage. In this architecture, the server still handles unified generative AI model invocation and diary data management, but distributed processing is introduced in the feature extraction and preprocessing stages, further improving the overall system's scalability and response speed.
[0101] Therefore, this invention constructs a complete data processing chain around generative artificial intelligence models and natural language processing technology through the collaborative work between servers, terminals, and users. This chain not only achieves the application goal of conversational diaries, but also provides a set of technical solutions in terms of algorithm design, data structure organization, and model control mechanisms within the computer, which can directly improve processing efficiency, generation accuracy, and data management capabilities. This makes the invention not limited to business process automation, but rather embodies an improvement on computer technology itself.
[0102] use Figure 11 The processing flow is explained.
[0103] Step 1: During startup, the server loads the configuration and model. The server's input includes the configuration file content and pre-trained model weight files, and its output includes a configuration object residing in memory, a database connection object, a natural language processing model instance, and a generative artificial intelligence model instance. The server reads parameters from the configuration file, such as the database address, port, user ID, password, and model path, and initializes the database connection pool accordingly. The server reads parameters from the storage device for the word segmentation model, sentiment classification model, and topic classification model, creating an inference session on the processor and optional graphics processing unit. The server loads the weights of the generative artificial intelligence model (e.g., a multi-layer Transformer network) from the storage device, building and warming up the model structure to enable efficient execution of subsequent text encoding and generation computations in memory.
[0104] Step 2: The terminal requests the server to start a conversation log session upon user interaction. The terminal's input is the user's action of clicking "Start Log" or a similar control on the interface, and the output is session initiation request data containing the user identifier, terminal identifier, and current time. The terminal generates a request message locally, writes the user identifier and timestamp into a data structure, and sends the request to the server using an encryption protocol via the network communication module. Before sending, the terminal encodes and compresses the request data to reduce network load.
[0105] Step 3: The server retrieves historical diary data based on the session initiation request and generates parsed information. The server's input is the session initiation request data sent by the terminal, and the output is a parsed information object representing the user's historical emotional state and areas of interest. The server reads the user identifier from the request, sends a query command to the database management system, and retrieves several recent response messages, sentiment tags, and topic tags. The server applies a tokenizer to the text fields in the query results, decomposing sentences into word sequences and mapping them to vector sequences using a text encoding model. The server performs aggregation operations on these vectors over time (e.g., weighted average or attention aggregation) to obtain the representation vector for each diary entry and the summary vector for the overall time window. The server uses a sentiment classification model to classify and infer these vectors, statistically analyzing the frequency and trends of each sentiment tag, and simultaneously uses a topic classification model or keyword matching algorithm to identify high-frequency areas of interest. The server organizes the statistical results, the current dominant sentiment, the main themes, and representative historical fragments into structured parsed information for subsequent prompt generation.
[0106] Step 4: The server constructs an initial prompt statement based on the parsed information and system strategy, and then calls the generative artificial intelligence model to generate the first query message. The server's input includes the parsed information object and the system's preset dialogue strategy parameters; the output is the first query message text for the user to answer. The server creates a prompt statement string in memory, concatenating the following elements according to a predetermined template: a description of the role the generative artificial intelligence model should play, a summary of the user's recent sentiment, the main areas of interest, and the length and tone control conditions of the query message. The server tokenizes the constructed prompt statement into a sequence of lexical units and inputs it into the encoding module of the generative artificial intelligence model. Within the model, the server calculates the correlation between the lexical units of the prompt statement using multi-head self-attention, generates a context representation vector, and then progressively generates the output lexical sequence through the decoding module. In each generation step, the server selects the highest probability lexical unit based on the probability distribution or employs a bundle search strategy to improve the generation quality. The server ultimately obtains the complete query message sentence text.
[0107] Step 5: The server sends the generated initial query to the terminal, which then displays it to the user. The server's input is the query text output by the generative AI model, and its output is a response data packet containing that text. The terminal's input is the server's response data packet, and its output is the natural language question displayed on the interface. The server writes the query information into a response structure, along with a session identifier and timestamp, and sends it to the terminal via a network interface. Upon receiving the response, the terminal parses the data structure, extracts the query text, and renders the text content in the display component. The terminal may optionally invoke a speech synthesis engine to convert the text into audio data and play it through a speaker, thereby driving the audio output device at the hardware level.
[0108] Step 6: After reading or listening to the inquiry, the user enters a response on the terminal. The user's input consists of natural language expressions of their daily experiences and emotions, while the output is edited text content displayed on the terminal interface. The user can bring up the soft keyboard via the touchscreen and input text character by character, or press the voice input button to speak the content. In voice input mode, the terminal invokes the voice recognition module to sample, extract features from, and recognize the received audio signal as a character sequence. The user checks and edits the recognition result on the terminal interface until satisfied, then clicks the send button.
[0109] Step 7: The terminal preprocesses the response information and sends it to the server. The terminal's input is the response text confirmed by the user on the interface, and its output is the response request data sent to the server. The terminal performs basic text cleaning operations, including removing leading and trailing spaces, replacing abnormal control characters, and limiting the maximum length. The terminal encapsulates the cleaned text, along with the user identifier, session identifier, and timestamp, into a data structure, and uses the network module to convert the data into a transmission format using encoding and encryption algorithms before sending it to the server. After sending, the terminal waits for the server's processing result or the next query.
[0110] Step 8: The server receives user response information and writes it to the database as a diary record. The server's input is the response request data sent by the terminal, and its output is the diary data record written to the database and the confirmation result. The server parses the request data, checking that all required fields are present and that the response text length is within the allowed range. The server constructs an insert command, writing the user identifier, session identifier, response text, and timestamp fields to the diary table, and reserving storage fields for sentiment status, areas of interest, and related query information. The server calls the database driver to execute the insert operation, ensuring the atomicity and consistency of the write through a transaction mechanism. After successful insertion, the server obtains a unique identifier for the new record from the database and caches this identifier for subsequent analysis and correlation processing.
[0111] Step 9: The server uses natural language processing (NLP) technology to perform sentiment and topic analysis on the response information and update the diary records. The server's input is the recently stored response text and its record identifier; the output is the updated diary record with added sentiment tags and domain-of-interest tags. The server reads the text field corresponding to the identifier from the database, performs word segmentation and tokenization on the text, converting it into a sequence of tokens. The server uses a pre-trained text encoding model to map the token sequence into a vector sequence, and obtains sentence-level representation vectors through pooling or attention-weighted operations. The server inputs this vector into a sentiment classification model, calculates the probability distribution across preset sentiment categories, and selects the sentiment tag based on the highest probability. Simultaneously, the server uses a topic classification model or a keyword-based classification algorithm to identify the topic of the text, generating one or more domain-of-interest tags. The server constructs an update instruction to write the sentiment tag and domain-of-interest tag into the corresponding fields of the diary record, achieving a structured expansion of the diary data.
[0112] Step 10: The server constructs new prompts based on the latest response information and historical parsing information, and calls a generative AI model to generate the next query. The server's input includes the current response text, its sentiment tag, domain of interest tag, and the parsing information constructed in previous steps; the output is the next query text. The server combines these data elements in memory to generate new prompts, including a verbatim summary of the current response, the system-inferred sentiment state, the domain of interest, and style constraints for the next question. The server tokenizes the prompts and inputs them into the generative AI model, which internally fuses the current response features with historical aggregated features using a self-attention mechanism. During the decoding phase, the server generates the next query word by word, selecting output tokens at each step using probability distributions, and can set temperature parameters or bundle search size to balance fluency and diversity. After generating a complete sentence, the server performs simple post-processing, such as removing extra spaces and checking for length constraints.
[0113] Step 11: The server sends the next query message to the terminal, which then displays and initiates a new round of dialogue. The server's input is the generated next query message text, and its output is a network response; the terminal's input is this response, and its output is a new question displayed on the interface. The server encodes the query message along with the updated session state into response data and returns it to the terminal via the network interface. After receiving the data, the terminal parses the query message text, inserts a new system message bubble into the dialogue interface, and displays the question to the user. The terminal can adjust the interface presentation based on the sentiment tags provided by the server, such as using soft colors or adding reassuring auxiliary text when the mood is low. After seeing the new question, the user can enter a response again, thus looping through the aforementioned response, storage, analysis, and generation steps to form a continuous conversational diary recording process.
[0114] Step 12: The server generates summary diaries or statistical information based on stored diary data and sends them to the terminal for display when needed. The server's input is a user identifier and a specified time range, and its output is summary text or statistical results. The server retrieves corresponding diary records and their sentiment and topic tags from the database based on the user identifier and time range, constructing a time-series data structure. The server applies a text summarization algorithm to cluster and extract data from multiple diary entries, generating concise text summaries reflecting major events and sentiment changes. The server simultaneously calculates the distribution of each sentiment state along the timeline and the frequency of occurrence of each topic, forming visualizeable statistical data. The server sends the summary text and statistical results to the terminal via the network. The terminal displays this information in the form of lists, charts, or timelines, allowing users to intuitively view their emotional trajectory and changes in focus.
[0115] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0116] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0117] In this invention, the server includes: means for communicating with a terminal including a display device and an input device via a communication device to present questions to a user in a conversational format and obtain user response information; means for performing natural language processing analysis on the obtained response information and external data obtained from external information sources to assess the user's emotional state and behavioral tendencies, and storing the response information and the external data as diary data; means for generating prompt statements input to a generative artificial intelligence model based on the diary data and the external data to instruct the generative artificial intelligence model to generate the next round of conversational questions; means for generating prompt statements input to a generative artificial intelligence model based on the diary data and the emotional state to instruct the generative artificial intelligence model to generate personalized information that conforms to the user's interests and emotions; means for obtaining the questions and personalized information output by the generative artificial intelligence model and sending them to the terminal via the communication device for presentation to the user on the terminal; and means for updating the content or structure of the prompt statements based on feedback information obtained from the user and the presentation history of the personalized information to adaptively change the model input conditions used for subsequent question generation and personalized information generation. This allows for the formation of a closed-loop data flow within the computer system, centered on diary data and external data. On one hand, the server automatically constructs and adjusts the prompts for the generative artificial intelligence model based on multi-source data fusion and sentiment analysis, thereby optimizing the model input conditions and improving the relevance and diversity of question generation and personalized information generation. On the other hand, the server adaptively updates the prompts based on user feedback and presentation history, enabling the human-computer interaction process and recommendation strategy to iterate automatically on the server side. This, in turn, enhances the intelligence of human-computer interaction, the effectiveness of model invocation, and the overall data processing efficiency and scalability of the system at the computer technology level.
[0118] "Terminal" refers to an electronic device that includes a display device and an input device, used to present and input information to users, and to send and receive data with a server through communication functions. It can be a mobile communication device, a computing device, or other information processing device with a human-computer interface.
[0119] A "server" refers to an electronic information processing device equipped with processing, storage, and communication devices, used to perform data reception, storage, analysis, generation, and information delivery to terminals. It can be a single physical device or a processing system composed of multiple devices.
[0120] "Communication device" refers to hardware and / or software modules used for sending and receiving data between a server and a terminal, or between a server and an external information source. It may include network interfaces, communication protocol stacks, and related control programs.
[0121] "Display device" refers to an output device used to present information such as problems, personalized information, and interface elements to users in a visual form on a terminal, and may include a display screen, virtual display interface, etc.
[0122] "Input device" refers to a hardware and / or software interface used to receive user input, including touch screens, keyboards, voice input interfaces, and other input modules capable of receiving text, voice, or operation commands.
[0123] "User response information" refers to natural language text, speech-to-text, or other response data that can be processed by the server when a user enters a question on the terminal in response to the question presented by the server.
[0124] "External data" refers to data obtained by a terminal or server from external information sources that is related to user behavior or environment, including location data, data published from electronic information services, activity record data, etc.
[0125] "External information source" refers to a service or device that is separately located from the system and can provide user-related data, including location service systems, electronic information service systems, social platform service systems and other network service systems.
[0126] "Location data" refers to numerical or symbolic data provided by a location acquisition device or location service system to represent a user's geographic location or location information, including coordinate information, location identification information, etc.
[0127] "Electronic information services" refer to service systems that provide users with information publishing, browsing, or interactive functions through the network, including social networking services, content sharing services, and messaging services.
[0128] "Published data" refers to data obtained from electronic information services that is related to users' activities on the service, including text data that can be used for analysis, such as text posted by users, text content of image descriptions, text comments or their summaries.
[0129] "Natural Language Processing Analysis" refers to the computational processing of natural language text in user responses and external data, including word segmentation, syntactic analysis, semantic analysis, sentiment analysis, and topic extraction, to obtain structured features or semantic tags.
[0130] "Emotional state" refers to the emotional characteristics of a user in a specific time or context, inferred from relevant text by natural language processing analysis. These characteristics include emotional polarity, emotional type, and intensity.
[0131] "Behavioral tendencies" refer to information derived from the analysis of external data and user historical data, reflecting users' preferences or habit patterns in terms of time, location, activity type, etc.
[0132] "Diary data" refers to a collection of information stored in a structured or semi-structured form by a server after associating user responses with external data. It represents a user's experiences, feelings, and context within a certain period of time.
[0133] "Generative AI models" refer to machine learning-based generative models that can automatically generate natural language text content based on input prompts and contextual information, including but not limited to deep learning-based language models.
[0134] "Prompt statements" refer to natural language or structured instruction text that is constructed by the server and input into the generative artificial intelligence model to instruct the model to generate questions, personalized information or other target text.
[0135] "Conversational questions" refer to questions presented in a natural language dialogue format, suitable for users to answer freely on the terminal, and used to guide users to describe their experiences, emotions, or needs.
[0136] "Personalized information" refers to text information generated based on diary data, external data, emotional state, and user behavior tendencies, which is adapted to the interests, emotions, or needs of a specific user. This includes suggestions, recommendations, and feedback information.
[0137] "Presentation history" refers to the historical data recorded in the system regarding the time, content, number of times personalized information is displayed to users on the terminal, and related interactions.
[0138] "Feedback information" refers to the user's response to questions or personalized information presented on the device, including likes, disinterest, clicks, dwell time, follow-up questions, or comments.
[0139] "Model input conditions" refer to the prompts, contextual data, parameter settings, and other input elements that affect the model's output when calling a generative artificial intelligence model.
[0140] "Adaptive change" refers to the process by which the server automatically adjusts the content, structure, or model input parameters of the prompt statements based on feedback information and presentation history, so that the behavior of subsequent model calls dynamically changes with the user's state and the system's performance.
[0141] "Additional input questions" refer to question texts generated by generative artificial intelligence models based on existing diary data and sentiment analysis results, used to further guide users to supplement information or engage in self-reflection.
[0142] "Behavioral suggestions" refer to suggestive text messages generated by generative artificial intelligence models based on users' emotional state, behavioral tendencies, and diary topics, used to guide users to take certain specific activities or adjust their behavior.
[0143] In this embodiment of the invention, the server, through a program running on general-purpose computer hardware, performs unified modeling, feature extraction, prompt statement construction, and generative artificial intelligence model invocation of user response information, external data, and diary data, thereby forming an adaptive human-computer interaction and personalized content generation mechanism within the computer. By improving data structure design, prompt statement generation algorithms, and model invocation processes, the server achieves comprehensive optimization of processing speed, recommendation accuracy, storage management efficiency, and communication load.
[0144] Servers can be deployed on computing devices including a central processing unit, main memory, non-volatile memory, and network interface. Servers can use general-purpose operating systems, such as Unix-like operating systems or general-purpose desktop operating systems. Applications can be implemented using general-purpose software frameworks, such as web frameworks based on general-purpose scripting languages or service frameworks based on object-oriented languages. Servers can use relational database systems, such as Structured Query Language databases, or non-relational database systems, such as document-oriented databases, to store user information, log data, external data, prompts, and generated results. Servers can communicate with terminals and external information sources over packet-switched networks via network interfaces.
[0145] The terminal can be a mobile communication device or a tablet computing device, and includes a display device, an input device, a location acquisition device, and a network communication module. The terminal can run a mobile operating system and use a mobile application framework to implement the user interface. The terminal acquires location data and sensor data through the application programming interface provided by the operating system, and accesses electronic information services through third-party service software development kits.
[0146] Users can view questions and personalized information sent by the server through the terminal's display device, and input answers, feedback, or additional questions through the input device. Users can authorize the terminal to access location data and electronic information service data as needed to enrich the data dimensions available to the system.
[0147] In terms of data storage, the server can construct a multi-table relational data structure for each user. For example, the server maintains a user master table, a journal table, a question-and-answer table, an external data table, and a recommendation history table in the database. The journal table can include date fields, journal text fields, sentiment tag fields, topic tag fields, and external data reference fields. The question-and-answer table records conversation questions and corresponding user answers, and is linked to the journal table via foreign keys. The external data table stores location data summaries, summaries of electronic information service publication data, and their extracted topic keywords. The recommendation history table records personalized information content presented to users, presentation time, click behavior, feedback results, etc., for subsequent adjustments to prompts and model input conditions.
[0148] In terms of feature extraction and analysis, the server can use a natural language processing module to process user responses and external text data. The server can utilize word segmentation algorithms, part-of-speech tagging algorithms, and dependency parsing algorithms to transform natural language text into word sequences, part-of-speech tag sequences, and dependency graphs. The server can use a sentiment classification model to predict sentence-level or document-level sentiment; this model can be a convolutional neural network, a recurrent neural network, or an attention-based classification network. The server can write the predicted sentiment polarity (positive, neutral, negative) and subcategories of sentiment (such as happy, anxious, tired, etc.) into a sentiment tag field. The server can also employ topic modeling algorithms or clustering methods based on pre-trained vectors to extract topic keywords from the text and store them as topic tags.
[0149] For generative AI models, the server can use a language model based on deep neural networks. This model can employ a structure with multi-layered self-attention encoders and decoders, or an autoregressive structure with only multi-layered self-attention encoders. During model training or fine-tuning, the server can use a large amount of dialogue and diary text as training data to minimize the cross-entropy loss function and update the model weights through backpropagation and gradient descent. During training, the server can employ batch normalization, residual connections, and multi-head attention mechanisms to improve the model's training stability and expressive power. Data augmentation techniques, such as dialogue rearrangement, synonym replacement, or noise injection, can be used during training to improve the model's robustness to inputs with different styles.
[0150] In constructing prompts, the server does not simply concatenate user text and input it into the model. Instead, it generates prompts according to specific structures and rules. The server can maintain different prompt templates for different tasks, such as a template for "generating the next round of conversation questions" and a template for "generating personalized information." When generating prompts, the server can retrieve the latest diary data, sentiment tags, topic tags, and recent external data summaries from the database and insert them into the template in a predetermined order to form structured prompts. This reduces redundant information and lowers the transmission and model computation load. Furthermore, it explicitly provides sentiment and topic signals to the model, improving the match between the output content and the user's state.
[0151] For example, when generating personalized content, the server can construct the following prompt statement: "Based on the following diary entries and user sentiment, please generate personalized content recommendations and insightful analyses for the user."
[0152] Diary entry: 'Today I went to a coffee shop with a friend and we talked a lot about the past. I felt happy but also a little sad.' Emotional tags: happy, slightly sentimental Tags: Friends, Cafe, Memories Please provide: 1) A brief sentiment analysis; 2) Three movie or book recommendations related to friendship and memories (with a brief reason); 3) Two questions to guide the user's self-reflection. When generating follow-up questions, the server can construct the following prompt statements: "You are a gentle and empathetic psychological support assistant. Based on the user's diary and location information, design three in-depth but not overly offensive follow-up questions for the user."
[0153] Diary entry: 'This week I've gone to the same coffee shop almost every day after get off work, just sitting there alone, lost in thought.' Location information: 'Café, 7:00 PM – 9:00 PM, repeated multiple times' Please output only three suitable Chinese questions that a chatbot can ask users. When processing logs generated from multiple rounds of question-and-answer sessions, the server can construct the following prompt statements: Please organize the following multiple rounds of questions and answers into a fluent diary entry for the day, using the first person and a natural tone, without adding fictional events.
[0154] Q&A Record: Q1: What did you do after get off work today? A1: 'I went to eat hot pot with my colleagues, and then went to a coffee shop by myself.' Q2: What do you mainly do when you're in a coffee shop? A2: 'I was listening to music alone, thinking about some recent things.' Q3: How are you feeling overall today? A3: 'I feel quite relaxed, but also a little worried about the future.' Please write a Chinese diary entry of no less than 200 words based on the above content. By constructing the aforementioned prompts in a structured manner, the server enables the generative AI model to receive filtered, labeled, and sorted key information at the input end, rather than raw, messy text. This improves the generation quality and shortens the inference time under the same model parameter scale.
[0155] The server can configure different parameters in the model invocation process to control the generation length, randomness, and diversity. For example, the server can set the generation temperature, sampling strategy, and maximum output length according to the task type, using higher diversity for session-type questions and lower diversity for summary-type diaries. The server can also maintain parameter configurations for different users to further personalize the model's running behavior.
[0156] In terms of external data collection, the terminal can call location service interfaces to obtain latitude and longitude coordinates, and obtain location category information (such as "coffee shop," "office," "park," etc.) locally or through a server by calling map service interfaces. The terminal can perform temporal clustering on repeatedly collected location information, marking frequently occurring locations as "frequently visited locations." When acquiring electronic information service data, the terminal can perform local summarization processing on recently published content, such as extracting only text titles and parts of the body, and filtering sensitive content on the terminal before sending it to the server. This reduces the amount of data transmitted over the network and lowers the load on communication devices.
[0157] In terms of multi-source data fusion, the server can create a list of references to external data records for each diary entry. During analysis, the server can quickly access relevant location data and e-information service summaries through these references and input them along with the diary text into the feature extraction module. In implementation, the server can map multi-source text features to a unified vector space, for example, using pre-trained word vector or sentence vector models to convert diary sentences, location descriptions, and social text into fixed-dimensional vectors, and then form a comprehensive semantic representation through vector concatenation or weighted summation. The server uses this comprehensive representation to determine sentiment and topic, thus achieving higher recognition accuracy than analyzing the diary text alone.
[0158] In terms of adaptive adjustment of personalized information generation, the server can construct evaluation metrics based on recommendation history and user feedback, such as click-through rate, dwell time, and positive feedback ratio. The server can use these metrics as additional input conditions each time a new prompt is generated, or adjust the instructions in the prompt based on the metric results, such as making the instructions "more concise," "more specific," or "emphasizing action suggestions." Compared to traditional static rules, this feedback-based adaptive prompt adjustment mechanism forms a closed-loop control within the computer, helping to reduce invalid generation, lower the computational load on the server side, and improve the overall recommendation effect.
[0159] In typical usage scenarios, users can engage in multi-turn dialogues with the system via a terminal. After viewing the question text generated by the server on the terminal, the user inputs their answer using an input device. The server combines this answer with the day's diary entries, historical sentiment tags, and external data summaries to generate new diary entries and personalized information. After viewing the personalized information on the terminal, the user can click "like" or "not helpful" on a particular piece of content. Based on this, the server updates its strategy for constructing relevant prompts for that user, such as reducing the appearance of certain types of recommendations or increasing the detail of other types of suggestions. Through this iterative process, the computer system can gradually "converge" to an interaction style and content structure that is more suitable for the user, demonstrating its technical adaptability to user preferences.
[0160] In alternative implementations, servers can employ different types of generative AI models. For example, servers can deploy lightweight models with smaller parameter sizes on edge servers to generate short questions in real time, while deploying large-parameter models on central servers to generate long diary entries or complex analysis reports. Servers can configure routing strategies between different models, selecting the appropriate model based on task complexity and latency requirements, thereby conserving computing resources while ensuring response speed.
[0161] In other implementations, the terminal can partially undertake natural language preprocessing or feature extraction tasks. For example, the terminal can use a local lightweight model to perform preliminary sentiment analysis and only send sentiment tags and compressed text summaries to the server, thereby reducing data transmission at the network communication level and reducing the computational burden of preliminary analysis on the server side. In this structure, the server mainly performs prompt generation and complex generation tasks, while the terminal undertakes front-end preprocessing and some user state estimation, realizing a collaborative structure between edge computing and central computing.
[0162] Because the server in this invention employs a specially designed data structure, feature extraction process, prompt generation rules, and feedback-based adaptive adjustment strategy, the generative artificial intelligence model is no longer simply regarded as a "black box answering tool," but is precisely manipulated within a specific computational framework. By controlling the selection, organization, and annotation of input information, the server alters the model's internal attention allocation and generation path, thereby improving output quality without increasing the model's parameter size and reducing the time consumption and storage footprint caused by meaningless generation. Thus, this invention not only automates the user interaction process but also achieves a comprehensive improvement in processing efficiency, data management capabilities, and generation accuracy at the computer technology level through improved data flow design and model invocation methods.
[0163] use Figure 12 The processing flow is explained.
[0164] Step 1: The server initializes the session environment.
[0165] Inputs: User ID, historical diary data from the database, historical Q&A records, historical sentiment tags, and external data summaries (location data summaries, electronic information service publication data summaries).
[0166] Output: Session context data structure.
[0167] The server reads the latest records from the database associated with the current user's identity, including the log table, question-and-answer table, external data table, and recommendation history table. It extracts fields such as date, sentiment tags, topic tags, location categories, and social summaries from these records and combines them into an in-memory session context object. The server cleans the read text fields (removing control characters and truncating excessively long text) and sorts the records according to timestamps. The server temporarily stores the session context in a session cache, providing a unified input for subsequent prompt construction and generative AI model invocation.
[0168] Step 2: The server generates a prompt message for the next round of sessions.
[0169] Input: Session context data structure (containing recent diary summaries, recent sentiment tags, and recent external data summaries).
[0170] Output: The text used to generate prompts for conversational questions.
[0171] The server selects one or more recent diary entries, their corresponding sentiment tags and topic tags, as well as recent location categories and social topic keywords from the session context, and maps this content to a predefined "question generation" template. The server inserts fields such as "diary content," "sentiment tags," "topic tags," and "external behavior summary" in a predetermined order according to the template, concatenating them into a natural language prompt, such as: "You are a conversational assistant that cares about users' daily lives and emotions. Based on the following information, generate one suitable Chinese question for the user as an opening..." During this process, the server performs string concatenation, placeholder replacement, and length control (truncating and summarizing when the text exceeds a set length), thus obtaining a structured prompt output.
[0172] Step 3: The problem of server-side calls to generative artificial intelligence models to generate session formats.
[0173] Input: Prompt text for generating conversational questions, model parameters (temperature, maximum length, sampling strategy).
[0174] Output: Question text in conversation format.
[0175] The server takes the prompt generated in step 2 as input, encapsulates it into a model request, and calls the generative AI model service deployed on the server cluster via a network interface. Internally, the model encodes the prompt using a multi-layer self-attention network, performs linear transformations, attention weight calculations, and residual connections on each word vector, and then predicts the probability distribution of the next word step by step using an autoregressive approach. After receiving the model output, the server removes the start and end markers from the returned sequence to obtain the final question text. The server performs sensitive word filtering and length checks on this question text, writes the valid question text along with the question ID and generation time into a question-and-answer table, and caches it in memory for association with subsequent user responses.
[0176] Step 4: Issues with the terminal receiving and displaying session formats.
[0177] Input: Question text in session format from the server, and question ID.
[0178] Output: The problem screen on the terminal display device.
[0179] The terminal receives the server's response via the network module, parses the question text and question ID from the data packet, and renders the question text in a designated area using a graphical interface component in the interface thread. The terminal adjusts the interface layout based on the question type; for example, it displays a multi-line text input box and a "Submit" button for open-ended questions. The terminal also activates the soft keyboard or provides a voice input button. The terminal caches the question ID locally to bind subsequent user answers to that question, ensuring correct association information is included during data upload.
[0180] Step 5: The user enters and confirms the answer.
[0181] Input: The problem interface displayed on the terminal display device.
[0182] Output: The user-confirmed response text.
[0183] After reading the question displayed on the terminal, the user enters their answer via the touchscreen keyboard, such as "I went to a new coffee shop with a friend after get off work today. The environment was great and very relaxing," or speaks their answer via the voice input button. In voice mode, the terminal uses the operating system's speech recognition interface to convert the audio signal into text, which is then displayed in the input box. The user can modify the recognition result. After confirming the answer is correct, the user clicks the "Submit" button. This action triggers the terminal to output the current input content as the answer text to the network module.
[0184] Step 6: The terminal packages the user's answer and sends it to the server.
[0185] Input: User-confirmed answer text, question ID, user identifier, timestamp.
[0186] Output: The response data packet sent to the server.
[0187] The terminal performs basic validation on the response text (checking for emptiness and length exceeding the limit), then combines the question ID, user identifier, current timestamp, and response text into structured data. The terminal can perform simple compression or encoding on the text to reduce data size. Subsequently, the terminal sends this data packet to the server's designated interface using a secure transmission protocol via the network communication module. The terminal temporarily stores this data copy locally, deleting the temporary record only after the server confirms successful reception, in case of network failure and retransmission.
[0188] Step 7: The server stores the answers and updates the log data.
[0189] Input: Response data packet from the terminal (response text, question ID, user identifier, timestamp).
[0190] Output: Updated diary entries and Q&A records.
[0191] The server first verifies the user ID and authentication token to confirm the legitimacy of the request source. Then, it writes the answer text along with the question ID into the question-and-answer table, recording it as a new answer record. The server searches for the current day's log entry based on the timestamp and user ID. If it doesn't exist, the server creates a new record in the log table, writing the answer text into the "Original Content" field. If it exists, the server appends the new answer to the "Question-and-Answer List" field of that log entry. The server also records the question ID associated with the answer to ensure that the question context is preserved when organizing the logs later. Through this associative storage method, the server achieves unified management of multi-turn question-and-answer sessions and single log entries.
[0192] Step 8: The server performs natural language processing analysis on the response text and external data.
[0193] Input: Updated diary entries (including new response text) and external data records associated with the date (location summary, electronic information service release summary).
[0194] Output: Updated sentiment tags, topic tags, and behavioral summaries.
[0195] The server extracts the latest response text from the diary records and merges it with associated external data text as analysis input. In its natural language processing module, the server performs word segmentation and part-of-speech tagging on this text, generating word sequences and part-of-speech sequences. The server calls a sentiment classification model, inputting the word vector sequences into a neural network structure, extracting features through multiple hidden layers, and finally outputting sentiment category and sentiment intensity score. Based on preset thresholds, the server labels sentiment polarity as positive, neutral, or negative, and writes sentiment subcategories (such as "happy," "tired," etc.) as sentiment tags into the diary records. Simultaneously, the server performs keyword extraction and topic clustering, mapping high-frequency nouns and verbs to predefined topic sets, such as "work," "friends," and "coffee shop," and writes the matching results into the topic tag field. For external data, the server counts the frequency of specific location categories over a period of time and records patterns such as "frequently appearing in coffee shops at night" as behavioral summaries for generating follow-up questions.
[0196] Step 9: The server generates prompts to organize the multi-round question-and-answer sessions into a coherent log.
[0197] Input: The original diary entry for the day, containing multiple question and answer records, along with relevant sentiment tags and topic tags.
[0198] Output: Prompt text used to generate coherent diary entries.
[0199] The server reads all questions and answers for the same date from the log table and the question-and-answer table, sorts them chronologically, and labels the questions as Q1, Q2, etc., and the answers as A1, A2, etc. The server combines these question-and-answer pairs into a "question-and-answer record" paragraph and appends the day's sentiment and topic tags to the prompt. For example, the server generates the following text: "Please organize the following multi-round question-and-answer content into a fluent diary entry for the day, using the first person, a natural tone, and without adding fictional events. Question-and-answer record: Q1:……A1:…… Q2:……A2:…… Please generate a Chinese diary entry of no less than 200 words based on the above content." During the generation process, the server performs string formatting and text concatenation operations to ensure that the output prompt meets length and structural constraints.
[0200] Step 10: The server invokes a generative artificial intelligence model to generate coherent diary text.
[0201] Input: Prompt text for generating coherent diary entries, and model inference parameters.
[0202] Output: The edited diary text for that day.
[0203] The server sends the prompts constructed in step 9 to the generative AI model service. Internally, the model vectorizes the prompts based on its encoding layer and generates output text word-by-word using an autoregressive strategy in its decoding layer. Externally, the server controls the generation process by setting a maximum generation length and termination conditions. After receiving the output, the server performs basic syntax checking and sensitive word filtering on the text, writes the generated coherent diary text into the "Edited Diary" field of the diary table, and keeps the original question-and-answer list unchanged, achieving a dual-storage structure of original data and generated text. Through this processing, the server transforms multi-round discrete question-and-answer sessions into structured diary content, providing a more stable input for subsequent personalized information generation.
[0204] Step 11: The server generates prompts for personalized information generation.
[0205] Input: The edited diary text of the day, sentiment tags, topic tags, and behavioral summaries.
[0206] Output: The text used to generate personalized prompts.
[0207] The server reads the "Edited Diary" field, along with sentiment and topic tags, from the diary entries, and then retrieves recent behavioral pattern information from the behavioral summary. The server inserts these elements into a predefined "Personalized Recommendation" prompt template, for example: "Based on the following diary content and analysis results, please generate personalized suggestions for the user: Diary content: ... Sentiment: ... Topic: ... Behavioral summary: ... Please output: 1) Brief sentiment feedback; 2) 3 content recommendations related to the above topics (with brief reasons); 3) 2 questions to guide the user to further self-reflect." During the construction process, the server trims and simplifies the fields to avoid lengthy text affecting the model's inference efficiency.
[0208] Step 12: The server calls a generative artificial intelligence model to generate personalized information.
[0209] Input: Prompt text used to generate personalized information, and model inference parameters.
[0210] Output: Structured personalized information data (emotional feedback, recommended items, self-reflection questions).
[0211] The server submits prompts to the generative AI model, requesting the generation of comprehensive text containing emotional feedback, recommended items, and self-reflection questions. Upon receiving the output, the server divides the overall text into multiple parts according to predefined text tags or natural language patterns, extracting the emotional feedback paragraphs, the sequence of recommended items, and the list of questions. The server generates a structured record for each recommended item, including a type field (e.g., "movies," "books," "events"), a title field, and a brief description field. When necessary, the server can call external content service interfaces to find specific content links based on keywords in the recommended items and write the link information into the record. The server then writes the final personalized information to a recommendation history table, providing a data foundation for subsequent display and feedback collection.
[0212] Step 13: The terminal receives personalized information and displays it.
[0213] Input: Structured, personalized information data from the server.
[0214] Output: A list of sentiment feedback and recommendations displayed on the terminal interface.
[0215] The terminal receives personalized information data packets sent by the server via the network module and parses the emotional feedback text, recommendation list, and self-reflection questions within them. The terminal constructs a page with multiple information areas at the interface layer: emotional feedback is displayed at the top, recommended items are displayed in a list or card format in the middle, and self-reflection questions and a "Continue to answer" button are displayed at the bottom. The terminal generates interactive controls for each recommended item, such as "View Details," "Like," and "Not Interested" buttons. The terminal uses a rendering engine to draw these components onto the display device, allowing users to visually view the generated content.
[0216] Step 14: Users can view personalized information and provide feedback or ask additional questions.
[0217] Input: The personalized information interface displayed on the terminal display device.
[0218] Output: User feedback data and / or appended question text.
[0219] When browsing emotion feedback and recommended items on the terminal, users can click on a recommendation to view a detailed description, or click the "Like" or "Disinterested" button. Users can also enter new questions or evaluations of recommended content in the input boxes at the bottom, such as "I don't really like this type of movie" or "Can you recommend some activities suitable for doing with friends?" After confirming the input, users click "Send," and the terminal outputs this feedback or additional questions as new data to the network module.
[0220] Step 15: The terminal packages and submits feedback and additional issues to the server.
[0221] Input: User feedback action (click, dwell time), appended question text, recommended item ID, user identifier, timestamp.
[0222] Output: Feedback data packets and appended question data packets sent to the server.
[0223] The terminal records user actions on the personalized information interface, such as the time of clicking recommended items, the duration of the interaction, and button selections (like, dislike), and associates these events with the corresponding recommended item IDs. The terminal encapsulates this action data along with the user identifier and timestamp into a feedback data packet. For user-inputted append questions, the terminal encapsulates them along with the current log date and the context of recent recommendations into an append question data packet. The terminal sends this data to the server via the network communication module for subsequent adaptive adjustments to prompts and generation of new responses.
[0224] Step 16: The server generates a strategy and a response based on the feedback update prompt statement.
[0225] Input: User feedback data packet, appended question data packet, recommendation history.
[0226] Output: Updated prompt statement generation parameters and new system response text.
[0227] The server analyzes feedback data to calculate click-through rates, like rates, and disinterest rates for various recommendation types, and then calculates user preference vectors for different content types. The server writes these statistics into a user preference table and adjusts the weight fields in the prompt template based on these preferences, such as adding prompts for high-preference categories and reducing recommendation requests for low-preference categories. For appended question text, the server constructs a new prompt combining the user's question with a recent diary summary, for example: "User Question: ... User's Recent Diary and Behavior Summary: ... Please provide three specific, actionable suggestions in Chinese, no more than 200 characters, using a caring and supportive tone." The server submits this prompt to a generative AI model and stores the model's response text in a recommendation history table before pushing it to the terminal. Through this feedback-based strategy update, the server continuously adjusts the prompt structure and model input conditions internally, thereby improving the relevance and efficiency of the generated results in subsequent processing.
[0228] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0229] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0230] In existing technologies, systems that use natural language processing and generative artificial intelligence models to generate questions for users typically only assess the user's emotional state based on text input at a specific point in time or short-term dialogue context, and then generate subsequent questions accordingly. This approach has the following technical limitations: First, information processing devices are mostly centered on instant input, lacking systematic collection and structured processing of multi-source activity data generated by users at different times and locations (such as images collected by the terminal, location information, and social activity records provided by external services). This makes it impossible to extract stable behavioral patterns and emotional tendencies from long-term behavioral trajectories. Second, the prompts received by generative artificial intelligence models are often manually or simply pieced together according to rules, failing to be automatically constructed based on data-analyzed activity information. Therefore, they are not precise enough in terms of time, location, and behavioral context, affecting the relevance and depth of the questions generated by the model. Third, user answer records are often stored in the form of simple text logs, without establishing a strict time-series correlation with specific activity information and analysis results. Information processing devices cannot use these historical records to continuously improve the logic for generating subsequent questions, thus limiting the system's continuous learning ability in personalized reflection and guidance.
[0231] Therefore, an improved information processing technology is needed. This technology standardizes, tabulates, and performs statistical / machine learning analysis on the server side of activity information from terminals and external information sources. It automatically generates high-quality prompts containing time, location, and behavioral information. It also uses generative artificial intelligence models to generate questions that help users deeply review their inner state. At the same time, it establishes a time-series correlation between user responses, activity information, and analysis results. This improves the overall quality of text generation and user introspection support in human-computer interaction, and achieves synergistic optimization of data analysis and natural language generation processes in computer technology.
[0232] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0233] In this invention, the server includes means for acquiring information related to user behavior from an external information processing device or information recording device via a communication link, and storing the acquired behavior-related information as activity information including time and location information; means for performing data parsing processing, including statistical processing or machine learning processing, on the stored activity information to generate parsing results representing user behavior tendencies or emotional tendencies; means for automatically generating prompt statements input to a generative artificial intelligence model based on the parsing results and the activity information, and inputting the generated prompt statements into the generative artificial intelligence model so that the generative artificial intelligence model generates questions to prompt the user to recall their internal state; means for sending the generated questions to a terminal device and having the terminal device present the questions to the user; and means for acquiring user response information to the questions from the terminal device, storing the response information as record information, and associating the record information with the activity information or the parsing results to form a time-series record. This enables the formation of an integrated processing flow within the server, encompassing multi-source activity data acquisition, tabular storage, statistical and sentiment analysis, automatic construction of prompt statements, and generative AI model-driven question generation and response record association. This allows the generative AI model to reason with precise temporal and spatial contextual prompts, thereby improving the relevance and contextual consistency of question generation. Furthermore, the time-series-based record structure provides computable historical features for subsequent data analysis and model prompt statement optimization, achieving an overall improvement in the performance of computer data processing and natural language generation technologies.
[0234] A "system" refers to an overall device or combination of devices consisting of multiple functional components, used to perform data acquisition, storage, analysis, and problem generation through electronic computing devices.
[0235] A "server" is an information processing device equipped with a processor, memory, and communication interface, used to centrally perform data storage, data parsing, prompt statement construction, and calling generative artificial intelligence models.
[0236] "Terminal device" refers to an information processing device carried or operated by a user for collecting user behavior information, presenting questions to the user, and receiving responses, including but not limited to smartphones, tablet computers, or other computing terminals.
[0237] A "communication link" refers to a physical or logical communication path used to transmit data between servers, terminal devices, and external information processing devices, including wired and wireless communication.
[0238] "External information processing device" refers to other information processing systems or service platforms that are connected to a server via a network and are used to store or provide data related to user behavior.
[0239] "Information recording device" refers to a device used to electronically record and store user behavior data, response data, or other related data, including local storage devices and remote storage devices.
[0240] "User behavior related information" refers to data that reflects a user's activities in the real world or a virtual environment, including but not limited to image data, text data, location information, and time information.
[0241] "Activity information" refers to structured data that is compiled from information related to user behavior and contains at least time and location information, used to represent user behavior at a specific time and place.
[0242] "Time information" refers to data used to identify when an event occurs, including specific dates, times, or time periods.
[0243] "Location information" refers to spatial data used to identify the location where an activity takes place, including geographic coordinates, area identifiers, or location names.
[0244] "Storage" refers to the process of keeping data in a reusable form through storage media, including writing, updating, and managing data to databases or file systems.
[0245] "Data parsing and processing" refers to the computational processing performed on stored data to extract features, generate statistical results, or infer user behavior and emotional tendencies.
[0246] "Statistical processing" refers to mathematical and statistical methods performed on activity information, including counting, summing, grouping and summarizing, and calculating averages, to obtain behavioral characteristics or distribution information.
[0247] "Machine learning processing" refers to the process of using trained models to perform pattern recognition or prediction on activity information, thereby inferring user behavior or emotional tendencies.
[0248] "Behavioral tendencies" refer to the characteristics of a user's behavioral patterns or preferences in time and space, obtained by analyzing activity information over a period of time.
[0249] "Emotional tendency" refers to the characteristics of a user's emotional state (positive, neutral, or negative, etc.) obtained by analyzing text, images, or other signals related to user behavior.
[0250] "Analysis results" refers to the result data obtained from activity information through data analysis and processing, including behavioral tendencies, emotional tendencies, and related characteristic quantities.
[0251] "Generative artificial intelligence models" refer to data processing models trained using machine learning methods that can automatically generate natural language text based on input prompts.
[0252] "Prompt statements" refer to natural language text constructed as input to generative artificial intelligence models, used to provide the model with contextual information and generation constraints required to generate questions.
[0253] "Automatic prompt statement generation" refers to the process where the server automatically constructs prompt statements based on activity information and parsing results through program logic, rather than manually writing prompt statements.
[0254] "Question" refers to a natural language query text generated by a generative artificial intelligence model based on prompts, used to prompt users to recall or express their internal state.
[0255] "Internal state" refers to non-explicit information such as a user's psychological state, emotional feelings, or subjective evaluations when a specific activity occurs.
[0256] "Sending" refers to the process of transmitting data from a server to a terminal device or other device via a communication link.
[0257] "Presentation" refers to the processing of outputting problems or information in a form that is perceptible to the user on the display interface, audio output, or other human-computer interface of a terminal device.
[0258] "Response information" refers to the natural language text, voice content, or other forms of response data that a user inputs or provides in response to a presented question.
[0259] "Record information" refers to a data set formed by storing response information in a reusable form, which can be used in conjunction with other information.
[0260] "Time series records" refer to a collection of recorded information organized and stored in chronological order, where each record establishes a temporal correspondence with the corresponding activity information or analysis results.
[0261] "Table data" refers to data that represents activity information in a structured form with rows and columns, where each row corresponds to an activity record and each column corresponds to a field attribute.
[0262] "Summary processing" refers to the process of performing aggregation operations on specified dimensions based on tabular data, used to obtain statistical results by time unit or geographical region unit.
[0263] "Emotion estimation processing" refers to the process of inferring the emotion category or intensity from text or other signals related to activity information.
[0264] "Date information" refers to time data used to identify the date an activity occurs, and is used in prompt statements to indicate the date the activity occurred.
[0265] "Behavioral content information" refers to summary data describing a user's activities at a specific time and place, including activity type, text summary, or image summary.
[0266] “Text information” refers to natural language data recorded in the form of characters or strings that is related to user activities or responses.
[0267] "Image information" refers to visual content related to user activities that is recorded in the form of image data, including photographs or other image resources.
[0268] An "abstract" refers to the result or processing of extracting the main content from long or complex data and presenting the core information in a concise form.
[0269] "Free-form questions" refer to questions that do not have fixed options or sentence structures, are generated in natural language by generative artificial intelligence models, and allow users to answer them in an open manner.
[0270] In one implementation, the server operates as a central information processing device. Server hardware includes a multi-core general-purpose processor (e.g., an x86 architecture CPU), a graphics processing unit (e.g., a general-purpose graphics processing unit), main memory, and non-volatile storage devices (e.g., solid-state storage devices), and communicates with terminals and external information processing devices via a network interface. The server operating system can be a Unix-like operating system, on which a script interpreter (e.g., a Python interpreter), data analysis libraries (e.g., data frame processing libraries), network communication libraries, and generative artificial intelligence model calling libraries run. The server can also be configured with a dedicated database management system (e.g., a relational database or a document-oriented database) to store activity information, parsing results, questions, and user response information.
[0271] In one embodiment, the terminal is a portable information processing device, such as a smart mobile device with a touch display unit, a positioning module, an image acquisition module, and a wireless communication module. The terminal's operating system can be a mobile operating system, and applications running on the terminal access the photo album, location services, and network communication interfaces through the system's application programming interface (API). With user authorization, the terminal application collects image, time, and location data related to user behavior and exchanges information with the server via an encrypted communication protocol.
[0272] Users are the primary agents in human-computer interaction within the system. Through a terminal application interface, users can browse questions sent by the server, input natural language responses, and optionally upload a range of activity records, such as recently taken photos and their metadata.
[0273] In one implementation, the server uses a database management system to store activity information in a structured manner. Each activity record is represented as a record unit, which includes at least the following fields: user identifier, timestamp, location coordinates, activity source identifier (terminal or external service), optional text content summary, and optional image feature vector. In its implementation, the server can utilize a data analysis library to load the activity information into a tabular data structure, where each row corresponds to one activity record and each column corresponds to a specific data field. At the database level, the server establishes a composite index for the user identifier and timestamp fields to improve retrieval efficiency when querying specific user activity data in chronological order, thereby reducing latency in subsequent parsing and prompt statement construction.
[0274] In one implementation, the server performs data parsing and processing on activity information. The server utilizes a data frame processing library to perform preprocessing operations on the tabular data, including time standardization, missing value imputation, and outlier removal. During time standardization, the server converts all different time formats into a unified timestamp format. In location processing, the server maps the raw latitude and longitude vectors to administrative region identifiers via geocoding services and caches this mapping. This avoids multiple remote calls when parsing repeated activity locations, thereby reducing communication load and shortening parsing time. In this way, the server transforms raw activity information into structured data that can be efficiently indexed and aggregated in both time and space dimensions, improving upon the low query efficiency and high memory consumption problems caused by storing behavioral data in unstructured log format in traditional systems.
[0275] In one implementation, the server uses machine learning to estimate user behavioral and emotional tendencies. The server can pre-train a multi-layer neural network model offline for emotion classification of active text or image features. During training, the server uses a training dataset labeled with emotion tags and defines a cross-entropy loss function as the error function. The server updates the neural network weights using a gradient descent-based optimization algorithm (e.g., adaptive moment estimation). Data augmentation techniques can be used during training, such as synonym substitution and word order perturbation for text, and slight rotation, cropping, or brightness adjustment for images, to improve the model's robustness to diverse inputs in real-world scenarios. After training, the server stores the model parameters as a weight file and deploys it in the inference environment for online emotion estimation.
[0276] In online parsing, the server extracts text content from activity information (such as social media post content and terminal notes) and inputs it into the sentiment classification model. The server first performs word segmentation, stop word removal, and word vectorization on the text, using pre-trained word vectors or sub-word encoding methods. The server then inputs the word vector sequence into a neural network, such as an encoder structure incorporating multi-head attention. The network outputs a probability distribution of the sentiment category for each piece of text. The server uses the category with the highest probability as the sentiment label and writes this label along with a confidence value into the activity information record. Through this processing, the server transforms unstructured text, which is difficult to use directly for statistical analysis, into feature fields with clear category labels. This allows for precise filtering at the data level, combining sentiment trends, during subsequent behavioral pattern analysis and prompt generation, thereby improving the matching accuracy between generated questions and the user's internal state.
[0277] In behavioral tendency analysis, servers can employ clustering algorithms or sequence analysis methods. For example, a server can model the activity location sequence of the same user over multiple days, mapping the daily set of visited locations into a vector representation and using clustering algorithms to segment typical behavioral patterns such as "weekday pattern," "travel pattern," and "stay-at-home pattern." The server can also calculate statistical characteristics such as nighttime activity frequency and inter-city travel frequency within a specific time period. When constructing the analysis results, the server combines behavioral tendency features with emotional tendency features to generate a generalized analysis record. Therefore, with lower computational overhead, the server can extract high-level features suitable as input conditions for generative artificial intelligence models from a large amount of raw activity records, achieving compression of data dimensionality and feature space, thereby improving overall computational efficiency.
[0278] In one implementation, the server generates a prompt statement based on the parsed results. The server uses a template generation module to input time information, location information, behavioral content information, and sentiment tags into a natural language template. For example, after detecting that a user traveled to a specific city on a specific date and has a positive sentiment, the server can generate the following prompt statement text: "Based on the following information, please generate a Chinese question to help the user recall their feelings at the time:" Date: October 5, 2023 Location: Tokyo, Japan Scenario: The user took many travel photos on this day and posted on social media, saying, "A day traveling in Tokyo, the weather is great."
[0279] Instructions: Please generate only one open-ended question; do not provide any answers. In another scenario, when the server detects that a user has been working late into the night multiple times within a week and is showing signs of fatigue, it can construct the following prompt: "Here is a summary of the user's activities over the past week: working late into the night at the workplace for many days and handling work-related matters on weekends. Based on this information, please generate a gentle yet thought-provoking question to help the user reflect on their sources of stress and needs." When generating prompts, the server doesn't simply concatenate text; instead, it selects different template types and content field combinations based on quantified features from the parsed results. For example, when it detects that an emotional fluctuation exceeds a certain threshold, the server prioritizes templates that guide users to reflect on their emotional changes; when it detects travel or special events, it selects templates that guide users to record important experiences. Through this combination of rules and data-driven approaches, the server enables the generation of prompts with fine-grained conditional control that is difficult to achieve through intuition alone. This provides generative AI models with a structured and information-dense input context, significantly improving the relevance and diversity of the generated prompts.
[0280] In one implementation, the server invokes a generative AI model. The server can deploy the generative AI model on computing nodes equipped with graphics processing units (GPUs), and the model architecture can be an encoder-decoder architecture with multi-layered self-attention networks. The server receives the prompt text using a model service interface, encodes it into a sequence of subwords, and inputs it into the model's encoder. The encoder extracts semantic representations using a multi-head attention mechanism and a feedforward network. The server then generates the output question text word-by-word on the decoder side using conditional language modeling. During the decoding process, the server can set temperature parameters and a maximum generation length, and uses a beam search algorithm to select the output with the highest score from multiple candidate sequences. In this way, the server controls the length and diversity of the generated questions while ensuring semantic coherence.
[0281] In its implementation, the server maintains dedicated parameter configurations for generative AI models. For example, the server can select different model sub-configurations based on the task type indicated in the prompt: using a smaller decoding length and lower temperature parameters when generating short questions; and using a slightly higher temperature to increase creativity when open-ended, deeply reflective questions are needed. The server can explicitly add control markers to the prompt, such as "Question Type: Reflective" and "Tone: Mild," which the model is conditionally trained to use during training. During inference, the server guides the model to generate questions according to a specific style by passing in these markers. Through this parameter and marker control mechanism, question generation does not rely on manual, sentence-by-sentence editing but on computable rules and parameters, thus constructing a stable and reproducible data-driven generation process that achieves better accuracy and efficiency than traditional manual writing or simple template replacement.
[0282] In one implementation, the server stores the generated questions in a database and establishes foreign key associations with the corresponding activity information and parsing results. When sending questions to the terminal, the server includes a question identifier and the time and location information of the relevant activities. Upon receiving the questions, the terminal displays the question text in a graphical interface, and can simultaneously display associated thumbnails or map locations. In this way, the terminal provides users with a multimodal context, helping them recall the situation and thus improving the completeness and authenticity of their answers. This association between activity information and questions is not only a presentation optimization at the user interface level, but also requires the terminal and server to adhere to a unified data structure and identifier mapping at the underlying level. This enables rapid alignment of user answers with corresponding activity records during time-series analysis on the server side, reducing matching errors.
[0283] After reading the question on the terminal interface, the user inputs an answer in natural language. The terminal stores the text content, the answer time, and the question identifier in a local data storage structure and uploads it to the server via an encrypted channel when the network is available. After receiving the response information, the server stores it in the response table and associates it with the question record and activity information record. The server can further perform sentiment estimation and keyword extraction on the response text and write the results into additional fields. In this way, the server gradually builds a time series record spanning multiple days and multiple events. When analyzing historical responses, the server can then calculate indicators such as changes in response length, smoothness of the sentiment trajectory, and frequency of keyword occurrences, which are used for future selection of prompt statement templates and adjustment of feature weights, forming a closed-loop optimization mechanism.
[0284] The server in this system not only implements the conventional processes of data acquisition, analysis, and display, but also achieves technical improvements at multiple levels through specific data structure design, machine learning parsing methods, and generative artificial intelligence model control methods. By standardizing activity information into tables and establishing multi-dimensional indexes, the server reduces query latency and saves memory usage. Through deep model sentiment estimation of text and images, the server transforms unstructured data into computable features, thereby achieving more accurate feature matching in the selection of question generation conditions and reducing the probability of generating irrelevant or repetitive questions. By using conditional control tags and parameter-adjustable generative artificial intelligence models, the server enables the question generation process to automatically adapt to different situations without requiring manual configuration, which brings substantial improvements in the rational allocation of computing resources and generation efficiency.
[0285] The collaboration between the server and the terminal enables the system to operate with low communication load in real-world environments. The server utilizes location resolution result caching and prompt template reuse mechanisms to reduce repeated remote service calls and long text transmissions. The terminal only needs to upload metadata and condensed text when necessary, without transmitting all the original media data. When constructing prompts and generating questions, the server can perform calculations based on stored features rather than the original large file, thereby reducing network bandwidth consumption and storage pressure while ensuring question quality.
[0286] In another implementation, the server can employ different generative AI model architectures, such as an autoregressive language model with an encoder layer or a multi-task learning model, enabling it to perform both sentiment estimation and question generation. In this case, the server can define different loss functions at the model's multi-task output layer for sentiment classification and text generation respectively, and use joint training to update the weights of the shared representation layer. This integrated model reduces the number of models deployed, lowers memory usage, and reduces latency caused by multiple model calls during online inference.
[0287] The terminal can also include a voice input module in different implementations. When a user answers a question by voice, the recorded audio signal is converted into text by a local or cloud-based voice recognition module and then uploaded to the server. The server performs parsing processing on this text in the same way as keyboard input. In its implementation, the server can also dynamically adjust the data upload frequency and compression strategy based on the device status information uploaded by the terminal (such as battery level and network type) to avoid burdening the terminal's user experience, thereby achieving a balance between performance and user experience at the overall system level.
[0288] The server utilizes the aforementioned technologies to achieve a complete technical chain, from multi-source activity data parsing, automatic construction of prompt statements, and generative artificial intelligence model-driven question generation to the recording of user response time series. By employing specific data structures, well-trained neural network models, controllable parameter configuration, and caching and indexing mechanisms internally, the server outperforms traditional systems that rely solely on real-time input and simple rule generation in terms of accuracy, speed, and resource consumption. This solves the technical challenge in computer science of efficiently transforming complex behavioral data into high-quality natural language and forming computable time series records.
[0289] use Figure 13 The processing flow is explained.
[0290] Step 1: The server obtains activity information from terminals and external information processing devices. Inputs include image metadata, location information, timestamps uploaded by the terminals, and social activity data (text, image links, posting locations, etc.) returned by external service interfaces. The server uses a network communication library to receive this data via encrypted communication protocols, parses the scattered raw data into record units of a unified format, and outputs a set of activity information containing user identifiers, time information, location information, and original content fields.
[0291] Step 2: The server preprocesses and stores the activity information in a structured manner. The input is the set of activity information output from step 1. The server uses a data frame processing library to standardize the format of the time field, fill in or remove missing values, and convert the latitude and longitude coordinates into administrative region names using a geocoding service. The processed results are then loaded into a tabular data structure. The server subsequently uses a database management system to write each record into a database table as a row, and creates an index between the user identifier and the time field. The output is a structured dataset of activity information that can be efficiently retrieved by user and time.
[0292] Step 3: The server performs statistical processing on the structured activity information. The input is the tabular activity information dataset obtained in step 2. The server calls statistical functions to summarize the number of activities by date or time period, count the number of trips by geographical region, and calculate features such as nighttime activity frequency and cross-regional movement frequency. Through these arithmetic operations, grouping, and aggregation, the server outputs a statistical feature vector containing each time unit or region, which is used for subsequent behavioral tendency analysis.
[0293] Step 4: The server performs sentiment estimation processing on activity-related text. The input is the text field from step 2 (e.g., social media content, terminal notes, etc.). The server first uses word segmentation tools and word vector encoding algorithms to convert the text into a vector sequence, and then inputs this vector sequence into a pre-trained neural network sentiment classification model. The server performs forward propagation within the model, calculates the probability distribution of each sentiment category, and selects the category with the highest probability as the sentiment label. The label and its confidence level are written to the corresponding activity record, and the output is a set of activity information with a sentiment tendency field.
[0294] Step 5: The server generates analytical results by combining statistical and sentiment features. The input consists of the statistical feature vector from step 3 and the sentiment labels from step 4. The server uses a rule engine or clustering algorithm to divide multi-day activities into different behavioral patterns, such as weekday patterns, travel patterns, or stay-at-home patterns, and generates summaries of behavioral and sentiment tendencies for each time period. The server stores these summaries as records in the analytical results table, and the output is a collection of analytical results organized by user and time segment.
[0295] Step 6: The server selects suitable activity segments as the subjects of questions. The input consists of the activity information from step 2 and the parsing results from step 5. The server filters candidate dates or events based on preset rules (e.g., emotional fluctuations exceeding a threshold, activity type being travel, excessive overtime, etc.). During the filtering process, the server performs conditional checks on the parsing results, selecting a set of activity record identifiers that meet the criteria, and outputting a list of target activities for which questions will be generated.
[0296] Step 7: The server constructs prompt statements based on the target activities and the parsing results. The input is the list of target activities output in step 6 and their corresponding parsing summaries. The server uses a template generation module to fill in date information, location information, behavioral content information, and sentiment into a natural language template. Internally, the server selects different sentence templates based on different activity types, performs string concatenation and formatting on each field, and generates prompt text such as "Based on the following information, please generate a Chinese question to help the user recall their feelings at the time: Date: ... Location: ... Situation: ... ". The output is a set of one or more prompt statements for each target activity.
[0297] Step 8: The server inputs the prompt text into the generative AI model to generate questions. The input is the prompt text generated in step 7. The server encodes the prompt text into a sequence of subwords through the model service interface and sends it to the generative AI model deployed on the graphics processing unit. The server sets parameters such as decoding length and temperature in the model, allowing the model to perform multi-layer self-attention operations and conditional language generation, progressively outputting candidate question texts. The server then uses beam search or a scoring function to select semantically reasonable and appropriately sized questions from multiple candidates, outputting a final set of question texts corresponding to each prompt text.
[0298] Step 9: The server stores the generated questions and prepares to send them to the terminals. The input is the set of question texts output from step 8, along with their corresponding activity identifiers and parsing result identifiers. The server writes the question content, relationships, and generation time into the question database table and generates a unique identifier for each question. Simultaneously, the server constructs a response data structure containing the question text, question identifier, and brief activity information, outputting a list of questions that can be obtained by the terminals via an interface.
[0299] Step 10: The terminal retrieves and displays the questions from the server. The input is the list of questions provided by the server in step 9. The terminal calls the server interface via network request to receive and parse the question data. The terminal arranges the question text along with associated date and location information, presenting it to the user on the screen as cards or lists. It can also simultaneously display relevant image thumbnails or map markers on the interface. Through graphics rendering and event listening mechanisms, the terminal allows the user to visually view each question, outputting a question screen visible on the terminal interface.
[0300] Step 11: The user reads the question and enters an answer on the terminal. The input consists of the question text and related background information displayed on the terminal in step 10. The user enters their answer to the question in natural language in the input area provided on the terminal via a touch keyboard or voice input interface. After confirming the content, the user operates the send or save control, and the output is one or more segments of answer text or voice data associated with a specific question identifier.
[0301] Step 12: The terminal processes and uploads the user's response. The input is the user's response text or voice data generated in step 11. The terminal converts the voice data into text locally (e.g., using a speech recognition module) and packages the question identifier, response text, and response time into structured data. The terminal sends this data to the server interface using an encrypted communication protocol. The data can be compressed before sending to reduce bandwidth usage. The output is a response data packet transmitted to the server over the network.
[0302] Step 13: The server stores the responses and establishes a time-series association with the activity information. The input is the response data packet uploaded in step 12. After validating the data packet's fields, the server writes the response content into the response database table, recording its question identifier and response time. The server traces the associated activity information and parsing results using the question identifier, sorts these records along the time dimension, and generates a time-series index, thus constructing a time-series chain of "activity—question—response." The server output is a set of response records that can be searched chronologically and includes behavioral and emotional characteristics.
[0303] Step 14: The server utilizes time-series records to optimize subsequent prompts and model parameters. The input consists of the time-series records generated in step 13, along with historical questions, answers, and parsing results. The server performs statistical calculations on metrics such as answer text length, sentiment change trajectory, and keyword frequency to form a long-term user interaction feature vector. Based on these feature vectors, the server adjusts the prompt selection rules and control parameters of the generative AI model, such as changing the temperature, adjusting question length preferences, or modifying template selection probabilities. The server saves the new rules and parameter configurations in a configuration store, and the output is an optimized prompt generation strategy and model invocation configuration, used in subsequent loops to improve the relevance of question generation and overall system performance.
[0304] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0305] The following technical problems exist in existing computer-based interactive recording and content generation technologies.
[0306] First, servers typically only store or analyze natural language text entered by users in a single instance, making it difficult to integrate behavioral data generated by users across multiple terminals and scenarios (such as captured images, location information, and online service records). This prevents the system from accurately modeling the user's real-world context and emotional state. Consequently, subsequent questions or text content generated based on this data often become disconnected from the user's current behavioral context, resulting in low system interaction efficiency and user engagement.
[0307] Second, even when existing systems introduce generative artificial intelligence models, they mostly use the user's original input directly as the model input. They lack a structured prompt generation mechanism that comprehensively processes user behavior information, environmental information, and emotional state. This makes it difficult for generative artificial intelligence models to make full use of contextual information, resulting in output results that are highly random, lack specificity, and are difficult to be stably used by downstream processing modules. As a result, the processing quality and controllability of the entire computer system in tasks such as automatic questioning, diary generation, and personalized content generation are limited.
[0308] Third, in many systems, user response data is simply saved as logs without being organically linked to user behavior information, environmental information, and sentiment analysis results at the data level. Servers lack a unified data structure and processing flow to jointly model and continuously utilize multi-source heterogeneous data, making it impossible to efficiently reuse existing records in subsequent rounds of interaction, and also making it difficult to perform efficient calculation and retrieval analysis in large-scale user scenarios, resulting in low utilization efficiency of storage and computing resources.
[0309] Fourth, traditional content generation processes are often linear pipelines of "front-end data collection - back-end simple processing - direct text generation," lacking a closed-loop optimization mechanism on the server side for "prompt statement construction - generative AI model invocation - structured saving of generated results - reconstructing prompt statements." This makes it difficult to develop scalable, reusable, and optimizable system-level algorithmic processes for complex tasks such as user experience recording, sentiment analysis, and personalized suggestion generation. As a result, the system needs to design independent program logic for different business scenarios (such as store feedback collection, personal diary generation, and emotion-driven personalized recommendations), increasing system complexity and maintenance costs.
[0310] Therefore, there is a need for a system-level technical solution that unifies the collection and association of user behavior information and environmental information on the server side, uses natural language processing and sentiment analysis to comprehensively analyze multi-source data, and drives the work of generative artificial intelligence models by constructing targeted prompts, so as to improve the processing efficiency, generation quality and resource utilization of computers in interactive content generation, data recording and personalized information provision.
[0311] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0312] In this invention, the server includes a device for acquiring user behavior information and environmental information; a device for constructing a prompt statement based on the acquired user behavior information and environmental information to generate a prompt statement for prompting the user, and instructing a generative artificial intelligence model to process the prompt statement as input; a device for sending the prompt statement output by the generative artificial intelligence model to the user terminal and associating and recording the answer information obtained from the user with the user behavior information and the environmental information; and a device for applying natural language processing technology and sentiment analysis to the recorded answer information and user behavior information, constructing a prompt statement based on the analysis result to automatically generate a record information representing user experience, instructing the generative artificial intelligence model to process the prompt statement as input, and saving the record information output by the generative artificial intelligence model to a recording medium for later review. This allows for the unified collection and association of multi-source behavioral and environmental data on the server side. By constructing structured prompts, generative artificial intelligence models can be efficiently scheduled, ensuring that the questions and recorded information generated by the models are highly matched with the user's actual situation and emotional state. This improves the processing performance, generation quality, and utilization efficiency of storage and computing resources in the computer system's interactive content generation and data recording tasks, and lays a unified data and algorithm foundation for subsequent personalized suggestions or information.
[0313] "User behavior information" refers to data related to user behavior obtained by computing devices based on user operations or activities, including but not limited to image data obtained by imaging devices, location data obtained by location detection devices, input data obtained by interactive interfaces, and usage data recorded by online services.
[0314] "Environmental information" refers to data used to represent the state of a user's environment, including but not limited to geographic location information, time information, location type information, network service information, and other contextual data related to the user's surroundings.
[0315] "Imaging device" refers to a hardware device used to acquire scene images or video data, including but not limited to camera modules, image sensors and their associated image acquisition circuits.
[0316] "Image information" refers to digital image data acquired by an imaging device and processed through encoding, including still image data and frame data extracted from a video stream.
[0317] "Location detection device" refers to a hardware device or combination thereof used to detect the spatial location information of a user or terminal, including but not limited to satellite positioning modules, base station positioning modules, short-range wireless positioning modules and software components used in conjunction with them.
[0318] "Location information" refers to data acquired by a location detection device to represent a geographic location, including but not limited to latitude and longitude coordinates, altitude, location accuracy information, and information corresponding to location identifiers.
[0319] "Information sharing services" refers to network service platforms that provide information publishing, storage, browsing, or exchange functions over communication networks, including but not limited to social networking services, media content sharing services, and online diary services.
[0320] "User terminal" refers to an electronic device that allows users to input information, browse or interact with it, including but not limited to mobile terminals, tablet terminals, portable computing devices and desktop computing devices.
[0321] "Prompt statements" refer to text data constructed by the server and provided as input to the generative AI model to instruct the generative AI model to perform specific generation tasks and limit the type and style of generated content.
[0322] "Generative artificial intelligence models" refer to data processing models built on machine learning algorithms that can generate natural language text or other content based on input prompts, including natural language generation models using deep neural network structures.
[0323] "Question content" refers to the query text generated by a generative artificial intelligence model based on prompts and used to guide users to answer specific events, experiences, or emotions.
[0324] "Response information" refers to the response data entered by the user in response to the questions prompted by the system, including the response entered in the form of natural language text and the time and context information associated with the response.
[0325] "Recorded information" refers to text data automatically generated based on user behavior information, environmental information, response information, and their parsing results, used to characterize user experience or status, including but not limited to diary text, summary text, and personalized content text.
[0326] Natural Language Processing (NLP) technology refers to algorithms or programs used to analyze and process natural language text, including but not limited to word segmentation, syntactic analysis, entity recognition, keyword extraction, and semantic representation construction.
[0327] "Sentiment analysis processing" refers to the analysis of user input text or related data to determine specific emotional states, including but not limited to algorithms or models for judging emotional polarity, emotional category, and emotional intensity.
[0328] "Personalized content" refers to text content that is customized for a specific user based on user behavior information, environmental information, response information, and emotional state, including but not limited to suggestions, product information, and service information.
[0329] "Suggestion information" refers to text content that the system automatically generates based on the user's emotional state or behavioral context, used to provide behavioral guidance or psychological comfort to the user.
[0330] "Product information" refers to descriptive data related to items available for purchase by users, including but not limited to product name, function description, usage scenarios, and reasons for recommendation.
[0331] "Service information" refers to descriptive data related to the service items available to users, including but not limited to service type, service content, applicable scenarios, and reasons for recommendation.
[0332] "Recording medium" refers to a data storage medium used to preserve recorded information in a reproducible form, including but not limited to semiconductor storage devices, magnetic storage devices, optical storage devices, and their logical representation in a network storage system.
[0333] In this invention, the server operates as the core data processing node, and the terminal operates as the data acquisition and presentation node. Users interact with the server through the terminal. The following provides a detailed description of the hardware structure, software modules, data structure, generative artificial intelligence model configuration, prompt statement construction method, and technical effects to enable those skilled in the art to implement this invention.
[0334] I. Overall System Composition The server includes at least one processor, main memory, non-volatile storage, a network interface, and optional hardware accelerators. The processor may be a multi-core general-purpose processor, or may further include a graphics processing unit (GPU) or a dedicated accelerator chip for neural network inference operations. The non-volatile storage is used to store application programs, model parameters, and log data, while the main memory is used to cache behavioral data, environmental data, and intermediate processing results during runtime. The network interface enables bidirectional communication with multiple terminals via a wireless or wired network.
[0335] The terminal includes a processor, memory, display device, imaging device, and location detection device. The imaging device is, for example, a camera integrated into the mobile terminal, and the location detection device is, for example, a satellite positioning module or a positioning module based on a base station or wireless local area network. The terminal executes a client program to collect image information, location information, and natural language text input by the user, and sends this data to the server via a communication network. Simultaneously, it receives and displays question content, recorded information, or personalized content generated by the server.
[0336] Users operate the device through a graphical user interface, including triggering shooting, authorizing access to external services, entering text responses, and browsing generated recorded information.
[0337] II. Server Software Modules and Data Structures The server's software structure includes a communication module, a data storage module, a preprocessing module, a feature extraction module, a sentiment analysis module, a prompt statement generation module, a generative artificial intelligence model interface module, and a record information generation and persistence module.
[0338] The server defines various data tables or logical collections at the data layer. The server records user behavior information as records belonging to the behavior data collection. Each record includes: behavior identifier, user identifier, timestamp, behavior type, image identifier, location information identifier, and association identifier with external service resources. The server records environmental information as records in the environmental data collection, including time information, location category, venue type, and indicators related to network service status. The server records user answer information as records in the answer data collection and establishes foreign key relationships between these records and the corresponding question content identifier and behavior identifier. The server records generated record information (such as diary text or personalized content) as records in the record information collection, which includes index fields and summary fields for retrieval.
[0339] This data structure enables the server to perform correlation queries and joint analysis on multi-source heterogeneous data at the database level, thereby improving data management efficiency and subsequent computing efficiency.
[0340] III. Acquisition and Preprocessing of User Behavior Information and Environmental Information When a user takes a photo, the terminal invokes the imaging device to capture image information and saves the image file locally. Simultaneously, the terminal invokes the location detection device to obtain the latitude, longitude, and accuracy information of the current location and records the time of the photo capture. The terminal further obtains time information and information about the currently connected network service, such as the access point identifier.
[0341] The terminal packages image data, location information, and time information into structured data via the communication module and sends it to the server. Upon receiving the data, the server performs basic image processing in its preprocessing module, such as thumbnail generation, image hash calculation, or simple scene classification; it also geocodes the location information and queries the location name and type, such as "shopping area" or "public green space." These processing steps provide structured semantic features for subsequent feature extraction and prompt construction.
[0342] IV. Natural Language Processing and Sentiment Analysis The server uses natural language processing (NLP) technology in its feature extraction module to process the user's input response information. The server uses a software library that includes functions such as word segmentation, part-of-speech tagging, named entity recognition, and dependency parsing. The server extracts time expressions, location expressions, names of people, emotion-related words (such as "happy," "tired," and "anxious"), and verb phrases describing activities from the response text to form a structured text feature vector.
[0343] The server uses a pre-trained sentiment classification model in its sentiment analysis module. This model can employ a bidirectional recurrent neural network or an attention-based encoder structure. The server inputs the word vector sequence obtained from the natural language processing module into the sentiment classification model, outputting a sentiment category (e.g., positive, neutral, negative, or more specific terms like "happy," "fatigued," "nervous," etc.) and a sentiment intensity score. The server can also calculate independent sentiment scores for different sentences, thus highlighting the emotional expression of key sentences when generating subsequent record information.
[0344] By employing structured natural language processing and sentiment analysis, the server can accurately capture implicit emotions in complex sentences and multi-sentence texts compared to simple keyword matching, thereby improving the relevance and accuracy of subsequent generated content.
[0345] V. Construction of Prompt Statements and Invocation of Generative Artificial Intelligence Models In the prompt generation module, the server combines user behavior information, environmental information, extracted text features, and sentiment analysis results to construct prompts suitable for driving generative artificial intelligence models. Instead of directly providing raw data to the model, the server employs specific data abstraction and templating strategies to make the generation task more controllable.
[0346] For example, the server constructs the following prompt statement when generating the problem content: "Please generate a question based on a product photo taken by a user at a physical location to guide the user in describing their subjective feelings about the product. The photo was taken at a shopping location at night." For example, the server might construct the following prompt statement when generating log information: "The user's experience today is as follows: 'I had a picnic with friends in the park today. The weather was great, and I felt very relaxed.' Please write a first-person diary entry of about 300 words, focusing on the emotional changes you experienced, and adding some environmental details and inner monologue." The server constructs the following prompt statement during the generation of emotion-driven personalized content: "The user says: 'I'm very tired today and just want to be alone in peace.' Please write a short text of 200 to 300 words in a gentle tone to help the user relax and provide three simple and actionable suggestions for rest." The server passes these prompts to the generative AI model interface module. The generative AI model can be a language generation model based on a transformer architecture, featuring multi-layered self-attention encoders and decoders. During training, it uses an autoregressive objective function and is trained on a large amount of text to learn the distribution of natural language. When calling the model, the server can set parameters such as temperature, sampling strategy, and maximum output length to control the determinism and diversity of the generated content.
[0347] The server encodes the input prompts and calls a generative AI model service deployed on dedicated computing resources via a network interface. The model generates question content or log information based on the prompts. After receiving the model output, the server performs basic format validation and length correction, and optionally filters sensitive content.
[0348] By incorporating structured features into prompts and adopting a unified format, the server can significantly reduce the unpredictability and noise of the generated results, making the generated content more consistent with the context and purpose of the target task, thereby improving the reasoning performance of the generative model in specific application scenarios.
[0349] VI. Generation and Storage of Recorded Information In the record information generation and persistence module, the server establishes a mapping relationship between the record information output by the generative artificial intelligence model and the corresponding behavioral information, environmental information, and sentiment analysis results. The server generates a unique identifier for each record and stores it in the record information collection. At the same time, it writes the summary field and index field into the retrieval engine to support fast retrieval based on timeline, location, or sentiment tags.
[0350] The server provides the recorded information to the terminal, which then presents it to the user in diary format on the display device. Users can query recorded information from any past time period through the terminal, and the server efficiently returns results based on the index structure, thereby improving search response speed.
[0351] VII. Integration with real-world technological applications and technological effects The server links physical world perception data, such as image and location information, with user text input data within the same data structure, directly binding the abstract language generation process to the actual environmental scene. This multimodal fusion processing is used to write to recording media, improving the fidelity of the recorded content to the real-world context. This has practical technical applications in areas such as collecting feedback from physical locations, environmental experience logs, and travel records.
[0352] The server employs a unified prompt statement construction strategy and a centralized generative AI model service, enabling the system to reuse model capabilities to handle various types of tasks, including question generation, log generation, and personalized suggestion generation, without altering the business front-end. This unified architecture reduces repetitive work in implementing natural language generation logic independently across different business lines, lowers overall system complexity, improves maintainability, and reduces communication load through centralized optimization of model inference bandwidth usage and caching strategies.
[0353] In terms of computational performance, the server performs feature compression and filtering before constructing the prompt statements, retaining only abstract features relevant to the generation task, rather than directly feeding lengthy raw data into the model. This approach reduces the length of the model input and the complexity of the context, shortens inference time, and reduces memory usage, thereby increasing the number of requests that can be processed per unit time under the same hardware conditions.
[0354] In terms of accuracy, the server performs structured analysis on the user's text in advance through the natural language processing module and the sentiment analysis module, and then explicitly encodes the analysis results into the prompt statement. Compared with the method of directly using the user's original text without analysis, this design makes emotional information and key entity information occupy a higher weight in the model input, which helps the generated results to be more consistent with the user's historical behavior and current emotions in terms of sentiment expression and detail selection, and reduces misunderstandings and semantic biases.
[0355] The sentiment analysis model used by the server employs a labeled dataset during training, updating parameters using the cross-entropy loss function or other loss functions suitable for classification tasks, and iteratively training using gradient descent and its variants. Data augmentation methods, such as synonym replacement and sentence perturbation, can be incorporated during training to improve the model's robustness to real user input. Through this training process, the server can obtain stable sentiment judgment results with low computational cost during the inference phase, providing high-quality features for constructing prompt statements.
[0356] VIII. Server-side Non-standard Processing Flows and Technical Improvements In this invention, the server does not operate in a simple linear "user input - model output" process, but rather employs a non-conventional architecture that includes multi-round feature processing and multi-level prompt statement construction. The server first performs structured modeling of behavioral and environmental information, and then constructs a first type of prompt statement based on this model for question content generation. Subsequently, the server persists the user's answers along with the behavioral and environmental data, and incorporates sentiment analysis results when constructing a second type of prompt statement to generate memorized recorded information or personalized content.
[0357] This two-stage process, which combines "behavior / environment-driven question generation" and "response / emotion-driven record generation," prevents generative AI models from working in isolation. Instead, they become part of the collaborative optimization of multiple server modules. Through this collaborative mechanism, the server achieves controllable guidance of model behavior, reducing the uncertainty and repetitiveness of traditional generative models in business scenarios. This, in turn, improves the efficiency and usability of content generation based on generative AI models at the system level.
[0358] IX. Multiple Implementation Forms and Variations In one implementation, the server can deploy the generative AI model on an external service platform and call it through a network interface; in another implementation, the server can deploy acceleration hardware locally and run a simplified version of the generative model locally, thereby reducing dependence on the external network and improving response speed.
[0359] In one variation, the server can employ a multi-model combination structure. For example, a smaller model can be used to quickly generate an initial draft, while a larger model can be used to refine the selected segment. The server can also select different sets of model parameters based on the request type. For example, a lower maximum output length and a lower temperature parameter can be used for short question generation tasks, while a higher maximum output length and an appropriate temperature parameter can be used for long diary generation tasks to balance generation speed and text richness.
[0360] In one embodiment, the terminal can be a smartphone; in another, it can be an in-vehicle terminal or a wearable device. As long as the terminal can collect at least one type of behavioral or environmental information and can communicate bidirectionally with the server, it can be adapted to the system of this invention.
[0361] Through the coordination of the aforementioned modules, data structures, and processes, the server, terminal, and user collaboratively complete the entire process from real-world data collection, feature processing, prompt statement construction, generative artificial intelligence model invocation, to the persistence of recorded information. This invention, by introducing structured prompt statements and a multi-level feature processing mechanism, achieves fine-grained scheduling and context enhancement of the generative artificial intelligence model within the computer. This technically improves the accuracy, efficiency, and resource utilization of content generation tasks, going beyond a simple automated replacement of manual recording or questioning processes; it represents a substantial improvement in computer language processing and data management capabilities.
[0362] use Figure 14 The processing flow is explained.
[0363] Step 1: The user initiates a recording operation on the terminal. The input is the user's operation event on the terminal interface (such as clicking the "Start Recording" button), and the output is a status flag indicating that the terminal has entered data acquisition mode. Based on the user's operation event, the terminal sets a session identifier in local memory and initializes a buffer for caching image data, location information, and text input, preparing for subsequent data acquisition.
[0364] Step 2: The user collects behavioral and environmental information on the terminal. Inputs include the user's specific actions (e.g., taking a photo, granting location access) and terminal sensor outputs (raw image data from the camera, raw position signal from the positioning module). Outputs are structured behavioral and environmental information. The terminal calls the imaging device to capture a single frame image, converts the raw image into a JPEG or PNG file using an encoding algorithm, and records the file path or URI. The terminal calls the location detection device to obtain latitude, longitude, precision, and timestamp, combining these values into a location information object. The terminal appends the shooting time, device ID, network status, etc., to this object, forming a data structure for behavioral and environmental information.
[0365] Step 3: The terminal sends the collected behavioral and environmental information to the server. The input consists of the image file, location information, and time information generated by the terminal in step 2; the output is a request data packet transmitted to the server over the network. The terminal constructs an HTTP or HTTPS request, serializes the image data (or its URL), location information (latitude and longitude, timestamp), user identifier, and session identifier into JSON or form format, and calls the network communication library to send it to the server's specified API endpoint via the network interface, thus uploading the raw multi-source data.
[0366] Step 4: The server receives and parses uploaded data from the terminal. The input is a network request data packet sent by the terminal (containing image data, location information, user identifier, etc.), and the output is a behavior data object and an environment data object generated in the server's memory. In the communication module, the server parses the request header and body, verifying data integrity and format; the server saves the image data to the file system or object storage, recording the image path; the server converts the location and time information into an internally unified coordinate format, combines it with the user identifier and image path to form a behavior data record, and writes it to the database or cache structure.
[0367] Step 5: The server preprocesses and extracts features from behavioral and environmental information. The input consists of the behavioral and environmental data objects obtained in step 4, and the output is an abstract feature set used to construct prompt statements. The server calls a geocoding service to convert latitude and longitude into human-readable location names and location type labels (e.g., "shopping mall" or "park"). The server calls an image analysis module (e.g., a pre-trained classification model) to obtain a rough scene category (e.g., "product shelf" or "scenery"). The server encodes location category, location type, and time period (e.g., "evening") into discrete features or text description strings, storing them in a feature structure as input for subsequent prompt statement generation.
[0368] Step 6: The server constructs a prompt statement to generate the question content. The input is the feature set obtained in step 5 (location type, time information, scene category, etc.), and the output is the prompt statement text describing the generation task in natural language. The server selects the corresponding template based on the location type, for example, for a shopping mall, the template is "Please generate a question based on a product-related scene photographed by the user at the shopping mall." The server fills the placeholders in the template with specific features to form the complete text, for example, "Please generate a question based on a product photo taken by the user at the shopping mall to guide the user in describing their subjective feelings about the product. The photo was taken at the shopping mall at night." This text is then cached as a prompt statement object.
[0369] Step 7: The server invokes a generative AI model to generate the question content. The input is the prompt text constructed in step 6, and the output is the question content text generated by the generative AI model. The server uses the generative AI model interface module, taking the prompt text as model input, setting parameters such as maximum output length, temperature, and sampling strategy, and then sending the request to the model inference service. The server receives the text sequence output by the model, parses it into a complete question, such as "What was your first impression when you saw this product?", and performs length and character set validation to obtain the final question content.
[0370] Step 8: The server sends the generated question content to the terminal. The input is the question text generated in step 7 and the corresponding behavior data identifier, and the output is a response message carrying the question content. The server constructs a response object in the communication module, packages the question text, session identifier, and question identifier together into a JSON structure, and returns it to the terminal that initiated the request; if a push method is used, the server calls the push service interface to send the question text as notification content to the specified terminal.
[0371] Step 9: The terminal receives and displays the question content to the user. The input is the response message returned by the server (containing the question text and question identifier), and the output is the question interface presented on the terminal display device. The terminal parses the JSON in the response, extracts the question text and related identifiers, displays the question text in the graphical interface, creates a text input box for the user to enter their answer, and records the question identifier in its local state for later reference when uploading an answer.
[0372] Step 10: Users input their answers to questions via the terminal. The input consists of the question text displayed on the terminal interface and the user's typed natural language response; the output is a temporarily stored answer string on the terminal. Users input their answers using either a virtual or physical keyboard, such as "This red dress looks very elegant and would be perfect for a party." The terminal locally stores the answer text along with the corresponding question identifier and timestamp in its session cache, preparing it for uploading to the server.
[0373] Step 11: The terminal sends the user's answer information to the server. The input consists of the answer text, question identifier, and session identifier cached in step 10, and the output is a network request containing the answer information. The terminal constructs an HTTP or HTTPS POST request, serializes the answer text, question identifier, user identifier, and time information into a JSON structure, and sends it to the server's specified answer receiving API through the network interface, thereby enabling the uploading of the user's natural language answer.
[0374] Step 12: The server receives and stores user response information. The input is the response request data packet uploaded by the terminal, and the output is a newly added response record in the database. The server parses the request body, extracts the response text, question identifier, user identifier, and timestamp, and generates a response data object. The server writes this object into the response data table and associates it with the behavior data record and question content record through foreign key fields, thus establishing a three-way relationship of "behavior-question-response" at the data level.
[0375] Step 13: The server performs natural language processing and sentiment analysis on the response information. The input is the response text obtained in step 12, and the output is structured text features and sentiment tags. The server calls a natural language processing library to perform word segmentation, part-of-speech tagging, and dependency parsing on the response text to extract keywords, subject-verb-object structures, and important phrases. The server maps the word segmentation results into word vectors or embedding representations, inputs them into the sentiment classification model, calculates the probability distribution of each sentiment category, determines the dominant sentiment category (e.g., "positive and happy") and the sentiment intensity value, and writes these results as sentiment analysis output into the extended fields of the response record.
[0376] Step 14: The server constructs prompt statements to generate record information. Inputs include the user's response text, behavioral information (such as location and time), environmental information, and the sentiment analysis results obtained in step 13. Outputs are the prompt statement text used to drive the generative AI model to generate record information. The server selects different templates based on the task type (such as generating a diary or generating personalized suggestions). For example, a diary generation template might be: "The user's experience today is as follows: '{response text}'. Please write a first-person diary entry of approximately 300 words, focusing on the emotional changes and adding some environmental details and inner monologue." The server inserts the actual response text along with location, time, and other information into the template to form a complete prompt statement and saves it as a new prompt statement object.
[0377] Step 15: The server invokes a generative AI model to generate recorded information. The input is the prompt text constructed in step 14, and the output is the recorded information text (such as diary content or a long experience record). The server sends the prompt text to the model service through the generative AI model interface, setting the output length and tone control parameters. The server receives the text generated by the model, performs basic formatting (removing extra blank lines and standardizing punctuation), and checks whether it contains complete emotional cues and scene descriptions, ultimately obtaining usable recorded information text.
[0378] Step 16: The server associates and saves the recorded information with the behavioral and response information. The input is the recorded information text generated in step 15, along with the associated behavioral and response identifiers. The output is a new record in the recorded information set. The server generates a recorded information ID and stores the recorded information text, along with the user identifier, behavioral identifier, response identifier, and emotion tag, into the recorded information table. Simultaneously, the server generates an index for the recorded information (e.g., an inverted index by date, location, and emotion tag) to facilitate faster subsequent retrieval and improve query efficiency.
[0379] Step 17: The server returns or pushes the generated log information to the terminal. The input is the log information identifier and log information text saved in step 16, and the output is the response message sent to the terminal. The server sends the log information text and related metadata to the user terminal via API response or message push, enabling the terminal to display the generated diary or experience record on the interface.
[0380] Step 18: The terminal displays recorded information to the user and supports subsequent access. The input is the recorded information text and metadata returned by the server, and the output is the recorded content interface presented on the terminal's display device. The terminal stores necessary summaries of the recorded information locally and updates the local list; when the user later selects a date or location, the terminal requests the corresponding recorded information from the server based on the metadata. The server uses the previously built index to quickly locate and return the result, thus achieving efficient access and browsing of historical records.
[0381] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0382] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0383] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0384] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0385] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0386] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0387] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0388] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0389] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0390] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0391] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0392] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0393] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0394] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0395] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0396] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0397] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0398] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0399] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0400] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0401] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0402] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0403] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0404] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0405] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0406] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0407] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0408] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0409] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0410] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0411] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0412] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0413] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0414] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0415] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0416] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0417] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0418] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0419] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0420] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0421] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0422] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0423] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0424] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0425] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0426] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0427] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0428] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0429] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0430] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0431] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0432] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0433] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0434] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0435] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0436] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0437] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0438] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0439] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0440] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0441] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0442] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0443] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0444] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0445] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0446] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0447] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0448] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0449] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0450] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0451] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0452] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0453] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0454] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0455] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0456] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0457] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0458] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0459] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0460] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0461] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0462] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that executes specific processes by executing software, i.e., a program. Furthermore, processors can include, for example, FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are dedicated circuits with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0463] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0464] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0465] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0466] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0467] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0468] In addition, the following notes are provided in response to the above explanation.
[0469] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for controlling information input and output, presenting a user with query information for writing a diary in a conversational format, and receiving response information from the user; An apparatus for parsing a set of past response information stored in a storage device using natural language processing techniques, extracting the user's emotional state and areas of interest from the information, and generating parsed information that summarizes the extraction results. A device for constructing a prompt statement for inputting into a generative artificial intelligence model based on the parsed information and the currently input response information, and for instructing the generative artificial intelligence model to generate the next query information to be presented to the user. A device for obtaining the next query information output from the generative artificial intelligence model and controlling the information input / output device to present the next query information to the user; A means for registering the response information and the correspondence between the emotional state and focus area extracted by the natural language processing technology and the information of the next inquiry as diary data in the storage device in chronological order; A device for generating historical information that can refer to the conversation history and emotional changes of each user, based on the diary data registered in the storage device.
[0470] (Note 2) The information processing system according to Appendix 1 is characterized in that, When constructing the prompt statement, control information is added to each emotional state and each area of interest contained in the parsed information to specify the tone, level of detail, and depth of the inquiry, and the next inquiry information generated based on the control information enables a personalized dialogue that adapts to the user's emotional state.
[0471] (Note 3) The information processing system according to Appendix 1 is characterized in that, A device for acquiring the diary data registered in the storage device, analyzing changes in emotional state and areas of interest within a predetermined time period using the natural language processing technology, and automatically generating summary diaries or statistical information to be presented to the user based on the analysis results.
[0472] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for communicating with a terminal, including a display device and an input device, via a communication device, thereby presenting a question to a user in a conversational manner and obtaining the user's response information; An apparatus for performing natural language processing analysis on acquired response information and external data obtained from external information sources to assess user emotional state and behavioral tendencies, and for storing the response information and the external data as diary data in association. A means for generating prompts based on the diary data and the external data, which are then input to a generative artificial intelligence model to instruct the generative artificial intelligence model to generate the next round of conversational questions; A device for generating prompts based on the diary data and the emotional state, which are then input to a generative artificial intelligence model to instruct the model to generate personalized information that aligns with the user's interests and emotions. A device for acquiring the question and the personalized information output by the generative artificial intelligence model, and sending them to the terminal via the communication device for presentation to the user on the terminal; An apparatus for updating the content or structure of a prompt statement based on feedback information obtained from the user and the presentation history of the personalized information, so as to adaptively change the input conditions of the model used for subsequent question generation and personalized information generation.
[0473] (Note 2) According to the information processing system described in Appendix 1, the terminal is configured to use a location acquisition device and communication functions with external services to send location data and publication data obtained from electronic information services as external data to the communication device, and to use the external data to generate prompt statements input to the generative artificial intelligence model.
[0474] (Note 3) According to the information processing system described in Appendix 1, the processing device in the system is configured to generate prompt statements input to the generative artificial intelligence model based on emotional states and topic information extracted from the diary data, and to generate additional input questions and personalized information containing suggestions for user behavior through the generative artificial intelligence model to prompt users to reflect on themselves.
[0475] Example 2 (Note 1) An information processing system, characterized in that it comprises: It is a means of acquiring information related to user behavior from external information processing devices or information recording devices through a communication link, and storing the acquired behavior-related information as activity information including time and location information; A means of performing data parsing processing, including statistical processing or machine learning processing, on stored activity information to generate analytical results that represent user behavioral or emotional tendencies; Based on the analysis results and the activity information, the system automatically generates prompts for input into the generative artificial intelligence model and inputs the generated prompts into the generative artificial intelligence model, so that the generative artificial intelligence model generates a means to prompt the user to recall their internal state. Means for sending a generated question to a terminal device and causing the terminal device to present the question to a user; A means for obtaining user response information to the question from the terminal device, storing the response information as record information, and associating the record information with the activity information or the parsing result to form a time-series record.
[0476] (Note 2) The information processing system according to Appendix 1 is characterized in that, The data parsing and processing includes: storing the activity information as tabular data, performing summary processing and sentiment estimation processing on the tabular data according to time units or geographical area units, and determining the date information, location information, and behavioral content information to be included in the prompt statement based on the processing results.
[0477] (Note 3) The information processing system according to Appendix 1 is characterized in that, The prompt statement includes the time and location of a specific user activity, as well as a summary of text or image information related to the activity, and the generative artificial intelligence model is configured to generate a question in a free-form descriptive form about the emotion or meaning attributed to the activity at that time based on the prompt statement.
[0478] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: Devices for acquiring user behavior information and environmental information; A device for constructing prompt statements for generating prompts to users based on acquired user behavior information and environmental information, and for instructing a generative artificial intelligence model to process the prompt statements as input; An apparatus for sending the question content output by a generative artificial intelligence model to a user terminal, and for associating and recording the answer information obtained from the user with the user behavior information and the environmental information; A device for applying natural language processing and sentiment analysis to the recorded answer information and the user behavior information, constructing prompt statements to automatically generate recorded information representing user experience based on the analysis results, and instructing a generative artificial intelligence model to process the prompt statements as input. An apparatus for saving the recorded information output by a generative artificial intelligence model to a recording medium for later viewing.
[0479] (Note 2) The information processing system according to Appendix 1 is characterized in that, The user behavior information and environmental information include at least one of image information acquired by an imaging device, location information acquired by a location detection device, and information acquired from an information sharing service, and the device for constructing the prompt statement is configured to construct the prompt statement based on the information so that the generative artificial intelligence model generates question content related to events that are presumed to be of high user interest.
[0480] (Note 3) The information processing system according to Appendix 1 is characterized in that, The apparatus for constructing prompt statements for automatically generating the recorded information is configured to construct prompt statements based on user response information and specific emotional states processed by the sentiment analysis, so as to instruct a generative artificial intelligence model to generate personalized content containing user-oriented suggestion information, product information, or service information.
Claims
1. An information processing system, characterized in that, include: processor; The processor is configured as follows: Provide an interface for receiving user information; Natural language processing technology is used to analyze the received information in order to assess the user's emotional state; Based on the results of sentiment analysis, prompt text is generated to instruct the generative AI model to generate the next question.
2. The information processing system according to claim 1, characterized in that, The processor is also configured to input the prompt text into the generative artificial intelligence model to generate a question for presentation to the user.
3. The information processing system according to claim 1, characterized in that, The processor is also configured to: record the user's answers to a database; and automatically generate a diary based on the content of the answers and the sentiment analysis results.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A