Information processing system
Patent Information
- Application Number
- CN202610234156.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-02-27
- Publication Date
- 2026-09-22
AI Technical Summary
例如,在办理地址变更手续时,用户往往需要在多个不同的网站或系统中重复填写姓名、联系方式、旧地址、新地址等信息,不仅操作繁琐、耗时耗力,而且容易因手动输入错误导致手续办理失败或延误
[0024] "Authentication reliability" refers to the system's ability to correctly identify genuine users and prevent impersonation or forgery during the identity authentication process, including the accuracy of authentication results, tamper resistance, and resistance to attacks or forgery.
Smart Images

Figure CN122802158A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot speech in response to the user's speech.
[0003] In existing technologies, when users use communication applications or information search tools to query information, conduct business, or generate documents, they often need to repeatedly and manually input large amounts of personal information and business-related information. For example, when processing address changes, users often need to repeatedly fill in their name, contact information, old address, and new address on multiple different websites or systems. This is not only cumbersome and time-consuming, but also prone to failure or delays due to manual input errors. Furthermore, although generative AI models have emerged to automatically generate text or answer user questions, there is a lack of a technical solution that can effectively utilize user identifiers in communication applications or information search tools to automatically complete user information and provide high-quality prompts for generative AI models. This results in AI-generated responses not fully matching actual user needs, limiting the level of intelligence. On the other hand, the security and reliability of identity authentication are particularly important when dealing with personal information and various online procedures. However, existing systems do not adequately utilize the linkage between user biometric information and public authentication systems, failing to significantly improve authentication reliability while ensuring user experience. Therefore, there is a need for a system and its implementation that can simplify user information input, improve the accuracy of generative AI responses, and simultaneously enhance authentication security. Summary of the Invention
[0004] To address the aforementioned issues, this invention provides an information processing system, including a processor; wherein the processor is configured to: provide an interface for receiving user input information, enabling users to input basic information related to a target business through communication applications, web pages, or other terminal interfaces; parse communication application identifiers or information search tool identifiers, obtain and complete the required personal and business-related information from user account data, historical search records, or other available data sources associated with the identifier, thereby automatically constructing a more complete input dataset even when the user only inputs a small amount of key content; generate prompt information to instruct a generative artificial intelligence model to generate a response, and input the prompt information, including the completion information, into the generative artificial intelligence model to generate response content or document content highly relevant to the user's needs.
[0005] In a preferred embodiment of the above system, the processor is further configured to: when the user inputs information related to address change, automatically generate various documents or electronic forms required for address change procedures based on the completed user information and address change-related information, including but not limited to address change application forms, various public utility account change applications, etc., thereby reducing the burden on users to repeatedly fill in information in different systems and improving the efficiency and accuracy of address change procedures.
[0006] In another preferred embodiment, the processor is configured to: link the user's biometric information (e.g., fingerprint information, facial image, voiceprint information, etc.) with a public authentication system; by associating and verifying the user's identity information in this system with the authentication results of the public authentication system, the security and reliability of the authentication process are significantly improved while maintaining user convenience, reducing the risk of account misuse and information leakage. Through the above technical means, this invention can achieve intelligent completion of user input information and automatic document generation, and combined with a highly reliable identity authentication mechanism, comprehensively solves the problems of cumbersome operation, repetitive information input, inaccurate response, and insufficient authentication security in existing technologies.
[0007] "System" refers to an integrated device that includes at least one processor, as well as related storage devices, communication modules, and software programs, for performing the functions described in this invention. It may be a single server, a cloud computing platform, a terminal device, or a combination thereof.
[0008] A processor is a hardware unit that can execute computer program instructions, process input data and output processing results, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a programmable logic device (FPGA) or any combination thereof.
[0009] An "interface" refers to a set of software and hardware provided by a system for data interaction with users or external systems, including graphical user interfaces, command-line interfaces, application programming interfaces (APIs), web forms, chat window interfaces, etc.
[0010] "User input information" refers to information actively entered or selected by the user through the above-mentioned interface, including natural language text, structured form data, option selection results, and other data that can be used to describe user needs or personal circumstances.
[0011] "Communication applications" refer to software applications used to enable text, audio, video, or multimedia information exchange in a network environment, including but not limited to instant messaging software, social media applications, enterprise communication tools, chatbot platforms, etc.
[0012] "Information search tools" refer to software or services used to perform keyword retrieval, semantic search, or other forms of information querying on the Internet or specific data sources, including search engines, site search tools, and browser built-in search functions.
[0013] "Communication application identifier" refers to the identification information used to uniquely identify a user, session, or account in a communication application, including username, user ID, session ID, device ID, or equivalent identification data.
[0014] "Information search tool identifier" refers to the identification information used in information search tools to identify a user's identity or usage session, including search account ID, login account, device identifier, session identifier, etc.
[0015] "Completing required information" refers to additional information inferred or obtained based on user input information and data associated with user identifiers to improve user profiles or business processes. This includes user name, contact information, address information, historical behavior records, and other information related to the target business.
[0016] "Generative artificial intelligence models" refer to artificial intelligence models based on technologies such as deep learning that can automatically generate text, code, or other content based on input prompts. These include large language models (LLM), multimodal generative models, and their variants.
[0017] "Prompt information" refers to the input content constructed to guide generative artificial intelligence models to generate specific types of output. It includes descriptions of task background, user requirements, constraints, output format, etc., and may contain natural language or structured data composed of user input information and completion information.
[0018] "Response" refers to the output of a generative artificial intelligence model after receiving prompts, including natural language text replies, draft documents, form field suggestions, instructions on processing steps, or other generated content related to user needs.
[0019] "Information related to address change" refers to the input information related to the change of a user's residential or contact address, including the old address, new address, relocation date, reason for relocation, information of the members involved, and other information necessary to complete the address change procedure.
[0020] "Documents required for address change procedures" refers to various documents or electronic forms that need to be submitted to complete the administrative or business procedures related to address change, including address change application form, address change declaration form, public utility account address change application, and relevant notices.
[0021] "Biometric information" refers to data related to the physiological or behavioral characteristics that can be used to identify or verify the identity of a specific natural person, including but not limited to fingerprint information, facial images, iris images, palm prints, voiceprints, gait features, and feature vectors extracted from the above features.
[0022] "Public authentication system" refers to an information system or platform operated by government agencies or authorized third-party organizations for online authentication or electronic signature authentication of user identities, including but not limited to electronic identity authentication platforms, unified government authentication systems, and similar systems such as "MyNumberPortal".
[0023] "Linkage" refers to the data exchange and functional collaboration between this system and external systems (such as public authentication systems) through interfaces, including a series of interactive processes such as initiating authentication requests, receiving and verifying authentication results, and associating user identity information.
[0024] "Authentication reliability" refers to the system's ability to correctly identify genuine users and prevent impersonation or forgery during the identity authentication process, including the accuracy of authentication results, tamper resistance, and resistance to attacks or forgery. Attached Figure Description
[0025] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0026] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0027] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0028] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0029] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0030] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0031] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0032] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0033] Figure 9 This represents an emotion map that maps multiple emotions.
[0034] Figure 10 This represents an emotion map that maps multiple emotions.
[0035] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0036] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0037] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0038] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0039] Hereinafter, an example of an implementation of the system to which the technology of this disclosure relates will be described with reference to the accompanying drawings.
[0040] First, let me explain the terminology used in the following instructions.
[0041] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose Computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0042] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0043] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0044] In the following embodiments, the communication I / F (interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0045] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0046] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0047] like Figure 1As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0048] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0049] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0050] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0051] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge-Coupled Device) image sensor.
[0052] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0053] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0054] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0055] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0056] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0057] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0058] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0059] In existing dialogue systems based on generative artificial intelligence models, users' natural language input is usually simply forwarded to the generative artificial intelligence model for response generation. The server lacks fine-grained identification and utilization of the front-end environment such as communication applications or information retrieval tools, and also lacks a mechanism for unified integration and parameter completion of user attributes, business rules, multi-turn dialogue context, and external data sources. The results are as follows: (1) The server has difficulty understanding the user's intent and key parameters in a timely and accurate manner, and the generated prompts are often too rough, resulting in the response information output by the generative artificial intelligence model not matching the actual business needs; (2) For scenarios involving procedures and document generation, the server cannot automatically generate high-quality electronic documents that can be directly submitted based on the structured business template, but requires the user to manually input and correct them multiple times, resulting in low interaction efficiency and easy errors; (3) In business scenarios that require high security, the server lacks a unified authentication framework that links biometric information with external authentication devices, and it is also unable to accurately embed the authentication status into the prompts to guide the generative artificial intelligence model to output content confirmation information that matches the authentication results, thus making it difficult to meet the high credibility of online processing requirements; (4) The server fails to systematically utilize historical dialogues and generated results to maintain the conversation context in multi-round dialogues, resulting in a lack of continuity in the construction of prompts, making it difficult for the generative artificial intelligence model to provide stable and consistent responses, which affects the human-computer interaction experience.
[0060] Therefore, how to perform structured parsing and completion of input information from user terminals and their associated application environment information on the server side using computer implementation, construct high-quality prompt statements containing various contextual information such as dialogue history, business rules, authentication status and display constraints, and deeply integrate with the calling process of generative artificial intelligence models, while achieving end-to-end automated processing in terms of automatic generation of procedural documents and authentication linkage guidance information output, has become a technical problem that urgently needs to be solved.
[0061] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0062] In this invention, the server includes: a device for receiving user input information and identification information of an information processing program with communication or information retrieval functions from a user terminal via a communication channel; a device for applying a natural language processing algorithm to the user input information to extract intent information and parameter information, and automatically completing missing items of some parameter information based on the identification information, user attribute information, past dialogue history information, and additional information obtained from external information sources; a device for constructing a prompt statement that issues a response generation processing instruction to a generative artificial intelligence model based on the completed intent information and parameter information, and including dialogue history information, business rule information, authentication status information, and display constraints in the prompt statement; a device for inputting the prompt statement into the generative artificial intelligence model to generate natural language response information, and performing content filtering, format shaping, and structured data processing on the response information to generate response data that can be sent to the user terminal; and a device for storing the response data in association with the user input information in a storage device, and updating the prompt statement using the stored information during subsequent multi-round dialogue processing. This allows for unified modeling and utilization of user input, front-end application environment, business rules, multi-turn dialogue context, and external authentication results on the server side through improved data structure design and processing flow. This generates structured and semantically rich prompts to drive generative artificial intelligence models, thereby significantly improving the accuracy of intent recognition and parameter completion. It also automatically generates electronic documents that conform to the procedure template and outputs content confirmation guidance information linked to the authentication status in scenarios requiring identity verification. Overall, this achieves improvements in computer technology for the reliability, security, and interaction efficiency of the conversational business processing system.
[0063] "System" refers to an entire system consisting of at least one information processing device, a storage device, and a communication device that communicates with a user terminal, which is a combination of hardware and software used to receive and process user input information, invoke generative artificial intelligence models, and output response data.
[0064] "Information processing device" refers to an electronic device with computing power and control functions, which may include a processor, memory and input / output interface, used to execute programs to perform processing such as parsing user input information, parameter completion, constructing prompt statements and generating response data.
[0065] "User terminal" refers to an electronic device operated by a user and connected to an information processing device via a network. It may include mobile terminals, fixed terminals, or other devices with communication functions, used to send user input information and receive response data returned by the server.
[0066] "User input information" refers to data information that represents user requests or instructions, which is entered by the user through the user terminal and sent to the server through the communication channel. It usually exists in the form of natural language text, but may also contain structured fields or biometric information.
[0067] "Identification information" refers to identification data used to identify information processing programs or front-end application environments with communication or information retrieval functions. It is used to indicate the source environment of user input information so that the server can perform differentiated processing accordingly.
[0068] "Natural Language Processing Algorithms" refer to a class of program algorithms that run on information processing devices and are used to process natural language text, such as word segmentation, part-of-speech tagging, syntactic analysis, intent recognition, and entity recognition. They are used to extract intent information and parameter information from user input.
[0069] "Intent information" refers to abstract information extracted from user input by natural language processing algorithms, representing the target task or request type that the user wants to complete, such as querying the weather, changing information, generating documents, or performing authentication.
[0070] "Parameter information" refers to various parameter data associated with intent information, used to specifically limit the content of user requests, including time, location, object, quantity, attribute values, etc., used to support subsequent business logic processing and response generation.
[0071] "User attribute information" refers to data describing user characteristics that are associated with user identifiers and stored in storage devices, including but not limited to geographic location preferences, language preferences, historical behavioral characteristics, account information, etc.
[0072] "Dialogue history information" refers to dialogue turn data recorded and saved by the server during multiple rounds of dialogue, including past user input and server responses, which is used to provide context when constructing subsequent prompt statements.
[0073] "External information sources" refer to data providers or service systems located outside the information processing device and accessible through networks or other interfaces, including external databases, third-party service interfaces, public information platforms, etc.
[0074] "Additional information" refers to extra data obtained from external information sources or internal systems to supplement or correct parameter information, including environmental information, configuration information, and data corresponding to business rules.
[0075] "Missing parameter information items" refer to parameter fields in the set of parameters required for the target task determined based on intent information that are not currently obtained from user input information or existing context, and need to be determined through a completion mechanism.
[0076] "Generative artificial intelligence models" refer to data processing models that can automatically generate natural language text or other forms of output based on input prompts, including but not limited to large-scale language models and sequence-to-sequence models.
[0077] "Prompt statements" refer to text or equivalent data structures composed of intent information, parameter information, dialogue history information, business rule information, authentication status information, and display constraints, which are used as input to generative artificial intelligence models to instruct the models to perform specific response generation tasks.
[0078] "Business rule information" refers to rule data that defines specific business processing flows, constraints, field requirements, and compliance requirements. It is used to guide the model to output results that meet business logic when constructing prompt statements and generating response data.
[0079] "Authentication status information" refers to status data that indicates the current user identity authentication result and its credibility, including whether authentication is complete, authentication level, authentication time, authentication credibility, etc. It is used to reflect the authentication status in the prompt statement and affect the model output content.
[0080] "Display constraints" refer to the limiting information related to the display capabilities of the user terminal, interface layout, or business scenario, including text length limits, language types, format requirements, etc., which are used to constrain the form of the output of generative artificial intelligence models.
[0081] "Natural language response information" refers to response text expressed in natural language, generated by generative artificial intelligence models based on prompt statements, used to answer user questions or complete business prompts.
[0082] "Content filtering" refers to the processing steps performed on natural language response information to delete, replace, or block information that does not conform to predetermined specifications, contains sensitive content, or is inappropriate.
[0083] "Formatting and shaping" refers to the process of adjusting the paragraph structure, punctuation, lists, and headings of natural language response information to conform to the predetermined display format or business format requirements.
[0084] "Structured data processing" refers to the process of converting natural language response information into a structured representation with field labels, hierarchical structure, or predefined format, so that the system can store, retrieve, or interface with other applications.
[0085] "Response data" refers to the output data that is directly sent to the user terminal for display or further processing after the natural language response information has undergone content filtering, format shaping, and structured data processing.
[0086] "Storage device" refers to a non-transitory information storage medium used to store user input information, response data, dialogue history information, user attribute information, business rule information, etc., including semiconductor memory, magnetic storage medium or optical storage medium, etc.
[0087] "Multi-turn dialogue processing" refers to the process by which a server processes multiple consecutive inputs from the same user and generates corresponding responses in a continuous session, where there is a contextual relationship between each round of input and response.
[0088] The “Procedure Style Information Storage Department” refers to the data storage unit used to store document style definition information related to various procedure processing, including field templates, layout information, and constraints.
[0089] "Style definition information" refers to the templated description information that is pre-set in the procedure style information storage department and is used to generate electronic documents, specifying the required fields, structure, and layout.
[0090] "Electronic document data" refers to document data in the form of electromagnetic records generated according to style definition information and used for online submission or archiving, including formatted text, form data and its metadata.
[0091] "Summary information" refers to a simplified summary of the key content in an electronic document, used to briefly present the main information of the document to the user on the confirmation screen.
[0092] "Acknowledgment information" refers to data information generated and sent to the server when a user confirms or agrees to a confirmation response on the terminal interface, indicating their acceptance of the document content or operation instructions.
[0093] "Biometric information" refers to data that is related to the physical or behavioral characteristics of an individual user and can be used for identity verification, including fingerprint features, facial features, iris features, voiceprint features, etc.
[0094] "Additional authentication information" refers to data other than biometric information that is used to assist in completing identity authentication, including passwords, one-time verification codes, hardware token data, etc.
[0095] "External authentication processing device" refers to an information processing device or service system independent of this system, which is used to perform identity authentication processing on received biometric information and additional authentication information and output authentication results.
[0096] "Authentication credibility information" refers to quantitative or graded data obtained by external authentication processing devices or this system based on the authentication process, which indicates the reliability of the authentication results and is used to indicate the credibility level of user identity authentication.
[0097] "Authentication-linked guidance information" refers to explanatory content that is dynamically generated in the response information output by the generative artificial intelligence model based on the authentication status information and authentication credibility information. It includes the user's confirmation result, content confirmation items, and suggestions for subsequent operations.
[0098] In one implementation, a server is located in a data center server room as an information processing device. The server includes a multi-core central processing unit, a graphics processing unit, high-speed semiconductor memory, and solid-state storage media. The server runs network server software, middleware, and applications on an operating system and interacts with terminals via a communication network. The terminal can be a mobile information processing device or a fixed information processing device equipped with a display device and an input device. Users interact with the server through communication applications or information retrieval tools on the terminal.
[0099] In this implementation, the server runs network server software (e.g., a web server) and an application runtime environment (e.g., a scripting language runtime or an application framework) on an operating system (e.g., a Unix-like operating system). The server maintains program modules in its storage device, including: an input parsing module, a natural language processing module, an intent and parameter extraction module, a parameter completion module, a prompt generation module, a generative artificial intelligence model interface module, a result post-processing module, a session management module, an authentication linkage module, and a document generation module, etc.
[0100] The server uses natural language processing (NLP) software libraries within the application, such as an open-source NLP library, a word segmentation library, or an inference component based on a pre-trained language model. When the server needs to invoke a generative AI model, it can do so via an application programming interface (API) by calling a large-scale language model deployed in an external inference service, or by calling a large-scale language model inference service deployed on a local graphics processing unit. In one implementation, the generative AI model can employ a multi-layered encoder-decoder architecture based on a self-attention mechanism, i.e., a transformer structure containing multiple attention sub-layers and feedforward sub-layers, with the number of model parameters exceeding one billion, and encodes the input text using a word segmentation algorithm.
[0101] In a more specific implementation, the server maintains a user attribute database, a business rules database, a dialogue history database, a procedure style information storage unit, and an authentication record storage unit in the storage device. The server stores the user session state in memory using a key-value structure, document structure, or relational table structure, including user identifiers, current session identifiers, recent dialogue texts, intent labels, parameter dictionaries, authentication status flags, etc. When parsing user input, the server loads the user input text into memory as a string buffer and stores the identification information of the communication application or information retrieval tool as an environment context field.
[0102] In the natural language processing module, the server performs word segmentation, part-of-speech tagging, and named entity recognition on the user's input text. Utilizing an intent classifier fine-tuned based on a pre-trained language model, the server inputs the segmented and encoded text vectors into several fully connected layers and normalization layers, outputting a probability distribution corresponding to a predefined intent category. The server determines the intent information based on the label with the highest probability. For example, when a user inputs "I want to change my address" through the terminal, the server encodes the input into a vector sequence, then feeds this vector sequence into the intent classification subnet. Based on the output, the server labels the intent information as "change of residence information."
[0103] The server uses sequence labeling models or attention-weighted span extraction models in its parameter extraction module to extract parameters such as time, location, document type, and personal identification number from user input and historical dialogues. When processing the input "What will the weather be like in Beijing tomorrow?", the server identifies "tomorrow" as a date parameter, "Beijing" as a location parameter, and "weather" as a query category. In address change scenarios, the server identifies the old address, new address, and effective date as key parameters.
[0104] In the parameter completion module, the server determines the list of required parameters based on intent information and business rule information. When the server detects missing parameters, it infers the missing parameters based on user attribute information, dialogue history information, and data returned from external information sources. For example, when a user enters "Check tomorrow's weather for me" without providing a city, the server reads the user's usual city from the user attribute information and uses that city as the missing location parameter. If user attribute information is missing, the server can determine the approximate region from the location information uploaded by the terminal or the network address resolution result and use that as the default location parameter. Internally, the server maintains a parameter dictionary in the form of a key-value mapping table and ensures that the key parameters corresponding to each type of intent are filled through a state machine or rule engine.
[0105] In the prompt generation module, the server organizes the completed intent and parameter information, along with dialogue history, business rules, authentication status, and display constraints, into structured text. Following a predefined template, the server concatenates this information into a natural language prompt, enabling the generative AI model to receive a complete task description and context from a single input. For example, in a weather query scenario, the server generates the following prompt: "The user wants to check the weather."
[0106] City: Beijing.
[0107] Date: Tomorrow.
[0108] Please answer tomorrow's weather forecast for Beijing in Simplified Chinese, including the temperature range, weather phenomena (e.g., sunny, cloudy, rain / snow), and brief travel advice. Your answer should not exceed 100 characters. In a document generation scenario, the server generates the following prompt: "Users need a formal leave request in Chinese."
[0109] Known information: Name: Zhang San Leave period: March 1, 2024 to March 5, 2024 Reason for leave: An urgent matter at home needs to be handled. Please draft a polite and well-formatted leave request for the user in Simplified Chinese, keeping it under 200 characters. The server generates the following prompt statement in authentication linkage scenarios: "The user has completed identity verification through an external authentication system, and the authentication credibility is high."
[0110] The following is a summary of the application documents generated by the system for the user: 1. Item to be processed: Change of residence information; 2. Effective Date: April 1, 2024; 3. New address: No. ××, ×× Road, ×× District, ×× City.
[0111] Please restate the above key terms to the user in Simplified Chinese, and remind the user to confirm their accuracy. If necessary, please prompt the user to make corrections. When generating the aforementioned prompts, the server encodes display constraints (such as maximum character count and whether to use list format) into text requirements or incorporates them into internal parameters. This guides the generative AI model to generate output that meets the terminal's display capabilities and business requirements. By pre-constructing prompts on the server side, the server centrally encodes contextual information, originally scattered across multiple system modules, into a single input sequence. This allows the model to comprehensively consider multi-source information in a single forward inference, thereby reducing the need for multiple network round trips and repetitive inference.
[0112] In the generative AI model interface module, the server inputs the prompt as a text sequence into the generative AI model. The server uses a transformer-based language model in its local or external inference service. This model includes an input embedding layer, multiple self-attention encoding layers, multiple self-attention decoding layers, a layer normalization layer, and an output projection layer. During inference, the server decomposes the prompt into sub-word units and maps them to high-dimensional vectors using an embedding matrix. The server performs matrix multiplication and normalization operations on the multiple self-attention mechanism in the graphics processing unit to calculate the relevance weights between different positions at each layer. At the output end, the server obtains the probability distribution of output words using a soft maximum function and gradually generates the response text using a greedy strategy or a sampling strategy.
[0113] During model training or retraining, the server can use the cross-entropy loss function as the error function and update the weight parameters through stochastic gradient descent and its variants (such as adaptive moment estimation). The server introduces an expanded dataset with explicit intent labels, parameter labels, and business constraint samples into the training data, enabling the model to learn specific output preferences in the business context through supervised learning or instruction fine-tuning. In some implementations, the server employs data augmentation methods, such as synonym replacement and template expansion, to generate diverse training samples, thereby improving the model's robustness to diverse natural language expressions.
[0114] In the post-processing module, the server performs a series of technical processes on the natural language response information output by the generative artificial intelligence model. First, the server performs content filtering on the response text, using keyword matching, regular expressions, and a classifier-based security filtering model to remove content that does not meet compliance requirements or exceeds predetermined boundaries. Next, the server performs formatting, inserting line breaks, list symbols, and header rows according to terminal display requirements to ensure clear presentation even on small-screen terminals. Simultaneously, the server performs structured data processing on fragments containing structured semantics. For example, in weather queries, it extracts fields such as maximum temperature, minimum temperature, and weather phenomena from the text; in procedural processing scenarios, it extracts fields such as name, date, and address from the text, storing them in a data structure in key-value pairs or a tree structure for use by subsequent business systems.
[0115] In the session management module, the server stores each user input, along with the corresponding response data and prompts, in a dialogue history database. During storage, the server adds a timestamp, intent tag, parameter snapshot, and authentication status flag to each record. In a new round of interaction, the server reads several recent dialogues associated with the current session identifier, extracts highly relevant content using filtering and summarizing algorithms, and inserts it into the "context" section of the new prompt. This selective context injection strategy reduces redundant text length, lowers the computational load and response latency of model inference, while maintaining consistency and continuity across multiple dialogue rounds.
[0116] In the authentication-related implementation, the server receives biometric information and additional authentication information uploaded by the terminal in the authentication linkage module. The server encapsulates this information into an authentication request and forwards it to an external authentication processing device. The external authentication processing device performs pattern matching, similarity calculation, or discriminative operations based on classification networks, and returns the authentication result and credibility score to the server. The server generates authentication status information based on the returned credibility score and explicitly indicates this status in subsequent prompts, such as "Authentication completed and credibility is high" or "Authentication not yet completed." When the generative artificial intelligence model receives a prompt containing authentication status information, it generates a response text containing the user's confirmation result and content confirmation items. The server then sends this text as confirmation response data to the terminal. After the terminal displays it, the user can express acceptance or modification instructions through interactive operations, and the server decides whether to confirm the electronic document data as the final version.
[0117] In the document generation implementation, the server uses pre-defined style information stored in the document generation module to fill parameter information into template fields. The server constructs a tree-like data structure in memory to represent the document structure, with nodes corresponding to paragraphs, sentences, table cells, etc., and leaf nodes corresponding to specific field values. During the filling process, the server performs field validity checks based on business rules, such as checking if the date is within the allowed range and if the address format meets predetermined rules. If necessary, the server calls the generative AI model again to refine or expand some descriptive fields, then converts the structured document into a submitable format (e.g., text with markup language or a formatted document), and outputs summary information to the user for confirmation.
[0118] Through the above methods, the server not only automates the human dialogue and document writing process, but also achieves multiple improvements at the technical level within the computer. By uniformly encoding application environment identification information, user attributes, and historical dialogue into prompt statements, the server effectively reduces the model's sensitivity to input ambiguity, thereby improving the accuracy of intent recognition and parameter selection. Furthermore, by performing parameter completion and context pruning on the server side, the server controls the input length of the generative artificial intelligence model, reduces the computational burden caused by irrelevant information, and achieves lower latency responses given hardware resources.
[0119] The server manages historical data in a structured manner through a session management module, reducing redundant calculations and external interface calls, thereby lowering network communication load. Through data augmentation and instruction fine-tuning during model training, the server ensures that the model's internal weights converge to a suitable parameter space for specific business domains, reducing the probability of inference output errors. Through explicit content filtering and format shaping processes, the server converts generative output into structured data that is easily processed by computers, allowing subsequent systems to directly verify, archive, or perform further calculations on this data.
[0120] Furthermore, in authentication-linked scenarios, the server directly incorporates authentication credibility information into the prompts and responses. This transforms the output of the generative AI model from a generic answer unrelated to identity status into a technical feedback tightly coupled with authentication progress and strength. This feedback not only enhances security during online processing but also indirectly improves overall system throughput by reducing the number of times users need to confirm. Because the server employs specific data structure design, model call order, and prompt composition rules throughout the process, the system develops a unique processing flow distinct from traditional rule engines and plain text forwarding methods, achieving comprehensive technical improvements in response time, response quality, and resource utilization.
[0121] In other alternative implementations, the server can employ generative AI models of varying scales, such as medium-parameter models or lightweight distillation models, to adapt to edge computing or resource-constrained environments. The server can also merge the natural language processing module with the generative AI model, simultaneously performing intent classification, parameter extraction, and response generation in a single model, or adopt a cascaded structure, using a smaller model for coarse intent classification and a larger model for complex response generation. The server can also dynamically adjust display constraints based on the terminal type, thereby generating a response layout more suitable for large or small screen devices. These variations, without departing from the technical concept defined in the claims, all fall within the scope of this invention.
[0122] use Figure 11 The processing flow is explained.
[0123] Step 1: The user enters a request in the terminal and sends it. Users open a communication application or information retrieval tool on their terminal, type natural language text in the input box, such as "I want to apply for an address change" or "What will the weather be like in Beijing tomorrow?", and send a request to the server by clicking the send button.
[0124] Input: Natural language text entered by the user in the terminal interface, as well as application identification information, user identifier, session identifier, etc. automatically attached by the terminal.
[0125] Output: The terminal packages the request data message, which includes user input text, application identification information, user identifier, session identifier, etc., and sends it to the server through the communication network.
[0126] Before sending, the terminal encodes the text and encapsulates the application layer data into transport layer and network layer data packets through the operating system's network protocol stack. The data packets are then transmitted to the server address via a wireless or wired network interface.
[0127] Step 2: The server receives the request and parses the basic fields. The server listens on a predetermined port on the network server software, receives request messages from the terminal, performs protocol parsing and decryption on the messages, and extracts the application layer payload from them.
[0128] Input: Network data packets sent by the terminal, which contain fields such as serialized user input text, application identification information, user identifier, and session identifier.
[0129] Output: A request object constructed in the server's memory, containing structured data such as raw text fields, application identification fields, user identification fields, and session identification fields.
[0130] The server calls the serialization parsing library to restore the request body from a byte sequence to a structured data format, and temporarily stores the structured object in the session management module for use by subsequent processing modules.
[0131] Step 3: The server performs natural language preprocessing on user input. In the natural language processing module, the server performs word segmentation, part-of-speech tagging, and basic cleaning operations on the user input text in the request object, including removing meaningless symbols and unifying the encoding format.
[0132] Input: A request object containing the original natural language text.
[0133] Output: A word sequence representation annotated with word boundaries and part-of-speech information, and a cleaned text version.
[0134] The server calls pre-trained models or rules through the loaded natural language processing library to segment the text by word boundaries, adds part-of-speech tags to each word, and stores the results as a list or tensor for use as input to subsequent intent recognition and parameter extraction algorithms.
[0135] Step 4: The server performs intent recognition and initial parameter extraction. In the intent and parameter extraction module, the server uses a classification sub-network fine-tuned based on a pre-trained language model to perform forward inference on the pre-processed text vectors, obtaining the probability distribution representing different intent categories, and selecting the label with the highest probability as the intent information. Simultaneously, the server uses a sequence labeling model to extract parameters such as time, location, and object from the text.
[0136] Input: preprocessed word sequence, corresponding word vector representation, and session identifier.
[0137] Output: Intent labels (e.g., “Change of residence information”, “Weather query”, “Document generation”, etc.) and an initial parameter dictionary containing several parameter key-value pairs.
[0138] The server maps the word sequence into a vector sequence and inputs it into a multi-layer fully connected network or a lightweight transformer structure. It then calculates the confidence scores for each category using matrix multiplication and activation functions. Simultaneously, it calculates the label for each word (such as time, location, name, etc.) in the sequence labeling path and constructs an initial parameter list based on this. This list is then written into the request context along with the intent label.
[0139] Step 5: The server checks parameter integrity based on business rules. In the parameter completion module, the server queries the business rules database based on the intent tag to obtain the list of parameters required for the corresponding intent. For example, "change of residence information" requires the old address, new address, and effective date; "weather query" requires the location and date. The server compares the key-value pairs in the initial parameter dictionary with the list of required parameters, marking which parameters already exist and which are missing.
[0140] Input: Intent label, initial parameter dictionary, and business rule information.
[0141] Output: A list of parameter requirements marked with missing items, and a dictionary of parameters marked with completeness indicators.
[0142] The server iterates through the set of required parameters, searching for each parameter name in the current parameter dictionary. If a parameter does not exist, it is added to the missing list. The server saves the results in the current session context to guide subsequent completion logic.
[0143] Step 6: The server uses user attributes and conversation history to complete missing parameters. In the parameter completion module, the server queries available information from the user attribute database, the dialogue history database, and external information sources based on the list of missing parameters.
[0144] Input: List of missing parameters, user ID, session ID, conversation history, and data returned by external information interfaces.
[0145] Output: The completed parameter dictionary, where the original parameters remain unchanged and missing parameters are filled with inferred or default values.
[0146] The server first queries user attribute records based on the user identifier, such as frequently used locations and default language. If the location parameter is missing, the server reads the user's city of residence from the user attributes and fills it in. If there is no data in the user attributes, the server searches for the most recent valid location information from the dialogue history. If it still cannot be determined, the server can obtain the terminal's location from an external geolocation service and fill in this region as a candidate value into the parameter dictionary. At the same time, the server adds a source marker to the parameter dictionary for the completed parameters, which is used for subsequent confirmation or correction.
[0147] Step 7: The server consolidates authentication status and updates the session context. In the authentication linkage module, the server checks the authentication records associated with the current session identifier to determine whether the user has completed identity authentication through an external authentication processing device and the authentication trust level.
[0148] Inputs: Session ID, User ID, External authentication result record, and authentication credibility information.
[0149] Output: Authentication status information (authenticated / unauthenticated, trust level, etc.) and write it to the session context.
[0150] The server determines whether to include authentication instructions and content confirmation requirements in subsequent prompts based on the authentication status. For example, it adds "User has not completed identity authentication" when authentication has not yet occurred, and adds "User has passed high-confidence authentication" when authentication has occurred and the credibility is high. The server stores this status information in the data structure of the session management module for reference by the prompt generation module.
[0151] Step 8: Prompt statements for the server to construct generative artificial intelligence models In the prompt generation module, the server converts various types of data into natural language prompts based on the completed intent information and parameter information, combined with dialogue history information, business rule information, authentication status information, and display constraints, according to a predefined template.
[0152] Inputs: intent label, completed parameter dictionary, session history summary, business rule entries, authentication status flags, and display constraint settings.
[0153] Output: Prompt text used as input to the generative artificial intelligence model.
[0154] The server summarizes historical conversations, selects several rounds of questions and answers highly relevant to the current intent, and places them at the beginning as "contextual descriptions." The server adds necessary output formatting requirements and precautions according to business rules. The server adds explicit instructions to the prompts based on display constraints (such as maximum character count and whether to use an item list). The server then concatenates all of the above information into a continuous text, for example: "The user wants to check the weather."
[0155] City: Beijing.
[0156] Date: Tomorrow.
[0157] Please answer tomorrow's weather forecast for Beijing in Simplified Chinese, including the temperature range, weather phenomena (e.g., sunny, cloudy, rain / snow), and brief travel advice. Your answer should not exceed 100 characters. Or generate in a document context: "Users need a formal leave request in Chinese."
[0158] Known information: Name: Zhang San Leave period: March 1, 2024 to March 5, 2024 Reason for leave: An urgent matter at home needs to be handled. Please draft a polite and well-formatted leave request for the user in Simplified Chinese, keeping it under 200 characters. Step 9: The server invokes a generative artificial intelligence model to generate response text. In the generative AI model interface module, the server sends prompts in text form to the generative AI model deployed locally or in an external inference service.
[0159] Input: The prompt text output by the prompt generation module.
[0160] Output: Natural language response text output by the generative artificial intelligence model.
[0161] The server segments the prompt into sub-words, encodes the resulting sequence into a vector, and sends this sequence to the graphics processing unit or remote inference service via a network interface. The generative AI model performs forward propagation of the prompt within a multi-layered self-attention structure, calculating the attention weights at each position and generating an output word probability distribution. The model progressively samples or selects the word with the highest probability to form the complete response text. The server receives the text sequence returned by the inference service and reconstructs it into a natural language string as the initial response result.
[0162] Step 10: The server performs content filtering and format shaping on the response text. In the post-processing module, the server performs security checks, compliance filtering, and formatting adjustments on the response text output by the generative artificial intelligence model.
[0163] Input: The original response text returned by the generative artificial intelligence model.
[0164] Output: Filtered and formatted response text, and possibly additional structured fields.
[0165] The server first uses a keyword list and regular expressions to detect whether the response contains sensitive words, illegal content, or information irrelevant to the business, deleting or replacing relevant segments if necessary. Then, it inserts line breaks, adds sequence numbers, or rearranges paragraphs according to terminal display requirements to meet the needs of small screen displays or standardized tone. If the response contains structured data, such as weather values, time ranges, or address fields, the server calls the parsing submodule to capture these values and strings, performs type conversion, and fills them into a structured data table.
[0166] Step 11: The server stores the response text in association with the user input. In the session management module, the server writes the current user input, prompts, generative AI model output, final response text, and structured data into the dialogue history database.
[0167] Input: Original user input text, prompt text, response text before and after filtering, structured parameter data, and current session identifier and timestamp.
[0168] Output: One or more dialogue records stored in persistent storage for subsequent rounds of dialogue and model optimization.
[0169] The server generates a unique identifier for each record and stores it in the storage medium, sorted by session identifier and time. During write operations, the server also updates the fast index structure to facilitate rapid retrieval of historical data by user, intent, or time range for use in constructing the next round of prompts.
[0170] Step 12: The server sends response data to the terminal, which is then displayed by the terminal. The server encapsulates the final response text and its necessary structured data into a response message in the network server software and returns it to the terminal through the communication network. After receiving the response, the terminal parses it and displays it to the user in the form of message bubbles or result areas on the interface.
[0171] Input: The response text and structured data generated by the server, as well as the network connection information of the terminal.
[0172] Output: The natural language response displayed on the terminal interface, as well as possible interactive controls (such as confirmation buttons, edit buttons, etc.).
[0173] The server serializes the response data into a response format and attaches a status code, then sends it to the terminal via the network protocol stack. Upon receiving the response, the terminal uses a parsing library to reconstruct it into text and fields, updates the user interface, displays the response text on the screen, and generates corresponding buttons or forms based on the structured data, allowing the user to proceed to the next step or confirm. After reading the response, the user can enter supplementary questions or confirmation commands on the terminal as needed; new input will trigger the cycle of the aforementioned steps again.
[0174] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0175] With the widespread application of generative AI models in content generation, information retrieval assistance, and other scenarios, related computer systems typically simply forward the user's natural language input directly to the generative AI model, which then generates the result and returns it to the user. This traditional approach suffers from the following technical problems: (1) In terms of constructing prompt statements, existing systems lack a standardized and structured processing mechanism based on user input features, which leads to unstable format of prompt statements input to generative artificial intelligence models and the presence of redundant noise information, thereby affecting the model's reasoning efficiency and generation quality, and increasing the consumption of ineffective computing resources.
[0176] (2) In terms of the utilization of the generated results feedback, the existing systems often only generate and display once, lacking the ability to systematically collect, store and analyze user evaluation information. They cannot dynamically map user feedback into the basis for adjusting prompt statement generation rules or model call parameters, which makes it impossible for the system to continuously adaptively optimize the generation behavior during operation, thus limiting the evolutionary ability of the overall performance of the computer system.
[0177] (3) In terms of prompt statement template management, existing technologies usually rely on hard coding or manual maintenance of prompt texts in different scenarios. There is a lack of a mechanism to automatically select and combine fixed and variable parts of the template based on user input type. The prompt statement construction logic is scattered in the application code, which increases the system maintenance cost. At the same time, it is not conducive to reusing existing experience and data between different topics and tasks, and reduces the scalability of the system in complex business scenarios.
[0178] (4) In terms of system-level resource utilization and response characteristics, existing solutions rarely differentiate the calling parameters (such as output length, generation diversity, etc.) of generative artificial intelligence models based on different themes or different quality evaluation results. This results in requests with low value or low user satisfaction still occupying the same level of computing and network resources, making it difficult for computer systems to achieve an optimal balance in terms of throughput, response time and resource utilization.
[0179] Therefore, a system is needed that can generate structured prompts from user input on the server side, perform adaptive parameter control by combining statistical user evaluation information, and build a continuous optimization closed loop through template management and thematic evaluation indicators. This system can improve the efficiency, stability, and controllability of generative artificial intelligence model invocation at the computer technology level, and improve the overall human-computer interaction quality and resource utilization.
[0180] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0181] In this invention, the server includes: a device for acquiring user input information via a communication network through a user terminal and storing the user input information as data; a device for generating a prompt statement containing the user input information based on the user input information and processing the length and character type of the prompt statement according to predetermined conditions to shape the prompt statement into an input format suitable for input into a generative artificial intelligence model; a device for sending a response generation request to the generative artificial intelligence model using the shaped prompt statement and converting the generated content obtained from the generative artificial intelligence model into distribution data for the user terminal and sending it; a device for acquiring evaluation information for the generated content from the user terminal and updating the generation rules of the prompt statement or the operating parameters of the generative artificial intelligence model based on the evaluation information to change the generation characteristics of subsequent generated content; a template management function for selecting a fixed part of the prompt statement based on the query information acquired as user input information and combining the fixed part with the user input information to form the prompt statement; and an adaptive control function for calculating evaluation indicators based on the statistical results of the evaluation information according to the theme of the generated content and automatically changing the constituent elements of the prompt statement or the output length control value of the generative artificial intelligence model according to the evaluation indicators. This allows for the standardization and structuring of user input prompts on the server side, as well as the dynamic adjustment of prompt generation rules and generative AI model calling parameters based on topic-based evaluation metrics. This forms a closed-loop mechanism that continuously optimizes generation behavior using user feedback, thereby improving the relevance and stability of the generated results, reducing unnecessary computational overhead, and improving the processing efficiency and resource utilization of computer systems in generative AI scenarios.
[0182] A "system" refers to a whole consisting of multiple electronic devices interconnected by a communication network and the programs running on them, which is a computer technology solution used to perform a series of collaborative processes such as user input acquisition, data processing, model invocation, and result output.
[0183] "User terminal" refers to an electronic device that allows users to input information, browse results, and evaluate performance. This includes, but is not limited to, mobile communication devices, computer devices, and display devices, which interact with servers via networks.
[0184] A "server" refers to an information processing device that runs on the network side and provides data processing, model calling, and result distribution functions. It typically includes a processor, memory, network interface, and operating system and applications running on it.
[0185] "Communication network" refers to wired or wireless communication infrastructure used to transmit data between user terminals and servers, including local area networks (LANs), wide area networks (WANs), mobile communication networks, and the Internet.
[0186] "User input information" refers to data related to content generation or query intent submitted by users through the input interface of user terminals and transmitted to the server via communication networks. It is usually natural language text, but may also include structured parameters or metadata.
[0187] "Data storage" refers to the process and state of saving user input information, generated content, evaluation information, etc. in a searchable form on the server side using storage media and their management programs.
[0188] "Prompt statements" refer to text data generated by the server based on user input information and organized in a predetermined format. These are used as input to generative artificial intelligence models to guide the models in generating target content according to the desired task and style.
[0189] "Handling of prompt statement length and character type" refers to the formatting operations performed by the server on the prompt statement, including trimming and padding the number of characters, and standardizing the character encoding or character type to meet the technical requirements of the generative artificial intelligence model interface for input.
[0190] "Generative artificial intelligence models" refer to artificial intelligence models that are trained through machine learning and can automatically generate text and other content based on input prompts. They typically employ neural network structures and are trained on large-scale datasets.
[0191] "Responding to a generation request" refers to the process by which a server sends a call instruction and a corresponding prompt statement to a generative artificial intelligence model, so that the model can perform inference operations and output generated content associated with the prompt statement.
[0192] "Generated content" refers to the result data automatically output by a generative artificial intelligence model after receiving a prompt statement through internal reasoning and calculation. It is usually natural language text, but can also be extended to text fragments containing structured information.
[0193] "Distribution data" refers to the data that a server encapsulates, encodes, and converts the generated content after obtaining it, in order to facilitate display or processing on user terminals, and is then used for transmission over the network.
[0194] "Evaluation information" refers to user feedback data on generated content, including but not limited to ratings, tags, option feedback, and natural language comments, which reflect users' subjective evaluation of the quality or relevance of the generated content.
[0195] "Prompt statement generation rules" refer to the set of predefined or dynamically updated logic used by the server when generating prompt statements, which specifies how to select templates, combine fixed and variable parts, and set additional descriptions or constraints based on user input information.
[0196] "Running parameters" refer to the control variables set by the server when calling the generative artificial intelligence model, including output length limits, randomness control parameters, sampling strategy parameters, etc., which are used to affect the style, length and diversity of the generated content.
[0197] "Fixed parts" refer to text fragments that are pre-defined in the prompt template and do not change with the specific user input, used to indicate the task type, output format, or style requirements.
[0198] The "template management function" refers to the program and data structure on the server side used to save, retrieve, select, and update prompt statement templates. This function selects the appropriate fixed part and combines it with the variable part to generate prompt statements based on the type of user input information or scenario.
[0199] "Topic" refers to a classification tag used to categorize generated content based on the domain, topic, or task type it involves. This tag is used for grouping when compiling evaluation information and adjusting system behavior.
[0200] "Evaluation metrics" refer to quantitative or qualitative measures calculated based on evaluation information by theme or other dimensions, used to reflect the overall performance of generated content in terms of user satisfaction, relevance, or usefulness.
[0201] The “components of a prompt statement” refer to the various parts that make up a prompt statement, including fixed parts, variable parts, supplementary explanations, constraints, formatting instructions, and other text units.
[0202] "Output length control value" refers to the parameter value used to constrain the maximum or desired length of the output content of the generative artificial intelligence model. By adjusting this value, the number of words or paragraph size of the generated content can be controlled.
[0203] "Adaptive control function" refers to the program function that automatically adjusts the constituent elements of prompt statements or the operating parameters of generative artificial intelligence models based on accumulated evaluation information and evaluation indicators, thereby continuously optimizing the generation behavior without human intervention.
[0204] In one embodiment of the present invention, the server, as an information processing device, includes a processor, main memory, non-volatile storage, and a network interface. The server runs an operating system (e.g., a general-purpose server operating system), a web service framework (e.g., an HTTP-based application server), a database management system, and a generative artificial intelligence model invocation module at the software level. The terminal, as a user-side device, can be a mobile communication device, a personal computing device, or an embedded display device. The terminal runs an application or browser for sending user input information to the server, receiving generated content, and sending evaluation information. The user performs input, reading, and evaluation operations through the terminal.
[0205] The server stores the executable program in non-volatile storage. When this program runs on the processor, it configures the server to perform functions such as prompt generation, generative AI model invocation, evaluation information statistics, and adaptive control. The server maintains data structures in main memory for each session, such as user input buffers, prompt objects, model invocation parameter objects, generated content records, and evaluation statistics records. Each object can be in the form of a key-value map or a structured record, with fields including user identifier, topic tags, timestamps, text content, and parameter values.
[0206] When generating prompt statements, the server performs specific data processing on the user input. The server stores the user input as a string in main memory and uses a string processing library to perform length checks, character set filtering, and formatting checks on that string. The server maintains multiple prompt statement templates in its template management function. Each template includes a fixed part and variable placeholders. The fixed part can be a task description in English or other languages, for example: Generatecontentabout: [User input information] When selecting a template, the server chooses an appropriate template based on the metadata of the user input information (such as input source, task type, and historical topic tags). For example, when the user inputs "the latest technology trends," the server combines the fixed and variable parts to produce the following prompt: Generatecontentabout: The latest technology trends In one variant, the server can use more complex templates, such as indicating structured output and depth requirements: This is a detailed, structured article about the latest technological trends, including recent developments, key technologies, and future directions. After generating the prompt statement, the server further encodes the prompt statement string according to the technical requirements of the generative artificial intelligence model interface (e.g., uniformly to UTF-8), controls the total number of characters to not exceed a preset limit, and deletes or replaces unacceptable control characters. Through the above processing, the server ensures that the prompt statement is semantically adapted to the task and formally meets the model input constraints, thereby reducing invalid computations caused by tokenization anomalies and truncation within the model.
[0207] In terms of the specific implementation of the generative artificial intelligence model, the server employs a neural network model based on the Transformer architecture. This model includes an embedding layer, a multi-layer self-attention encoder-decoder structure, a feedforward network, and an output probability distribution calculation module. The server uses a deep learning framework (such as a general tensor computation framework) on a graphics processing device to perform matrix multiplication and attention weight calculations during training or invocation.
[0208] During model training, the server uses a large-scale text corpus to pre-train the model parameters. The server sets specific loss functions in the training pipeline, such as cross-entropy loss, to measure the difference between the model's output token probability distribution and the target token. In each training iteration, the server calculates the gradients of the network weights and bias parameters of each layer using the backpropagation algorithm and updates the parameter values using optimization algorithms (such as adaptive learning rate optimization methods with momentum). The server can perform data augmentation on the training data, such as random truncation, random masking, and sentence rearrangement, to improve the model's robustness to different prompt formats. The generative AI model obtained by the server in the above manner encodes the prompts during the inference phase and progressively predicts the conditional probability of each token in the output sequence.
[0209] During online inference, the server inputs the prompts into the generative AI model interface. The model first performs tokenization and embedding on the prompts, generating a high-dimensional vector representation. Then, it calculates the dependencies between tokens using a multi-head self-attention layer. Next, it outputs the probability distribution of the next token through a feedforward network and a normalization layer. The server selects the next token at the model's output layer based on predefined parameters (e.g., temperature, top-p, top-k). The server appends the selected token to the current generated sequence and repeats the inference steps until the maximum length is reached or a termination flag is encountered. Finally, the server reconstructs the generated token sequence into natural language text, which serves as the generated content.
[0210] In this invention, the server does not simply repeat human editing behavior, but drives model invocation through a set of explicit, non-human intuitive control rules. The server adopts an adaptive control strategy for the components of the prompt statements and the model running parameters: the server accumulates evaluation information for each generated content in the database, and the evaluation information is stored in the form of structured records, including content identifiers, topic tags, user ratings, usefulness markers, and text feedback summaries.
[0211] In the background statistics module, the server performs aggregation operations on the evaluation information. The server can group evaluation records by topic and calculate the average score, score variance, and number of valid evaluations for each topic. Based on these evaluation metrics, the server automatically adjusts the prompt template. For example, when the average score of a topic remains consistently low, the server can automatically switch the template associated with that topic from "simple description" to "structured detailed description," or add requirements for examples and step-by-step instructions to the template. The server can also dynamically adjust the output length control value of the generative AI model according to evaluation metrics: for topics with low ratings and user feedback indicating "too long content," the server automatically shortens `max_tokens`; for topics with low ratings and user feedback indicating "shallow content," the server extends the output length and increases the level of detail.
[0212] The server, through the aforementioned adaptive control functions, modifies the components of the prompt statements and model parameters, and stores the new rules in the configuration store. The server automatically applies the updated rules in subsequent sessions without requiring manual modifications to the program logic by developers. Through this feedback loop, the server continuously optimizes the generation behavior internally, improving processing efficiency and generation accuracy. Since the structure of the prompt statements and model parameters have a substantial impact on the neural network inference path and token selection distribution, the server can reduce invalid inference steps in the evaluation-driven adaptive process, thereby reducing overall computational resource consumption and improving response speed.
[0213] In one embodiment of the present invention, the terminal provides an interactive interface module. The terminal displays an input area and a result area on a display device. The user types natural language content in the input area, such as "the latest technology trends," "the application of artificial intelligence in the medical field," or "explaining blockchain technology to non-specialist readers." The terminal sends the text as user input to the server, and after receiving the generated content returned by the server, displays the generated paragraphs, titles, and explanatory text in the result area. The terminal also provides a rating control (e.g., a star rating button) and a feedback text input box below the result, allowing the user to evaluate the generated content.
[0214] In this invention, users provide only simple, high-level natural language input and express their subjective evaluations through ratings or short text feedback. In contrast, the server performs fine-grained data structure operations, parameter optimization, and model control in the background. Therefore, it can be seen that the technical contribution of this invention lies not in the business process itself, but in the server's technical processing methods for the structure of prompt statements, model operating parameters, and feedback data.
[0215] In one optional implementation, the server can also dynamically adjust its generation strategy based on runtime environments such as network bandwidth. For example, when the server detects high bandwidth usage on the network interface, it can reduce the maximum length of the generated content or adopt a higher compression encoding scheme to reduce the size of the distributed data. This mechanism, combined with evaluation-driven template adjustment, can reduce communication load while ensuring content availability, achieving end-to-end resource optimization.
[0216] In another implementation, the server can use the generated content for downstream device control. For example, in a smart home scenario, when a user inputs "create a comfortable reading scene for the living room lighting," the server generates a structured description, including brightness, color temperature, and device grouping, using a generative artificial intelligence model. The server parses this description into device control commands and sends them to the lighting control device via the local network, thereby actually changing the state of the lighting fixtures. During this process, the server explicitly requests the model to output text conforming to a predefined control format using prompts, such as: Generate a JSON-like lighting configuration for: Creating a comfortable reading scene for living room lighting, including brightness (0-100), color temperature (K), and target device group. The server then performs syntax checking and field mapping on the generated text to generate the final control commands. In this way, the output of the generative AI model is no longer limited to abstract text but directly drives physical devices, realizing the conversion from natural language instructions to specific control signals. The prompt statement shaping and evaluation-driven parameter control mechanism proposed in this invention can significantly reduce the model output format error rate, thereby reducing device control errors and achieving higher control reliability and response speed.
[0217] In another implementation, the server can be deployed on an edge computing device, located on the same local area network as the terminal. The server can choose to use a simplified generative AI model based on local computing resources, and achieve low-latency local content generation through the same prompt generation and adaptive control logic. In this case, template management and evaluation statistics on the server side can still be performed in a local database without accessing a remote cloud. By introducing the prompts and parameter optimization mechanism of this invention at the edge, inference efficiency can be improved and output quality guaranteed in environments with limited computing resources.
[0218] The server demonstrates through the aforementioned implementation methods that the system of this invention achieves its effects technically through the following mechanisms: On the one hand, by controlling the structured generation and length / character type of prompt statements, it reduces abnormal situations and redundant calculations in the encoding stage of generative artificial intelligence models, thereby improving inference speed and stability; on the other hand, by utilizing evaluation information to construct thematic evaluation indicators, it adaptively controls the constituent elements of prompt statements and model operating parameters, enabling the system to automatically optimize generation quality and resource allocation during operation. Since these control logics strictly rely on the computer's internal data structures, statistical algorithms, and parameter update rules, rather than a simple simulation of human workflows, this invention substantially improves the computer's processing capabilities in generative artificial intelligence scenarios, achieving technical effects such as increased accuracy, accelerated response, improved resource utilization, and reduced communication load.
[0219] use Figure 12 The processing flow is explained.
[0220] Step 1: The terminal receives user input information and generates request data.
[0221] The terminal presents a text input box and a submit control to the user on the display interface. The user enters natural language text in the input box, such as "the latest technology trends" or "the application of artificial intelligence in the medical field." The terminal takes the user's current session identifier and timestamp as input, reads the character sequence entered by the user into memory, performs basic checks (such as whether it is empty or whether the length exceeds a preset limit), and confirms the text encoding format (e.g., uniformly UTF-8). Based on the above input, the terminal constructs a request data structure, adding user identifier, session identifier, and input text fields to form a logical "user input information" message object. The terminal serializes this message object into a network transmission format (e.g., text in an HTTP request body) and sends it as output to the server via the communication network.
[0222] Step 2: The server receives user input and parses it to generate internal data structures.
[0223] The server takes network requests sent by the terminal as input and receives messages containing user input information through the network interface. Using a web framework, the server parses the encoding method, content length, and JSON or form data from the HTTP request header and body, and extracts the user input text field. The server performs a length truncation operation on this text (keeping the first N characters if it exceeds the model's maximum allowed length) and performs character type filtering (removing control characters or unsupported characters). The server organizes the parsed text, user identifier, timestamp, and other fields into an internal data structure, such as a key-value mapping record, and stores it in main memory. Simultaneously, it passes this internal record as output to the subsequent prompt generation module.
[0224] Step 3: The server generates prompts based on user input and templates.
[0225] The server takes the internal data structure generated in step 2 and the set of prompt message templates stored in the template management module as input. First, the server selects the most suitable template from the template set based on the content characteristics of the user input (e.g., whether it contains specific keywords, the business type, or historical topic tags). The template includes a fixed part and placeholders for inserting user input. Then, the server performs string concatenation on the user input text, joining the fixed part string with the user input text to generate a complete prompt message. For example, when the user inputs "latest technology trends" and the selected template's fixed part is "Generatecontentabout:", the server outputs the following prompt message: Generatecontentabout: The latest technology trends The server further checks the length and character set of the prompt statement, truncating excessively long parts and replacing characters that do not conform to the encoding standards if necessary, thereby outputting the final prompt statement string that meets the requirements of the generative artificial intelligence model interface, which is then used by the model calling module.
[0226] Step 4: The server encodes and binds parameters to the prompt statement, preparing for the generative artificial intelligence model invocation request.
[0227] The server takes the prompt statement output in step 3 and the preset model call parameters of the current session as input. The server writes the prompt statement into the request object according to the field names required by the generative AI model, and binds runtime parameters such as model name, maximum output length, temperature, and top-p. The server performs a token estimation operation on the prompt statement (e.g., estimating the number of tokens based on the model's word segmentation rules). If the sum of the estimated token value and the preset `max_tokens` exceeds the model's upper limit, the server performs a further truncation operation to ensure the total number of tokens remains within the allowed range. The server packages the final prompt statement and parameters into a model call request structure and passes it as output to the calling interface module for sending inference requests to the generative AI model.
[0228] Step 5: The server invokes the generative artificial intelligence model and receives the generated content.
[0229] The server takes the model call request generated in step 4 as input and sends it to the generative AI model server via an HTTP(S) client library or a local model interface. Upon receiving the prompt, the generative AI model performs Transformer-based inference operations internally, including tokenization, embedding encoding, self-attention calculation, and feedforward network operations, gradually generating an output token sequence. After receiving the response data from the model via the network interface, the server extracts the text fields from the response as the raw generated content. The server performs trimming operations on this raw text (removing leading and trailing whitespace, deleting redundant line breaks), and truncates the length if necessary. After processing, the server outputs the resulting natural language text as the generated content object, which is then passed to the subsequent formatting and storage modules.
[0230] Step 6: The server formats the generated content and generates distribution data for the target terminal.
[0231] The server takes the generated content text output from step 5, along with associated user and topic information, as input. The server performs paragraph segmentation on the generated content (breaking the text into multiple paragraphs based on line breaks or semantic segmentation rules), and can generate titles, summaries, or key point lists as needed. For example, the server can extract the first sentence as a summary, or select several key sentences based on a simple sentence scoring algorithm. The server organizes the generated content and this additional information into a structured record, adding metadata such as content identifiers and generation time. Subsequently, the server serializes this structured record into a transmission format suitable for terminal parsing (e.g., a structure containing field names and corresponding strings), outputting it as distribution data and sending it to the terminal via the communication network.
[0232] Step 7: The terminal receives the distribution data and displays the generated content on the interface.
[0233] The terminal takes the distribution data sent by the server as input, receives the response data through the network stack, and parses its content according to the protocol header. The terminal deserializes the received data into an internal object, from which it reads the generated content text, titles, paragraph list, and possible summary information. The terminal uses graphical interface components for layout based on device type and screen size, placing the titles at the top and displaying the paragraphs sequentially in a scrollable area. The terminal may adjust the font size or pagination strategy based on the content length. The terminal's output is readable text content presented on a human-computer interface, allowing users to browse the results generated by the generative artificial intelligence model.
[0234] Step 8: Users read the generated content and enter their evaluation information.
[0235] Users take the generated content displayed on the terminal as input and decide whether to provide a rating based on their subjective feelings after reading it. Users can select a rating level (e.g., 1-5 stars) in the rating control provided on the terminal, or click the "Helpful" or "Not Helpful" button, and enter natural language comments in the text evaluation box, such as "The content is relatively comprehensive" or "Lacking specific examples." The terminal takes these user actions as input, summarizes the rating, selection results, and text comments into an evaluation information object. After applying length limits and filtering illegal characters to the evaluation text, the terminal sends the evaluation information object as output to the server via the network.
[0236] Step 9: The server receives the evaluation information and stores it as an evaluation record.
[0237] The server takes the evaluation information object uploaded by the terminal as input, receives data through the network interface, and parses out the content identifier, rating value, tag Boolean value, and evaluation text. The server performs a validity check on the rating value (limiting it to a predetermined range) and performs necessary de-identification and length truncation on the evaluation text. Subsequently, the server writes these fields, along with the corresponding generated content identifier and user identifier, into the database, forming a new evaluation record. Each evaluation record in the database includes a primary key, content ID, user ID, topic tag, rating value, evaluation tag, evaluation text, and timestamp. The server outputs a message indicating successful insertion of the record for subsequent statistical analysis.
[0238] Step 10: The server calculates evaluation metrics based on evaluation records and updates the rules for generating prompt statements.
[0239] The server takes accumulated review records from the database as input and executes the statistical analysis module periodically or on an event-triggered basis. The server groups and aggregates review records according to content themes, calculating evaluation metrics such as average score, score variance, number of valid reviews, and "helpful" percentage for each theme. Based on these metrics, the server performs rule-based judgment operations: for example, when the average score of a theme is below a threshold and the keyword "shallow content" appears frequently in the review text, the server updates the prompt template for that theme to "requires more details and structured output," and increases the output length control value of the generative AI model; conversely, when user reviews focus on "content too long," the server reduces max_tokens or shortens the task requirement description in the template. The server writes the updated template and runtime parameters to the configuration storage area, forming new prompt generation rules and model call parameter configurations, and uses these updated results as output for direct application in subsequent sessions.
[0240] Step 11: The server adaptively adjusts the behavior of subsequent generative artificial intelligence model calls based on the updated configuration.
[0241] The server takes the latest prompt generation rules and runtime parameter configurations output in step 10 as input. When it receives new user input, it prioritizes the updated template selection strategy and parameter binding logic. During the prompt generation phase, the server automatically selects a more suitable template based on the new rules, and during the model invocation phase, it uses parameter values such as max_tokens, temperature, and top-p, which have been optimized by evaluation metrics. Because these updates are automatically completed internally by the server based on evaluation data and statistical metrics, the server achieves continuous optimization of the generative AI model invocation behavior without increasing the user's operational burden. Through this output behavior, the server improves the relevance, length suitability, and user satisfaction of subsequently generated content, while reducing the computational and communication overhead caused by redundant generation.
[0242] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0243] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0244] In existing human-computer interaction technologies, servers typically construct fixed-format query statements or simple commands based solely on the user's current natural language input and send them to a generative AI model to obtain a response. In this model, servers often fail to fully utilize user-specific attribute information and behavioral history stored long-term in data storage devices, instead treating the generative AI model as a general question-answering tool. This technical approach suffers from the following problems: First, when constructing inputs for generative AI models, the server lacks structured analysis and modeling of users' historical behavior, resulting in prompts that contain almost no personalized interests, preferences, or current context. As a result, generative AI models can only provide answers with a high degree of generalization and low personalization.
[0245] Second, when processing requests from communication applications or information retrieval tools, servers typically only transmit the surface text input by the user. They do not uniformly associate and extract user identification information, individual attribute information, and behavioral history information scattered across multiple data sources. This prevents the system from effectively modeling the user's state at the infrastructure level, and computing resources are wasted on repetitive and undifferentiated generation tasks.
[0246] Third, there is a lack of standardized and computable mechanism for constructing prompts between the server and the generative AI model. Existing technologies mostly rely on manually written static prompts, lacking template-based programmatic generation methods, and do not explicitly embed the user's authentication status and business constraints in the prompts. This makes it difficult to reliably apply generative AI models in security-sensitive scenarios (such as address changes and identity authentication-related businesses).
[0247] Fourth, in scenarios involving document generation or sensitive operations (such as address change procedures), servers often treat the results of generative artificial intelligence models only as "reference texts," lacking a mechanism to deeply integrate user attribute information, business rules, and the model generation process. This makes it impossible to guarantee the consistency and compliance of the output results with the backend records along the computation path.
[0248] Therefore, a new computer implementation is needed to enable servers to automatically acquire user identification information at the system level, combine it with individual attribute information and behavioral history information stored in multi-source data storage devices, perform statistical or classification processing to extract user interest information and user status information, and construct prompt statements for generative artificial intelligence models based on template mechanisms. Key states such as authentication results are embedded as computable features in the prompts, thereby improving the personalization, business relevance, and overall system processing efficiency of the output results of generative artificial intelligence models while ensuring security, thus achieving an improvement in computer technology itself.
[0249] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0250] In this invention, the server includes a device for acquiring user identification information and user request information in an information processing device; a device for retrieving individual attribute information and behavioral history information stored in an information storage device based on the user identification information to obtain related information associated with the user; a device for performing statistical processing or classification processing on the related information and the user request information to extract user interest information and user status information; a device for selecting template information based on the user interest information and the user status information and generating a prompt statement for inputting into a generative artificial intelligence model by filling the variable portion of the template information with the related information and the user request information; a device for inputting the prompt statement into the generative artificial intelligence model and obtaining response information from the generative artificial intelligence model; and a device for performing display format conversion processing or security determination processing on the response information and sending the processed response information to a terminal device. This allows for the formation of a computational path within the server, indexed by user identification information, based on data storage devices, and centered on statistical analysis and templated prompt statement construction. This enables the input of generative artificial intelligence models to include structured user interest information, user status information, and authentication status information. Consequently, at the computer technology level, this enhances the expressiveness and computability of prompt statements, improves the personalization and business adaptability of model output results, and reduces the burden on terminals and external applications through unified security judgment processing, thereby improving the overall processing efficiency and reliability of the human-computer interaction system.
[0251] A “system” refers to a set of hardware and software that work together to perform user information processing, data analysis, and generative artificial intelligence model invocation, consisting of at least one information processing device, at least one information storage device, and at least one terminal device connected through a communication network.
[0252] "Information processing device" refers to an electronic device equipped with a processor, memory and communication interface, used to execute program instructions, process user-related data and interact with information storage devices and terminal devices.
[0253] "Information storage device" refers to data storage resources used to store individual attribute information, behavioral history information and other related information in a searchable manner, including database systems, file storage systems or other persistent storage media.
[0254] "Terminal device" refers to an electronic device operated by a user and interacting with an information processing device through a communication network, including but not limited to mobile terminals, computing terminals, or other devices with display and input functions.
[0255] "User identification information" refers to the identification information used to uniquely or quasi-uniquely distinguish different users within the system, including account identifiers, login identifiers, device identifiers, or combinations thereof.
[0256] "User request information" refers to natural language information or structured request data that is input by the user through a terminal device and sent to the information processing device to represent the user's current needs or problems.
[0257] "Individual attribute information" refers to static or slowly changing characteristic information related to a specific user, including attribute data recorded in information storage devices such as age, residential area, language preference, interest tags, and identity status.
[0258] "Behavioral history information" refers to information that represents a user's operation or interaction records within a certain time range, including search records, browsing records, click records, conversation records, or other behavioral log data.
[0259] "Associated information" refers to a comprehensive set of information related to a user that is retrieved and associated from an information storage device based on the user's identification information. It includes at least individual attribute information and behavioral history information.
[0260] "Statistical processing" refers to the quantitative operations performed by information processing devices on related information or user request information, such as counting, summarizing, frequency analysis, and time distribution analysis, to extract patterns or features.
[0261] "Classification processing" refers to the process by which an information processing device divides related information or user request information into one or more categories according to predefined rules or models, including rule-based classification and machine learning-based classification.
[0262] "User interest information" refers to information inferred from historical behavioral information and related data through statistical processing or classification, used to represent users' preferred topics or areas of interest.
[0263] "User status information" refers to information determined based on associated information and user request information, used to represent the user's current status or context, including current problem type, time conditions, budget preference, authentication status, etc.
[0264] "Template information" refers to a text or data structure with a fixed structure and replaceable variable parts that is predefined and stored in an information processing device, used to generate prompt statements for generative artificial intelligence models.
[0265] The “variable section” refers to the placeholder area in the template information that is used to be replaced by dynamic data. This area is filled with associated information or user request information when the prompt statement is generated.
[0266] "Prompt statements" refer to natural language text or equivalent representations generated by an information processing device based on template information, related information, and user request information, and used as input to generative artificial intelligence models.
[0267] "Generative AI models" refer to AI models that can automatically generate text, code, or other data content based on input prompts, including but not limited to language models based on deep learning.
[0268] "Response information" refers to the generated data that a generative artificial intelligence model outputs after receiving a prompt, representing the answer or recommendation result, including text or structured data.
[0269] "Display format conversion processing" refers to the process by which an information processing device adjusts the content structure or presentation format of response information to adapt to the display requirements of the terminal device, including operations such as segmentation, formatting, summarization, and structuring.
[0270] "Security assessment and processing" refers to the processing performed by the information processing device on the response information, such as content inspection, sensitive information filtering, compliance verification, and risk assessment, in order to prevent inappropriate information or information that does not comply with security policies from being sent to the terminal device.
[0271] "Authentication infrastructure" refers to external or internal authentication systems used to perform identity verification processing, including biometric authentication systems, account authentication systems, or unified identity authentication platforms.
[0272] "Biometric information" refers to user physiological or behavioral characteristic data used for identity recognition, including fingerprints, facial features, voiceprints, iris features, or other biometric data that can be used for identity verification.
[0273] In various embodiments of this invention, the server, terminal, and user collaboratively process information around a generative artificial intelligence model and prompts to provide personalized information based on user identification information and historical behavior information. The system of this invention makes specific limitations on hardware structure, software module division, and data structure design, thereby achieving improvements in computer technology itself.
[0274] In one implementation, the server includes at least one processor, at least one memory, and a network interface. At the operating system layer, the server can use a general-purpose operating system, such as a Unix-like operating system. At the application layer, the server can deploy application service frameworks, such as a scripting language-based web framework. At the data layer, the server can connect to a relational database management system, such as a relational database, for storing structured data such as user attribute tables and configuration tables; the server can also connect to a document-oriented database system, such as a document database, for storing semi-structured data such as user behavior log documents and conversation record documents. The server can also access generative artificial intelligence model inference services deployed in the same data center or a remote computing environment via the network interface, such as a language model service built on a deep learning framework.
[0275] In one embodiment, the terminal includes a mobile computing device or a desktop computing device. At the operating system layer, the terminal can use a mobile operating system or a desktop operating system. At the application layer, the terminal can run communication applications or information retrieval tools. The user inputs a natural language request through the terminal's graphical user interface, and the terminal combines this request with the login identifier and device identifier obtained from the local graphical interface to form user identification information and user request information.
[0276] In one implementation, the server uses program instructions stored in memory to process user identification information and user request information from the terminal. In its data access module, the server accesses relational and document-oriented databases via a database driver. The server maintains a user attribute table in the relational database, which may include fields for storing unique user identifiers and fields for storing attributes such as age, residential region, language preference, and interest tags. The server maintains a collection of user behavior logs in the document-oriented database, with each document containing fields such as user identifier, timestamp, behavior type, behavior content (e.g., search keywords, clicked object identifiers), and terminal type.
[0277] In one implementation, the server performs feature extraction and statistical calculations on the associated information composed of individual attribute information and behavioral history information retrieved from the data storage device. In its statistics module, the server can use algorithms such as counting, frequency statistics, and time window statistics to perform word frequency statistics on keyword fields in the behavioral history information and frequency statistics on behavior type fields, thereby obtaining the user's interest distribution within a predetermined time window. For example, if the server analyzes the search keywords "beach, diving, island" in the behavior log and finds that the frequency of these keywords is higher than a preset threshold, then the server assigns a higher weight to the "island vacation" dimension in the interest vector.
[0278] In the classification module, the server can use either rule-based classification or machine learning classification. In rule-based classification, the server maps "beach, diving, island" to the "island vacation" interest category based on a pre-built keyword-to-category mapping table, and "family trip, amusement park" to the "family travel" interest category. In machine learning classification, the server can use vector space-based classifiers, such as linear classifiers or tree models, to transform behavioral logs into high-dimensional feature vectors and output interest category labels through a trained model. The server combines the obtained interest categories, interest intensity, and temporal distribution to form user interest information.
[0279] In one implementation, the server synthesizes user status information by combining user request information with user interest information. In its natural language processing module, the server uses word segmentation and syntactic analysis algorithms to process the user request text, identifying request type, time conditions, budget conditions, etc. For example, if a user enters "Where would you recommend for your next vacation?", the server identifies "vacation" as the time condition and "travel" as the request category. The server then combines this request category with the high-weighted "island vacation" from the user's interest information, labeling the user status information as "vacation travel request + island vacation preference + no explicit budget condition or inferred medium budget based on attribute information".
[0280] In one implementation, the server embeds user interest information, user status information, and individual attribute information as features into the prompt statement generation process. The server pre-stores multiple prompt statement templates in its template management module; each template consists of a fixed text section and variable placeholders. The server selects the appropriate template based on the user request type. For example, in a travel recommendation scenario, the server selects a template that includes a "user history behavior description," a "current problem description," and an "output requirement" section.
[0281] In one implementation, the server generates a prompt statement by filling in the associated information and user request information into the variable section of a template. For example, in a travel scenario, the server generates the following prompt statement: "The user has previously searched for content related to beach vacations and diving multiple times and read travel guides for Phuket and Bali. Now the user is asking: 'Where would you recommend for my next vacation?' Please recommend 3 to 5 beach vacation destinations suitable for a week-long trip within a medium budget, explaining the best time to visit each destination, recommended activities, and approximate cost per person." In one implementation, the server uses the aforementioned prompt as input to the generative AI model. Within its model interface module, the server invokes the generative AI model service deployed on another computing node via a network request. In this implementation, the generative AI model employs a neural network language model based on a transformer architecture. During training, the model uses a large-scale text corpus, encodes the input sequence through a self-attention mechanism, and performs feature transformation through a multi-layer feedforward network. During training, the model uses a cross-entropy loss function as the error function and updates the network parameters using a gradient descent-based optimization algorithm (e.g., an adaptive learning rate optimization method). During training, the model can initialize parameters through pre-training tasks such as masked language modeling and next-sentence prediction, and can be fine-tuned using domain data to enhance the generation quality in specific application scenarios.
[0282] In one implementation, the server leverages structured information from prompts to improve the computational efficiency and output accuracy of the model. By explicitly providing user interest information (e.g., "beach vacation and diving related content") and output constraints (e.g., "recommend 3 to 5 destinations, specifying the best season, recommended activities, and budget") in the prompts, the server reduces the search range of the model in the output space. This allows the model to converge more quickly to candidate sequences highly relevant to user needs during inference, thereby shortening generation time and reducing the probability of generating irrelevant content.
[0283] In one implementation, the server can also use a similar mechanism in scenarios related to address changes. When the server detects that a user's request is related to an address change, it automatically generates the document information structure required for the address change procedure using existing fields such as address and identity information in the individual's attribute information. This includes, for example, a list of required fields and format constraints. The server embeds these document information generation conditions as constraint information into a prompt statement. The prompt statement could be, for example,: "The user needs to change their address. The user's existing records include name, old address, and contact information. Please generate an address change application template that includes a new address field, a date field, and a user signature field, according to the address change requirements, and ensure that the field names are clear and the format is standardized." In one implementation, in scenarios involving identity authentication, the server uses user identification information, including biometric data, to interact with the authentication infrastructure. After communicating with the authentication infrastructure, the server writes the authentication result (e.g., authentication successful or failed, authentication timestamp, authentication level) as an authentication status field into the associated information. When generating a prompt statement, the server embeds the authentication status information, for example: "The user has passed advanced identity verification within the last 24 hours. Please provide only a summary of sensitive information that matches the verified user's own records in your response, and do not include any information from other users." In one implementation, the server performs formatting and security checks on the response information returned by the generative artificial intelligence model. In the formatting module, the server can segment long text based on the terminal type and display area, number list items, or separate the destination name and description into a key-value structure to adapt to the terminal's interface layout. In the security checks module, the server can maintain a list of sensitive words and a set of rules, and can use a lightweight classification model to detect inappropriate content in the response. When a potential risk is detected, the server can replace, block, or require the model to regenerate the corresponding segment.
[0284] In one implementation, after receiving the response information from the server, the terminal displays it according to its own application logic. In a communication application scenario, the terminal displays the response text as a message bubble; in an information retrieval tool scenario, the response text is combined with structured items into a list or card format. The user then proceeds to the next step based on the results displayed on the terminal, such as asking more detailed questions or performing subsequent operations like booking or submitting.
[0285] The system of this invention is not limited to the single implementation described above. In another embodiment, the server can decentralize the template information and prompt generation logic to edge nodes to reduce the load on the central node and network latency. In this embodiment, the server can distribute the template management module and statistical analysis module across multiple nodes, reducing repeated access to the backend database through an intermediate caching mechanism, thereby reducing communication load and improving response speed.
[0286] In another implementation, the server can use different types of generative AI models, such as language models with varying parameter sizes, retrieval-enhanced generative models combined with retrieval modules, or domain models fine-tuned for specific fields (such as law or medicine). The server can dynamically select models of different sizes based on the complexity of user requests and response latency requirements to balance accuracy and computational resource consumption.
[0287] The system of this invention introduces a multi-source data fusion, statistical analysis, and templated prompt generation mechanism driven by user identification information within the server. This allows generative artificial intelligence models to no longer directly face raw, unstructured user request text, but instead receive structured and constraint-enhanced prompts. Because the prompts contain user interest information, user status information, and authentication status information, the model can more quickly focus on user-related feature dimensions when performing calculations based on attention mechanisms and parameter weights, reducing useless computation and error path exploration. Therefore, it can achieve higher generation speed and lower error rate with the same hardware resources.
[0288] These technical features of the present invention are not merely simple automation of manual operations, but rather optimization and reconstruction of the information processing flow inside the computer at three levels: data structure design, prompt construction algorithm, and model calling strategy. This results in verifiable technical effects in terms of processing speed, generation accuracy, data management, and communication resource utilization.
[0289] use Figure 13 The processing flow is explained.
[0290] Step 1: The user initiates a request on the terminal.
[0291] Users enter natural language request text in the communication application or information retrieval tool interface of the terminal and submit the request by clicking the send or search button.
[0292] Input: The request text that the user types into the input box, such as "Where would you recommend to travel for the next holiday?".
[0293] Output: A user input event containing the requested text is generated internally within the terminal.
[0294] Step 2: The terminal generates and sends user identification information and user request information.
[0295] The terminal reads data such as account identifier and device identifier from the local session, combines them to form user identification information, and encapsulates it together with the user's input request text into a request data packet, which is then sent to the server via network protocol.
[0296] Input: The user request text from step 1 and the account identifier and device identifier stored locally on the terminal.
[0297] Specific data processing: The terminal concatenates or packages the account identifier and device identifier into a structured field, writes the request text into the request body, and attaches metadata such as timestamp and application identifier.
[0298] Output: The request data packet sent to the server over the network, which includes at least user identification information and user request information.
[0299] Step 3: The server parses the request data packet and verifies its basic validity.
[0300] The server receives request data packets from the terminal in the network interface module, and parses out metadata such as user identification information, user request information and timestamps at the application layer, while performing basic format verification and identity verification.
[0301] Input: The request data packet sent to the server in step 2.
[0302] Specific data processing: The server decodes the requested data, checks whether the required fields exist, verifies whether the user identification information conforms to the predetermined format, and determines whether the request comes from a valid session based on the session token or signature.
[0303] Output: A formatted user identification information object and a user request information object. If the verification fails, an error status will be output.
[0304] Step 4: The server retrieves related information from the information storage device.
[0305] The server uses user identification information objects as retrieval keys to access relational databases and document databases, retrieve individual attribute information and behavioral history information corresponding to the user, and merge them into related information.
[0306] Input: The user identification information object obtained in step 3.
[0307] Specific data processing: The server executes queries in the relational database, retrieving fields such as age, residential area, and interest tags from the user attribute table; it also executes conditional queries in the document database, retrieving recent search records, click records, and browsing categories from the behavior log collection. The server then joins and merges multiple query results based on user identifiers, removing duplicate and expired records.
[0308] Output: An associated information object containing individual attribute information and behavioral history information.
[0309] Step 5: The server performs statistical analysis on historical behavioral information and extracts preliminary interest features.
[0310] In the statistics module, the server performs word frequency statistics, behavior category counting, and time distribution analysis on the behavioral history information in the associated information to extract the user's high-frequency interest topics.
[0311] Input: The behavior history portion of the associated information object output in step 4.
[0312] Specific data processing: The server segments the keyword field in the behavior log, maps each keyword to a standard term, and counts the number of times each term appears and the most recent appearance time; the server categorizes and summarizes the behavior types (such as "search: travel", "browse: travel guide"), calculates the access frequency and time decay weight of different categories, and forms a preliminary interest vector by weighted summation.
[0313] Output: A preliminary interest feature vector or a list of interest keywords representing the distribution of user interests.
[0314] Step 6: The server classifies preliminary interest features and generates user interest information.
[0315] In the classification module, the server maps initial interest features to a predefined set of interest categories based on rules or machine learning models, thereby obtaining user interest information that represents the user's long-term preferences.
[0316] Input: The preliminary interest feature vector or list of interest keywords output in step 5.
[0317] Specific data processing: The server uses a keyword-to-category mapping table or classification model to map keywords such as "beach, diving, island" to the "island vacation" category, and keywords such as "family travel, amusement park" to the "family trip" category; the server normalizes the weights of each category, selects several categories with weights higher than the threshold as the user's main interests, and records the most recent activity time for each category.
[0318] Output: A user interest information object containing one or more interest categories, their weights, and time information.
[0319] Step 7: The server parses the user request information and extracts the request type and context conditions.
[0320] In the natural language parsing module, the server performs word segmentation, part-of-speech tagging, and intent recognition on the user request text to determine contextual elements such as request type, time conditions, and budget conditions, forming part of the user status information.
[0321] Input: The user request information object output from step 3.
[0322] Specific data processing: The server detects keywords through rules or models, such as identifying "holiday" and "weekend" as time-related words, and "travel," "study plan," and "address change" as request type words; the server analyzes the sentence intent based on specific phrase structure, marks "Where do you recommend to travel next holiday?" as a "travel destination recommendation" request type, and checks whether there is an average spending level field in the user attribute information under the default budget to infer the budget category.
[0323] Output: A request parsing result object containing fields such as request type, time conditions, and budget conditions.
[0324] Step 8: The server synthesizes user status information.
[0325] The server combines user interest information objects, request parsing result objects, and individual attribute information to form user status information describing the current session context.
[0326] Input: Individual attribute information from step 4, user interest information object from step 6, and request parsing result object from step 7.
[0327] Specific data processing: In the logical combination module, the server determines the interest category most relevant to the current request based on the correlation between the request type and the interest category. For example, it associates the "travel destination recommendation" category with the "island vacation" category. The server attaches attributes such as age and city of residence to the status information and marks metadata such as the current session time and request priority.
[0328] Output: A user status information object including primary interests, secondary interests, request type, time scenario, budget inference, and basic user attributes.
[0329] Step 9: Server selection prompt statement template.
[0330] In the template management module, the server selects a matching prompt template based on the request type and user status information, and prepares to populate the variable section.
[0331] Input: The user status information object output from step 8 and the pre-stored template set.
[0332] Specific data processing: The server calculates the applicability of each template and the degree of matching with the current request. For example, if the request type is "travel recommendation", then the travel template that includes "historical behavior description + current problem description + output format requirements" will be selected first. The server selects the optimal template from multiple candidate templates based on priority and suitability.
[0333] Output: The selected template object containing fixed text portions and variable placeholders.
[0334] Step 10: The server fills in the template and generates the prompt message.
[0335] The server fills in the template variables with information such as association information, user interest information, and user status information, and generates complete natural language prompts for input into the generative artificial intelligence model.
[0336] Input: the associated information object from step 4, the user interest information object from step 6, the user status information object from step 8, and the selected template object from step 9.
[0337] Specific data processing: The server will fill in the history description position of the template with a summary of the user's past behavior, such as "the user has searched for content related to beach vacations and diving multiple times in the past and read travel guides for Phuket and Bali", fill in the current question position with "Where do you recommend to travel next vacation?", and fill in the output requirement position with "Recommend 3 to 5 beach vacation destinations suitable for a week trip, and explain the best season, recommended activities and budget".
[0338] Output: Complete prompts that the generative AI model can directly receive, such as: "The user has previously searched for content related to beach vacations and diving multiple times and read travel guides for Phuket and Bali. Now the user is asking: 'Where would you recommend for my next vacation?' Please recommend 3 to 5 beach vacation destinations suitable for a week-long trip within a medium budget, explaining the best time to visit each destination, recommended activities, and approximate cost per person." Step 11: The server invokes the generative artificial intelligence model and obtains the response information.
[0339] In the model interface module, the server sends the prompt statement as input to the generative artificial intelligence model service, and receives the generated response text after the model completes inference.
[0340] Input: The prompt message output in step 10.
[0341] Specific data processing: The server encodes the prompt statement into request parameters and submits them to the language model deployed on the inference server via a network protocol. Internally, the language model performs vectorization encoding and sequence generation based on a transformer structure, while the server externally only waits for and receives the generated results. Upon receiving the data, the server decodes it to extract the main answer text or a set of candidate answers.
[0342] Output: A response information object containing the text of the model-generated answer.
[0343] Step 12: The server formats and performs security checks on the response information.
[0344] In the result processing module, the server performs structured segmentation and format optimization on the generated text, and checks whether it contains sensitive content or expressions that do not comply with the policy through the security judgment module.
[0345] Input: The response information object output in step 11.
[0346] Specific data processing: The server divides a long text into multiple suggestions according to predefined segmentation rules, such as "1.", "2.", etc., and extracts key fragments such as destination name and budget description to build a list structure; the server scans the text using a sensitive word list and classification model, and replaces, truncates, or re-requests the generation of inappropriate content if it is found.
[0347] Output: The final response data object that has passed security checks and is adapted to the terminal display requirements, including formatted text and an optional list of structured items.
[0348] Step 13: The server then sends the final response to the terminal.
[0349] The server encapsulates the final response data object into a response message through the network interface and sends it to the terminal that originally made the request.
[0350] Input: The final response data object output from step 12.
[0351] Specific data processing: The server sets adaptation information in the response header according to the terminal type, and encodes the text content and structured data in a unified format to ensure that the terminal can parse it without ambiguity.
[0352] Output: The response message sent to the terminal.
[0353] Step 14: The terminal parses and displays the response content.
[0354] The terminal receives the server's response message in the communication module, parses out the text content and optional structured items, and displays them on the user interface in the form of messages or lists.
[0355] Input: The response message received from the server in step 13.
[0356] Specific data processing: The terminal decodes the message, displays the main text as a dialogue message or result area, renders structured items as cards or list items with titles and summaries, and automatically adjusts the layout according to the screen size and interface template.
[0357] Output: Personalized recommendations or answers displayed on the terminal screen, which users can perceive and use for subsequent decision-making.
[0358] Step 15: Users can then interact with the displayed results.
[0359] After reading the response displayed on the terminal, users can make decisions based on the information they have obtained, or issue more detailed new requests, thereby triggering a new round of request processing.
[0360] Input: The result information displayed on the terminal in step 14.
[0361] Specific data processing: At the cognitive level, users compare and filter the candidate suggestions, and choose to continue asking questions or perform a specific action (such as clicking on a destination card). This action is then converted into a new user request event in the terminal.
[0362] Output: New user input behavior or business operation instructions, serving as the starting point for the next round of steps 1 and 2.
[0363] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0364] In existing technologies, computer-based personalized service systems primarily employ fixed rules or single machine learning models to analyze users' historical behavior offline and generate recommendations or document content. These systems typically suffer from the following problems: (1) The system cannot simultaneously utilize user input information, communication application or information retrieval tool identifiers, historical behavior information and user emotional state under the same technical framework, resulting in insufficient understanding of user context and limited personalization of recommendation results and automatically generated document content.
[0365] (2) When calling generative artificial intelligence models, the system often uses fixed and static prompt texts and fails to dynamically construct prompt statements based on the user's real-time behavioral characteristics and emotional state. This results in a low degree of matching between the generated results and the user's current needs and psychological state, thereby reducing the overall interaction efficiency.
[0366] (3) The system usually implements the recommendation module, document generation module and user feedback collection module separately, lacking an online update mechanism with behavior history as a closed loop. It cannot use the user's actual operation behavior on the output results for subsequent feature extraction and candidate information selection dynamic optimization, resulting in low utilization of computing resources and difficulty in timely convergence to a better model strategy.
[0367] (4) In electronic procedures or online transaction scenarios, existing systems mostly only realize the generation of surface template-filled documents, without making full use of generative artificial intelligence models to automatically construct complex document structures. At the same time, when linked with external authentication systems, there is a lack of computer implementation methods to integrate the authentication results with the document content confirmation process, which increases the complexity of the process and the probability of errors.
[0368] (5) In scenarios with legal or administrative constraints, such as address changes, traditional systems often require users to manually understand the regulations and procedures, select forms and fill in content themselves. The computer is only a passive form container and fails to provide integrated intelligent support based on data processing and generative model reasoning, resulting in high human-computer interaction costs, time-consuming processes and easy omissions or errors.
[0369] Based on the above problems, the technical challenge to be solved by this invention is to provide a unified technical solution implemented at the computer system level, enabling the server to use data processing programs and learning programs to jointly model input information from user terminals and their associated behavioral history and emotional state. By constructing context-adaptive prompts, the generative artificial intelligence model is driven to perform inference processing. In the process of recommending goods or services, automatically generating document data, and handling online procedures, an efficient, dynamic, and self-learning integrated processing flow is achieved, thereby substantially improving the technical performance and resource utilization efficiency of computers in personalized information processing, document generation, and authentication process management.
[0370] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0371] In this invention, the server includes: means for acquiring past user behavior information and attribute information from a data storage device based on user input information from a user terminal and the identifier of a communication application or information retrieval tool, and for completing missing information in the user input information; means for performing preprocessing and feature extraction on the completed information and user input information using a data processing program and a learning program, thereby determining user preference information and user emotional state within the computer; means for selecting candidate information for goods or services and / or generating candidate information for document data or input data based on user preference information and user emotional state; means for generating prompt statements containing candidate information, user preference information and / or user emotional state, and using the prompt statements as prompt input to a generative artificial intelligence model to instruct the generative artificial intelligence model to perform inference processing; means for determining recommended information for goods or services and / or document data or input data in a processor based on the output results obtained from the generative artificial intelligence model, and generating it as output information for the terminal; and means for acquiring user operation information related to the output information from the terminal, storing the operation information as a behavior history in the data storage device, and using the behavior history to update feature extraction and the selection and / or generation of candidate information. This allows for joint data processing and online learning of user input, behavioral history, and emotional state within a unified server architecture. By dynamically constructing prompts, generative artificial intelligence models can be driven to generate more context-appropriate recommendation results and document content. Furthermore, the candidate information selection and feature modeling process can be continuously optimized based on user feedback, thereby improving the computational efficiency, response speed, and result quality of the computer system in the integrated processing of personalized recommendations, automatic document generation, and authentication processes.
[0372] "User input information" refers to text information, option information, form information, control commands, and various data contents related to the current task that users provide to the system through the terminal.
[0373] "Communication applications" refers to applications used for sending and receiving messages over a network, including but not limited to instant messaging applications, social networking applications, and client programs with message interaction capabilities.
[0374] "Information retrieval tool identifier" refers to the tag information used to uniquely identify the information retrieval tool or search service used by a user, which enables the server to distinguish different information retrieval environments and associate them with corresponding historical behavior data.
[0375] "Data storage device" refers to storage resources used to store user behavior information, attribute information, candidate information, model parameters and other related data in a computer system, including local storage devices and network-based storage resources.
[0376] "Past behavior information" refers to historical data generated by a user's interaction with a terminal or service within a predetermined time frame, including search history, browsing history, click history, purchase history, and other operation records.
[0377] "Attribute information" refers to static or semi-static feature information related to users, including users' basic information, preference features, group identifiers, and other feature data related to personalization.
[0378] "Completing missing information" refers to generating and filling in fields that are missing or incomplete in the input information based on user input, past behavior information, and attribute information through data processing or inference.
[0379] "Data processing program" refers to a software module that runs on a processor and performs data cleaning, format conversion, aggregation statistics, feature construction and other data preprocessing operations.
[0380] "Learning program" refers to a program module that runs on a processor for training, updating, and inferring machine learning models, including algorithm implementations for feature extraction, classification, clustering, regression, and recommendation.
[0381] "Preprocessing" refers to the process of cleaning, normalizing, encoding, standardizing, and other preparatory treatments of raw data before performing model inference or data analysis.
[0382] "Feature extraction" refers to the process of extracting numerical or symbolic features from user input information, completed information, and behavioral history to describe user state, preferences, and context.
[0383] "User preference information" refers to parameterized information that reflects users' long-term or short-term preferences in terms of goods, services, content types, price ranges, etc., determined based on feature extraction results.
[0384] "User emotional state" refers to the category and intensity of a user's current emotion inferred by a computer through analysis of user data such as text, voice, images, and behavioral patterns.
[0385] "Candidate information" refers to a set of product information, service information, or document fragments that the system pre-selects based on user preference information and / or user emotional state before generating the final recommendation information or document data.
[0386] "Candidate information for goods or services" refers to the intermediate set of goods or services used to provide recommendations to users, including their identifiers, categories, attribute parameters, and user-related rating information.
[0387] "Candidate information for document data or input data" refers to the set of text fragments, field contents and their structured representations used to constitute the target document, form or data record, which are in a candidate state before final determination.
[0388] "Prompt statements" refer to instructional or explanatory texts automatically generated by the server based on user input, completed information, user preferences, and user emotional state. These texts are used as input to generative artificial intelligence models to guide them in performing specific reasoning tasks.
[0389] "Generative artificial intelligence models" refer to artificial intelligence models that can automatically generate text, code, or other data content based on input prompts, including natural language generation models based on deep learning.
[0390] "Prompt input" refers to providing prompt statements as input parameters to a generative artificial intelligence model to trigger the model to perform corresponding reasoning and generation operations.
[0391] "Inference processing" refers to the computational process by which a generative artificial intelligence model, after receiving prompts, analyzes the input based on its internal parameters and learned knowledge, and generates output results.
[0392] "Recommended information" refers to the product or service suggestions determined and presented to users based on candidate information, the results of generative artificial intelligence models, and business rules.
[0393] "Output information" refers to the structured or unstructured data that the server constructs to provide display and interaction to the terminal after completing the recommendation, document generation or data generation, including text, lists, tags and metadata.
[0394] "User action information" refers to the interactive behavior data of users on the terminal in response to output information, including clicks, selections, editing, confirmations, rejections, ratings, and other interaction records.
[0395] "Behavioral history" refers to a time-series data set accumulated from user operation information and other interaction records, which is used for subsequent analysis, modeling, and system adaptive updates.
[0396] "Update feature extraction and candidate information selection and / or generation" refers to adjusting the feature extraction method, model parameters, or candidate information generation strategy by utilizing newly obtained behavioral history in order to improve subsequent data processing and result generation.
[0397] "Address information" refers to structured or unstructured data that represents geographical location, including elements such as country, region, city, street, house number, and postal code.
[0398] "Document data for address change procedures" refers to electronic document data used to process address change procedures with administrative agencies, business agencies or other entities, including table field content and document text content.
[0399] "Display data" refers to the visual data generated for presentation on the terminal interface, including text, control properties, layout information, and other parameters related to the display.
[0400] "Biometric information" refers to physiological or behavioral characteristics used for individual identification, including fingerprint features, facial image features, voiceprint features, and other characteristic data that can be used for identity verification.
[0401] "Authentication information" refers to various types of information used to complete identity verification, including but not limited to account identifiers, passwords, one-time verification codes, ID numbers, and encrypted credentials.
[0402] "Communication methods for authentication" refers to the communication protocols and security mechanisms used to transmit authentication information between the server and external authentication systems, including secure transmission channels, authentication protocols, and encryption methods.
[0403] "External authentication system" refers to a computer system or service platform outside of the present invention used to verify the identity of users.
[0404] "Personal verification" refers to an automated verification process that confirms the authenticity of a user's identity through an external authentication system based on biometric information and / or authentication information.
[0405] "Content confirmation result" refers to the status information regarding whether the user has viewed, understood, and agreed to the output information content, including confirmation marks, modification records, and related time information.
[0406] Embodiments of the present invention will be described in conjunction with the system architecture defined in the appendix. In the following description, the terms "server," "terminal," and "user" are used exclusively.
[0407] In one implementation, the server is deployed on a computing platform based on general-purpose computing hardware. The server hardware includes a multi-core processor, main memory, persistent storage, and network interface circuitry. The server software includes an operating system, database management program, web service program, middleware, and multiple application modules. In its implementation, the server can use a general-purpose operating system as its runtime environment, a relational database management system as its data storage device, a service program based on the Hypertext Transfer Protocol (HTTP) as its network interface, and a scripting environment as its application logic execution environment. For data processing, the server uses data analysis libraries, machine learning libraries, and natural language processing libraries as data processing and learning programs.
[0408] In one embodiment, the terminal is a smart mobile terminal or a general-purpose computing terminal. Hardware-wise, the terminal includes a processor, main memory, display panel, input devices (touchscreen, keyboard, etc.), image acquisition devices, and audio acquisition devices. Software-wise, the terminal includes a terminal operating system, a browser program or dedicated application program, and a communication module for communicating with a server.
[0409] In one implementation, the user provides user input information to the server through a terminal interface. The user enters text information, address information, or control commands on the terminal and allows the terminal to access the user's biometric information or authentication information so that the server can perform identity authentication.
[0410] In terms of data structure, the server divides user-related data into multiple tables or collections. In one implementation, the server stores at least the following data structures in the data storage device: The server stores static or semi-static fields such as user identifiers and attribute information (e.g., age group codes, regional codes, user group labels) in the user basic information table.
[0411] The server stores past behavior information in the user behavior history table, including timestamps, interaction types (enumerated values of search, click, purchase, edit, etc.), resource identifiers (product identifiers, document identifiers, page identifiers), quantitative attributes (duration of stay, number of operations, amount, etc.), and source application identifiers (identifiers of communication applications or information retrieval tools).
[0412] The server stores the context information of the current session in the session state table, including the most recent user input information, the most recently inferred user sentiment state, the current task type (recommendation, document generation, address change, etc.), and the identifier of the document being edited.
[0413] The server stores candidate information for goods or services, as well as candidate information for document data or input data such as document fragments and template fragments, in the candidate information table, and supports fast querying through index structures (such as inverted indexes or multi-field indexes).
[0414] In one implementation, the server processes the behavioral history using DataFrame objects from a data analysis library. After loading user behavior records, the server adds derived fields to each record, such as the location code of the behavior time within a week, whether it is a holiday, and the time difference with the current session time. During preprocessing of this data, the server performs operations such as missing value imputation, outlier filtering, and standardization to construct a data matrix suitable for feature extraction and learning programs.
[0415] In one implementation, the server employs a vector-representation-based recommendation model and a sentiment recognition model for its learning process. When constructing the recommendation model, the server uses matrix factorization algorithms or vector-embedded collaborative filtering algorithms to map users and items / services into a low-dimensional continuous vector space. During model training, the server minimizes a loss function based on squared error or cross-entropy and updates the weight parameters using gradient descent or variant optimization algorithms (such as adaptive learning rate-based optimization methods). The server selects interaction records from user behavior history as positive samples in the training data and generates non-interacted candidate items as negative samples through negative sampling to improve training efficiency.
[0416] In one implementation, the server uses a deep neural network-based text sentiment classifier for sentiment state recognition. This classifier employs an embedding layer to map words to vectors, uses a bidirectional recurrent neural network or a self-attention-based neural network to encode sequences, and then outputs the sentiment category probability distribution through a fully connected layer and a softmax output layer. During training, the server uses cross-entropy as the loss function and updates the model weights through backpropagation and gradient descent. During the online inference phase, the server uses the trained model to encode the user's latest input text and a portion of historical text to obtain the user's current sentiment state, such as "anxious," "expectant," or "confused," and provides the corresponding confidence level.
[0417] In one implementation, during the candidate information generation phase, the server uses the feature extraction results to calculate the similarity between the user feature vector and the candidate information feature vector. The server uses vector dot product or cosine similarity to calculate the similarity and generates a candidate set based on the results. When selecting candidate information for goods or services, the server considers multiple features such as user preference characteristics, price preference, and historical conversion rate, balancing relevance and diversity through a multi-objective ranking strategy. When selecting candidate information for document data or input data, the server encodes task type, required fields, and legal constraints as features and selects suitable template fragments and clause fragments as candidate elements.
[0418] In one implementation, the server converts structured features and candidate information into natural language input by constructing prompt statements. The server explicitly includes the following information in the prompt statements: a summary of the user's recent behavior, a summary of user preferences, the user's emotional state, a list of candidate information and their key attributes, the current task objective, and output format constraints. In a specific example, the server generates the following prompt statement: "The user recently searched for: 'camping tent', 'camping light', and 'sleeping bag'. Their historical purchase preference is for mid-priced camping gear, and their current mood is 'anxious'. The following are candidate items: 1. Waterproof sleeping bag, price XXX; 2. Lightweight camping light, price XXX; 3. Folding chair, price XXX. Please generate 3 personalized recommendations for this user, each no more than 80 characters, in a friendly tone, and also aiming to soothe their mood." In a specific instance of an address change scenario, the server generates the following prompt: "The user wishes to change their address. The original address is... and the new address is... Please generate a draft address change application for government departments, including the reason for the change, the original address, the new address, and the declaration clauses. Use formal written language and keep it within 1,000 words." In a specific example within a loan or tax scenario, the server generates the following prompt: "Please generate a standard housing loan application based on the following information, in formal Chinese: Applicant's name: Zhang San; ID number: XXXX; Current address: Haidian District, Beijing...; Loan amount: 800,000 yuan; Loan purpose: Purchase of first home; Repayment method: Equal principal and interest payments. Please include the applicant's basic information, explanation of loan purpose, explanation of repayment ability, and declaration clauses." In a specific instance of document correction or supplementary explanation, the server generates the following prompt: "The user is currently on the payment page, exhibiting rapid clicking and a short dwell time, suggesting an anxious mood. Please generate a clear and concise instruction of no more than 40 words, informing the user that they only need to click the 'Pay Now' button to complete the transaction." When invoking the generative AI model, the server calls the model endpoint over the network. During the call, the server uses a prompt statement as input text and sets maximum output length, temperature parameters, penalty parameters, etc., to control the length, randomness, and repetition of the output text. After receiving the output from the generative AI model, the server performs format validation, sensitive field consistency checks, and necessary rule filtering. The rule filtering uses a combination of whitelists and blacklists, regular expression matching, and simple syntax checks to ensure that the generated content conforms to business constraints and is technically parsable.
[0419] In the scenario of automatically generating documents for address changes, after completing output validation, the server combines structured fields with the generated document text into a unified data representation before providing it to the terminal for display and editing. The terminal presents the generated document as text areas or editable segments in the user interface, allowing users to make partial modifications. After the user completes the modifications, the terminal sends the revised content back to the server as new input. Upon receiving the modifications, the server updates the document version and can store the modified paragraphs as new training samples in the training dataset for subsequent optimization of generative artificial intelligence models or post-processing rules.
[0420] In one implementation, the server is configured with an authentication interface module for communicating with external authentication systems. Upon receiving the user's biometric data or authentication credentials forwarded by the terminal, the server sends them to the external authentication system via a secure communication channel. After receiving the authentication result and authentication identifier returned by the external authentication system, the server associates and stores the authentication result with the corresponding document or recommended output. The server only allows specific document data to be submitted to external business systems or written to protected data storage areas upon successful authentication.
[0421] In one embodiment, the terminal displays the output information generated by the system of this invention as interface elements to the user when the user makes online payments, signs contracts, or changes their address. In payment scenarios, the terminal displays generated guidance instructions, simplifying the user's decision-making process in complex interfaces. In contract signing scenarios, the terminal highlights key terms, allowing users to focus on verifying important content and reducing human error.
[0422] In this invention, the server quantifies user behavior history and emotional state as features and incorporates them into the construction of prompt statements. This represents a significant improvement over simply using fixed templates to call generative AI models. In recommendation scenarios, the server embeds users and candidate items into a unified vector space using a matrix factorization-based recommendation model. This allows similarity calculations to be performed efficiently using linear algebra in the vector space, reducing query time complexity. When introducing emotional state as a feature, the server adds a one-dimensional or multi-dimensional emotional vector to the model and includes a regularization term for emotional consistency in the loss function. This improves the degree of emotional matching while satisfying historical preference matching, thereby reducing recommendations that are "inappropriate in tone" or "inappropriate in timing" for the user.
[0423] In document generation scenarios, the server explicitly encodes legal rules or structural constraints into prompt statements. Combined with rule-based validation in the post-processing stage, this makes the output of generative AI models easier to parse and reuse, reducing the difficulty of secondary structuring. When training sentiment classification and recommendation models, the server uses data augmentation and negative sampling techniques to improve the models' generalization ability and training efficiency, reducing the required training sample size and training time. During runtime, the server caches features and candidate sets from recent sessions, preventing repeated queries and feature extractions for consecutive requests within the same session. This reduces storage access and network interactions, thereby alleviating communication load and improving response speed.
[0424] When working with the terminal, the server transmits only concise output information relevant to the current interface rendering, while keeping large-scale features and behavioral history centrally processed on the server side, avoiding the transmission of redundant intermediate data between the terminal and the server. This architecture allows computationally intensive feature extraction and model inference to be completed on the server side, while the terminal only handles lightweight display and local input collection, thereby improving the overall system's computational resource utilization efficiency and reducing system latency under poor network conditions.
[0425] In this invention, the server does not rely on a predefined set of static business rules to determine recommended content or document text. Instead, it automatically learns the mapping relationship between features and target output from large-scale historical data during the training phase through a learning program, and uses a generative artificial intelligence model for text generation during the runtime phase. In the candidate information selection and generation phase, the server employs a nonlinear mapping and vector space approximation method, different from traditional rule systems. This enables the system to perform rapid similarity calculation and clustering in a high-dimensional feature space, thereby achieving effective screening and combination of candidate information in situations where traditional manual logic is insufficient for exhaustive searching.
[0426] In alternative implementations, the server can employ different generative artificial intelligence model structures. For example, the server can use a sequence generation model based on an encoder-decoder architecture, taking prompts as input to the encoder and autoregressively generating document text from the decoder. In another implementation, the server can use a graph-based model to model the document structure, treating paragraphs and clauses as graph nodes, and optimizing the document structure through graph neural networks to make the generated text more consistent with the expected structural constraints.
[0427] In alternative implementations, the terminal can implement some emotion recognition functions locally. For example, it can perform preliminary classification of facial expression images using a local lightweight model and only send the classification results to the server to further reduce the amount of original image data transmitted over the network, thereby reducing privacy risks and bandwidth consumption.
[0428] In another implementation, users can provide input information to the terminal via voice input. After converting the speech into text, the terminal sends the text along with an acoustic feature summary to the server. In this implementation, the server can use a joint model to perform multimodal emotion recognition on the text and acoustic features, thereby obtaining a more accurate emotion state estimate and demonstrating stronger context sensitivity in prompt construction and candidate information selection.
[0429] Through the aforementioned hardware and software configuration, data structure, and algorithm flow, the server in this invention achieves joint modeling of user input information, behavioral history, and emotional state. It then drives a generative artificial intelligence model to perform inference processing by constructing prompts containing multi-dimensional features. After obtaining the generated results, the server continuously optimizes the feature extraction and candidate information generation process through structured post-processing and a closed-loop behavioral update mechanism. Therefore, this invention not only achieves automation at the business level but also improves processing speed, result accuracy, and resource utilization efficiency from the perspectives of the computer's internal data processing structure, model structure, and operating mechanism, constituting a substantial improvement to computer technology itself.
[0430] use Figure 14 The processing flow is explained.
[0431] Step 1: The user enters initial information on the terminal. Users enter text content and structured information, such as search keywords, address information, and application type, into the terminal interface.
[0432] Input: Text content entered by the user on the terminal (such as "camping tent recommendation" or "residence change application"), form fields (name, address, etc.), and user and device identifiers automatically attached by the terminal.
[0433] The terminal performs basic validation on the input locally (e.g., checking whether required fields are empty or whether the phone number format is correct), and packages the validated data together with the identifier of the communication application or information retrieval tool into a request data structure (e.g., a JSON object).
[0434] Output: The terminal generates request data containing fields such as user input information, user identifier, application identifier, and timestamp, and sends it to the server over the network.
[0435] Step 2: The server receives and parses the request data. The server receives request data from the terminal through the network interface.
[0436] Input: Request data sent by the terminal, including user input information, user identifier, communication application or information retrieval tool identifier, timestamp, etc.
[0437] The server uses a parser to perform syntax parsing on the request data, extracting user input text, structured fields, user identifiers, and application identifiers, and performing format checks (such as checking whether the field types and lengths are valid).
[0438] The server writes the parsed data into a temporary session context structure for subsequent processing modules to access.
[0439] Output: The server generates a standardized session context data structure, which includes user input text, structured fields, user identifier, application identifier, and session identifier.
[0440] Step 3: The server retrieves past behavior information and attribute information from the data storage device. The server accesses the data storage device (such as a relational database or key-value store system) based on the user identifier and application identifier in the session context.
[0441] Input: User ID and application ID in the session context.
[0442] The server performs a query operation, reading attribute information (such as region code and historical preference tags) from the user's basic information table and past behavior information (search history, browsing history, click history, purchase history, etc.) from the user behavior history table.
[0443] The server merges and deduplicates the query results, and sorts the records according to time order or weight.
[0444] Output: The server generates a set of user attribute information and past behavior information sorted by time or weight, and appends it to the session context.
[0445] Step 4: The server completes missing information in the user input. The server analyzes the relationship between the user's current input information and the obtained attribute information and behavioral history.
[0446] Input: User input information (such as incomplete addresses or vague descriptions), user attribute information, and a collection of past behavior information.
[0447] The server uses data processing programs and rule sets or statistical models to complete missing fields. For example, the server infers the name of a city that is not currently specified based on historical address records, or looks up the administrative division based on the postal code.
[0448] The server uses algorithms such as string matching, pattern recognition, and geocoding table lookup at the data level to fill in missing information, and generates multiple candidate values when necessary and sorts them by confidence level.
[0449] Output: The server outputs a completed user input data structure, in which missing fields are filled with specific values or candidate value lists, and updates the session context.
[0450] Step 5: The server preprocesses the information and extracts features. The server transforms the completed information and past behavior data into a data format suitable for machine learning processing.
[0451] Input: Completed user input information, user attribute information, and a collection of past behavior information.
[0452] The server uses a data analysis library to build data frames, standardize numerical fields, encode categorical fields (e.g., one-hot encoding), and perform feature processing on time fields (e.g., extracting hours, days of the week, etc.).
[0453] The server converts the text input into a vector representation through word segmentation, stop word removal, and word embedding; it generates feature vectors from the behavior sequence through statistics (such as the number of clicks, the time difference of the most recent operation) and aggregation (such as statistics by category).
[0454] Output: The server generates user feature vectors, session context feature vectors, and candidate information feature vectors, providing input for subsequent preference inference and sentiment recognition.
[0455] Step 6: The server infers user preference information. The server uses the recommendation model in the learning program to infer users' long-term and short-term preferences.
[0456] Input: User feature vector, behavioral history feature vector, attribute information.
[0457] The server calls a pre-trained recommendation model (e.g., a model based on matrix factorization or embedding representation) to calculate the similarity or rating prediction between user vectors and item vectors.
[0458] The server sorts the items based on the predicted scores and extracts high-scoring items in different categories and price ranges to summarize user preference information (such as preference for camping equipment or preference for the mid-range price range).
[0459] Output: The server outputs user preference information in a structured form, including preference category weights, price range preferences, brand or feature preferences, etc., and appends it to the session context.
[0460] Step 7: Server identifies user's emotional state The server uses an emotion recognition model to infer the user's current emotional state.
[0461] Input: The user's current input text (including several historical input texts if necessary), optional acoustic feature summary and behavioral pattern features (such as frequent modifications, long dwell times).
[0462] The server embeds the text input into a word or sub-word vector space, and performs forward propagation through a sentiment classification network (such as a bidirectional recurrent network or a self-attention-based network) to obtain the probability distribution of each sentiment category.
[0463] The server outputs the emotion category with the highest probability or that meets the threshold condition as the user's emotion state (such as "anxious", "expectant", "confused") and records the corresponding confidence level.
[0464] Output: The server outputs the user's sentiment state label and its confidence level, and writes this information into the session context.
[0465] Step 8: The server generates candidate information. The server selects candidate information for goods or services or generates candidate information for document data based on user preference information and emotional state.
[0466] Input: User preference information, user sentiment state, available items in the candidate information table, and current task type (e.g., recommendation, document generation).
[0467] In product recommendation scenarios, the server filters products that match the category and price range based on preference weights, sorts them using similarity scores, and selects the top few as a candidate product set.
[0468] In document generation scenarios, the server selects structural templates and clause fragments from the document template library based on task type and regulatory constraints, and combines them into a set of candidate document fragments.
[0469] Output: The server outputs a structured set of candidate information, including the identifier, main attributes and preliminary score of each candidate, as the basis for constructing the prompt statement.
[0470] Step 9: Server construct prompt statement The server organizes key information from the session context into prompts in natural language.
[0471] Input: The user's most recently entered text, a summary of the user's preference information, the user's sentiment state, a set of candidate information, the current task objective, and the output requirements.
[0472] The server, following a predefined prompt template, concatenates the above information into well-organized natural language text. For example: "The user recently searched for: 'camping tent', 'camping light', and 'sleeping bag'. Their historical purchase preference is for mid-priced camping gear, and their current mood is 'anxious'. The following are candidate items: 1. Waterproof sleeping bag, price XXX; 2. Lightweight camping light, price XXX; 3. Folding chair, price XXX. Please generate 3 personalized recommendations for this user, each no more than 80 characters, in a friendly tone, and also aiming to soothe their mood." The server generates a notification message in the event of an address change, for example: "The user wishes to change their address. The original address is... and the new address is... Please generate a draft address change application for government departments, including the reason for the change, the original address, the new address, and the declaration clauses. Use formal written language and keep it within 1,000 words." Output: The server generates prompt text for the input of the generative artificial intelligence model and records it in the session context.
[0473] Step 10: The server invokes a generative artificial intelligence model to perform inference processing. The server sends the prompt statement to the inference service corresponding to the generative artificial intelligence model via a network interface.
[0474] Input: Prompt text, generation parameters (maximum output length, temperature value, repetition penalty parameter, etc.).
[0475] The server sends a request to the interface of the generative AI model, with the request body containing prompts and parameters. Upon receiving the request, the generative AI model performs forward propagation of its neural network, encodes the prompts, and gradually generates a sequence of output text.
[0476] While waiting for a response, the server monitors timeout times and error status, and executes retry or degradation strategies for abnormal situations.
[0477] Output: The server receives the output text returned by the generative artificial intelligence model, which includes recommendations, document drafts, or guidance statements.
[0478] Step 11: The server verifies and processes the generated results. The server performs technical verification and formatting on the text output by the generative artificial intelligence model.
[0479] Input: Raw text output by the generative artificial intelligence model, raw structured data (user information, candidate information, regulatory constraints).
[0480] The server uses regular expressions and string matching to check whether key fields (such as name, amount, address) match the input data, preventing the model from modifying key data without authorization.
[0481] The server applies rules to filter based on task type, deleting paragraphs that do not meet business tone or format requirements, and rearranging or truncating certain paragraphs to meet length or structural limitations.
[0482] The server combines the validated text with the original structured fields into a unified data object, such as a recommendation list or document object, and adds tags to each field to indicate which parts are available for user editing.
[0483] Output: The server generates a validated and formatted output information object, ready to be sent to the terminal.
[0484] Step 12: The server sends output information to the terminal. The server sends the output information to the terminal through the network interface.
[0485] Input: Formatted output information objects (including a list of recommended results, document text, guidance text, etc.).
[0486] The server encapsulates the output information into a response data structure, adds a status code, error message (if none), and session identifier, and returns it to the terminal via a network transmission protocol.
[0487] Output: The server outputs response messages to the terminal, which the terminal can use to render the interface and present interactive content.
[0488] Step 13: The terminal displays output information and collects user operation information. The terminal receives the output information returned by the server and renders it on the interface.
[0489] Input: Recommendation list, document text, instructions, and other output information returned by the server.
[0490] In recommendation scenarios, the terminal displays goods or services in the form of lists or cards, along with recommendation descriptions generated by generative artificial intelligence models; in document scenarios, it displays application forms or document text in an editable area.
[0491] Users can perform actions such as clicking, selecting, editing, confirming, and rejecting on the terminal. The terminal translates these actions into user operation information, such as "Click on product A", "Modify paragraph 3 of the document", and "Confirm submission".
[0492] Output: The terminal generates structured user operation information and sends it back to the server at appropriate times (e.g., when the operation is completed or at a set time).
[0493] Step 14: The server records user operation information and updates the behavior history. The server receives user operation information from the terminal and updates the behavior history.
[0494] Input: User operation information sent by the terminal (operation type, target object identifier, timestamp, session identifier, etc.).
[0495] The server writes user actions into a user behavior history table, and merges or aggregates multiple operations within a short period of time when necessary to reduce storage and computing overhead.
[0496] The server updates the user feature vector and session feature vector based on new behavior records, such as increasing the interest weight of a certain type of product, or recording the user's preference for a certain document structure.
[0497] Output: The server generates updated behavioral history data and feature representations to provide more accurate input for subsequent requests.
[0498] Step 15: The server uses the updated behavioral history and characteristics to improve subsequent processing. When the server receives a new request, it prioritizes processing the request using the updated behavioral history and characteristics.
[0499] Input: Updated behavior history, user feature vector, session feature vector.
[0500] The server refreshes its internal cache with the latest data in the online portions of the recommendation and sentiment recognition models, and performs incremental training or parameter fine-tuning on the models in batch tasks when necessary.
[0501] When constructing the next round of prompts, the server incorporates the latest behavioral information and user feedback signals, enabling the prompts to more accurately reflect the current user status and actual preferences.
[0502] Output: The server outputs candidate information and prompts based on more refined features and the latest behavior, thereby obtaining higher quality generation results when the generative artificial intelligence model is called in the future, improving the overall system's recommendation accuracy, document generation quality and response speed.
[0503] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0504] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0505] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0506] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0507] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0508] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0509] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0510] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0511] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0512] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0513] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0514] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0515] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0516] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0517] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0518] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0519] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0520] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0521] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0522] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0523] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0524] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0525] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0526] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0527] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0528] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0529] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0530] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0531] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0532] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0533] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0534] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0535] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0536] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0537] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0538] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0539] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0540] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0541] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0542] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0543] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0544] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0545] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0546] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0547] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0548] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0549] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0550] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0551] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0552] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0553] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0554] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0555] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge-Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0556] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0557] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0558] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0559] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0560] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0561] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0562] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0563] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0564] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0565] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0566] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0567] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0568] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0569] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0570] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0571] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0572] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0573] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0574] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0575] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0576] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0577] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0578] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0579] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0580] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0581] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0582] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0583] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0584] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0585] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0586] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0587] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0588] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0589] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0590] In addition, the following notes are provided in response to the above explanation.
[0591] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for receiving user input information sent by a user terminal and identification information of an information processing program having communication or information retrieval functions from a user terminal via a communication channel in an information processing apparatus. An apparatus for applying natural language processing algorithms to the user input information to extract intent information and parameter information, and for automatically completing missing items of some parameter information based on the identification information, user attribute information, past dialogue history information and additional information obtained from external information sources. A device for constructing a prompt statement that generates a response to a generative artificial intelligence model based on the completed intent information and parameter information, and including dialogue history information, business rule information, authentication status information and display constraints in the prompt statement; An apparatus for inputting the prompt statement into the generative artificial intelligence model to generate natural language response information, and performing content filtering, format shaping and structured data processing on the response information to generate response data that can be sent to the user terminal; An apparatus for storing the response data in association with the user input information in a storage device, and for updating the prompt statement using the stored information during subsequent multi-turn dialogue processing.
[0592] (Note 2) The information processing system according to Appendix 1 is characterized in that, When the intent information extracted from the user input information is determined to be related to procedures for changing residence information, the information processing device automatically fills in the required items to generate electronic document data for procedures based on the completed parameter information and the style definition information stored in the procedure style information storage unit. It then prompts the user terminal with confirmation response data containing the summary information of the electronic document data and performs confirmation processing on the electronic document data based on the user's confirmation information.
[0593] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device will compare the biometric information and additional authentication information obtained from the user terminal with the external authentication processing device, and reflect the authentication credibility information obtained therefrom as the authentication status information in the prompt statement, so that the response information generated by the generative artificial intelligence model includes authentication linkage guidance information of self-confirmation result and content confirmation matters.
[0594] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring user input information via a communication network through a user terminal and storing the user input information as data; An apparatus for generating a prompt statement containing the user input information based on the user input information, and processing the length and character type of the prompt statement according to predetermined conditions, thereby shaping the prompt statement into an input format suitable for input into a generative artificial intelligence model; An apparatus for sending a response generation request to the generative artificial intelligence model using the shaped prompt statement, and for converting the generated content obtained from the generative artificial intelligence model into distribution data for the user terminal and sending it. An apparatus for obtaining evaluation information on the generated content from the user terminal, and updating the generation rules of the prompt statement or the running parameters of the generative artificial intelligence model based on the evaluation information, thereby changing the generation characteristics of subsequent generated content.
[0595] (Note 2) The information processing system according to Appendix 1 is characterized in that, The information processing device further includes a template management function for selecting a fixed portion of a prompt statement based on the query information obtained as user input information, and combining the fixed portion with the user input information to form the prompt statement.
[0596] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device further includes: an adaptive control function for calculating evaluation indicators based on the statistical results of the evaluation information according to the theme of the generated content, and automatically changing the constituent elements of the prompt statement or the output length control value of the generative artificial intelligence model according to the evaluation indicators.
[0597] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring user identification information and user request information in an information processing apparatus; A device for retrieving individual attribute information and behavioral history information stored in an information storage device based on the user identification information, thereby obtaining related information about the user; A device for performing statistical or classification processing on the associated information and the user request information to extract user interest information and user status information; An apparatus for selecting template information based on the user interest information and the user status information, and for generating prompt statements for input to a generative artificial intelligence model by filling the variable portion of the template information with the associated information and the user request information; A device for inputting the prompt statement into the generative artificial intelligence model and obtaining response information from the generative artificial intelligence model; A device for performing display format conversion processing or security determination processing on the response information, and sending the processed response information to a terminal device.
[0598] (Note 2) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to: when the user request information is related to address change, automatically generate the document information required for address change procedures based on the user attribute information contained in the associated information, and include conditions for the generation of the document information in the prompt statement for the generative artificial intelligence model.
[0599] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to: use user identification information containing biometric information to link with the authentication infrastructure to perform identity verification processing, and include the obtained authentication result as the associated information in the prompt statement, thereby enabling the generative artificial intelligence model to generate response information corresponding to the authentication status.
[0600] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: An apparatus for obtaining past behavior information and attribute information of a user from a data storage device based on user input information from a user terminal and the identifier of a communication application or information retrieval tool, and for completing any missing information in the user input information; An apparatus for preprocessing and extracting features from the completed information and the user input information using a data processing program and a learning program, thereby determining user preference information and user emotional state; A means for selecting candidate information for goods or services and / or generating candidate information for document data or input data based on the user preference information and the user emotional state; A device for generating prompt statements containing the candidate information, user preference information, and / or user emotional state, using the prompt statements as prompt input to a generative artificial intelligence model, and instructing the generative artificial intelligence model to perform inference processing; A means for determining recommendation information for the goods or services and / or the document data or input data based on the output results obtained from the generative artificial intelligence model, and generating them as output information for the terminal; A device for obtaining user operation information for the output information from the terminal, storing the operation information as a behavior history in the data storage device, and using the behavior history to update the feature extraction and the selection and / or generation of the candidate information.
[0601] (Note 2) According to the information processing system described in Appendix 1, the candidate information is procedural document data related to address change generated based on address information contained in user input information and the completed information, and the system is configured to automatically generate the procedural document data related to address change based on the output result generated by the inference process, and output it to the terminal as display data.
[0602] (Note 3) According to the information processing system described in Appendix 1, the system is configured to obtain user biometric information and / or authentication information from the terminal based on the user operation information and the behavioral information contained in the user input information, perform personal authentication in conjunction with an external authentication system using an authentication communication method, and determine the output information based on the personal authentication result, and prompt the terminal with the authentication result and the content confirmation result of the output information.
Claims
1. An information processing system, characterized in that, include: processor; The processor is configured to: provide an interface for receiving user input information; parse communication application identifiers or information search tool identifiers to complete the required information; and generate and input prompt information to instruct the generative artificial intelligence model to generate a response.
2. The information processing system according to claim 1, characterized in that, The processor is configured to automatically generate the documents required for address change procedures based on the user input information related to address change.
3. The information processing system according to claim 1, characterized in that, The processor is configured to use the user's biometric information to link with a public authentication system in order to improve the reliability of authentication.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A