Information processing system

CN122797477APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610242462.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-02-28
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

通过上述结构与处理流程,本发明的系统能够在无需用户额外切换应用或进行复杂输入操作的情况下,实现词典数据库释义与生成式人工智能扩展应答的结合,同时通过字符数限制机制保证展示结果简洁明了,从而有效解决现有技术中存在的操作繁琐、信息展示冗长以及难以兼顾基础释义与上下文扩展解释的技术问题

Benefits of technology

服务器在模型训练阶段利用大规模文本语料,采用数据扩展手段,如随机掩码、句子顺序打乱、同义替换等,以增强模型对多样输入的鲁棒性。服务器采用小批量训练方式,在每个批次中计算预测输出与真实下一个词之间的交叉熵损失,并通过误差反向传播更新网络权重。服务器可以利用梯度裁剪、学习率退火等方法提高训练稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797477A_ABST
    Figure CN122797477A_ABST
Patent Text Reader

Abstract

The application provides an information processing system. An information processing system, characterized by comprising: a processor; wherein the processor is configured to: query a dictionary database with a sentence selected by a user on a terminal to obtain meaning information of the sentence; send the obtained meaning information to the terminal of the user and display it on a display screen; generate prompt text based on the selected sentence, the prompt text being used to instruct a generative artificial intelligence model to generate a specific response; send the generated prompt text to the generative artificial intelligence model, and receive a response from the generative artificial intelligence model; apply a character number limit to the received response for adjustment, and send the adjusted response to the terminal of the user for display on the display screen.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot speech in response to the user's speech.

[0003] In existing messaging applications or text reading environments, when users want to understand the meaning of a selected statement, they typically need to manually copy the statement, switch to a separate dictionary application or browser, and actively input or paste the statement for retrieval. This method involves multiple steps and fragmented processes, reducing the continuity and efficiency of the user's conversation or reading process. Furthermore, traditional dictionary searches only provide fixed dictionary definitions and struggle to generate more targeted explanations or extended information based on the user's current context. While the development of generative artificial intelligence models can provide users with more flexible and richer text responses, directly displaying the complete response output by these models often results in lengthy content and a lack of emphasis on key information. This is not conducive to presentation within the limited display area of ​​a terminal and hinders users from quickly obtaining core information in conversational scenarios. Therefore, it is necessary to provide an information processing system that, without requiring the user to leave the current terminal interface, provides basic semantic information for the selected statement based on a dictionary database and automatically generates specific responses related to the selected statement using generative artificial intelligence models. The system should also impose character limits on the generated responses to ensure information richness while avoiding lengthy displays, thereby improving user convenience and information acquisition efficiency in actual use. Summary of the Invention

[0004] To address the aforementioned problems, this invention provides an information processing system comprising a processor configured to: query a dictionary database for the meaning information of a statement selected by a user on a terminal, and send the obtained meaning information to the user's terminal for display on a screen, thereby allowing the user to obtain a basic definition within the current application interface; the processor is further configured to: generate prompt text based on the selected statement to instruct a generative artificial intelligence model to generate a specific response, the prompt text preferably including the selected statement itself and instruction information related to the generation target; the processor sends the generated prompt text to the generative artificial intelligence model and receives a response from the generative artificial intelligence model. To avoid excessively long responses, the processor is further configured to: adjust the response received from the generative artificial intelligence model according to a pre-set character limit, controlling the response within a target character range through truncation or other compression methods, and send the adjusted response to the user's terminal for display on a screen. Through the above structure and processing flow, the system of the present invention can combine dictionary database definitions with generative artificial intelligence extended responses without requiring users to switch applications or perform complex input operations. At the same time, the character number limitation mechanism ensures that the displayed results are concise and clear, thereby effectively solving the technical problems of cumbersome operation, lengthy information display, and difficulty in taking into account basic definitions and contextual extended explanations in the prior art.

[0005] "System" refers to an overall device or collection of devices consisting of hardware and / or software, used to perform the various functional processes described in this invention, including at least one processor and connected storage devices, communication interfaces and user terminals, etc.

[0006] A processor is a computing unit that can execute program instructions to perform data processing and control flow. It can be a single CPU, multiple CPUs, GPU, ASIC, FPGA or any combination thereof, or a processing circuit integrated in a server, cloud platform or other electronic device.

[0007] "User" refers to a natural person who uses a terminal device to select statements, view meaning information, and generate AI responses on the interface, without being limited to a specific user identity or permission level.

[0008] "Terminal" refers to an electronic device that is operated by a user and used to display information, including but not limited to smartphones, tablets, personal computers, wearable devices, or other devices with display interfaces and network communication capabilities.

[0009] A “statement” is a string of one or more characters, words or phrases that can be selected as a whole in text. It can be a word, phrase, sentence or part thereof, and is not limited to a specific type of natural language.

[0010] "Selected statement" refers to the target statement that the user selects from the text content on the terminal's display screen through operations such as clicking, long-pressing, and dragging. This statement serves as the object of query and response generation.

[0011] A "dictionary database" is a collection of data used to store the meaning information, usage instructions and related metadata of words, phrases or sentences. It can be implemented in the form of a local database, a server database or a cloud database.

[0012] "Semantic information" refers to the semantic explanations, conceptual descriptions, usage examples, or other structured or unstructured textual information corresponding to the selected statement that helps users understand the meaning of the statement.

[0013] "Display screen" refers to the user interface area presented on the display screen of a terminal device, used to display chat content, query results, meaningful information, and responses from generative artificial intelligence models.

[0014] "Generative artificial intelligence models" refer to models that process input prompt text and automatically generate corresponding natural language responses or content based on machine learning and deep learning technologies, including but not limited to large-scale language models and dialogue generation models.

[0015] "Prompt text" refers to the text content generated by the processor and sent to the generative artificial intelligence model to instruct the generative artificial intelligence model to perform a specific response generation task based on the selected statement. It usually includes the selected statement itself and the instruction statements related to the task.

[0016] "Specific response" refers to the answer content or explanatory text generated by the generative artificial intelligence model based on the prompt text and associated with the selected statement. Its content may include explanations, extended explanations, examples or other forms of textual information.

[0017] "Response" refers to the text output of a generative artificial intelligence model after receiving a prompt text. It can be a single sentence or multiple sentences of natural language description, used to respond to the generation task indicated by the prompt text.

[0018] "Character limit" refers to the upper limit constraint on the number of characters in the response output by the generative artificial intelligence model. It is used to control the length of the response so that it does not exceed the preset maximum number of characters.

[0019] "Adjusting according to the character limit" refers to the process by which the processor detects and processes the length of the received response text, including truncation, deletion or other compression operations when the response length exceeds the preset character limit, in order to generate an adjusted response that meets the character limit. Attached Figure Description

[0020] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0021] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0022] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0023] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0024] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0025] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.

[0026] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0027] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0028] Figure 9 This represents an emotion map that maps multiple emotions.

[0029] Figure 10 This represents an emotion map that maps multiple emotions.

[0030] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0031] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0032] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0033] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0034] Hereinafter, an example of an implementation of the system to which the technology of this disclosure relates will be described with reference to the accompanying drawings.

[0035] First, let me explain the terminology used in the following instructions.

[0036] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose Computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0037] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0038] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0039] In the following embodiments, the communication I / F (interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0040] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0041] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0042] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0043] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0044] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0045] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0046] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge-Coupled Device) image sensor.

[0047] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0048] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0049] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0050] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0051] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0052] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0053] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0054] In existing information processing technologies, when users encounter unfamiliar words or phrases in dialogue interfaces or other application interfaces on a terminal, they typically need to manually switch to a separate search application or online search service to input the word or phrase and obtain its meaning. This method requires multiple interface switching and repetitive input operations, which not only reduces the efficiency of human-computer interaction but also easily interrupts the user's workflow in the current application. Furthermore, with the widespread application of generative artificial intelligence models in the field of natural language processing, although these models can generate explanations for words or phrases, current technologies often simply forward the user's raw input to the generative artificial intelligence model without structured control over parameters such as word count and explanation style. This results in lengthy and unrelevant generated results, making it difficult to provide users with concise and context-sensitive explanations in a timely manner.

[0055] Furthermore, existing systems often neglect the systematic recording and analysis of user interaction history and model response history data. They are unable to automatically optimize prompt templates and word limit parameters for generative AI models based on this historical data, resulting in the system's inability to adaptively improve response quality and human-computer interaction experience over long-term operation. Simultaneously, many implementations are limited to isolated interpretations of user-selected words, failing to fully utilize the contextual information of those words in the dialogue, thus making it difficult to obtain interpretations that match the specific context.

[0056] Therefore, how to integrate dictionary retrieval and generative artificial intelligence model invocation on the server side with an improved program processing flow, without increasing the local computing burden on the terminal, automatically generate prompt statements with structured control conditions, dynamically adjust prompt strategies and character limits based on user settings and historical usage data, and support the generation of context-dependent interpretation results based on dialogue context, thereby improving the overall processing efficiency, response quality, and user experience of computer systems in semantic query and result presentation, has become a technical issue that urgently needs to be solved in this field.

[0057] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0058] In this invention, the server includes a processing unit for receiving a character sequence selected by a user on the display screen of an information processing device from a terminal, retrieving semantic information corresponding to the character sequence from a data storage device storing vocabulary information, and sending the semantic information to the terminal for display on the screen; a processing unit for generating a prompt statement for a generative artificial intelligence model based on the selected character sequence received from the terminal, wherein the prompt statement includes not only the selected character sequence but also conditional information constraining the word count limit and description style of the response, and sending the prompt statement to the generative artificial intelligence model; and a server for receiving response information generated based on the prompt statement from the generative artificial intelligence model, determining and trimming the response information according to a predetermined word count limit to generate adjusted response information, and sending the adjusted response information to the terminal for display on the screen. The display screen shows a processing unit associated with the selected character sequence; it includes a parameter management unit for storing various template information for generating prompt statements and user settings related to answer character limits and description styles, and automatically selecting template information based on user settings to dynamically change the prompt statement content and answer character limit parameters on the server side; a historical analysis unit for recording historical information such as the selected character sequence, corresponding prompt statements, response information of the generative artificial intelligence model, and the application's answer character limit in chronological order, and updating and optimizing template information and answer character limits based on this historical information; and a context processing unit for receiving data from the terminal containing historical dialogue information surrounding the selected character sequence, generating extended prompt statements containing contextual content, and sending them to the generative artificial intelligence model to obtain response information matching the specific context. This allows for unified control and optimization of the semantic query process on the server side, achieving end-to-end automated workflow from user selection to semantic presentation. It reduces the computational and control burden on the terminal side. Through structured prompt generation, word limit control, dynamic adjustment of templates and parameters, and adaptive optimization based on historical data and contextual information, it improves the compactness and relevance of generative artificial intelligence model responses, thereby enhancing the overall performance and user experience of the computer system in natural language semantic query and display processing.

[0059] A "system" refers to a combination of devices consisting of one or more servers, terminals, and corresponding software programs, used to perform functions such as information reception, processing, storage, and output.

[0060] "Terminal" refers to an electronic device that is operated by a user and has the functions of displaying screens and network communication, including but not limited to information processing devices such as smartphones, tablets, and personal computers.

[0061] "Information processing device" refers to an electronic device capable of processing, storing and outputting input data, including its hardware such as processor, memory and input / output interfaces, and the software system running on it.

[0062] "Display screen" refers to the display area presented by a display device, used to visually output text, images and interface elements to the user, including dialog interfaces, application interfaces, etc.

[0063] A "character sequence" refers to a text data unit composed of one or more characters arranged in a certain order, which can be a word, phrase, sentence, or other text fragment.

[0064] "Vocabulary information" refers to dictionary data related to a character sequence, including semantic information such as the meaning, part of speech, and usage examples of the character sequence.

[0065] "Data storage device" refers to a storage component used to store vocabulary information, template information, user settings information and historical information, which can be a database system, file storage system or other non-volatile storage medium.

[0066] "Semantic information" refers to structured or unstructured data obtained by retrieving lexical information or processing character sequences, which is used to express the meaning or interpretation of the character sequence.

[0067] "Generative AI models" refer to AI models that model input text or other data and automatically generate natural language responses or other content based on machine learning and deep learning algorithms.

[0068] "Prompt statements" are text instructions sent to generative artificial intelligence models to describe the generation task, constrain the content and style of the response, and may include selected character sequences and restrictions.

[0069] "Conditional information" refers to the control parameters used in the prompt statements to constrain the output of generative artificial intelligence models, including information such as word limits for responses, description styles, and tone requirements.

[0070] "Response information" refers to the output content generated and returned by the generative artificial intelligence model based on the prompt statement, which is used to explain or supplement the selected character sequence.

[0071] "Answer character limit" refers to the maximum number of characters that can be output in the response information, or its corresponding length parameter, used to control the length of the generated content.

[0072] "Template information" refers to pre-defined text or structured rules used to generate prompt statements, including placeholders and fixed expressions, to unify the format and style of prompt statements.

[0073] "User settings information" refers to the preference parameters specified by the user regarding the number of characters in the answer, the style of the description, the language style, etc., which are used to personalize the generation of prompts and the adjustment of response information.

[0074] "Historical information" refers to data related to character sequences, prompts, response information, and word limits recorded in chronological order during system operation, which is used for subsequent analysis and optimization.

[0075] The "dialogue display area" refers to the part of the screen that displays the interactive content between the user and the system, usually in the form of dialog bubbles, message lists, etc.

[0076] A "display unit" refers to an independent interface element in the display screen used to carry and display a semantic message or response message, such as a single speech bubble or a single list item.

[0077] "Dialogue history information" refers to the preceding and following message content that has been presented in the dialogue display area around the position of the selected character sequence, including user messages and system messages.

[0078] "Extended prompt statements" refer to text that includes relevant dialogue history information in addition to the selected character sequence, and are used to provide context to generative artificial intelligence models to generate context-relevant response information.

[0079] A "processing unit" refers to a functional module consisting of a processor and the program logic running on it, used to perform specific processing operations such as receiving, generating, adjusting, or outputting data.

[0080] The "Parameter Management Unit" refers to a functional module used to store and manage template information and user settings information, and to select or update parameters for prompt statements and answer word limits based on this information.

[0081] The "historical analysis unit" refers to a functional module used to statistically analyze recorded historical information and adjust parameters such as template information or word limits for answers accordingly to optimize the generated results.

[0082] The “context processing unit” refers to a functional module that receives dialogue history information, combines the information with the selected character sequence to generate extended prompt statements, and obtains context-dependent response information based on the output of a generative artificial intelligence model.

[0083] In this embodiment of the invention, the server, terminal, and user each assume different functional roles, working collaboratively to achieve semantic interpretation of the character sequence selected by the user in the dialogue interface, as well as the generation and presentation of responses based on a generative artificial intelligence model. The specific implementation of this invention will be described below in conjunction with its hardware structure, software modules, data structure, and algorithm processing.

[0084] I. Overall System Structure In one embodiment, the server is deployed on a general-purpose processor-based computer device, which may be a computer system with a multi-core CPU (e.g., an x86_64 architecture general-purpose processor), main memory, and non-volatile memory. The server runs on an operating system (e.g., Linux) and can deploy backend services via container management software (e.g., container runtime environments and orchestration systems). The server can use web server software (e.g., reverse proxy servers) and application server software (e.g., Python-based web frameworks or JavaScript-based web frameworks) to construct a network service environment.

[0085] In one embodiment, the terminal is a mobile terminal with an instant messaging application installed, such as a smartphone or tablet running a mobile operating system. The terminal includes a touchscreen display, a processor, memory, a wireless communication module (such as a Wi-Fi or cellular communication module), and a graphical user interface component. The terminal runs a chat application on the mobile operating system, which uses system-provided text display components (such as a text view component) and text selection interfaces to detect user selections in conversation messages.

[0086] In this invention, users operate through a terminal, including selecting character sequences on the display screen, issuing query commands, and viewing semantic information and response information returned by the server.

[0087] II. Server-side functional modules and data structures In one specific implementation, a server includes multiple functional units, which can be implemented by program modules executed by one or more processors and stored in a non-transitory computer-readable storage medium.

[0088] 1. Server retrieval processing unit After receiving the selected character sequence from the terminal, the server accesses the vocabulary information data structure in the data storage device. In one embodiment, the server uses a relational database management system (e.g., a general-purpose database system such as a relational database) to store the vocabulary information, which can be organized in the following form: - Glossary: ​​Fields include glossary ID, character sequence text, part-of-speech tag, basic meaning description, extended definition, usage examples, etc.

[0089] - Index structure: The server uses B+ tree indexes or full-text search indexes to accelerate the retrieval of term records based on character sequences.

[0090] During retrieval, the server uses the received character sequence as a key for exact or fuzzy matching, and constructs a semantic information object in memory containing the meaning and brief description of the terms. After encoding this semantic information, the server sends it to the terminal for presentation as an independent display unit in the dialog display area.

[0091] 2. Server prompt statement generation unit In another implementation, the server generates a prompt statement based on template information stored by the terminal or itself. The server maintains a template information table in a data storage device, which may contain: - Template ID; - Template text content (including placeholders {character sequence}, {word count}, {style}, etc.); - Applicable scenario tags (such as "brief explanation", "for beginners", "technical description").

[0092] After receiving the selected character sequence and user settings (such as default character limits and description style), the server reads the template record that matches the user settings from the template information table. The server performs a string replacement operation in memory, replacing placeholders with actual values ​​to generate the final prompt message. For example, the server can generate the following prompt message text: "Please explain the meaning of 'generative artificial intelligence model' in no more than 100 Chinese characters. It should be concise and to the point." or: "Please briefly explain the meaning and role of 'prompt statements' in interacting with generative artificial intelligence models, within 80 characters." By using this structured template generation method, the server explicitly encodes parameters such as word count limits, descriptive style, and tone in the prompt statements. This allows the output range of the generative artificial intelligence model to be precisely controlled at the request stage, thereby creating predictable constraints on the model's inference results within the computer, rather than simply forwarding natural language problems.

[0093] 3. Generative AI Invocation Unit of the Server In one implementation, the server uses a generative artificial intelligence model interface to generate natural language responses. The server can call external model services through a dedicated SDK or HTTP client library. The generative artificial intelligence model can be a deep neural network model based on the Transformer architecture. This model typically includes a multi-layer self-attention encoder and decoder structure, and the internal parameter dimensions (such as the hidden layer dimension, the number of attention heads, and the number of layers) are fixed during the training phase.

[0094] The server sends the prompt statement as a vectorized sequence of input to the model interface during the call. The server specifies the following parameters in the request: - max_tokens: The maximum number of tokens corresponding to the character limit of the answer; - Temperature: A real-valued parameter that controls the variety of output parameters; -top_p: A parameter that controls the quality of the sampling probability.

[0095] By explicitly setting these parameters, the server enables the model to perform restricted autoregressive sampling during the decoding phase. The probability distribution of candidate words at each step is truncated with top-p and temperature-scaled, and then the next token is determined using a greedy or random sampling strategy. This decoding process relies on the model's encoding state of the prefix sequence at each step, thus creating an efficient iterative inference process between the server and the model.

[0096] The server internally sets timeout periods and retry policies for model calls to ensure the overall system response stability even when the network is unstable.

[0097] 4. Server response adjustment unit After receiving the response information from the generative artificial intelligence model, the server performs structured processing on the response text in memory. First, the server calculates the number of characters or length information in the response text, which can be obtained through UTF-8 encoding length statistics or Unicode character counting functions. The server compares this length with the character limit specified in the prompt statement. If the limit is exceeded, the server prunes the response text at character or word boundaries to ensure that the text returned to the terminal meets the length convention.

[0098] In one implementation, the server can also perform rule-based filtering operations, such as removing redundant opening phrases (e.g., "Okay, let me explain:") and extra blank lines, to further improve the conciseness of the response. This unified processing on the server side avoids repeatedly implementing complex text post-processing logic on the terminal side, thereby reducing the computational burden on the terminal and improving the consistency of the overall system.

[0099] 5. Server parameter management unit and historical analysis unit The server records template information, user settings information, and historical information in the data storage device. Historical information may include: - The selected character sequence text; - The content of the prompt statement used; - The response text output by the generative artificial intelligence model; - The application has a word limit for responses; - Signals of user behavior, such as whether to perform secondary queries on the results.

[0100] In one implementation, the server periodically performs statistical analysis algorithms on this historical information. The server can calculate metrics such as the average response length, truncation rate, and number of repeated queries by users for different templates in actual use. Based on the statistical results, the server automatically adjusts the word limit parameters and wording in the template text. For example, it lowers the default word limit for templates that frequently produce excessively long responses, or adds constraints such as "output only one or two sentences" to the templates.

[0101] Through this historical data-driven parameter tuning mechanism, this invention achieves adaptive control of the generative artificial intelligence model on the server side, improving the overall response quality and efficiency under fixed computing resource conditions. This technique differs from simple human rule setting; instead, it utilizes accumulated data to iteratively optimize internal system parameters, representing an improvement in the management and performance tuning of computer system parameters.

[0102] 6. Server Context Processing Unit In another implementation, the server receives dialogue history information sent by the terminal. This information can be a text set containing several messages before and after the message containing the selected character sequence. The server concatenates these dialogue messages chronologically to construct a context string, and embeds it along with the selected character sequence into an extended prompt statement. For example, the server can generate the following prompt statement: In the following dialogue, what does 'model' refer to? Please explain in no more than 80 words. Dialogue content: ... (dialogue text) ... By explicitly incorporating dialogue history information into the prompts, the server enables generative AI models to acquire richer contextual vectors during the encoding phase, thus providing context-specific interpretations during the inference phase. This context-enhanced prompting process is implemented internally by the computer through specific operations such as string concatenation, length truncation, and preprocessing before encoding, and is uniformly controlled on the server side. This facilitates the stable provision of context-sensitive interpretation services in multi-user environments.

[0103] III. Terminal-side functions and graphical interface control In one implementation, the terminal runs a chat application that uses the operating system's text component to display conversation messages. When the user long-presses or drags to select text, the terminal calls the system-provided text selection callback interface to obtain the start and end indices of the selected area, and uses these indices to extract the selected character sequence from the message text.

[0104] After capturing the character sequence, the terminal can display an operation menu on the local user interface, allowing the user to select "Query Meaning" or a similar command. Upon receiving the user's command, the terminal packages the selected character sequence along with metadata such as the current session ID and message ID into structured data and sends it to the server via the network communication module using an encrypted communication protocol.

[0105] After receiving the semantic and response information from the server, the terminal creates a new independent display unit (such as a message bubble) in the dialogue display area, and presents the semantic information and the response information from the generative artificial intelligence model separately. For example, the terminal can display a brief definition in the dictionary definition bubble and display explanatory text generated by the generative artificial intelligence model in the AI ​​definition bubble. The terminal can also display a "Searching" loading indicator during server processing, thereby providing feedback to the user on the current system status.

[0106] IV. User Operation Method In a typical use case, a user encounters the unfamiliar term "generative artificial intelligence model" in a chat interface. The user long-presses the term on the device and selects "Query Meaning" from the pop-up menu. The device then extracts the term as a character sequence and sends it to the server.

[0107] The server first performs a dictionary search to obtain the basic meaning of the word and returns it to the terminal. Then, the server generates the following prompt statement based on the template: "Please explain the meaning of 'generative artificial intelligence model' in no more than 100 Chinese characters. It should be concise and to the point." The server sends the prompt to a generative artificial intelligence model. The model encodes the prompt based on its internal Transformer network structure and gradually generates a short explanation during the decoding phase, for example: "Generative artificial intelligence models are a type of algorithm system that uses deep learning to learn patterns from large amounts of data and automatically generate content such as text and images." The server performs a length check and trims the response (if necessary) and returns it to the terminal. The terminal inserts a new AI-generated explanatory bubble in the chat display area to present the explanatory text, allowing the user to understand the terminology without leaving the current chat application.

[0108] V. Structure and Training Instructions of Generative Artificial Intelligence Models In this invention, the generative artificial intelligence model invoked by the server is typically an autoregressive language model based on a multi-layer Transformer. This model uses subwords or characters as basic units, maps discrete tokens to vector representations using embedding matrices, models long-distance dependencies through multi-head self-attention layers, and performs feature transformation through a feedforward network. Model parameters are learned during pre-training by minimizing the cross-entropy loss function. During training, a gradient descent-based optimization algorithm (such as Adam or its variants) is employed, and unsupervised learning is performed using a large-scale text corpus.

[0109] During training, the model predicts the conditional probability of the next token in the sequence using an autoregressive approach, with the loss function being the negative log-likelihood of the actual tokens. Model parameters are updated via backpropagation. The server keeps the parameters fixed during inference, controlling the decoding strategy (such as beamsearch or top-p sampling) only during the decoding phase based on prompts. The server explicitly controls `max_tokens` using constraints in the prompts, indirectly constraining the number of decoding steps, thus achieving a balance between computational resource consumption and output quality. This approach significantly reduces average inference time and communication data volume compared to unconstrained generation.

[0110] VI. Technical Effects and Causal Relationships By internally maintaining template information, user settings, and historical information, and utilizing this information during prompt generation and response adjustment, the server improves parameter management and text control capabilities during the invocation of generative artificial intelligence models, thereby achieving the following technical effects: 1. By using structured coding of character count limits and style constraint parameters, the server imposes constraints on the output space before the model is called, reducing irrelevant and verbose content, improving the compactness of response information, reducing the amount of data transmitted, and achieving communication load reduction.

[0111] 2. By statistically optimizing the template and word count limit parameters through historical analysis, the server can automatically adjust the prompt strategy and length control, thereby improving the average response speed and response availability under the same hardware conditions. This is a performance improvement brought about by computer system parameter tuning, rather than a simple replacement of human labor.

[0112] 3. By introducing dialogue history information through the context processing unit, the server enables the generative artificial intelligence model to encode a more complete context in the vector space, thereby improving the accuracy and relevance of semantic interpretation, reducing the number of times users need to make secondary queries, and achieving error reduction and improved interaction efficiency.

[0113] 4. By centrally performing length statistics, text trimming, and redundancy removal on the server, the terminal does not need to implement complex post-processing logic. It can still run the chat application smoothly with a lower-performance processor and limited memory. From the perspective of the system as a whole, this reduces the computing burden on the terminal and optimizes resource allocation.

[0114] This invention, through the aforementioned hardware and software co-design, combines the natural language generation capabilities of generative artificial intelligence models with structured prompts, server-side parameter control, and historical data-driven optimization. This enables the computer system to exhibit efficient, stable, and controllable technical characteristics in semantic querying and result presentation, rather than simply automatically executing human queries and interpretations. This implementation method possesses clear technical characteristics and reproducible deployment in terms of data structure, module division, and algorithm flow, providing a universal implementation framework for different types of terminals and application scenarios.

[0115] use Figure 11 The processing flow is explained.

[0116] Step 1: Users select a character sequence on the terminal's display screen.

[0117] Users browse conversations in a chat application running on their device. They can select a piece of text, such as "generative artificial intelligence model" or "prompt statement," by long-pressing or dragging their finger.

[0118] Input: The complete dialogue text displayed on the terminal screen.

[0119] Output: The character sequence selected by the user and its start and end positions in the message.

[0120] The terminal obtains the selection range index based on the text selection interface provided by the operating system, performs substring extraction on the original message string, and copies the character sequence of the corresponding range into a variable in memory for subsequent data processing.

[0121] Step 2: The terminal will select a character sequence, structure it, and send it to the server.

[0122] After obtaining the selected character sequence, the terminal creates a structured data object locally containing the session ID, message ID, user ID, and the selected character sequence.

[0123] Input: The selected character sequence and its position information, and metadata related to the current session.

[0124] Output: A request data packet encapsulating the above information is sent to the server over the network.

[0125] The terminal uses a JSON serialization library to encode the structured object into a JSON string, and then sends the JSON as a request body to the specified interface of the server through a network communication module (such as an HTTP client), thus realizing the raw selection data transmission from the terminal to the server.

[0126] Step 3: The server retrieves lexical information from the data storage device and generates semantic information.

[0127] After receiving a request from the terminal, the server parses the selected character sequence from the request body. The server then uses this character sequence as a search key to access the vocabulary database and performs a search operation in the database index to match the corresponding word record.

[0128] Input: JSON request data containing the selected character sequence.

[0129] Output: A semantic information object corresponding to the character sequence (including basic definition, brief description, etc.).

[0130] The server constructs a semantic information data structure based on the search results, organizes the term meaning field, example field, etc. into a short text, and encapsulates the semantic information into a response JSON and returns it to the terminal, realizing data processing from dictionary data to displayable semantic text.

[0131] Step 4: The server generates prompt statements based on the template and user settings.

[0132] After completing the vocabulary search, the server reads the prompt template record corresponding to the user's settings from the database, such as "Please explain the meaning of '{character sequence}' in no more than {number of characters} Chinese characters, and keep it concise."

[0133] Input: Selected character sequence, user's word limit and description style settings, template information table.

[0134] Output: Specific prompt text for generative artificial intelligence models.

[0135] The server performs a placeholder replacement operation on the template string in memory, replacing the placeholder {character sequence} with the selected text and {number of characters} with the user-defined character limit, thereby generating a specific prompt statement, such as: "Please explain the meaning of 'generative artificial intelligence model' in no more than 100 Chinese characters, and keep it concise."

[0136] Step 5: The server constructs a model call request and invokes the generative artificial intelligence model.

[0137] After receiving the prompt, the server calculates the corresponding max_tokens parameter based on the character limit in the prompt, and constructs a generative artificial intelligence model invocation request by combining it with preset parameters such as temperature and top_p.

[0138] Input: Prompt text, character limit parameters, and model call configuration parameters.

[0139] Output: Request data conforming to the model interface format is sent to the generative artificial intelligence model server, and the response text output by the model is received.

[0140] The server uses an HTTP client library to encapsulate the prompt statement in the request body and sends it to the model service interface. The Transformer network inside the model service vectorizes the prompt statement and performs autoregressive decoding. After receiving the JSON response returned by the model, the server parses out the response text field to provide raw data for the next step of pruning and adjustment.

[0141] Step 6: The server performs length determination and pruning on the response information of the generative artificial intelligence model.

[0142] After the server receives the response text, it counts the number of characters in the text in memory, calculates the actual length using a traversal counting method for Unicode characters, and compares it with the preset character limit.

[0143] Input: The original response text returned by the model, and the word count limit parameter.

[0144] Output: Adjusted response text that meets the length constraint.

[0145] When the response text exceeds the character limit, the server truncates the text at character or word boundaries, removing the excess. Simultaneously, the server can perform rule-based filtering to remove redundant polite phrases, blank lines, etc. Through this data processing, the output text is compressed into a compact and controlled format.

[0146] Step 7: The server sends semantic information and adjusted response information to the terminal.

[0147] After the server completes the generation of lexical semantic information and the adjustment of the model response, it merges the two types of results into a unified response structure, which is divided into two parts: "lexical semantic information" and "AI response information".

[0148] Input: Lexical semantic information object, and the cropped response text.

[0149] Output: JSON response data containing two interpretations, returned to the terminal over the network.

[0150] The server uses JSON serialization to encode the above data into a string and sends it to the terminal via an HTTP response, enabling the terminal to obtain both traditional dictionary explanations and generative artificial intelligence model explanations in a single response.

[0151] Step 8: The terminal parses the server response and generates a display unit in the dialog interface.

[0152] After receiving the server's response, the terminal reads the semantic information and AI response information fields through a JSON parsing library and generates the corresponding display data structures in memory.

[0153] Input: The JSON string returned by the server.

[0154] Output: Semantic display units and AI response display units for interface rendering.

[0155] The terminal creates two new message bubbles or display blocks in the dialogue display area, displays dictionary semantic information in one bubble and AI response information in the other bubble, and inserts these two bubbles near the original selected character sequence in the message through UI layout logic, realizing the mapping and rendering from data fields to specific interface elements.

[0156] Step 9: The terminal displays progress and error messages based on the processing status.

[0157] After the terminal sends a request to the server, it marks the query as "in progress" in its local state machine and displays a loading indicator (such as a rotation icon or the text "Querying...") on the interface.

[0158] Input: Query start event, server response, or error status.

[0159] Output: The corresponding progress icon or error message display status.

[0160] When the terminal receives a normal response, it updates the status to "Completed" and hides the loading indicator. If no response is received within a preset time or a response containing error information is received, the terminal generates an error message display unit in the interface, such as "Query failed, please try again later". The UI display is controlled by conditional judgment to realize visual feedback on network and server anomalies.

[0161] Step 10: The server records historical information and updates templates and parameters.

[0162] After each processing step, the server records the selected character sequence, the prompt statement used, the model response text, the word limit parameter, and the flag indicating whether pruning occurred in the history information table.

[0163] Input: All key fields in this interaction process (character sequence, prompt statement, response text, character limit and clipping flag, etc.).

[0164] Output: A new record stored in the historical information database, as well as statistical results and updated templates or parameters based on multiple records.

[0165] The server periodically reads historical information and performs statistical calculations on factors such as the average response length, cropping ratio, and repeated user query behavior across different templates. Based on these statistical results, it adjusts the default character limit in the templates or modifies the wording of prompts, thereby adopting the new template configuration in subsequent request processing. This data analysis and parameter update process is completed internally by the server, enabling adaptive optimization of the system's prompt strategy and generated length control.

[0166] Step 11: The server uses the conversation history to generate extended prompts to obtain context-appropriate responses.

[0167] In an implementation where the server sends the dialogue history near the selected character sequence along with the terminal, the server parses several context message texts from the request and concatenates them into a dialogue context string in chronological order.

[0168] Input: The selected character sequence and the corresponding set of dialogue history text.

[0169] Output: The text of the extended suggestion statement containing context information, and the text of the context-dependent response generated based on the extended suggestion statement.

[0170] The server will select a character sequence and embed it with a context string into a unified prompt statement, such as: "In the following dialogue, what does 'model' refer to? Please explain in 80 words or less. Dialogue content: ...". Then, it will invoke the generative AI model and adjust the response as in steps 5 and 6. Because the dialogue context is added to the prompt statement, the model can construct a more accurate semantic representation during the internal encoding stage. Therefore, the server's output response has a higher contextual matching degree and reduces the occurrence of ambiguous interpretations.

[0171] Step 12: Users can view the results on the terminal and initiate further queries.

[0172] Users read dictionary semantic information bubbles and AI response information bubbles in the terminal dialogue interface, and users can judge whether they have understood the meaning of the character sequence as needed.

[0173] Input: Semantic information and AI response information displayed on the terminal screen.

[0174] Output: The user's instructions on whether to continue the query or select a new character sequence.

[0175] If the user still has questions, they can select a new character sequence or manually enter a more complex question as a new prompt on the same interface. The terminal will then re-trigger the processing flow after step 1, thereby achieving multi-round, efficient semantic query interaction.

[0176] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0177] In modern human-computer interaction technology, with the popularization of wearable information display devices, generative artificial intelligence models and network services, users' demand for high-quality explanatory information related to specific statements in real time in real-world environments is constantly increasing. However, existing technologies typically suffer from the following problems: First, when users wear smart glasses or other information display devices, the system struggles to locate and extract the selected statement from the user's field of vision in a timely and accurate manner, leading to the need for additional user input or multiple confirmation steps in subsequent processing, resulting in low interaction efficiency. Second, even if the selected statement can be sent to the server, the server-side calls to generative AI models often employ fixed or simple free text prompts, lacking an automatic generation mechanism for structured prompt statements based on scenarios and display modes, making it difficult to consistently obtain responses suitable for terminal display constraints and business contexts. Third, existing systems often prioritize long text when calling generative AI models, failing to closely link with the character limit of downstream display devices and lacking the ability to perform comprehensive length control and summarization at the server level and sentence level. This results in generated responses either being simply truncated at the terminal, causing semantic incompleteness, or being displayed cluttered, reducing readability and user experience. Fourth, existing systems generally lack structured records and storage of the correlation between "selected statement - prompt statement - generated response," failing to provide basic data support for subsequent template optimization, service quality assessment, and adaptive prompt generation, thus limiting the system's maintainability and scalability.

[0178] Therefore, the technical challenge this invention aims to address is how to collaboratively improve computer technology on both the server and terminal sides to achieve the following: automatically acquiring and recognizing sentences selected by users through wearable information display devices; generating structured prompts on the server by combining sentence information and service type information; intelligently pruning or extracting summaries on the server side based on character limits after obtaining answers using generative artificial intelligence models; and storing sentence information, prompts, and response information in association, thereby comprehensively improving the computer processing capabilities and user experience of prompt construction and result presentation in generative artificial intelligence applications.

[0179] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0180] In this invention, the server includes: a receiving means for acquiring statement information based on a user's statement selection via an information display device and sending it to the server via a terminal; a meaning information acquisition and sending means for retrieving the statement information in a vocabulary information storage unit to obtain meaning information and sending the meaning information to the terminal for display on the information display device; a prompt statement generation means for selecting template information from a set containing multiple template information according to a display mode based on the statement information and service type information, and automatically generating a prompt statement for specifying the content to be queried by a generative artificial intelligence model by embedding the statement information into a predetermined position in the template information; a model interaction means for sending the automatically generated prompt statement to the generative artificial intelligence model and receiving the response information returned by the model; a response information processing means for determining the length of the response information according to a predetermined character limit, and extracting a predetermined number of sentences or extracting sentences with higher importance through extractive summarization processing when the number of characters exceeds the upper limit, thereby generating response information adapted to the terminal display constraints for display; and a recording means for associating and storing at least a portion of the statement information, prompt statement, and response information in a recording area for subsequent analysis and template optimization. This enables structured control and unified management of the generative AI model invocation process on the server side. While ensuring that the response content is highly relevant to the user context, it automatically generates high-quality prompts and performs intelligent summarization and formatting on the model response, taking into account the display capabilities and character limits of the terminal. This improves the overall efficiency and effectiveness of data acquisition, natural language generation and invocation, and result presentation in wearable information display environments, enhancing the human-computer interaction experience for users in real-time information query scenarios.

[0181] A "system" refers to a collection of integrated information processing devices consisting of multiple interconnected hardware and software components, used to acquire, process, send, and display user-selected statements.

[0182] "User" refers to an individual or organization that operates a terminal and selects statements, views meaningful information, and receives responses from generative artificial intelligence models through an information display device.

[0183] "Terminal" refers to an electronic device carried or worn by a user for data communication with a server and for presenting information on an information display device, including but not limited to wearable devices, mobile devices or other computing devices.

[0184] "Information display device" refers to a display component used to present text or image information to a user in a visual manner, including but not limited to head-mounted displays, smart glasses displays, flat panel displays or other image display modules.

[0185] A “statement” refers to a sequence of text that a user can select on an information display device, including words, phrases, or sentences, used to express a specific meaning or information.

[0186] "Field of view image" refers to image data acquired by the terminal through the imaging device that reflects the user's current observation scene, and may include areas containing statements.

[0187] "Character information" refers to text data extracted from the visual field image through character recognition processing and represented in character encoding form, which is used for subsequent sentence recognition and processing.

[0188] "Statement information" refers to the text data and its additional attributes related to the statement selected by the user, including the character information of the statement itself and the structured or semi-structured data related to the statement's identifier, position, context, etc.

[0189] A “vocabulary information storage department” refers to a data storage resource used to store vocabulary entries, definitions, attributes, and other information related to sentences. It can consist of a database, a file system, or other data storage media.

[0190] "Meaning information" refers to explanatory data used to explain the meaning of a statement, obtained by retrieving statement information from the vocabulary information storage department. This includes definitions, usage instructions, examples, etc.

[0191] "Service type information" refers to control information that indicates the current service category or application scenario type. It is used to indicate the target mode for generating prompt statements or handling responses, such as terminology explanation mode, product description mode, or usage method mode.

[0192] "Generative artificial intelligence models" refer to machine learning or artificial intelligence models that can automatically generate natural language response information based on input prompts, including but not limited to language models based on deep learning.

[0193] "Prompt statements" refer to natural language text used to provide query instructions or context to generative artificial intelligence models, specifying the content, style, or constraints of the response information generated by the model.

[0194] "Template information" refers to a predefined text or data structure containing fixed structure and insertable statement placeholders, which is used to fill and combine when generating prompt statements.

[0195] "Display mode" refers to the mode settings used to control how information is displayed, including font size, display layout, and level of detail of information, which affect template selection and prompt generation strategies.

[0196] "Response information" refers to the natural language text or other forms of response data output by generative artificial intelligence models based on prompt statements.

[0197] "Character count limit" refers to the maximum number of characters set for the response information, which is used to control the length of the response information to adapt to the display capabilities of the terminal or information display device.

[0198] "Response information for display" refers to response text data that is suitable for direct display on the terminal's information display device after the response information output by the generative artificial intelligence model has been processed according to character limit and summary rules.

[0199] "Length determination" refers to the process of calculating the number of characters in the response information and comparing it with a preset upper limit to determine whether the response information needs to be truncated or digested.

[0200] "Extractive summarization" refers to a processing method that evaluates the importance of sentences in the response information and selects sentences with higher importance to form a summary without generating or rewriting the sentences themselves.

[0201] "Overlay display" refers to a display method on an information display device that displays response information on top of the user's field of vision image or other basic screen in a way that includes overlaying, floating, or augmented reality, allowing the user to perceive both the basic screen and the response information simultaneously.

[0202] "Record area" refers to the data storage space used to persistently store statement information, prompt statements, response information and their associated relationships, which can be composed of database, storage server or other data storage media.

[0203] The embodiments of this invention will combine the collaborative work of the server, terminal and user to provide a detailed description of the system's hardware structure, software modules, data structure and the internal processing methods of the generative artificial intelligence model, so that those skilled in the art can implement this invention and understand the computer technology improvement effects it brings.

[0204] I. System Overall Structure In this invention, the server is deployed on a data processing device on the network side, preferably a computing device with a multi-core processor and large memory capacity, such as a network server running a general-purpose operating system (e.g., a Unix-like operating system). The server runs a web application framework (e.g., a Python-based web framework or a JavaScript-based web server framework) and interacts with external generative artificial intelligence model services and terminals through secure communication protocols.

[0205] In this invention, the terminal is a wearable or portable computing device, such as smart glasses or other head-mounted display devices running a mobile operating system. The terminal includes a camera module, a display module, a touch or button input module, a wireless communication module, and a local processor. The terminal runs a dedicated application program for this invention, used to perform processing such as image acquisition, character recognition, sentence information construction, prompt statement generation, and display control.

[0206] In this invention, users wear a terminal and observe product labels, explanatory text, or other information carriers in a real-world setting. Users select target statements through physical operations on the terminal (touch, buttons, gestures, etc.) and view the meaning information returned by the server and the response information from the generative artificial intelligence model displayed on the terminal.

[0207] II. Server-side software modules and data structures The server in this invention includes multiple logical modules, which can be implemented by one or more programs on the same physical device or different physical devices. The main modules include: a request receiving module, a statement information parsing module, a vocabulary information retrieval module, a prompt statement generation module, a generative artificial intelligence model interface module, a response information processing module, a display information generation module, and a recording module.

[0208] The server receives statement information from the terminal in the statement information parsing module. The statement information is preferably represented in structured data format, which the server can parse into a data structure containing the following fields in memory: statement text, statement language, context category, service type identifier, timestamp, and terminal identifier. The server performs field validation and filtering on this data structure to ensure data quality for subsequent processing from the outset, reducing the impact of abnormal data on system performance and accuracy.

[0209] The server accesses the vocabulary information storage unit within the vocabulary information retrieval module. This storage unit can be implemented using a relational database. The server pre-stores multiple vocabulary records in the database, each containing at least fields such as word form, part of speech, basic definition, extended definition, and typical use cases. Upon receiving the sentence information, the server standardizes the text (e.g., case conversion, removal of redundant symbols), generates a search key, and executes the query on the index structure. Because the server uses an index structure optimized for high-frequency queries (e.g., a B+ tree or inverted index structure), it still achieves high retrieval speed even with large-scale vocabulary data. This structured retrieval improves both performance and accuracy compared to simple string matching.

[0210] The server maintains a set of template information in its prompt generation module. Template information can be stored in a configuration database or file system. Each template contains a fixed text portion and one or more placeholders. When generating prompts, the server does not directly concatenate free text, but performs the following technical processing: based on the service type identifier and display mode parameters contained in the prompt information, it selects a specific template from the template set; then, it inserts the prompt text and context information into the predetermined positions within the template. For example, the server can use templates of the following form: Please explain the meaning of the statement '{statement}' and give a simple example. Your answer should not exceed 100 words. Please describe the product features and applicable scenarios related to '{statement}', with a maximum of 150 words. "Please explain the key points involved in '{statement}' to the user in plain language from the perspective of a service worker, and keep your answer concise." The server generates prompts using this template-based information filling method. Compared to completely free text construction, this improves the structural stability and predictability of the prompts, making the output of subsequent generative AI models easier to control and facilitating the achievement of desired style and length responses within character limits. This template-driven prompt construction represents a technical improvement over traditional natural language interfaces.

[0211] III. Structure and Training of Generative Artificial Intelligence Models The server interfaces with externally deployed language models in the generative artificial intelligence model interface module. This model is preferably a deep learning-based generative artificial intelligence model, which can internally employ a multi-layered self-attention neural network architecture, such as a multi-layered Transformer encoder-decoder or an autoregressive language model containing only a decoder structure.

[0212] During the training phase, the server pre-trains the model using a large-scale text corpus. Training data can include general text, technical documents, product descriptions, etc. The server implements the following technical details in its training control program: During forward propagation, the server encodes prompts into vector sequences and models the internal dependencies of the sequences using a multi-head self-attention mechanism; during the loss calculation phase, the server uses a cross-entropy loss function to compare the probability distribution of the next word generated by the model with the true target sequence and calculate the error; during the backpropagation phase, the server updates the weight parameters of each layer of the network based on gradient descent and its variants (such as adaptive learning rate optimization algorithms). The server can also apply data augmentation strategies to the training data, such as synonym rewriting and sentence transformation, to enhance the model's robustness to different expressions.

[0213] During the inference phase, when the server generates response information based on the prompts, it executes a decoding algorithm based on probability sampling or beam search. The server controls the diversity and length of the generated text by setting parameters such as decoding temperature, maximum generation length, and penalty coefficient, thus working in conjunction with the character count limit in this invention. The setting of these decoding parameters is a technical means that directly affects the model's output characteristics and computational resource consumption.

[0214] IV. Response Information Processing and Character Count Limit Control The server performs fine-grained processing on the response information returned by the generative artificial intelligence model in the response information processing module. The server first performs character counting and sentence segmentation on the response text. The server can divide sentences according to pre-defined delimiters (such as periods, question marks, exclamation marks, etc.) and build sentence arrays in memory.

[0215] When the server detects that the response text exceeds the character limit, it doesn't simply truncate it. Instead, it performs one or a combination of two technical processing methods: first, it retains the first few sentences in sentence order; second, it executes an extractive summarization algorithm. In the extractive summarization, the server calculates the weight of each sentence. Weight calculation can be based on simple features such as word frequency, position, and keyword matching, or further employ lightweight vector representations and similarity metrics. The server selects several sentences according to their weights to construct the response information for display. Through this character-limit-oriented summarization process, the server avoids simple hard truncation on the client side, which can lead to semantic breaks. This improves the integrity and readability of information delivery, reduces the number of user re-requests, and indirectly reduces communication load and server computational load.

[0216] When processing response information, the server can also adapt its processing based on the display mode. For example, in "concise mode," the maximum character limit can be reduced and the summarization can be strengthened, while in "detailed mode," more contextual information can be retained. This length control and summarization strategy, which is centrally implemented on the server side, enables the system to manage and optimize uniformly for different terminals and scenarios, representing a technical improvement over the traditional terminal-based truncation solution.

[0217] V. Image Acquisition and Character Recognition Processing on the Terminal Side The terminal uses its camera hardware in the image acquisition module to periodically or on demand capture images of the user's current field of view. When the terminal receives a selection action from the user (such as a tap on the side touchpad), it locks the image frame at that moment as the target field of view image. The terminal stores this image in its memory and initiates character recognition processing.

[0218] The terminal calls upon optical character recognition software libraries in its character recognition module, such as running an open-source character recognition engine or an integrated machine learning inference library on the terminal's local processor. The terminal preprocesses the target image, including grayscale conversion, binarization, noise filtering, and geometric correction, to improve the accuracy and stability of character recognition. After obtaining the character recognition results, the terminal segments the recognized text by lines and words, and infers the actual sentence selected by the user based on the user's gaze position or touch position. This processing enables the terminal to efficiently extract target sentences from complex backgrounds and significantly reduces manual input operations by the user.

[0219] When constructing statement information, the terminal packages the identified statement text along with basic metadata such as language, display context, and timestamps into structured data. The terminal can perform simple statement cleaning locally, such as removing leading and trailing spaces and standardizing encoding formats, to ensure the consistency and parsability of the data sent to the server. This terminal-side preprocessing helps reduce the server's parsing burden and improves the overall system response speed.

[0220] VI. Terminal-side prompt message filling and display control The terminal can also implement partial or complete template filling logic during the prompt generation process. The terminal can store frequently used templates locally, for example: Please explain the meaning of the word '{statement}' and give a simple example. Your answer should not exceed 100 words. Please provide product information related to '{statement}', including key features and applicable scenarios. Your answer should be within 150 characters. Please explain the basic usage of '{statement}' step by step, using concise language. After receiving the user's selected statement, the terminal chooses the appropriate template based on the current scenario and inserts the statement into the template to form a preliminary prompt. This prompt, along with other information, is then sent to the server. The server can then directly use the prompt or make further adjustments based on the server-side template. This collaborative division of labor between the terminal and server in prompt generation allows for a flexible balance between network latency and computational load, improving the overall system's response performance.

[0221] The terminal presents the semantic information returned by the server and the response information used for display in the display control module. Based on parameters such as the display hardware resolution and field of view, the terminal selects an appropriate font size and line-wrapping strategy, and overlays the text onto the unobstructed area of ​​the user's current field of view. The terminal can use a local graphics API during rendering to draw the text onto a transparent background layer, allowing the user to see the real scene while simultaneously reading related instructions. This overlay display method directly applies the computer's internal processing to the real-world visual output device, realizing a technical link from abstract data to concrete physical display control.

[0222] VII. Records and Subsequent Optimization Support The server stores the mapping relationships between statement information, prompt statements, response information, and related metadata in the database within the recording module. The server employs a relational table structure for this purpose, including statement tables, prompt tables, response tables, and their associated tables. The server creates or updates corresponding records after each interaction for subsequent analysis.

[0223] When the server performs offline analysis on these records, it can calculate statistical indicators such as the usage frequency of different templates, response truncation ratio, and user re-query rate to objectively evaluate the effectiveness of the prompt message templates and summary strategies. Based on these statistical results, the server iteratively updates the template set and summary rules, thereby continuously optimizing the prompt generation and response processing logic during system operation. This method of optimizing model call configuration using structured records substantially improves the computer system's adaptive capabilities in natural language interfaces, representing a data-driven, system-level technological improvement.

[0224] VIII. Technical Effects and Causal Relationships Through the aforementioned structured prompt generation, character-limited summary processing, and vocabulary retrieval and recording mechanisms, combined with the terminal's image acquisition and character recognition, this invention achieves improvements in computer technology in several aspects: The server uses template selection and placeholder filling when generating prompts, making the input distribution of the generative AI model more stable, reducing invalid or redundant outputs, and improving response quality and average inference efficiency. Because the model is more likely to produce high-confidence outputs on normalized inputs, the overall semantic accuracy of the responses is improved.

[0225] When processing responses based on character limits, the server employs sentence-level and digest-level algorithms for intelligent pruning. This avoids semantic incompleteness and duplicate requests caused by simple truncation at the terminal, directly reducing network communication frequency and data transmission volume, thereby reducing communication load and server computing resource consumption.

[0226] The terminal performs image preprocessing and character recognition locally, which can complete high-cost visual computing close to the data source, reduce the processing pressure on the server for visual data, and compress the amount of data transmitted by uploading only structured text data, thereby improving the overall processing speed.

[0227] The server establishes a correlation between statement information, prompt statements, and response information in the recording module, providing a computable basis for subsequent automatic analysis and rule optimization. This forms a feedback loop for prompt construction and response length control, enabling the system to continuously adjust parameters and templates based on actual operating conditions, thereby improving stability and adaptability under long-term operation.

[0228] With the above configuration, the system proposed in this invention is not merely a simple automation of manual query work, but also makes improvements at multiple internal computer technology levels, such as prompt statement construction, generative artificial intelligence model invocation, response length control, image-to-text conversion, and data recording and feedback, thereby achieving comprehensive technical effects in terms of processing speed, response quality, resource utilization, and user experience.

[0229] use Figure 12 The processing flow is explained.

[0230] Step 1: Users wear the terminal in real-world scenarios and select statements through the information display device.

[0231] Input: Real-world text information (such as text on product labels), user's gaze position, and physical operations (touch, button, or gesture).

[0232] Output: Indicates the user's selection command for the target area and the corresponding visual field trigger signal at that moment.

[0233] When a user focuses their gaze on a word on a product label (such as "materials") and lightly touches the touchpad on the side of the terminal or presses a button, the terminal determines that the user wants to select the text in the center of their field of vision and triggers the subsequent image acquisition and processing process.

[0234] Step 2: The terminal acquires images of the user's field of view and locates candidate text regions.

[0235] Input: Raw image frame data from the camera, as well as the trigger signal and line of sight / indication position generated in step 1.

[0236] Output: Image of the cropped candidate text region.

[0237] The terminal calls the camera hardware interface to acquire one or more high-resolution color images and caches them in memory. The terminal calculates the pixel coordinates of the candidate regions according to preset geometric rules (e.g., cropping a rectangular area of ​​fixed width and height centered on the image center region or the estimated line of sight), performs cropping data operations on the original image, and generates a local image containing only candidate characters for subsequent character recognition.

[0238] Step 3: The terminal performs character recognition on the candidate regions and extracts the text of the statements.

[0239] Input: Candidate region image obtained in step 2.

[0240] Output: The identified text string of the statement and its line and word boundary information.

[0241] The terminal calls optical character recognition software (such as a locally running OCR engine) to convert the candidate region image into a grayscale image. It then performs image preprocessing operations such as denoising, binarization, and tilt correction. Next, it performs character segmentation and pattern recognition to obtain the encoded value of each character. The terminal combines the characters line by line and word by word to generate structured text results. Based on the user-selected location, it filters out the corresponding target sentence (such as "material") and outputs the sentence text along with its location index.

[0242] Step 4: The terminal constructs statement information and performs local preprocessing.

[0243] Input: The statement text, row and column position information, and terminal internal state (language type, scene type, etc.) output from step 3.

[0244] Output: Structured statement information data object.

[0245] The terminal combines the statement text with metadata (such as language, timestamp, terminal ID, and context category) into a data structure. During this process, the terminal performs string cleaning operations (removing leading and trailing spaces, standardizing encoding, and deleting meaningless symbols) and standardizes the statement text (such as unifying capitalization) to obtain structured statement information that is easy for the server to parse.

[0246] Step 5: The terminal generates an initial prompt statement based on the statement information and display mode.

[0247] Input: The statement information generated in step 4 and the current display mode / service type settings.

[0248] Output: A preliminary prompt statement string containing the statement text.

[0249] The terminal reads various prompt templates from local storage, such as: Please explain the meaning of the word '{statement}' and give a simple example. Your answer should not exceed 100 words. Please describe the product features and applicable scenarios related to '{statement}', with a maximum of 150 words. The terminal selects a template based on the service type (such as "terminology explanation" or "product description") and display mode (concise / detailed), replaces the placeholders in the template with the statement text, and generates a specific prompt statement through string concatenation and formatting operations, such as: "Please explain the meaning of the word 'material' and give a simple example. The answer should not exceed 100 words." Step 6: The terminal packages the request and sends statement information and prompts to the server.

[0250] Input: the statement information from step 4, the initial prompt statement from step 5, and the network connection parameters.

[0251] Output: The request message sent to the server over the network.

[0252] The terminal constructs a request data structure (such as key-value pairs) containing statement information and prompts, and converts it into a byte stream suitable for network transmission through serialization. The terminal invokes the communication module, establishes a connection with the server using a secure transmission protocol, and sends the request message to the server's designated interface, thereby reporting the user's selection result and prompts to the server.

[0253] Step 7: The server receives and parses the terminal request.

[0254] Input: Byte stream of request message from the terminal.

[0255] Output: The server's internal statement information object and the prompt statement string.

[0256] The server receives request messages in its network listening module, performs decoding operations on the messages, and restores them from byte streams to structured data. The server checks the integrity of the message header and fields, performs syntax and length checks on statement information and prompts, and stores them in a data structure in memory as input for subsequent processing.

[0257] Step 8: The server retrieves the meaning information of the sentences from the vocabulary information storage department.

[0258] Input: The text and information of the statement obtained from step 7.

[0259] Output: Text containing the meaning of the corresponding statement.

[0260] The server performs standardization processing on the statement text (such as converting it to a unified character set, applying simple word segmentation or stemming algorithms) to generate search keywords. In the vocabulary information storage unit (such as a relational database or key-value store), the server performs retrieval operations based on the index structure (e.g., a B+ tree index), searching for entries that match the statement, reading their definition fields, usage description fields, etc., and combining these fields into semantic information text, which is then output as the search result.

[0261] Step 9: The server sends the meaning information to the terminal for immediate display.

[0262] Input: The meaning information text obtained in step 8 and the terminal identifier.

[0263] Output: A response message containing the meaning of the message sent to the terminal.

[0264] The server constructs meaningful response data, encapsulating the meaningful text and statement identifiers together in the response message. After serialization and encoding, it sends the message to the corresponding terminal over the network. At this stage, the server does not rely on generative artificial intelligence models but directly utilizes its local vocabulary to achieve a fast response, enabling the terminal to display the basic meaning in a short time.

[0265] Step 10: The server selects a template based on the statement information and service type, and generates the final prompt statement.

[0266] Input: Statement information from step 7, initial prompt statement (if generated by the terminal), and server-side template set and service type information.

[0267] Output: The final prompt string sent to the generative AI model.

[0268] The server searches the template set for template entries that match the service type and display mode, comparing them with the initial prompt statements uploaded by the terminal. If the terminal does not provide one or requires adjustment, the server performs a template selection operation and fills the statement text and its context information into the template placeholders to generate the final prompt statement. If the terminal has already generated a suitable prompt statement, the server can directly reuse it or make minor modifications (such as adding a word limit explanation). Through this template-based filling, the server constructs unstructured natural language into prompt statements with a stable format for subsequent processing by the generative artificial intelligence model.

[0269] Step 11: The server invokes the generative artificial intelligence model and obtains the response information.

[0270] Input: The final prompt statement generated in step 10.

[0271] Output: The raw response text returned by the generative artificial intelligence model.

[0272] The server, through a model interface module, feeds the prompt statement as an input sequence into a pre-trained generative AI model. Internally, the model employs a multi-layered self-attention network structure. During inference, the server performs the following operations: segments the prompt statement into words and embeds them into a vector space, inputting this into a multi-layered Transformer decoder. Each layer calculates attention weights and intermediate feature vectors through matrix multiplication and normalization operations. Then, the output layer performs a linear transformation and softmax operation to generate the probability distribution of the next word. The server executes a decoding algorithm (such as beam search) based on set parameters such as temperature, maximum length, and penalty coefficient, selecting the next word probabilistically, iterating until a termination condition is met. Finally, the server concatenates the predicted word sequence into a complete response text, outputting it as the original response information.

[0273] Step 12: The server performs character count limits and sentence-level segmentation on the response information.

[0274] Input: The original response text obtained in step 11 and the preset character limit value.

[0275] Output: The structure of the sentence array and the result of whether it exceeds the limit.

[0276] The server performs a character count on the original response text, comparing its length with a preset upper limit to determine if it exceeds the character limit. Simultaneously, the server segments the text into ordered sentence arrays based on punctuation marks such as periods, question marks, and exclamation marks. The server then uses these sentence arrays and the over-limit flag as input for subsequent summary processing.

[0277] Step 13: The server performs an extractive digest or truncation on the out-of-limit response to generate response information for display.

[0278] Input: The sentence array generated in step 12 and the result of the over-limit judgment.

[0279] Output: Response text for display that meets the character limit.

[0280] When the server detects that the response length exceeds the limit, it executes a sentence-level selection algorithm: On one hand, the server can simply add sentences sequentially from the beginning of the sentence array until it approaches but does not exceed the character limit, thus constructing a relatively complete first half of the content; on the other hand, the server can calculate the keyword density, positional weight, and other feature values ​​of each sentence, score and rank the sentences, and select the sentences with higher scores to combine into a summary. The server combines the selected sentences into a new response text through string concatenation operations, and adds ellipses when necessary to indicate that the content is compressed. Ultimately, the server ensures that the number of characters in the output text does not exceed the limit and that the semantics are relatively complete.

[0281] Step 14: The server records the relationship between statement information, prompt statements, and response information.

[0282] Input: the statement information in step 7, the final prompt statement in step 10, the response text to be displayed in step 13, and the original response text.

[0283] Output: Multiple database records or log entries written to the record area.

[0284] The server constructs a record data structure, combining fields such as statement identifier, statement text, prompt statement, original response text, response text for display, and timestamp, and performs database insert or update operations to write this data into the record area. Through this structured record, the server provides calculable historical data for subsequent statistical analysis and template optimization.

[0285] Step 15: The server sends the response information for display to the terminal.

[0286] Input: The response text output from step 13 for display and the terminal identifier.

[0287] Output: A server-to-terminal response message containing the response text.

[0288] In the information generation module, the server packages the response text and statement identifiers for display into a response data structure, serializes and encodes it, and then sends it to the corresponding terminal over the network. At this stage, the server can add display mode prompts based on the terminal type to allow the terminal to perform appropriate layout.

[0289] Step 16: The terminal receives the meaning information and the response information for display, and displays them overlaid on the information display device.

[0290] Input: The response message containing the meaning information returned in step 9 and the response message containing the response information returned in step 15 for display.

[0291] Output: The semantic information and model response text image superimposed on the user's field of vision.

[0292] The terminal receives two types of responses from the server in its communication module, parsing out the semantic text and the response text for display. The terminal performs layout calculations based on the display device resolution and the user-set display mode (e.g., concise / detailed), including text area size, number of lines, font size, and screen coordinates. In its graphics rendering module, the terminal calls a drawing interface to draw the text line by line on a semi-transparent background layer and overlays this layer onto the real-world image. When wearing the terminal, the user can simultaneously see the original scene and the server-processed semantic information and generative responses in their field of vision, enabling immediate and structured explanation of selected statements.

[0293] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0294] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0295] In existing human-computer interaction technologies, when a user selects a sentence on a terminal, the meaning can usually only be looked up through a fixed vocabulary database, or a generative artificial intelligence model can directly output long and redundant explanatory information. Firstly, traditional dictionary-based data query methods cannot dynamically generate explanatory content based on context and the user's current reading scenario, resulting in insufficient relevance and readability of the information. Secondly, existing systems utilizing generative artificial intelligence models often simply take the user-selected sentence as input, lacking structured design and adaptive generation of prompts, easily leading to output results that are not focused on the key points the user truly cares about. Thirdly, the output length of generative artificial intelligence models is usually not finely controlled by the front-end display environment and terminal type, resulting in problems such as line breaks, truncated key information, and the need for repeated scrolling on small screens or in limited display areas, thus reducing the user's efficiency in using the information processing system. Fourthly, existing systems lack systematic utilization of user historical interaction records and model generation results, failing to dynamically optimize the generation conditions and output length strategies of prompts based on recorded information, making it difficult to achieve continuous performance optimization for different users and different content scenarios.

[0296] Therefore, under the premise of the basic interactive action of users selecting sentences on the terminal, how to improve the collaborative processing flow between the server side and the terminal side to achieve: highly relevant response generation based on structured prompts, dual-channel information provision combined with lexical data query, adaptive character number control adapted to the display environment, and dynamic adjustment mechanism based on recorded information, thereby improving the overall processing efficiency and resource utilization of computers in natural language information acquisition and display tasks, has become an urgent technical issue to be solved.

[0297] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0298] In this invention, the server includes: a device for detecting and acquiring a selected statement based on user operations on a terminal display interface; a device for generating a prompt string containing the selected statement based on a predetermined string template, and using the prompt string as an input prompt statement for a generative artificial intelligence model; a device for sending the prompt string to an information processing device or a model server via a communication path, and receiving response information generated by the generative artificial intelligence model; a device for determining the length of the received response information according to a preset or dynamically adjusted character limit, and simplifying the response information by deleting excess parts and / or adding ellipsis markers; a device for retrieving semantic information from vocabulary data for the selected statement, and sending the semantic information together with the simplified response information to the terminal for combined display on the display interface; and a device for storing the processing content corresponding to the prompt string and response information as record information, and updating the generation conditions and / or character limit of the prompt string based on the record information. This allows for integrated control of prompt statement construction, model invocation, and result post-processing on the server side. Combined with vocabulary data query results and record-driven adaptive strategies, the output of generative artificial intelligence models is optimized in terms of relevance, length control, and display adaptability. This improves the processing efficiency, interface presentation, and resource utilization of computer systems in natural language information generation and display tasks.

[0299] "System" refers to a collection of technical solutions consisting of one or more information processing devices, terminal devices and communication connections between them, used to perform functions such as statement detection, prompt statement generation, model invocation, result processing and display control.

[0300] "User" refers to a human user who operates a terminal device, selects statements on the display interface, browses or utilizes information output provided by the system.

[0301] "Terminal device" refers to an electronic device used to present an interface to a user and receive user input, including but not limited to mobile communication terminals, portable computing devices, desktop computing devices, or other devices with display and network communication capabilities.

[0302] "Display device" refers to an output device that forms part of or is connected to a terminal device for presenting information to a user in graphic or textual form, including but not limited to liquid crystal displays, organic light-emitting displays, or electronic paper displays.

[0303] "Display interface" refers to the interface screen presented by the display device of a terminal device for displaying text, images and interactive controls, which includes statement content that can be selected by the user.

[0304] A “statement” refers to a sequence of text that is presented in text form on the display interface and can be selected by the user as a whole. It includes words, phrases, sentences, or other text units with semantic meaning.

[0305] "Selection statement" refers to the target statement that the user selects on the display interface through clicking, long-pressing, dragging, or other interactive operations. This statement serves as the input object for subsequent prompt statement generation and semantic information retrieval.

[0306] "Vocabulary data" refers to a collection of data used to store entries, sentences, and their related meanings, usage instructions, etc., which can be implemented in the form of databases, dictionary files, or other data structures.

[0307] "Semantic information" refers to the semantically related structured or unstructured information obtained by retrieving selected statements based on lexical data, such as meaning, definition, explanation, or usage description.

[0308] A "string template" is a pre-defined string structure that contains a fixed text portion and can be used to insert variable placeholders. It is used to generate a prompt string after inserting the selected statement.

[0309] "Prompt string" refers to a character sequence generated based on a string template, containing the selected statement, and used as input to a generative artificial intelligence model to instruct the model to generate response information related to that statement.

[0310] "Prompt statements" refer to the text content that prompts are presented as input to a generative artificial intelligence model. Their purpose is to clarify to the model the expected topic, scope, or style of the response.

[0311] "Generative AI models" refer to AI models that can automatically generate response text or other forms of output based on input text, including large-scale language models or other machine learning-based generative models.

[0312] "Information processing device" refers to a computing device that has components such as processor, memory and communication interface, and is used to execute programs to perform functions such as prompt statement processing, model calling and result generation. It may include servers or cloud computing nodes.

[0313] "Response information" refers to the text content output by the generative artificial intelligence model based on the prompt statement, which is used to explain, clarify, or expand the selected statement.

[0314] "Communication path" refers to the physical or logical connection used to transmit data between a terminal device and an information processing device, including wired networks, wireless networks, and the communication protocol stack running on them.

[0315] "Maximum character count" refers to the maximum number of characters that can be preset or dynamically adjusted for response information and / or displayed text. It is used to control the length of the output content to adapt to the display environment and user reading needs.

[0316] "Simplification processing" refers to length control operations performed on response information based on the maximum number of characters, including deleting excess parts, adding ellipses, and other processing methods used to compress text length.

[0317] "Recorded information" refers to the data formed by storing prompt strings, response information, and related context parameters during system operation, which is used to reflect historical interactions and processing results.

[0318] "Generation conditions" refers to the set of parameters or rules used to generate the prompt string, including string template selection rules, formatting rules for the inserted content, and context-related control parameters.

[0319] "Terminal device type" refers to terminal classification information that distinguishes different hardware forms or operating environments, such as mobile terminal, tablet terminal, desktop terminal or other display terminal. Its type will affect the upper limit of the number of characters and the display layout strategy.

[0320] "Size of display area" refers to the range of pixels or logical dimensions of the specific area in the terminal display interface used to present response information, which is used to determine the text layout method and character count control strategy.

[0321] The embodiments of this invention will be described in conjunction with the hardware structure, software modules, data structure, and the internal working mechanism of the generative artificial intelligence model. In the following description, the subject is limited to "server," "terminal," or "user."

[0322] I. Overview of System Overall Structure A server consists of a processor, main memory, non-volatile storage media, and a network interface. The server runs an operating system (such as a general-purpose server operating system) and on it run application services, network communication modules, and generative artificial intelligence model inference modules. The server stores lexical data, recorded information data, and parameter configurations for model invocation.

[0323] The terminal consists of a processor, memory, a touchscreen display, a wireless communication module (such as a Wi-Fi or cellular communication module), and a graphics processing unit. The terminal runs a terminal operating system (such as a mobile operating system or a desktop operating system) and runs user interface applications on it. The terminal communicates with the server via a network interface.

[0324] Users browse text content through the terminal's display interface and select target sentences via touch. The terminal generates prompts based on the user's selections and interacts with the server to obtain response information based on a generative artificial intelligence model, as well as semantic information from lexical data.

[0325] II. Implementation Modes of Terminal-Side Program Processing 1. Terminal user interface and statement retrieval The terminal displays text content through a display device, which can be a liquid crystal display (LCD) or an organic light-emitting diode (OLED). The terminal enables text selection functionality on the text component, utilizing the text selection event interface provided by the operating system to obtain the sentence selected by the user on the display interface.

[0326] The terminal stores the selected statement as a string in memory and uses it as input data for generating subsequent prompts and constructing requests. The terminal can store this string in a structured object, which may include fields such as the original text, start position, end position, and context paragraph identifier.

[0327] 2. Generation of terminal prompts The terminal stores multiple string templates locally, which constitute a rule base for generating prompt statements. The terminal selects a template from the multiple string templates based on the category of the selected statement (such as technical terms, common nouns, phrases, etc.) and the attributes of the currently displayed content (such as news, technical documents, or educational materials).

[0328] The terminal uses string concatenation and placeholder replacement operations to insert the selected statement into the corresponding template. For example, the terminal can generate the following prompt statement: Please briefly explain in no more than 80 words: What is a 'generative artificial intelligence model'? Please explain in no more than 80 words: What is the role of 'prompt statements' when using generative artificial intelligence models? Please describe the basic concepts and a typical application of "blockchain" in no more than 100 words. When generating prompts, the terminal can preprocess the selected statement, including removing extra spaces, removing control characters, and limiting the maximum length, to avoid increasing the model load due to excessively long input.

[0329] 3. Terminal data packaging and transmission The terminal uses the network communication module to construct request data. The terminal combines information such as prompts, terminal type (e.g., mobile phone, tablet, desktop computer), and display area size (e.g., width and height pixel values) into structured data and sends it to the server through the network protocol stack.

[0330] While sending the request, the terminal displays a "Request processing" indicator on the screen to inform the user that background processing is in progress. The terminal does not generate a local interpretation before receiving the server's response; instead, it waits for the server's result, which is a result of comprehensive processing based on a generative artificial intelligence model and vocabulary data.

[0331] 4. Terminal response display and character count control After receiving the response information from the server, the terminal performs a simple character count check locally. Considering the actual available space in the current display area, it may adjust the length of the displayed string again. For example, on smaller devices, the terminal may compress the display length to 50-80 characters, while larger devices may allow for longer text to be displayed.

[0332] The terminal displays the selected statement as the title, followed by semantic information derived from lexical data and response information generated by a generative artificial intelligence model. The terminal uses graphical interface controls to layout the text, ensuring that key parts are displayed preferentially within the visible area. This layout and length control constitute a technological improvement over the traditional "full-text scrolling display" method, thereby reducing the amount of irrelevant content users need to browse and decreasing the number of graphics rendering and scrolling operations.

[0333] III. Implementation Forms of Server-Side Program Processing 1. Server request reception and parsing The server receives request data from the terminal via a network interface. Within the application, the server uses a parsing module to extract the following elements from the request data: prompt message, terminal type information, display area size, and selected message identifier. The server stores these elements in a request context object in memory, which may contain fields such as: session identifier, timestamp, prompt text, and terminal characteristic parameters.

[0334] The server uses this context object as the basic data unit for subsequent data transfer between modules, avoiding repeated parsing and data copying, thereby improving processing efficiency.

[0335] 2. Server-side vocabulary data retrieval The server stores vocabulary data on storage media, which can be implemented using a relational database, key-value database, or in-memory database. Based on the selected statement, the server performs a query operation on the vocabulary data to retrieve information such as meaning entries, definition text, and usage examples related to that statement.

[0336] The server caches search results in memory in a structured format, storing them in fields such as: term identifier, brief definition text, and detailed definition text. The server can simplify the search results, extracting only the brief definition for quick display.

[0337] 3. Server-side generative artificial intelligence model structure and invocation The server loads a generative artificial intelligence model, which can employ a deep neural network based on a transformer architecture. The model loaded by the server consecutively contains multiple self-attention layers, feedforward network layers, and normalization layers, and is connected to a language modeling head at the output to predict the probability distribution of the next word or character.

[0338] The server employs a pre-trained model with fixed parameters during inference, the parameters of which are obtained during the pre-training phase. The pre-training phase utilizes large-scale text data and an autoregressive language modeling objective function, namely minimizing the cross-entropy loss for predicting the next word. During training, the server uses gradient descent and its variants (such as Adam) to update the weight and bias parameters in the network through backpropagation. After pre-training, the server can be fine-tuned for domain-specific corpora. This fine-tuning phase also uses cross-entropy loss as the error function and updates parameters with a small learning rate to improve the model's accuracy in specific domains.

[0339] During inference, the server converts the prompts into a discrete token sequence, mapping each token to a vector representation using a vocabulary and embedding matrix. In each self-attention module, the server calculates attention weights based on the query, key, and value vectors, obtains the context representation using a weighted sum, and then performs a non-linear transformation through a feedforward layer. The server achieves high-level semantic abstraction through multi-layer stacking, ultimately generating the word probability distribution for each position at the output layer. The server then progressively generates the response text according to a predefined sampling strategy (e.g., greedy decoding or beam search decoding).

[0340] When a server invokes a generative AI model, it uses prompts as input and can set parameters such as temperature, maximum generation length, and penalty to control the randomness and concentration of the output. For example, the server can set a lower temperature and a shorter maximum generation length to make the responses more stable and concise.

[0341] 4. Calculation and simplification of server character limit The server dynamically calculates the maximum number of characters based on the terminal type and display area size. The server can preset a set of mapping rules, such as limiting 80 characters for mobile devices, 120 characters for tablets, and 160 characters for desktop devices, and further fine-tune them according to the specific pixel width of the display area and the font size.

[0342] The server stores the character limit in the request context object. After generating the response text, the server compares the response text length with the character limit. If the response text length exceeds the limit, the server performs a truncation operation, retaining the first few characters and adding "..." as an ellipsis marker at the end. The server can prioritize truncation at word boundaries to avoid creating incomprehensible breakpoints, thereby improving readability.

[0343] After simplifying the text, the server packages the simplified response text along with the concise semantic information obtained from the vocabulary data to form an output object. The server then sends this object to the terminal. Because the server performs length control and truncation during the model output stage, the terminal does not need to perform complex scrolling and pagination on extremely long texts, thereby reducing the rendering burden and storage usage on the terminal and improving the overall response speed.

[0344] 5. Server record information storage and adaptive strategy updates The server stores the prompts, responses, character limits, terminal type, and simple user feedback (such as whether to request "more information" again) for each interaction as log information. The server can periodically analyze this log information in the background to determine the most suitable length distribution for different terminal types and the effectiveness of different prompt templates.

[0345] Based on the analysis results, the server updates the conditions for generating prompts. For example, it adjusts the template content to more explicitly limit the number of characters or specify the answer format (such as "listing two application scenarios in bullet points"), and resets the character limit range for different terminal types. Through this adaptive strategy, the server makes the subsequent generation process more suitable for different device conditions and user reading habits, thus achieving continuous optimization of computer system performance, rather than simply automating human behavior.

[0346] IV. Training and Technical Effects of Generative Artificial Intelligence Models During model training, the server utilizes a large-scale text corpus and employs data augmentation techniques such as random masking, sentence shuffling, and synonym substitution to enhance the model's robustness to diverse inputs. The server uses a mini-batch training approach, calculating the cross-entropy loss between the predicted output and the next true word in each batch, and updating network weights through error backpropagation. The server can leverage techniques such as gradient pruning and learning rate annealing to improve training stability.

[0347] This generative AI model, based on a deep transformer structure and large-scale pre-training, far surpasses traditional rule-based or simple statistical models in its semantic understanding and generation capabilities. In this invention's system, the server inputs structured prompts, allowing the model to more explicitly focus on constraints such as "definition," "brief description," and "word count" in a high-dimensional vector space. This results in more compact and relevant responses with the same computational resources, improving computational efficiency and output quality.

[0348] V. Explanation of Technical Improvements and Causal Relationships The collaborative processing between the server and the terminal in this invention does not merely replace manual queries, but rather brings about specific improvements in computer technology through the following technical means: 1. By generating prompts through templates on the terminal, the server receives structured input that clearly indicates the range and length of the answer, thereby reducing unnecessary exploration of the model in the search space and improving computational efficiency in the inference phase.

[0349] 2. The server dynamically calculates the maximum number of characters based on the terminal type and display area, and performs truncation and simplification processing on the server side. This reduces the burden on the terminal side for scrolling rendering and cache management of long texts, reduces graphics rendering and memory usage, and thus improves the overall response speed.

[0350] 3. While performing model inference, the server also performs vocabulary data retrieval, resulting in a combined output of "dictionary definition + model-generated explanation". Through this approach, the server can cover key semantics and auxiliary explanations with shorter texts, indirectly reducing redundant content generated by the model and the amount of data transmitted over the network, thus lowering the communication load.

[0351] 4. The server uses the recorded information to adaptively update the conditions for generating prompt statements and the upper limit of the number of characters. This allows the system to gradually select more effective templates and length strategies after multiple interactions, thereby reducing repeated attempts and invalid generation, and optimizing the allocation of computing resources.

[0352] 5. The server employs a model training approach that combines pre-training and fine-tuning, utilizing data expansion, error function optimization, and weight update strategies to achieve higher language understanding and generation accuracy with limited computing resources. The model internally constructs contextual dependencies through a self-attention mechanism, exhibiting higher information compression and reasoning capabilities compared to traditional methods based on fixed windows or manual rules.

[0353] Therefore, through comprehensive improvements to various aspects such as terminal and server program control flow, data structure design, internal structure and training methods of generative artificial intelligence models, and prompt statement generation and length control strategies, the present invention enables the entire system to achieve improved accuracy, faster response, optimized resource utilization, and reduced communication and rendering load in natural language information generation and display tasks, demonstrating a substantial improvement to computer technology itself.

[0354] use Figure 13 The processing flow is explained.

[0355] Step 1: Users select statements on the terminal's display interface.

[0356] Users can select a text as the target sentence in the terminal's text display area by clicking, long-pressing, or dragging on the touchscreen.

[0357] Input: The complete text content displayed in the interface.

[0358] Output: The selected target statement text, along with its start and end positions in the original text.

[0359] Based on the text selection interface provided by the operating system, the terminal obtains the offset of the selected range, extracts the corresponding substring from the original text buffer in memory, stores the substring as the target statement string in an internal variable, and records the context paragraph identifier corresponding to the statement.

[0360] Step 2: The terminal preprocesses the selected statement.

[0361] The terminal performs cleaning and normalization processing on the target statements obtained in step 1.

[0362] Input: target statement text, start and end positions, context paragraph identifiers.

[0363] Output: Preprocessed standardized statement text.

[0364] The terminal performs string operations on the statement text, such as removing leading and trailing spaces, removing redundant newline characters and control characters, and compressing consecutive spaces into single spaces; the terminal judges the length of the statement, and if it exceeds the predefined maximum length (e.g., 50 characters), the terminal truncates it to a specified length according to word boundaries to prevent subsequent prompts from being too long; the terminal saves the processed standardized statement in a new variable as input to the prompt generation module.

[0365] Step 3: The terminal selects a string template based on the statement type and interface attributes.

[0366] The terminal analyzes the type of the selected statement and the attributes of the currently displayed content, and selects a suitable prompt statement template from the local template library.

[0367] Input: Standardized statement text, content type (e.g., technical documents, news, teaching materials), and terminal type information.

[0368] Output: The selected string template.

[0369] The terminal can determine whether a statement contains technical terms or is a proper noun through simple rules or classification tags, and select from multiple templates based on this determination, such as "definition class template" or "function description template". The terminal reads the corresponding template text from memory, such as "Please briefly explain in no more than 80 words: What is '{statement}'?", and uses this template as the basis for subsequent string concatenation.

[0370] Step 4: The terminal generates a prompt message.

[0371] The terminal uses string concatenation and placeholder replacement to insert standardized statements into the selected template, generating prompts for generative artificial intelligence models.

[0372] Input: Standardized statement text, selected string template.

[0373] Output: The complete prompt text.

[0374] The terminal locates the placeholder in the template within the program, replaces the placeholder with a standardized statement, for example, replacing "{statement}" with "generative artificial intelligence model", and generates a specific prompt statement: Please briefly explain in no more than 80 words: What is a 'generative artificial intelligence model'? or Please explain in no more than 80 words: What is the role of 'prompt statements' when using generative artificial intelligence models? The terminal stores the generated prompt statement in the request data structure, ready to send it to the server.

[0375] Step 5: The terminal constructs and sends a request to the server.

[0376] The terminal packages the prompt message and terminal-related parameters and sends them to the server via the network interface.

[0377] Input: Prompt text, terminal type, display area size, session identifier.

[0378] Output: The request data packet sent to the server.

[0379] The terminal creates a request object, adding fields such as the prompt statement, terminal device type (e.g., mobile phone, tablet), and display area width and height to the request data structure; the terminal uses the network protocol stack to send the request to the server's specified service address via a wireless communication module (e.g., Wi-Fi or mobile network); the terminal records the request time when sending and displays a loading indicator on the interface to prompt the user that the system is processing the request.

[0380] Step 6: The server receives the request and parses the parameters.

[0381] The server receives data sent by the terminal from the communication interface and parses out the prompts and related parameters.

[0382] Input: A request data message from the terminal.

[0383] Output: A request context object containing prompts, terminal information, and display area parameters.

[0384] The server calls the parsing module in the application to parse the request message and extract fields such as the prompt text, terminal type, display area size, and session identifier. The server organizes these fields into an internal request context object and stores it in the main memory for use by the subsequent vocabulary retrieval and model calling modules.

[0385] Step 7: The server retrieves semantic information from the vocabulary data.

[0386] The server performs a query in the vocabulary data based on the selected statement or related identifier to obtain semantic information.

[0387] Input: The text of the selected statement or its identifier, and vocabulary data storage.

[0388] Output: A brief definition, extended explanation, and other semantic information corresponding to this statement.

[0389] The server uses the selected statement (or its encoding) contained in the request context as the query condition to perform a retrieval operation in a relational database or key-value database; the server extracts the main definition field and auxiliary description field from the matching record and encapsulates them into a semantic information object; the server can perform an initial truncation of lengthy definitions, retaining only the main definition sentence to reduce the burden of subsequent transmission and display.

[0390] Step 8: The server calculates the maximum number of characters and generates length control parameters.

[0391] The server determines the maximum number of characters in the response message based on the terminal type and the size of the display area.

[0392] Input: Terminal type, display area size, preset mapping rules.

[0393] Output: Maximum number of characters in the response message.

[0394] The server follows preset rules, such as setting an 80-character limit for mobile terminals, a 120-character limit for tablet terminals, and a 160-character limit for desktop terminals. The server further fine-tunes this limit value based on the display area width and font size to obtain the precise maximum number of characters. The server writes this character limit into the request context object as a constraint parameter for subsequent response simplification processing.

[0395] Step 9: The server calls a generative artificial intelligence model to generate response information.

[0396] The server takes the prompt as input, calls a generative artificial intelligence model based on a transformer structure to perform reasoning, and generates a response text.

[0397] Input: Prompt text, model parameters (including temperature, maximum generation step size, etc.), and maximum number of characters.

[0398] Output: The original response text generated by the model.

[0399] The server first segments the prompt statement into words or sub-words, converting the text into a token sequence. The server then maps these tokens to vector representations using an embedding matrix, inputting these vectors into a multi-layer self-attention network. In each layer, the server calculates the query, key, and value vectors, determines the attention weights, and obtains the context representation. After passing through a feedforward network and a normalization layer, the output is the next layer's input. In the output layer, the server sequentially selects or samples the next token based on a probability distribution, repeating this process until the maximum step size is met or a termination token is encountered. Finally, the server converts the generated token sequence back into text, forming the complete response content.

[0400] Step 10: The server performs simplification and truncation on the response information.

[0401] The server determines the length and simplifies the original response text generated by the model based on the maximum character limit.

[0402] Input: Original response text, maximum number of characters.

[0403] Output: Simplified response text.

[0404] The server counts the number of characters in the original response text. If it is less than or equal to the maximum number of characters, it is directly output. If it exceeds the maximum number of characters, the server truncates the text from the beginning to near the maximum number of characters without breaking word boundaries, and adds “…” as an ellipsis at the end. The server thus obtains a short and complete response text, ensuring that there is no excessive scrolling or line breaks in the terminal display area.

[0405] Step 11: The server combines semantic information with response information and generates response data.

[0406] The server combines the semantic information retrieved from the vocabulary data with the simplified response text and packages it into response data.

[0407] Input: Semantic information object, simplified response text, request context object.

[0408] Output: The response data packet sent to the terminal.

[0409] The server creates a response structure, with semantic information as the first part and the response text as the second part, and adds metadata such as character limit and terminal type if necessary. The server serializes the structure into a transmission format and sends it to the requesting terminal through the network interface. The server also saves the prompt statement, response text and terminal-related information in the record storage to provide a data foundation for subsequent adaptive strategy updates.

[0410] Step 12: The terminal receives the response and parses and displays the content.

[0411] The terminal receives response data from the server, parses out semantic information and response text, and prepares to display it on the interface.

[0412] Input: The response data packet returned by the server.

[0413] Output: Semantic text and response text for display.

[0414] The terminal uses the parsing module to extract semantic information text and response text from the response data and stores them in the interface state object. The terminal reconfirms the text length based on its current display area and performs a very small amount of truncation if necessary to accommodate special fonts or user-defined font size settings. The terminal determines the text layout order, uses the selected statement as the title, and arranges the semantic information and response information according to the set format.

[0415] Step 13: The terminal displays the results on the screen for the user to read.

[0416] The terminal renders the processed semantic information and response text onto the display device for the user to view.

[0417] Input: semantic text, response text, selected statement text, display layout parameters.

[0418] Output: An explanation interface displayed on the terminal screen.

[0419] The terminal invokes the graphical user interface rendering component to display the selected statement as a highlight or title, followed by a brief explanation from the vocabulary data, and then a simplified response output by the generative artificial intelligence model. The terminal controls line breaks and font size based on the number of characters, terminal orientation (vertical or horizontal), and screen size to ensure that key content is within the current visible area. By reading this interface, users can quickly obtain a dual-channel explanation of the selected statement.

[0420] Step 14: Users can then engage in further interactions as needed.

[0421] After reading the results, users can choose to close, copy, or request more detailed information.

[0422] Input: The explanatory interface and interactive controls displayed on the screen.

[0423] Output: New user commands (e.g., "More information", "Regenerate", etc.).

[0424] If users require more detailed information, they can click the "More Information" button. The terminal captures this interaction event and generates a new prompt based on the original selected statement, indicating a more detailed answer, such as: "Please explain in detail: the principles of generative artificial intelligence models and three typical applications, in no more than 300 words." The terminal then constructs a request again and sends it to the server, repeating the above processing flow. This allows for hierarchical and controllable length interactive information retrieval without changing the overall data structure and control logic.

[0425] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0426] In existing human-computer interaction systems, when users encounter unfamiliar words or phrases on the terminal's display interface, they typically need to manually switch to a separate search application or online search service, requiring multiple inputs and redirects to obtain explanatory information. This approach not only increases the number of interaction steps and reduces the continuity of conversations or transactions, but also makes it difficult to provide users with aggregated, multi-level explanatory information in a timely manner within the same interface.

[0427] Furthermore, traditional systems often rely solely on static dictionary data or fixed templates for interpretation, lacking the ability to dynamically generate natural language responses based on context using generative artificial intelligence models. They also fail to tailor and focus the explanatory content according to the user's interface type (e.g., communication interface, transaction interface) and current operating state, resulting in mismatches between the generated explanatory information and the actual task, thus affecting information acquisition efficiency.

[0428] Furthermore, existing solutions typically ignore the user's emotional state during the conversation and fail to perform sentiment analysis on the content of the user's messages on the communication or transaction interface. As a result, they cannot automatically adjust the style and depth of the prompts based on the user's emotions (such as anxiety, confusion, fatigue, etc.). Consequently, the generated responses are either overly technical and difficult to understand, or verbose, information-overloaded, and unable to adapt to the cognitive load and emotional state of different users.

[0429] Meanwhile, for long text responses returned by generative AI models, traditional systems often use simple truncation for length control, failing to implement structured summaries and segmented display strategies based on character limits on the server side. This prevents hierarchical organization and interactive expansion between "summary information" and "detailed information." This not only leads to a chaotic information layout on the display interface but also increases the burden on terminal rendering and user cognition, reducing overall human-computer interaction performance.

[0430] In summary, how can we achieve this within the same display interface without interrupting the current application scenario? (1) Automatically combine the user-selected string with the vocabulary information in the storage device and the generative artificial intelligence model for processing; (2) Perform emotion recognition on the user's speech on the server side, and dynamically adjust the prompts for the generative artificial intelligence model accordingly; (3) Implement character-limit-based summary and segmentation display control on the generated results. This will improve information retrieval efficiency and user experience at both the system architecture and interaction process levels, becoming a pressing computer technology problem that needs to be solved in this field.

[0431] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0432] In this invention, the server includes: a device for retrieving meaning information from a string selected by a user on a terminal display interface based on vocabulary information stored in a storage device connected to an information processing device; a device for encapsulating the meaning information into display information and sending it to the terminal for direct presentation on the current display interface; a device for generating a prompt statement to instruct a generative artificial intelligence model to generate a natural language response based on the selected string and additional information about the terminal's operating status; a device for sending the prompt statement to the generative artificial intelligence model and receiving response information related to the string; a device for parsing user speech content obtained on a communication interface or transaction interface to determine the user's emotional state; a device for adjusting the style and explanatory hierarchy of the prompt statement according to the determined emotional state, thereby changing the content and style of the response information output by the generative artificial intelligence model; and a device for controlling the length of the response information and removing irrelevant information according to a preset character limit, generating display response information optionally divided into summary information and detailed information, and sending it to the terminal for layered display on the display interface. This allows for integrated processing on the server side, encompassing vocabulary retrieval, sentiment recognition, prompt generation, response information summarization, and segmented control. Without interrupting the current application scenario, it provides the terminal with short text explanations and expandable detailed descriptions that adaptively match the interface type and user emotions. This reduces the interaction overhead of users searching across applications, lowers the rendering and computational burden on the terminal side, and improves the overall computer technology performance of the human-computer interaction system in terms of information acquisition efficiency and usability.

[0433] "System" refers to an entirety consisting of one or more information processing devices, terminals, and communication networks therebetween, and is a collection of technical devices used to perform the various functional processes of this invention.

[0434] "Information processing device" refers to an electronic device, including a processor and a storage device, used for acquiring, storing, calculating, controlling, and outputting data.

[0435] "Terminal" refers to an electronic device operated by a user for displaying information and receiving user input, including but not limited to smartphones, tablets, computers, or other devices with display interfaces and communication functions.

[0436] "Display interface" refers to the user interface area on a terminal that presents text, images, or other visual information, including communication interfaces, transaction interfaces, and other application interfaces.

[0437] A "communication interface" refers to an interface used to display and input messages and dialogue content, and is used for text or multimedia interaction between users or between users and the system.

[0438] "Transaction interface" refers to the interface used to display product information, transaction information, payment information, etc. during electronic transactions or online services.

[0439] A "string" is a unit of text consisting of one or more consecutive characters, including words, phrases, or sentences.

[0440] "Vocabulary information" refers to the semantic information, definition information, usage information, or other data related to the meaning of a string.

[0441] "Storage device" refers to a storage medium used to store vocabulary information, user data, model configurations, and program instructions, including semiconductor memory, magnetic memory, optical storage media, or cloud storage resources.

[0442] "Meaning information" refers to the definition, description, or example information obtained by retrieving vocabulary information, which is used to explain the meaning of the string selected by the user.

[0443] "Display information" refers to formatted information that is processed by the server and used to present to the user on the terminal display interface, including semantic information and response information.

[0444] "Generative artificial intelligence models" refer to artificial intelligence models that are based on statistical learning or deep learning techniques and can automatically generate natural language text and other outputs based on input prompts.

[0445] "Prompt statements" refer to the natural language input text sent to a generative artificial intelligence model, which instructs the model to generate corresponding response information for a specific object or task.

[0446] "Running status" refers to the operating environment and context information of the terminal at a specific point in time, including the type of application currently running, the type of interface, network status, or session status.

[0447] "Additional information" refers to environmental or contextual information associated with the string selected by the user, including interface type, application scenario, time information, or user interaction history.

[0448] "Response information" refers to the natural language text output by the generative artificial intelligence model based on the prompt statement, which is used to answer questions or explain strings.

[0449] "User-generated content" refers to the text content that users input or send in the communication or transaction interface to express their needs, emotions, or other information.

[0450] "Emotional state" refers to the category and intensity of a user's emotions inferred from the user's spoken content or behavioral data, including but not limited to states such as happy, sad, angry, tense, confused, or neutral.

[0451] "Style" refers to the expressive style or tone characteristics of natural language text, including linguistic style attributes such as formality, friendliness, comfort, and persuasion.

[0452] "Explanation level" refers to the level of detail or abstraction used when explaining a target object, including different granularities such as overview level explanation, basic level explanation, and detailed level explanation.

[0453] "Character limit" refers to the maximum number of characters that can be constrained in the length of text to avoid displaying excessively long or redundant content.

[0454] "Irrelevant information" refers to content in the generated response that is not directly related to the user's current query goal or task, including redundant background information, repetitive expressions, or paragraphs unrelated to the current scenario.

[0455] "Response information for display" refers to text information that is suitable for presentation on the terminal display interface after the response information output by the generative artificial intelligence model has been processed by length control, filtering and formatting.

[0456] "Summary information" refers to a simplified description extracted from the response information that covers the main meaning and key points in a small number of characters.

[0457] "Detailed information" refers to supplementary explanations that go beyond the summary information, providing a more comprehensive and detailed elaboration on related concepts, reasons, examples, etc.

[0458] "Phase-based display" refers to a display control method that initially displays only summary information, and then gradually displays corresponding detailed information based on additional user actions.

[0459] In the following embodiments, the server, terminal, and user are described as the main entities. This invention is not limited to a specific hardware or software platform; any electronic device capable of implementing the described functional modules can be used.

[0460] I. Overall System Composition The server can be a computer device configured with a multi-core central processing unit, graphics processing unit, main memory, and large-capacity storage device, and the operating system can be a general-purpose server operating system. The server communicates bidirectionally with multiple terminals through a network interface. The server internally includes: a vocabulary retrieval module, a sentiment recognition module, a prompt statement generation module, a generative artificial intelligence model invocation module, a response summarization and segmentation module, a communication management module, and a logging and monitoring module, etc.

[0461] The terminal can be a smartphone, tablet, or personal computer, running a general-purpose operating system (such as a mobile or desktop operating system). Applications running on the terminal can be browsing, messaging, or transaction applications, used to present content on the display interface and receive user selections and inputs. The terminal communicates with the server via cellular networks, wireless LANs, or wired networks.

[0462] Users interact with the terminal via touchscreen, keyboard, or pointing device, selecting unfamiliar strings on reading pages, communication interfaces, or transaction interfaces, or entering messages to express their current feelings and needs.

[0463] II. Server-side module structure and data structure The server divides its functions into multiple logical modules and stores them in the form of program instructions and configuration data in the storage device. The server uses a message-oriented inter-module communication method, and the modules exchange data through memory queues or remote procedure calls.

[0464] 1. Vocabulary Search Module The server stores a vocabulary dataset in its storage device. This dataset can be stored as a relational database table or a key-value store structure, where each record includes fields such as "lexicon identifier," "string," "basic definition," and "example sentence." After receiving a user-selected string uploaded by the terminal, the server uses that string as a search key to perform a query in the vocabulary dataset and returns the corresponding meaning information.

[0465] The server internally standardizes the search results, such as by standardizing encoding formats, removing control characters, and regulating punctuation. The server then encapsulates the processed definition text as part of the information to be displayed and subsequently sends it to the terminal.

[0466] 2. Emotion Recognition Module The server stores users' most recent messages in a log database or session storage. When it needs to identify sentiment states, the server extracts several recent texts from the session records and concatenates them into an input sequence. The server internally deploys a sentiment classification model, which can use bidirectional encoding to represent the structure of the network and classification layers, or it can use other deep learning models.

[0467] During the training phase, the server uses a large number of text samples with sentiment labels to perform supervised learning on the sentiment classification model. The server employs cross-entropy loss as the error function and updates the model parameters using backpropagation and stochastic gradient descent or its variants. To improve robustness, the server can perform data augmentation during training, such as synonym replacement, random word dropping, or sequential perturbation of the input text, thereby enhancing the model's generalization ability to different expressions.

[0468] During the inference phase, the server encodes the user's speech sequence into a vector representation, which is then input into the sentiment classification model to obtain probability distributions for multiple sentiment categories. The server determines the user's sentiment state based on the highest probability, such as "sad," "angry," "confused," "excited," or "neutral." The server can also determine the intensity of the sentiment based on the probability value, which can be used to fine-tune the style of prompts.

[0469] 3. Prompt Statement Generation Module The server stores several prompt statement templates in its storage device. These templates exist in the form of structured text, and each template contains placeholder fields such as "target string," "interface type," "emotional constraint," and "word count constraint." After receiving the string information from the terminal and the emotional state output by the emotion recognition module, the server selects a suitable template from the template set and replaces the placeholders.

[0470] When generating prompts, the server incorporates contextual information into the natural language description. For example, when the emotional state is "sad," the server adds a description like "gentle tone, comforting" to the template; when the interface type is "transaction interface," the server adds a conditional description like "electronic transaction in progress" to the template, thus enabling the generative AI model to generate content in a targeted manner.

[0471] The server can generate different types of prompts based on different objectives. For example, to obtain a definition explanation, the server uses the following type of prompt: "Please explain the meaning of 'cryptocurrency' to users who are conducting electronic transactions, using no more than 100 Chinese characters, and the language should be easy to understand." To provide comforting explanations when users are feeling down, servers can use the following types of prompts: "The user is very tired today. Please use a gentle and comforting tone to explain the basic meaning of 'cryptocurrency' to him in no more than 80 Chinese characters, and try to avoid using technical terms." To help users analyze their emotions, the server can use the following types of prompts: "When a user says 'I'm really angry,' please use a gentle tone to help the user analyze 'why I'm angry' within no more than 120 Chinese characters, and briefly explain with 2-3 common reasons." When generating prompts, the server doesn't simply concatenate fixed text. Instead, it selects rules and adjusts weights based on multiple features such as emotional state, interface type, and string length. For example, the server can maintain a set of weights within the prompt generation module, representing the priority of the explanation level under different emotions. When the user is fatigued, the server increases the weight of "brief explanation," prioritizing short templates and avoiding generating complex explanations, thereby reducing the display burden on the terminal and the reading burden on the user.

[0472] 4. Generative Artificial Intelligence Model Calling Module In this embodiment, the server can connect to external large-scale natural language model services via an interface, or deploy generative artificial intelligence models on local graphics processing hardware. The model can employ a multi-layered self-attention network structure, including embedding layers, self-attention encoding layers, and decoding layers. The model parameters are pre-trained using a large-scale corpus and can then be further fine-tuned using instruction data to adapt it to prompt-driven response generation tasks.

[0473] When the server invokes the generative AI model, it encodes the prompts into an input sequence and specifies the maximum output length, randomness control parameters, etc. After receiving the output, the server uses the model-generated response as the raw result for subsequent modules to process. Because the prompts are pre-constrained by the server within a specific word count and expression style, the model-generated results are more controllable in terms of length and topic focus, thereby reducing the number of times the server truncates the data, reducing the amount of communication data, and improving overall processing efficiency.

[0474] In another implementation, the server can employ a smaller generative language model for edge deployment. In this case, the server reduces model complexity by decreasing the number of layers or hidden dimensions, and compresses parameters using methods such as quantization or pruning to accelerate inference in resource-constrained environments, thereby reducing response time while maintaining sufficient interpretation quality.

[0475] 5. Response Summary and Segmentation Module After receiving the response information output by the generative artificial intelligence model, the server performs a length check on the text. The server calculates the text length based on a pre-set character limit; if the length exceeds the limit, it performs a summary calculation. The server can use an extractive summarization algorithm, such as vectorizing sentences and selecting several key sentences as summary information based on similarity and importance scores; or it can use a summarization model based on an encoder-decoder structure to generate short text.

[0476] When generating summary and detailed information, the server maintains corresponding data structures, marking the summary information as "Level 1 content" and the supplementary information as "Level 2 content." When responding to the terminal, the server only marks the summary information as to be displayed by default, while marking the detailed information as to be displayed later. This segmentation strategy allows the terminal to render only a limited number of characters in the initial stage, reducing the layout complexity and rendering time of the display interface. When the user performs additional operations on the terminal (such as clicking "Expand More"), the terminal then requests or expands the detailed information, thus achieving on-demand loading and phased display.

[0477] The server employs the aforementioned summarization and segmentation mechanism, which reduces the amount of data transmitted per transaction and the initial rendering load on the terminal screen in high-concurrency request scenarios, thereby improving overall performance in both computing and communication aspects.

[0478] III. Terminal-side display and interaction In this invention, the terminal is responsible for collecting user input and displaying results. The terminal captures the string selected by the user on the display interface through the selection interface provided by the operating system, encapsulates it, and sends it to the server. After receiving the display information returned by the server, the terminal displays summary information on the current interface as a floating box, bottom drawer, or embedded area, and provides interactive controls for the user to request or expand detailed information.

[0479] During the display process, the terminal can adjust the interface style based on the tagging information returned by the server. For example, when the server indicates that the response is of the emotion-adaptive type, the terminal can use soft colors or prompt icons to help users distinguish between general definition descriptions and emotion-care descriptions. When displaying a large number of responses, the terminal can use a local caching strategy to save recent descriptions, reducing repeated requests for the same string, thereby further reducing server pressure and network load.

[0480] In another implementation, the terminal can run a lightweight emotion recognition component locally to quickly and roughly assess the user's input, then send the assessment result to the server. The server then combines this assessment with a more refined emotion model to make a comprehensive decision. This distributed emotion processing architecture can reduce the computational burden on the server side and shorten the overall latency of emotion state estimation.

[0481] IV. User-side usage examples When a user browses products in a shopping app, they see the phrase "Supports cryptocurrency payment" below the product details. Unfamiliar with this term, the user selects it by long-pressing the string. The terminal captures this selection event and sends the selected string along with the interface type information to the server.

[0482] The server retrieves the basic definition of "cryptocurrency" from the vocabulary database, sends the definition text back to the terminal as semantic information, and simultaneously determines the user's emotion as "tired / sad" based on the interface being a "trading interface" and the user's recent message "I'm really tired today," and generates prompts with tone requirements and character limits accordingly, such as: "The user is very tired today. Please use a gentle and comforting tone to explain the basic meaning of 'cryptocurrency' to him in no more than 80 Chinese characters, and try to avoid using technical terms." The server inputs the above prompt into the generative AI model and obtains a brief description. If the generated content is slightly longer, the server compresses it to the character limit using a summarization and truncation mechanism. For example, the final description would be: "A cryptocurrency is a digital currency that exists only on a network. It uses distributed technology to record transactions, does not rely on bank intermediaries, and is often used for online payments or transfers." The server returns this explanation as a display response to the terminal. The terminal displays this explanation in a pop-up window on the current transaction interface, allowing the user to understand the relevant terminology without leaving the current page. Because the server has streamlined the explanation's objective when generating the prompt, the text generated by the generative AI model is more focused in length and topic, thus reducing irrelevant information and improving the readability and accuracy of the explanation.

[0483] V. Technical Effects and Causal Relationships Because the server comprehensively analyzes user sentiment, interface type, and string characteristics before generating prompts, and uses structured templates to dynamically control the level and tone of the instructions, generative AI models can generate responses within a smaller search space after receiving prompts. This reduces redundant content and improves the relevance and stability of the generated results. This strategy of pre-constraining prompts differs from the traditional approach of simply conveying the user's original question. It optimizes the computational path at the model inference level and reduces the amount of text requiring post-processing.

[0484] The server employs an emotion recognition model that uses a deep learning architecture to vectorize and classify user speech into multiple categories. It is trained using cross-entropy loss and backpropagation algorithms, resulting in high accuracy in emotion state estimation. When generating prompts, the server uses emotion state as a modulating parameter, allowing the explanatory text to adapt to the user's state in terms of tone and depth. This improves the subjective satisfaction and comprehension efficiency of the output information without requiring additional user intervention.

[0485] The server-deployed summarization and segmentation module employs an algorithm based on text vector representation and importance ranking to divide the long text output by the generative AI model into summary information and detailed information. This phased display reduces the number of characters initially rendered and the number of bytes transmitted over the network. Since the summary information prioritizes core content, users can complete their tasks by simply browsing the summary in most cases, reducing unnecessary reading and scrolling. This structured presentation method is technically significantly different from simple truncation or full-segment display, improving the efficiency of data flow utilization between the server and the terminal.

[0486] By integrating the collaboration among the aforementioned modules, the server in this invention not only automates human word lookup and interpretation processes, but also introduces an emotion-driven prompt generation mechanism, model-based summarization and segmentation strategies, and contextual control through multi-source data fusion at the computational structure level. This allows generative artificial intelligence models to leverage their advantages under controlled conditions, thereby achieving more efficient data processing, more accurate content matching, and better resource utilization. These technical features and implementation methods collectively constitute an improvement to computer technology itself, rather than merely automating business processes.

[0487] use Figure 14 The processing flow is explained.

[0488] Step 1: The user selects a string on the terminal.

[0489] Input: The text content displayed on the terminal screen.

[0490] Output: The string selected by the user and its position information.

[0491] When users browse text on the terminal's communication or transaction interface, and encounter unfamiliar words or phrases, they can select the target string on the touchscreen by long-pressing, double-clicking, or dragging. The terminal then calls the text selection interface provided by the operating system to obtain the range of characters selected by the user, extracts the text within that range into a string, and records contextual information such as the interface type and application identifier of the string.

[0492] Step 2: The terminal sends the selected string and context to the server.

[0493] Input: The string selected by the user, interface type information, application identifier, and user identifier.

[0494] Output: The request data packet sent to the server.

[0495] After obtaining the selected string, the terminal encapsulates it together with the current interface type (e.g., communication interface or transaction interface), application identifier, and user identifier into structured data. The terminal calls the network communication library to serialize the data into a request message and sends it to the server via the communication network using a secure protocol. Data processing includes encoding the string into a unified character encoding format and adding a timestamp and session identifier so that the server can perform session association and security verification.

[0496] Step 3: The server parses and preprocesses the terminal request.

[0497] Input: A request data message from the terminal.

[0498] Output: Standardized strings, interface types, user identifiers, and other internal representations.

[0499] The server receives request messages from terminals via a network interface and uses an application-layer protocol parsing library to parse the message header and body, extracting string fields, interface type fields, and user identifiers. The server performs data cleaning on the received strings, including removing leading and trailing whitespace, standardizing punctuation, and checking for illegal characters, and stores the results in an in-memory data structure to provide standardized input for subsequent module calls.

[0500] Step 4: The server performs a word search to obtain meaning information.

[0501] Input: The normalized string.

[0502] Output: The meaning information corresponding to this string.

[0503] The server uses strings as search keys to access a collection of vocabulary data in storage. Based on the string length and character type, the server selects an appropriate indexing strategy, performs a matching query in the vocabulary, and retrieves the basic definitions and example sentences corresponding to the string. Data computation includes hash-based or tree-structured index lookups, converting the search results into semantic information objects in a uniform format. The server then truncates or formats the results as needed for clear display on the terminal.

[0504] Step 5: The server returns basic information to the terminal.

[0505] Input: The retrieved information object and the original string.

[0506] Output: Basic explanatory data sent to the terminal for display.

[0507] The server constructs response data, which includes the original string and its corresponding semantic information, indicating that it is a basic dictionary definition. The server encodes this data into a response message via its network communication module and sends it back to the requesting terminal. Before sending, the server can check the length of the semantic information to ensure it does not exceed the preset maximum transmission length, thus avoiding invalid data consuming bandwidth.

[0508] Step 6: The terminal displays basic information on the current screen.

[0509] Input: Basic explanation data returned by the server.

[0510] Output: Basic explanatory text displayed on the terminal interface.

[0511] The terminal receives the server's response message and parses out the raw string and its meaning. The terminal creates a floating window or bottom information bar on the current communication or transaction interface, presenting the meaning to the user with an appropriate font and layout. Before displaying the text, the terminal can automatically wrap the text according to the screen size and control the maximum number of lines; any lines exceeding this limit are guided to the user via an "expand" button.

[0512] Step 7: The server retrieves users' most recent posts for sentiment analysis.

[0513] Input: Session logs or spoken text associated with the user ID.

[0514] Output: The integrated sentiment analysis input text.

[0515] The server retrieves several recent conversation records from the user's session database or cache, based on the user's identifier, such as chat messages, comments, or input. The server concatenates these multiple texts into a continuous text sequence in chronological order, truncating the text as needed to ensure the text length remains within the maximum input length of the sentiment recognition model. Data processing includes removing labels and system prompts, allowing sentiment analysis to focus on the user's actual content.

[0516] Step 8: The server uses an emotion recognition model to determine the user's emotional state.

[0517] Input: The integrated sentiment analysis input text.

[0518] Output: User's emotional state category and corresponding confidence level.

[0519] The server inputs the integrated text into a pre-trained sentiment recognition model. This model employs a deep neural network structure, converting the input text into embedding vectors, extracting semantic features through a multi-layer coding network, and making probability predictions for multiple sentiment categories at the output layer. The server receives the probability distribution output by the model and selects the category with the highest probability as the sentiment category, while retaining the confidence value for each category. The server treats the sentiment category and confidence value as a structured sentiment state object for subsequent modules to use.

[0520] Step 9: The server generates prompts based on strings, interface type, and emotional state.

[0521] Input: Standardized string, interface type, emotion state object.

[0522] Output: Prompt text for generative artificial intelligence models.

[0523] The server selects a template from the prompt template library that matches the current interface type and task objective, such as a "definition explanation template" or an "emotional reassurance template." The server adjusts the tone and description depth fields in the template based on the emotional state, replacing standardized strings with placeholders within the template. For example, when the interface is a trading interface and the emotional state is fatigued, the server generates a prompt similar to, "The user is very tired today. Please explain the basic meaning of 'cryptocurrency' to them in a gentle, comforting tone, within no more than 80 Chinese characters, avoiding technical jargon as much as possible." Data processing includes string concatenation, placeholder replacement, and selecting an appropriate description level based on emotional weights to form a complete natural language input text.

[0524] Step 10: The server sends the prompt to the generative artificial intelligence model and receives the response information.

[0525] Input: Prompt text.

[0526] Output: The original response text output by the generative artificial intelligence model.

[0527] The server submits the prompt statement as input to the generative AI model via a model call interface. The server specifies generation parameters, such as the maximum output length and randomness control coefficients, to constrain the generation process. Internally, the generative AI model performs encoding and decoding operations to generate response text based on the prompt statement. After the model's inference is complete, the server receives the complete output string and saves it in memory as the original response information.

[0528] Step 11: The server performs character limit and digest processing on the response information.

[0529] Input: Original response text, preset character limit.

[0530] Output: Display response information (summary information and optional detailed information) that meets the display requirements.

[0531] The server calculates the character count of the original response text. If the length is less than or equal to the character limit, the text is directly used as the summary information. If the length exceeds the limit, the server initiates a summarization algorithm to segment the text into sentences and uses importance scoring to select several key sentences to form the summary information, while the remaining part is used as detailed information. Data processing includes vectorizing sentences, calculating their relevance to the overall text, and ranking them. The server packages the summary information and detailed information into structured objects and marks the summary information as the primary display content.

[0532] Step 12: The server will send a response message to the terminal.

[0533] Input: Display response information object (containing summary information and optional details), raw string.

[0534] Output: Response display data sent to the terminal.

[0535] The server constructs a response message, encapsulating summary and detailed information together, and appending the original string and sentiment matching tags. The server sends this message back to the terminal via the network interface. When constructing the message, the server can choose whether to delay or send the detailed information all at once, depending on the network conditions; the terminal then decides whether to display it immediately.

[0536] Step 13: The terminal displays summary information and provides detailed information expansion controls on the display interface.

[0537] Input: The response data returned by the server.

[0538] Output: A summary description text and user-interactive expand controls displayed on the terminal interface.

[0539] The terminal parses the server response, reading summary information, detailed information, and the original string. The terminal displays the summary information as a floating layer or embedded area in the current interface, binding the detailed information to a "Expand More" or similar control. The user first sees the summary; for further information, they can click the control to view the detailed information. The terminal can adjust the interface style based on emotion-adaptation markers to better match the user's current state.

[0540] Step 14: Users can read the response information and expand on the details as needed.

[0541] Input: A summary description and expand controls displayed on the terminal.

[0542] Output: User behavior feedback (e.g., continue operation, expand details, close window).

[0543] Users read the summary instructions on the terminal. If they understand it sufficiently, they can close the instruction window and continue their current task. If they still have questions, they can expand the controls to view more details. These user actions can be recorded by the terminal and fed back to the server in subsequent sessions to evaluate the sufficiency of the summary information and optimize character limit parameters.

[0544] Step 15: The terminal responds to user requests or displays detailed information.

[0545] Input: Expand action triggered by the user, and details received or pending.

[0546] Output: A detailed description text displayed on the terminal interface.

[0547] When the terminal detects that a user has clicked on controls such as "Expand More," it will directly expand and display the detailed information on the interface if the information is already included in the current data. If the detailed information has not been transmitted, it will send a request to the server to obtain the corresponding detailed description and display it upon receipt. Data processing includes adjusting the text area height and rearranging interface elements to ensure that the summary and detailed information are clearly readable on the screen.

[0548] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0549] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.

[0550] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0551] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0552] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0553] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0554] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0555] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0556] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0557] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0558] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge-Coupled Device) image sensor, to capture images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0559] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0560] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0561] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0562] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0563] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0564] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0565] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0566] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0567] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0568] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0569] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0570] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.

[0571] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0572] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0573] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0574] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0575] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0576] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0577] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0578] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0579] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge-Coupled Device) image sensor, to capture images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0580] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0581] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0582] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0583] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0584] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0585] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0586] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0587] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0588] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0589] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0590] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0591] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.

[0592] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0593] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0594] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0595] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0596] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0597] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0598] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0599] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0600] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge-Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0601] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0602] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0603] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0604] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0605] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0606] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0607] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0608] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0609] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0610] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0611] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0612] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0613] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.

[0614] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0615] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0616] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0617] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0618] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0619] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0620] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0621] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0622] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0623] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0624] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0625] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0626] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0627] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0628] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0629] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0630] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0631] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0632] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0633] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0634] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0635] In addition, the following notes are provided in response to the above explanation.

[0636] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for receiving a character sequence selected by a user on the display screen of an information processing device from a terminal, retrieving the character sequence from a data storage device storing vocabulary information to obtain semantic information, and sending the obtained semantic information to the terminal for display on the display screen. A means for generating a prompt statement for a generative artificial intelligence model based on the selected character sequence received from the terminal, the prompt statement including the selected character sequence and conditional information for limiting the number of words in the response and the style of the description, and a means for sending the prompt statement to the generative artificial intelligence model; An apparatus for receiving response information generated based on the prompt statement from the generative artificial intelligence model, determining the number of characters in the response information according to the limit on the number of characters in the response, and adjusting the response information by deleting parts that exceed a predetermined upper limit; A means for sending the adjusted response information to the terminal and displaying the response information on the display screen in a manner associated with the selected character sequence; A device for maintaining multiple template information for generating the prompt statement and user settings information related to the answer character limit and description style, and for selecting the template information according to the user settings information, thereby dynamically changing the content of the prompt statement and the answer character limit; An apparatus for recording, in chronological order, the selected character sequence, the prompt statement, the response information, and historical information related to the response character limit, and for updating the template information or the response character limit based on the recorded historical information.

[0637] (Note 2) The information processing system according to Appendix 1 is characterized in that, The terminal is used to detect the user's selection operation of the character sequence on the display screen, extract the corresponding character sequence according to the detected selection range and send the extracted character sequence to the system. At the same time, according to the query processing status corresponding to the prompt statement, visual information indicating the processing progress is displayed on the display screen. In the dialogue display area of ​​the display screen, the semantic information received from the system and the adjusted response information are arranged in mutually independent display units.

[0638] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system is used to receive dialogue history information surrounding the selected character sequence from the terminal, generate an extended prompt statement containing the dialogue history information, send the extended prompt statement to the generative artificial intelligence model, and obtain response information generated by the generative artificial intelligence model after performing contextual dependency analysis on the selected character sequence based on the dialogue history information.

[0639] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: Means for obtaining the statement based on the statement selection made by the user through an information display device; Extraction means for acquiring a visual field image containing the acquired statement and extracting character information of the statement from the visual field image through character recognition processing; A means for receiving statement information containing the extracted character information from a terminal; Means for obtaining and sending meaning information by retrieving the sentence information in the vocabulary information storage unit and sending the meaning information to the terminal for display on the display device; A means for automatically generating prompt statements based on the statement information and service type information to specify the prompt statements for querying generative artificial intelligence models. Model interaction means for sending the automatically generated prompts to the generative artificial intelligence model and receiving response information from the generative artificial intelligence model; A response information processing means for determining the length of the response information based on a character limit, and generating response information for display by deleting part of the content or performing summary processing based on the length determination result; A display control means for sending the response information to be displayed to the terminal and for the information display device to overlay the response information to be displayed; Recording means for associating and storing at least a portion of the statement information, the prompt statement, and the response information in a recording area.

[0640] (Note 2) The information processing system according to Appendix 1 is characterized in that, The prompt statement generation method is configured to, when automatically generating the prompt statement, select template information from a set of multiple template information containing information into which the statement information can be inserted, according to a display mode, and generate the prompt statement by embedding the statement information into a predetermined position of the template information.

[0641] (Note 3) The information processing system according to Appendix 1 is characterized in that, When generating the response information for display based on the character count limit, the response information processing means is configured to, if the character count of the response information exceeds a predetermined upper limit, extract a predetermined number of sentences from the beginning in units of sentences, or extract sentences with higher importance through extractive summarization processing to constitute the response information for display.

[0642] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for detecting and obtaining the selected statement on a display device based on user operations; Apparatus for generating a prompt string containing the selected statement based on a predetermined string template, and using the prompt string as an input prompt statement for a generative artificial intelligence model; A device for sending the prompt as a string to an information processing device via a communication path, and for receiving response information generated by the generative artificial intelligence model from the information processing device; A device for determining the length of received response information according to a predetermined character limit, and simplifying the response information by deleting parts exceeding the character limit and / or adding ellipsis marks; A means for arranging and displaying simplified response information and semantic information related to the selected statement on a display interface of a user-operated terminal device; A device for retrieving the selected statement from vocabulary data to obtain the semantic information, and for sending the semantic information to the terminal device; An apparatus for storing processing content corresponding to the prompt string and the response information as record information, and for updating the generation conditions of the prompt string and / or the upper limit of the number of characters based on the record information.

[0643] (Note 2) The information processing system according to Appendix 1 is characterized in that, The apparatus for generating a prompt string is configured to select a string template from a plurality of string templates based on the category of the user-selected statement and / or the attributes of the displayed content, and to insert the selected statement into the selected string template to generate the prompt string.

[0644] (Note 3) The information processing system according to Appendix 1 is characterized in that, The apparatus for simplifying the response information is configured to, while performing reduction processing based on the character limit, adjust the character limit according to the size of the display area and / or the type of terminal device, and regenerate the response string for display based on the adjusted character limit.

[0645] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for retrieving meaning information from a string selected by a user on a terminal display interface based on vocabulary information stored in a storage device connected to an information processing device. A device for including the acquired meaning information in display information and sending the display information to the terminal for display on the display interface; A means for generating prompt statements to instruct a generative artificial intelligence model to generate natural language responses based on the selected string and additional information about the operating status of the terminal; A means for sending the prompt statement to the generative artificial intelligence model and obtaining response information related to the selected string from the generative artificial intelligence model; A device for analyzing user speech content obtained from the communication interface or transaction interface of the terminal to determine the user's emotional state. An apparatus for modifying the style or descriptive level of the prompt statement based on the determined emotional state, thereby adjusting the content of the response information generated by the generative artificial intelligence model; A device for controlling the length of the response information obtained from the generative artificial intelligence model according to a predetermined character limit and removing irrelevant information to generate response information for display; A device for sending the display response information to the terminal and displaying the display response information on the display interface.

[0646] (Note 2) The information processing system according to Appendix 1 is characterized in that, The device for generating the prompt statement is configured to switch between generating a prompt statement for defining and explaining the vocabulary of the explanatory object and a prompt statement for obtaining additional information related to the vocabulary of the explanatory object, based on the selected string, the type of interface in which the selected string is displayed, and the emotional state.

[0647] (Note 3) The information processing system according to Appendix 1 is characterized in that, The apparatus for generating the response information for display is configured to: summarize the response information obtained from the generative artificial intelligence model according to the character count limit; when the amount of information after the summary exceeds a predetermined threshold, divide the response information into summary information and detailed information; display the summary information on the display interface in real time; and display the detailed information in stages according to additional user operations.

Claims

1. An information processing system, characterized in that, include: processor; The processor is configured as follows: The system queries the dictionary database to obtain the meaning information of the statement selected by the user on the terminal. The acquired semantic information is sent to the user's terminal and displayed on the screen. Based on the selected statement, generate prompt text to instruct the generative artificial intelligence model to generate a specific response; The generated prompt text is sent to the generative artificial intelligence model, and a response is received from the generative artificial intelligence model. The character limit for the received response is adjusted, and the adjusted response is sent to the user's terminal for display on the screen.

2. The information processing system according to claim 1, characterized in that, The processor is configured to generate prompt text containing the selected statement.

3. The information processing system according to claim 1, characterized in that, The processor is configured to adjust the character limit for responses received from the generative artificial intelligence model.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A