Information processing system

CN122797718APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610268073.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-06
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]本发明要解决的主要课题在于:在利用生成式人工智能模型为用户提供各类智能服务(例如专业咨询、翻译、解释说明等)的过程中,现有技术存在以下问题:首先,现有系统通常要求用户以自然语言直接编写复杂的提示文本,普通用户缺乏提示设计能力,难以构建高质量且可复用的提示,从而导致生成式人工智能模型的输出质量不稳定;其次,现有系统在与大规模语言模型交互时,提示文本往往较长,令牌使用量高,造成计算资源浪费与运行成本上升,却缺乏针对提示文本进行系统性压缩与令牌优化的机制;再次,现有系统在与用户交互时,通常未充分考虑用户的情绪状态,无法根据用户当前情绪动态调整提示文本的内容和压缩策略,因而难以提升交互体验与响应的个性化程度;此外,现有基于块状拼装式语言或可视化编程界面的系统,多侧重教育编程或逻辑控制,对生成式人工智能模型的提示构建支持不足,缺乏通过可视化积木组合生成适配不同任务与角色设定的提示文本的统一框架

Benefits of technology

[0004]To address the aforementioned issues, this invention provides an information processing system comprising a processor configured to execute a series of collaborative functional modules to achieve visualization of prompt construction, token optimization of prompt content, and adaptive emotion adjustment. Specifically, the processor is first configured to distribute specific templates to users. These templates predefine various block-based language building blocks, including blocks for defining roles in generative AI models, blocks for describing tasks or operational instructions, and translation blocks for multilingual conversion of prompt text. Through template distribution, users do not need to write prompts from scratch; instead, they can select appropriate blocks based on the distributed template on the terminal interface and drag and drop to assemble them, thereby visually constructing the instruction structure for the generative AI model. Secondly, the processor is configured to automatically generate corresponding prompt text based on the user's assembly result of the selected template using block-based language. This text instructs the generative AI model to perform specific operations. During the generation process, the processor combines the semantics of multiple blocks into a structured, understandable, and model-compliant complete prompt text according to the predefined text templates of each block and user-defined parameters. Furthermore, the processor is configured to compress the generated prompt text to reduce token usage. Compression strategies include, but are not limited to, deleting redundant words and repeated instructions, merging semantically similar sentences, and replacing long sentences with internal short codes. This significantly reduces the number of tokens while ensuring the semantic integrity and interpretability of the prompts, thereby improving the overall processing efficiency of the system and reducing operating costs. In addition, the processor is configured to recognize user emotions and dynamically adjust the generation method and compression rate of the prompt text based on the recognized emotions. For example, when the user is tense or depressed, it adds explanatory and reassuring instructions and appropriately reduces the compression intensity to maintain the subtlety of expression; while when the user is calm or has high efficiency requirements, it increases the compression rate to reduce interaction latency. Finally, the processor is configured to send the adjusted prompt text to a generative artificial intelligence model and receive the response returned by the model, providing the response result to the user. This constructs an integrated intelligent interaction system encompassing template distribution, prompt assembly, compression optimization, emotion perception, and model invocation. Through the above methods, the present invention simplifies and improves the quality of prompt construction, effectively reduces token usage, and optimizes the emotion-adaptive interaction process, fundamentally solving the related technical problems existing in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797718A_ABST
    Figure CN122797718A_ABST
Patent Text Reader

Abstract

This invention provides an information processing system. The information processing system is characterized by comprising: a processor; wherein the processor is configured to: distribute a specific template to a user; generate prompt text for instructing a generative artificial intelligence model to perform a specific operation based on the user's assembly result of the selected template using a block-based language; compress the generated prompt text to reduce token usage; identify the user's emotion and adjust the generation method and compression rate of the prompt text according to the identified emotion; and send the prompt text to the generative artificial intelligence model and provide the user with the response returned by the generative artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] The main problem this invention aims to solve is that existing technologies have the following issues when using generative artificial intelligence models to provide users with various intelligent services (such as professional consultation, translation, and explanation): First, existing systems typically require users to directly write complex prompt text in natural language. Ordinary users lack the ability to design prompts, making it difficult to construct high-quality and reusable prompts, thus leading to unstable output quality of generative artificial intelligence models. Second, when interacting with large-scale language models, existing systems often produce long prompt texts and high token usage, resulting in wasted computing resources and increased operating costs, yet lack a mechanism for systematically compressing prompt texts and optimizing tokens. Third, existing systems typically do not fully consider the user's emotional state when interacting with them, and cannot dynamically adjust the content and compression strategy of prompt texts based on the user's current mood, thus making it difficult to improve the personalization of the interactive experience and response. In addition, existing systems based on block-based language or visual programming interfaces focus more on educational programming or logic control, providing insufficient support for prompt construction by generative artificial intelligence models, and lacking a unified framework for generating prompt texts adapted to different tasks and roles through visual block combinations. Therefore, the technical challenge that this invention aims to address is to provide a system that allows users to construct high-quality prompts using a block-based, modular language, effectively compresses the prompt text to reduce token usage, and dynamically adjusts the system based on user emotions. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides an information processing system comprising a processor configured to execute a series of collaborative functional modules to achieve visualization of prompt construction, token optimization of prompt content, and adaptive emotion adjustment. Specifically, the processor is first configured to distribute specific templates to users. These templates predefine various block-based language building blocks, including blocks for defining roles in generative AI models, blocks for describing tasks or operational instructions, and translation blocks for multilingual conversion of prompt text. Through template distribution, users do not need to write prompts from scratch; instead, they can select appropriate blocks based on the distributed template on the terminal interface and drag and drop to assemble them, thereby visually constructing the instruction structure for the generative AI model. Secondly, the processor is configured to automatically generate corresponding prompt text based on the user's assembly result of the selected template using block-based language. This text instructs the generative AI model to perform specific operations. During the generation process, the processor combines the semantics of multiple blocks into a structured, understandable, and model-compliant complete prompt text according to the predefined text templates of each block and user-defined parameters. Furthermore, the processor is configured to compress the generated prompt text to reduce token usage. Compression strategies include, but are not limited to, deleting redundant words and repeated instructions, merging semantically similar sentences, and replacing long sentences with internal short codes. This significantly reduces the number of tokens while ensuring the semantic integrity and interpretability of the prompts, thereby improving the overall processing efficiency of the system and reducing operating costs. In addition, the processor is configured to recognize user emotions and dynamically adjust the generation method and compression rate of the prompt text based on the recognized emotions. For example, when the user is tense or depressed, it adds explanatory and reassuring instructions and appropriately reduces the compression intensity to maintain the subtlety of expression; while when the user is calm or has high efficiency requirements, it increases the compression rate to reduce interaction latency. Finally, the processor is configured to send the adjusted prompt text to a generative artificial intelligence model and receive the response returned by the model, providing the response result to the user. This constructs an integrated intelligent interaction system encompassing template distribution, prompt assembly, compression optimization, emotion perception, and model invocation. Through the above methods, the present invention simplifies and improves the quality of prompt construction, effectively reduces token usage, and optimizes the emotion-adaptive interaction process, fundamentally solving the related technical problems existing in the prior art.

[0005] "System" refers to an overall device or set of devices including at least one processor and connected storage devices, communication interfaces and / or user interaction interfaces, used to perform the functions of template distribution, prompt text generation, compression processing, emotion recognition and generative artificial intelligence model invocation described in this invention.

[0006] A processor is a hardware unit that can execute program instructions to perform functions such as data processing, logical operations, control flow and communication management, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and processing modules or processing clusters composed of the above components.

[0007] A "template" refers to a set of structured configurations predefined by the server and distributed to users for building prompt text using a block-based assembly language. These include definitions of multiple building blocks, type information of the blocks, text templates, and parameter specifications, which constrain and guide users in constructing prompts.

[0008] "Block-based modular language" refers to a descriptive language form that uses visual building blocks or modules as basic elements, allowing users to combine different functional blocks by dragging, splicing, nesting, etc. to construct logical instructions or prompt text. This language focuses on graphical, modular, and structured expression, rather than pure natural language input.

[0009] "Building blocks" refer to functional units in block-based modular languages ​​that can be dragged and combined by users. Each building block corresponds to a specific logical or semantic function and has a predefined text template and parameter interface, which are used to replace specific natural language fragments when prompt text is generated.

[0010] "Prompt text" refers to natural language or structured text that is automatically generated by the processor based on the block-based language assembly results created by the user. It is used to issue task instructions or contextual constraints to generative artificial intelligence models, including role settings, task descriptions, input and output requirements, etc.

[0011] "Compression" refers to the process of controlling the length and redundancy of the generated prompt text, including but not limited to deleting redundant characters or repeated sentences, merging semantically similar expressions, and replacing long sentences with short codes or abbreviations, in order to reduce the amount of tokens used in the process of calling generative artificial intelligence models.

[0012] "Token utilization" refers to the amount of tokens used by the model to represent prompt text and / or user input when interacting with a generative artificial intelligence model, relative to the total available token resources. This metric reflects the amount of computational resources consumed by the text for the model.

[0013] "User emotions" refers to the psychological state or emotional tendency exhibited by users during interaction with the system, including but not limited to tension, anxiety, pleasure, calmness, and confusion, which are identified and judged by the processor through user input text, voice, facial expressions, or interactive behavior characteristics.

[0014] "Generative artificial intelligence model" refers to an artificial intelligence model trained based on machine learning and / or deep learning algorithms that can automatically generate text, code, images or other content based on input prompts. In this invention, it mainly refers to a generative text model based on a large-scale language model.

[0015] "Specific expert roles" refers to the professional identities or domain backgrounds preset for generative artificial intelligence models in the prompt text, such as medical experts, legal advisors, English teachers, etc., to guide the models to adopt specific professional knowledge and communication styles when responding.

[0016] A "translation block" is a predefined building block unit in a template used to translate prompt text or user input from one language to another. The block contains parameters such as the source language and the target language, and is converted into corresponding translation instructions when the prompt text is generated.

[0017] "Compression ratio" refers to the length comparison of the prompt text before and after compression. It is usually expressed as the ratio of the number of tokens after compression to the number of tokens before compression, and is used to measure the compression effect and the degree of token saving. Attached Figure Description

[0018] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0019] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0020] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0021] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0022] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0023] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.

[0024] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0025] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0026] Figure 9 This represents an emotion map that maps multiple emotions.

[0027] Figure 10 This represents an emotion map that maps multiple emotions.

[0028] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0029] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0030] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0031] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0032] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0033] First, let me explain the terminology used in the following instructions.

[0034] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0035] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0036] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0037] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0038] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0039] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0040] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0041] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0042] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0043] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0044] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0045] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0046] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0047] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0048] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0049] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0050] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0051] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0052] With the widespread application of generative AI models in text generation, dialogue interaction, and content creation, end users typically need to input high-quality prompts into these models to obtain the desired output. However, in existing technologies, the construction of prompts mainly relies on users freely writing them in natural language, which leads to the following technical problems: (1) The process of constructing prompt statements lacks structured guidance. Ordinary users find it difficult to understand the sensitivity of generative artificial intelligence models to elements such as context, role setting, and output format. They often adjust the text through repeated trial and error, resulting in a waste of a lot of computing resources and time, and low overall system processing efficiency.

[0053] (2) Prompt statements often contain a lot of redundant information or repetitive expressions when expressing the same intention, resulting in an excessive number of labels input into the generative artificial intelligence model, which consumes communication bandwidth and computing resources, increases inference latency, and is not conducive to deployment and operation in environments with limited computing resources.

[0054] (3) Traditional compression methods focus on general data compression or simple truncation, but lack fine-grained processing of the semantic structure of prompt statements. This can easily destroy the expression of key semantics such as roles, tasks, and constraints in prompt statements, thereby affecting the response quality of generative artificial intelligence models.

[0055] (4) The generation and compression of prompt statements are usually performed linearly on the same text plane. The visual block structure is not fully utilized to constrain the construction process of prompt statements. There is a lack of reversible or semi-reversible mapping mechanism between block sequence structure and natural language text, making it difficult to automatically optimize prompt statements while maintaining editability.

[0056] Therefore, how can we provide a new system and method that enables: Users are guided to construct structured prompt statements through templated and block-based visual editing methods; On the server side, conversion rules between block sequence data and natural language text are used to compress and optimize prompt statements specifically for generative artificial intelligence models, effectively reducing the number of tags while maintaining semantic integrity as much as possible. It also supports the re-editing and dynamic optimization of prompts during multi-round interactions. This invention aims to substantially improve the efficiency of computational resource utilization, response latency, and overall system performance during the invocation of generative artificial intelligence models, which is the technical challenge it seeks to address.

[0057] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0058] In this invention, the server includes: a component for reading multiple template data stored in a hierarchical structure from an information storage device and sending the template data to an information display device via a network; a component for receiving block sequence data generated by the information display device based on the user's selection, arrangement, and parameter input operations on multiple block elements, and parsing the identifiers and corresponding parameter values ​​of each block element contained in the block sequence data; a component for expanding the block sequence data into natural language text according to predetermined conversion rules, thereby generating original data for prompt statements for a generative artificial intelligence model; a component for performing text conversion processing on the original prompt statement data, including merging synonyms, deleting redundant expressions, and shortening sentence structure, to generate compressed candidate data, and calculating and comparing the number of tags on the original prompt statement data and the compressed candidate data respectively, thereby selecting the final prompt statement data to be adopted according to predetermined conditions; and a component for sending the final prompt statement data to an inference processing device of a generative artificial intelligence model, obtaining response data from the inference processing device, and returning it to the information display device. This allows for the effective reduction of the number of tags input into generative AI models, reducing network transmission and inference computation load, and shortening response time, while ensuring the semantic integrity and editability of the prompt statements. It also enables dynamic optimization of the prompt statements through multi-round interactions, thereby improving the overall processing efficiency and resource utilization of the computer system during the invocation of generative AI models.

[0059] "Information processing device" refers to an electronic device with computing and control functions, used to execute program instructions for data reading, parsing, conversion and communication control, and may include one or more processors, memory and communication interfaces.

[0060] The "processing unit" refers to a logical functional unit in an information processing device, consisting of a processor and the program it executes. It is used to control, calculate, and manage input data in order to achieve predetermined data processing functions.

[0061] "Information storage device" refers to a storage component used to save template data, block element data, prompt statement data and related configuration information in a readable form, which may include semiconductor memory, magnetic storage medium or optical storage medium.

[0062] "Information display device" refers to a terminal device used to present a graphical user interface, text information and response content to a user, and to receive user input. It may include a display screen, an input device and a computing unit that executes the interface program.

[0063] "Template data" refers to a set of configuration data stored in a predefined hierarchical structure for building prompt statements. It contains multiple reusable block elements and their text templates, parameter definitions, and structural relationships.

[0064] "Hierarchical structure information" refers to a data structure form organized through parent-child or nested relationships, used to represent multi-level subordinate or combined relationships between templates, block elements, and their parameters.

[0065] A "block element" refers to a structured component defined in template data that serves as the basic unit for constructing prompt statements. Each block element has a predefined text pattern and one or more fillable parameters.

[0066] "Block sequence data" refers to structured data formed by combining multiple block elements in sequence based on the user's selection, arrangement, and parameter input operations on the information display device. It is used to represent the composition order and content settings of prompt statements.

[0067] An "identifier" is a symbol or string used to uniquely mark each block element or related data item, and is used to reference, parse, and convert block elements within an information processing device.

[0068] "Parameter value" refers to the specific content that the user enters or selects at the parameter position corresponding to the block feature. It is used to replace the placeholder in the block feature text template to generate text with specific meaning.

[0069] "Pre-defined transformation rules" refer to a set of logical rules for text generation and transformation set in advance during the system design phase. These rules are used to convert block sequence data into natural language text and to standardize or simplify sentence structure and word usage while preserving semantics.

[0070] "Natural language text" refers to a continuous sequence of characters written in human natural language, used to express instructions, constraints, or questions for generative artificial intelligence models.

[0071] "Prompt statements" refer to the input text for generative artificial intelligence models, which indicate the role the model plays, the tasks it performs, the format of the output, and other constraints to guide the model in generating the desired response.

[0072] "Prompt statement raw data" refers to the uncompressed or unoptimized natural language text data of the prompt statement obtained by the information processing device based on the first expansion of block sequence data.

[0073] "Final prompt data" refers to the prompt data that is determined to be sent to the generative artificial intelligence model after text conversion processing, tag count comparison, and condition selection.

[0074] "Compressed candidate data" refers to one or more candidate text versions obtained after performing text compression on the original prompt statement data. These candidate text versions are used to compare with the original data to select the final prompt statement.

[0075] "Text transformation processing" refers to a set of semantically preserving transformation operations performed on natural language text, including synonym merging, redundant expression deletion, sentence structure shortening, and other processing methods used to reduce text length and optimize expression.

[0076] "Synonymous expression merging" refers to the process of combining multiple expressions with similar or identical meanings into a shorter and more unified expression without changing the overall semantics.

[0077] "Redundant expression removal" refers to the process of identifying and removing words or phrases that contribute little to the main semantics of the prompt statement and have repetitive or redundant features.

[0078] "Sentence shortening" refers to the process of reorganizing or simplifying sentences while maintaining the essential semantic points, in order to reduce the complexity of vocabulary and structure.

[0079] "Number of tags" refers to the minimum number of processing units obtained after segmenting the input text according to the word segmentation or encoding method adopted by the target generative artificial intelligence model. It is used to measure the scale of the model input.

[0080] "Generative artificial intelligence models" refer to artificial intelligence models trained based on machine learning and deep learning techniques that can automatically generate text responses or other output results based on input prompts.

[0081] "Inference processing device" refers to a collection of computing resources and program components that perform inference operations on generative artificial intelligence models, and is used to receive the final data of prompt statements and output response data.

[0082] "Response data" refers to the output data generated by a generative artificial intelligence model through inference and calculation after receiving the final data of the prompt statement. It is usually natural language text, but may also contain structured information.

[0083] "Visual interface elements" refer to interface components that are presented graphically on information display devices and allow users to operate them by clicking, dragging, or inputting, and are used to represent and edit block elements.

[0084] "Re-editing instruction" refers to the instruction information issued by the user through interface operation after reading the response data, which is used to modify block sequence data, adjust the content or structure of prompt statements.

[0085] "Control information" refers to data signals generated by the information display device and sent to the information processing device for indicating changes in block sequence data, regeneration of prompt statements, or control of related processing flows.

[0086] This invention will illustrate a system for generating and compressing prompt statements using a server, a terminal, and a generative artificial intelligence model working collaboratively. This invention is not limited to the specific embodiments described below; various modifications and variations are possible without departing from the technical concept defined in the claims.

[0087] I. Overall System Composition Servers run on computing devices with general-purpose processors and network interfaces, such as rack-mounted computers with multi-core CPUs, virtual machines, or cloud computing instances. Servers can utilize any combination of hardware and software environments, including but not limited to: Server hardware example: a computing device with a multi-core CPU (such as a general-purpose multi-core processor), at least tens of GB of memory, and solid-state storage; Example of a server operating system: a Unix-like operating system, such as a Linux distribution; Server application framework example: The backend uses an interpreted or compiled program runtime environment, such as a web framework based on a scripting language or a web framework based on a general programming language; Server database example: Relational database management system or document database management system, used to store template data, block elements and related records of prompt statements.

[0088] The terminal can be any combination of information display and input devices, such as smartphones, tablets, personal computers, and all-in-one terminals. The terminal can run common operating systems (such as mobile or desktop operating systems) and communicate with the server through a browser or a dedicated client.

[0089] Generative AI models can be deployed on a server locally or on other computing nodes, such as inference servers equipped with graphics processing units or cloud-based model services. Generative AI models are preferably deep learning-based autoregressive language models, such as neural networks with multi-layer transformer structures.

[0090] II. Server-side modules and data structures The server can store the following main data structures in the storage device: Template data: includes multiple template entries, each template consisting of several elements; Block feature data: Each block feature has a block identifier, text pattern, parameter list, and hierarchical relationship information; Transformation rule data: a set of rules for converting block sequence data into natural language text, and a set of rules for performing text compression; Logs and statistics: Used to record the number of tags in the original and compressed versions of prompt statements, the response time of generative artificial intelligence models, etc.

[0091] The server can store template data and block feature data in a hierarchical structure in the database, such as through the association of template tables, block tables, and parameter tables, representing a multi-level structure of "template-block-parameter". The text pattern of a block feature can contain placeholders for the terminal to fill in parameter values ​​in a later stage. For example, the text pattern of a block feature could be: "You (AI) are an expert in the domain". The server can load transformation rule data into memory. These transformation rules can include: Block-to-text expansion rules: Define how to convert a block identifier and its corresponding parameter values ​​into natural language text; Synonym merging rules: For example, rewrite "You are a history expert. Please explain in detail..." as "As a history expert, please explain..."; Redundant expression identification rules: For example, detect and delete repeated honorifics or irrelevant small talk; Sentence shortening rules: such as breaking complex sentences into shorter sentences or using more compact grammatical structures.

[0092] The server can also implement a tag counting module, which segments or encodes the input text according to the word segmentation or encoding method used by the target generative AI model and returns the number of tags. The server can use this module to quantitatively compare the original prompt statement with the compressed candidate text.

[0093] III. Terminal Interface and Data Representation The terminal can execute a front-end program in a browser or client, which can be implemented using a common scripting language and interface framework. After receiving template data from the server, the terminal can visually present a list of block elements on a display device.

[0094] The terminal can render each block element as a UI component that displays the parameterless portion of the text mode and provides input controls for each parameter placeholder, such as text input boxes, dropdown selections, and checkboxes. The terminal allows users to change the order of block elements by dragging, thus specifying the order of the various parts of the prompt statement.

[0095] The terminal can internally represent the user-constructed block sequence using structured data, for example, representing each block as a set of records containing block identifiers and parameter value fields. The terminal does not need to store specific transformation rules; it only needs to send the block identifiers and corresponding parameter values ​​to the server, which then performs a uniform rule transformation.

[0096] IV. User Operation Modes on the Terminal Users can perform the following typical operations on the terminal interface: The user selects the "Character Setting Block" from the block list and enters "History" in the parameter field, thus generating the text pattern: "You (AI) are an expert in history." The user selects "Task Description Block" from the block list and enters "Causes and Major Battles of World War II" in the topic parameter. The corresponding text pattern can be: "Please describe the following topic in detail: Causes and Major Battles of World War II." Users select "Tone and Audience Block" from the block list and specify "easy to understand, suitable for middle school students" in the parameters. The corresponding text mode can be: "Please explain in easy to understand language that is suitable for middle school students to comprehend." The terminal can arrange the above blocks in the user-specified order on the interface and generate a preview of the prompt statements in real time. For example, the terminal can display the following preview of the prompt statements to the user: "You are a history expert. Please explain in detail the causes and major battles of World War II using language that is easy to understand and suitable for middle school students." V. Server-side prompt statement expansion and compression After receiving the block sequence data sent by the terminal, the server can generate the raw data for the prompt statement as follows: The server searches for the text pattern of the corresponding block element in the template data based on the block identifier; The server replaces the placeholders in the text pattern of the block element with the parameter values ​​provided by the user to obtain the complete sentence; The server connects the sentences according to the block sequence order, using set delimiters (such as periods and newlines), to generate the original text of the prompt statement.

[0097] In this embodiment, the raw data for the prompt statement generated by the server can be: "You are a history expert. Please explain in detail the causes and major battles of World War II using language that is easy to understand and suitable for middle school students." The server can invoke the marker counting module to encode and count the text, thereby obtaining the initial marker count. Subsequently, the server can activate the text conversion processing module to perform a specialized compression operation on the original text. For example, the server can perform the transformation according to the following rules: The server can recognize the sentence structure "You are an expert in X" and standardize it into "as an expert in X" according to the synonym merging rules; The server can simplify "Please use language that is easy to understand and suitable for middle school students" to "Use language suitable for middle school students"; The server can replace "detailed description" with "system explanation" or "explanation" to reduce the number of characters while keeping the meaning the same.

[0098] Based on the above rules, the server can generate a compressed candidate text, for example: "As a history expert, I will explain the causes and major battles of World War II in language suitable for middle school students." The server can call the tag counting module again to calculate the number of tags for the compressed candidate text and compare it with the number of tags in the original text. If the number of tags in the compressed candidate text is significantly reduced and no semantic loss rule constraints are triggered (e.g., key fields such as "history", "World War II", and "cause and major battles" are still retained), then the server can determine the compressed candidate text as the final data for the prompt statement.

[0099] The server performs a specific compression process tailored to the structure and semantics of the prompt statements, rather than simple general text compression. This compression utilizes block structures and predefined transformation rules to reduce redundant text without altering the task's semantics, thereby reducing the number of labels input into generative AI models.

[0100] VI. Examples of Generative Artificial Intelligence Model Structures and Training Methods The generative artificial intelligence model that the server can invoke is preferably an autoregressive language model based on a transformer architecture. This model can contain multiple layers of encoding units, each consisting of a multi-head self-attention sublayer and a feedforward neural network sublayer. Each attention sublayer calculates attention weights using the dot product and normalization between the query vector, key vector, and value vector, while the feedforward network sublayer applies a non-linear transformation to each labeled position.

[0101] Generative AI models can use large-scale text data for unsupervised or self-supervised learning during the pre-training phase. During training, the server can use cross-entropy loss as the error function, comparing the predicted distribution of the current label with the true next label, and updating the network weights through backpropagation. Optimization algorithms (such as gradient descent methods with momentum) can be employed during training, using techniques like learning rate decay, batch normalization, or layer normalization to accelerate convergence and stabilize training.

[0102] After receiving the final data of the prompt statement during the inference phase, the server encodes the text into a sequence of tags, which are then input into the model one by one. At each step, the model calculates the probability distribution of the next tag based on the previous tag sequence and selects the tag with the highest probability or samples it under certain temperature parameters until the termination condition is met. The server can adjust the diversity and stability of the generated content by controlling the maximum generation length, sampling strategy, and temperature parameters.

[0103] VII. Interface Design between Server and Generative Artificial Intelligence Model Servers can interact with generative AI models through in-process calls, remote procedure calls, or network APIs. When a generative AI model is deployed on a standalone inference server, the server can encapsulate the final data of the prompt statement into a request message over the network, send it to the model service interface, and parse the response data into natural language text format after receiving the response.

[0104] The server can add internal model statistics to the response data, such as the number of input labels, the number of output labels, and inference time. This information can be recorded by the server and used for subsequent optimization analysis. The server can also dynamically adjust the activation level or priority of compression rules based on this statistical information to achieve a balance between response quality and resource consumption.

[0105] VIII. Technical Effects and Causal Relationships The server employs a block-sequence-driven prompt generation and specialized text compression, achieving a technical solution significantly different from traditional free text input. Because the prompts are constructed based on templates and block elements, the server can semantically decompose and reassemble each block, thereby conditionally applying targeted compression rules. This structured processing enables the server to more accurately identify redundant expressions and mergeable synonymous structures, thus significantly reducing the number of tags in most cases.

[0106] Reducing the number of labels directly decreases the input length of generative AI models, reducing the matrix operations required for multi-head self-attention computation in the transformer network, thus shortening inference time and reducing computational resource consumption. Since the time complexity of multi-head self-attention is quadratic with the sequence length, moderate compression of the number of labels can significantly improve inference speed. Simultaneously, because compression follows specific semantic preservation rules, the server can avoid the loss of task semantics caused by simple truncation, thereby maintaining or improving response quality.

[0107] Furthermore, by recording the original number of tags, the compressed number of tags, and the response time for each interaction, the server can subsequently adjust text conversion rules and compression strategies. Based on statistical results, the server can apply certain compression modes to more scenarios or reduce compression for specific prompt types, dynamically balancing performance and quality under different load and latency targets. This feedback-based rule adjustment mechanism constitutes adaptive optimization of the computer system's internal processing flow, enabling continuous improvement in computational efficiency compared to purely manual adjustments or static templates.

[0108] IX. Multiple Implementation Methods and Variations The server can support multiple types of block elements. For example: The server can define various blocks such as "role setting block", "task description block", "input / output format block", "translation block", and "length control block" to achieve fine-grained combination.

[0109] The server can support multilingual templates, configuring different language versions of text modes under the same identifier block, and the terminal can automatically select the appropriate version based on the user interface language.

[0110] The server can introduce a rule-based two-stage compression process, which first performs semantic compression based on block structure, and then performs local sentence simplification based on statistical language models to further reduce the number of tags.

[0111] The terminal can adopt different presentation formats in different scenarios: On the desktop, the terminal can display template hierarchy by dragging panels and tree structures; On mobile devices, the terminal can use paginated lists and collapsible menus to adapt to small screen operation.

[0112] The specific implementation of generative artificial intelligence models can also be varied: The server can use transformer models of different sizes, or a sequence-to-sequence model with an encoder-decoder structure; During the fine-tuning phase, the server can perform supervised training for specific application scenarios (such as translation, educational Q&A, and professional consultation) to further improve the model's response quality under corresponding prompts.

[0113] Through the above implementation forms and variations, a collaborative system is formed between the server, terminal, and generative artificial intelligence model. By utilizing structured block sequences, rule-based text conversion, and label quantity control, substantial improvements are achieved in the internal processing flow of the computer during the invocation of the generative artificial intelligence model. This results in causal technical effects in terms of processing speed, resource utilization, and response quality, rather than simply automating the writing of human prompts.

[0114] use Figure 11 The processing flow is explained.

[0115] Step 1: The server reads template data from the information storage device and sends it to the terminal. The input is a set of template records stored in the database, and the output is a template data message sent to the terminal via the network. The server calls the database query interface, performs retrieval operations based on conditions such as template type and language settings, assembles the query results into a hierarchical data structure (including template ID, block element ID, text pattern, parameter list, etc.), serializes this data structure, generates message data that can be transmitted over the network, and then sends the message to the terminal through the communication interface.

[0116] Step 2: The terminal receives and parses template data, displaying block elements on the interface. The input is the template data message sent by the server, and the output is a list of block elements stored in the terminal's memory and a visual interface presented on the display device. The terminal receives messages through the network module, deserializes the message content, maps the hierarchical structure to front-end objects (such as an array of block objects), and generates corresponding interface components based on the text pattern and parameter definitions of each block. Simultaneously, it draws draggable block areas and parameter input areas on the display device.

[0117] Step 3: Users select block features and enter parameter values ​​on the terminal. The input consists of the currently displayed block feature interface and user actions (clicks, drags, text input). The output is a block sequence structure containing the user's selection order and parameter value settings. Users select several block features on the terminal interface using a pointer device or touchscreen, drag them to the editing area, adjust their order, and simultaneously enter specific content in each parameter input box (e.g., entering "History" in the "Domain" parameter and "Causes and Major Battles of World War II" in the "Theme" parameter). The terminal converts each action event into an update operation on the block objects, records the block ID, sequence index, and parameter values, and maintains the current block sequence state in memory.

[0118] Step 4: The terminal generates a preview text for a prompt statement based on the block sequence structure. The input is the block sequence structure in the terminal's memory (including block IDs and parameter values), and the output is the preview string of the prompt statement displayed on the interface. The terminal iterates through the block sequence, reading the text pattern of each block in block order, replacing placeholders with their corresponding parameter values, and concatenating the text blocks using preset delimiters to form continuous natural language text. The terminal stores this text in a local variable for later transmission to the server, and simultaneously updates the preview area on the display device for user confirmation and further modification.

[0119] Step 5: The user confirms the prompt and instructs the terminal to send a request. The input consists of the preview text of the prompt displayed on the terminal and the user's confirmation action (e.g., clicking the "Send" button). The output is the request data packet to be sent to the server. After confirming the preview text, the user issues the send command. The terminal uses the current preview text as the original prompt field, packages it together with additional information such as user identification and language settings into a request data structure, performs serialization processing on this data structure to generate a network request body, and registers the target server address and interface path in the communication module, preparing for network transmission.

[0120] Step 6: The terminal sends a request containing the original prompt statement to the server. The input is the request data packet generated in step 5, and the output is an HTTP or other protocol request transmitted to the server over the network. The terminal invokes the network stack to encapsulate the request data packet in application layer protocol and transport layer protocol messages, performs byte stream transmission, sends the original prompt statement and related metadata as message content to the predetermined server interface address, and records the request time and request ID locally for subsequent correlation with the response.

[0121] Step 7: The server receives the request and parses the original prompt statement. The input is the request message sent by the terminal, and the output is the original prompt statement string and related parameters stored in the server's memory. The server receives the message through the network interface, performs protocol parsing and message decoding, and extracts fields from the request body, including the user ID, the original prompt statement text, and language settings. The server loads the original prompt statement into a memory buffer to prepare for subsequent tag counting and text conversion operations, and records basic information about this request in the log system.

[0122] Step 8: The server performs a tag count on the original prompt statement. The input is the original prompt statement string, and the output is the number of tags for the original prompt statement. The server calls a tagging function compatible with the target generative AI model to convert the original prompt statement into a tag sequence, and calculates the number of tags by traversing the sequence or directly from the result returned by the function. The server stores this tag count in a record associated with this request, serving as the basis for subsequent compression effect comparisons and resource statistics.

[0123] Step 9: The server performs text conversion on the original prompt statement to generate compressed candidate text. The input is the original prompt statement string, and the output is one or more compressed candidate text strings. In the text conversion module, the server sequentially applies predetermined conversion rules: through pattern matching and substitution operations, it rewrites sentences like "You are an expert in X" into "As an expert in X"; it removes repeated honorifics and redundant phrases that do not affect semantics using regular expressions; and it compresses long sentences into more compact sentences using syntactic simplification rules. After completing one round of conversion, the server generates compressed candidate text and can generate multiple versions with different degrees of compression as needed. These versions are stored in memory as a list for use in tag counting and selection operations.

[0124] Step 10: The server counts the tags on the candidate compressed texts and compares them with the original prompt statement. The input consists of one or more candidate compressed text strings and the original tag count. The output is the tag count for each candidate text and the selection result of the best candidate text. The server repeats the tagging and counting operations in step 8 for each candidate compressed text to obtain the corresponding tag count, then performs a comparison operation to calculate the tag reduction amount of each candidate text relative to the original text, and checks whether key semantic segments are completely preserved (e.g., fields such as "history," "World War II," "cause and major battles," etc.). The server selects an optimal compressed version based on predetermined conditions (such as a reduction ratio threshold and semantic integrity check results). If no candidate meets the conditions, it reverts to using the original prompt statement.

[0125] Step 11: The server determines the final data of the prompt statement and prepares to send it to the generative AI model. The input is the comparison result between the original prompt statement and candidate texts; the output is the final data string of the prompt statement and the model request structure. Based on the selection result in step 10, the server sets the optimal text as the final data of the prompt statement and combines it with inference control parameters such as model identifier, maximum generation length, and sampling parameters to form the model input structure. The server encodes this structure to form a request content suitable for sending to the generative AI model service, and internally records the compression strategy information used in this instance.

[0126] Step 12: The server sends the final data of the prompt statement to the generative AI model and receives the response. The input is the model request structure (containing the final data of the prompt statement), and the output is the response data returned by the generative AI model. The server sends the request to the inference processing unit where the model is deployed via a local process call or a remote interface call. The inference processing unit internally performs the forward propagation operation of the neural network and generates the response text based on the final data of the prompt statement. After receiving the response, the server parses the text fields and statistical fields (such as the number of model input / output tags and inference time), stores the response text in memory, and attaches relevant statistical information for subsequent processing and recording.

[0127] Step 13: The server performs optional post-processing on the response data and sends it to the terminal. The input is the response text and statistical information from the generative AI model; the output is the response message sent to the terminal. The server can perform basic formatting operations on the response text, including segmenting by periods or newlines, inserting paragraph separators, etc., and can also check the content using simple sensitive word detection rules. The server encapsulates the processed response text along with information such as the original token count and the compressed token count into a response structure, serializes it, and sends it to the terminal via the network interface.

[0128] Step 14: The terminal receives and displays response data. The input is the response message sent by the server, and the output is the response text displayed on the screen, along with an interface state that allows for further user interaction. The terminal receives and parses messages via the network module, extracts the response text and relevant metrics, renders the response text in the results display area, and presents it to the user in paragraph or list format. It can also display simple performance information on the interface (e.g., "How many tags were reduced in this compression"). The terminal retains the block sequence structure used in this round of interaction in memory so that the user can edit it again if needed.

[0129] Step 15: The user re-edits the response on the terminal, triggering a new round of processing. The input consists of the response text displayed on the terminal, the current block sequence state, and the user's editing actions. The output is the updated block sequence structure and the triggering of a new request. After reading the response, if the user deems the explanation too long or unclear, they can add or modify block elements in the block editing interface, such as adding blocks like "Please keep it under 500 words" or "Please use more colloquial examples." The terminal records and updates these operations, generates a new block sequence structure, and repeats steps 4 through 14. This, in conjunction with the server and the generative artificial intelligence model, enables multi-round optimization of the prompt statements and gradual improvement of response quality.

[0130] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0131] With generative AI models widely used in text generation and content creation, users typically need to write their own prompts to enable these models to perform specific tasks. However, existing technologies suffer from the following problems: First, the construction of prompts highly depends on the user's language skills and understanding of the model's behavior, lacking structured guidance. This leads to unstable prompt quality and poor controllability of the generated results, thus limiting the usability and ease of use of generative AI models on general-purpose terminal devices. Second, to accurately describe complex needs, prompts are often written very long, resulting in an excessive number of symbolic units (such as labeled units) input into the generative AI model, causing high computational resource consumption, increased response latency, and higher call costs. Third, existing systems mostly use simple string concatenation to directly transmit prompts to the generative AI model, lacking technical means to finely control and compress data length at the prompt level, which is detrimental to improving overall processing efficiency at the system level. Fourth, existing user interfaces often only provide text input boxes, failing to assist users in constructing prompts through visual and modular methods. This makes it difficult to standardize and streamline the prompt construction process, and users cannot intuitively control key parameters such as role settings and language types.

[0132] Therefore, it is necessary to provide a system for generative artificial intelligence models and a scheme for generating and processing prompt statements. By coordinating the design of template information, constituent element information, compression processing, and symbol unit quantity control mechanisms between the server and the terminal, the system can improve the structure and compactness of prompt statement input from the perspective of computer architecture and data processing flow, reduce symbol unit consumption, improve the overall processing efficiency and response performance of generative artificial intelligence model invocation process, and enhance the visualization and controllability of prompt statement construction process.

[0133] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0134] In this invention, the server includes means for abstracting and structuring user-oriented information into various types of template information in an information processing device and sending the template information to an external device; means for obtaining a selection result of one type of the template information from the external device and multiple component information arranged in a visual manner, and constructing a prompt statement representing the instruction content of a generative artificial intelligence model according to the order of the component information and their respective input values; means for performing compression processing on the prompt statement to reduce the length of the symbol sequence, and controlling the compression method or compression rate in the compression processing according to the structure or usage of the prompt statement to reduce the number of symbol units used in the generative artificial intelligence model; means for sending the compressed prompt statement or its decompression result to the generative artificial intelligence model and sending the generated content information obtained from the generative artificial intelligence model to the external device; and means for enabling the external device to display the template information and the component information in the form of a visually arrangeable editing interface and receiving the addition, deletion or order change of the component information according to the user's operation. This allows for the formation of a complete data processing chain within the computer, from template-driven construction of structured prompt statements to optimization of symbolic units based on compression processing, and control over the input length and number of constituent elements of generative artificial intelligence models. This systematically reduces the number of symbolic units used in generative artificial intelligence models without relying on advanced user skills, improves resource utilization and response speed during model invocation, and enhances the controllability and consistency of prompt statement construction through a visual constituent element editing interface, thus achieving an overall improvement in the mechanism for prompt statement generation and processing in computer technology.

[0135] "Information processing device" refers to a computing device with computing and storage capabilities, used to execute programs and process data, including but not limited to server devices, terminal devices, or computing systems composed of multiple computing nodes.

[0136] "External device" refers to a device that sends and receives data with an information processing device via a communication line. It is usually a terminal device operated by the user, such as a client device with display and input functions.

[0137] "Template information" refers to a preliminary data set that is pre-set and structured by an information processing device to guide the construction of prompt statements. This data set defines the basic framework, field structure, and recommended component types of the prompt statements.

[0138] "Component information" refers to the unit of the prompt statement based on template information. Each unit describes a specific functional part or semantic fragment in the prompt statement, including block type, parameter items and their values, etc., and is used to form a complete prompt statement by combination.

[0139] "Visual layout" refers to presenting the information of constituent elements in the form of graphical, modular, or block-shaped elements on a display screen with a graphical user interface, and allowing users to change their position and order through dragging, clicking, and other operations.

[0140] The "editing interface" refers to an interactive screen displayed on an external device for viewing, selecting, inputting parameters, and adjusting the order of template information and component information. This screen supports visual arrangement of components and preview of prompt statements.

[0141] "Prompt statements" refer to text data input into a generative artificial intelligence model to instruct the model to perform a specific task or generate specific content. They are usually natural language strings composed of multiple constituent information elements.

[0142] "Generated content information" refers to the output data of a generative artificial intelligence model based on prompts, which includes at least text content and, if necessary, structured information or metadata related to that text content.

[0143] "Compression processing" refers to the process of encoding or compressing prompt statements to reduce the length of symbol sequences or data. Specifically, it can use general compression algorithms or special encoding methods.

[0144] "Symbol sequence length" refers to the length of the sequence of symbol units used to represent prompt statements, including but not limited to the number of characters, bytes, or tokens generated by a particular tokenizer.

[0145] "Symbol unit" refers to an element used to represent the smallest unit of text processing in a generative artificial intelligence model or coding algorithm, including but not limited to characters, word fragments, tags, or other basic symbols defined by the model.

[0146] "Compression method" refers to the specific technical solution or algorithm category used when compressing the prompt statement, such as different encoding rules, compression algorithm types, or parameter configurations.

[0147] "Compression ratio" refers to the ratio of the length of the compressed data to the length of the uncompressed data, or a related metric, used to characterize the strength of the compression effect.

[0148] "Usage status" refers to operational information related to how prompt statements are used in the system, including but not limited to call frequency, prompt statement length, target task type, and resource consumption when interacting with generative artificial intelligence models.

[0149] "Role setting information" refers to the constituent elements of information used in prompt statements to specify the functional role or professional identity played by the generative artificial intelligence model, thereby limiting the knowledge domain and expression style of the model's output.

[0150] "Translation instruction information" refers to the constituent information in prompt statements used to instruct generative artificial intelligence models to perform language conversion on text content or output results in a specific language, including at least the target language or language equivalent parameters.

[0151] "Specialty domain information" refers to an abstract description used to indicate the scope of knowledge or application area that the prompt statement focuses on, such as tourism, healthcare, and finance, in order to limit the output range of generative artificial intelligence models.

[0152] "Language type information" refers to the parameter information used to specify the category of natural language used for inputting prompt statements or generating output content, such as Chinese, English and other language identifiers.

[0153] "User interface processing" refers to a series of processes performed on an external device that are related to displaying template information and component information, as well as receiving user input operations, including interface drawing, event response, and interactive logic control.

[0154] "Generative artificial intelligence models" refer to model systems trained through machine learning methods that can automatically generate text or other data content based on input prompts. They typically employ a multi-layer neural network structure and use sequences of symbolic units as the basic processing objects.

[0155] In one embodiment, the server operates as an information processing device on a computer system. The server may employ a hardware platform including a multi-core central processing unit and an optional graphics processing unit, such as a rack-mount computer with a multi-core processor and a graphics accelerator card, and run a general-purpose operating system, such as a Unix-like operating system. The server can communicate with multiple terminals through network server software (e.g., general-purpose HTTP server software) and execute the various functional modules of this invention through an application server framework (e.g., a general-purpose web application framework).

[0156] In this implementation, the server pre-stores template information and component definition information in a storage device (e.g., a hard disk drive or solid-state storage device). The server can use a relational database management system (e.g., general-purpose relational database software) or a non-relational database management system (e.g., general-purpose document-oriented database software) to store the following data structures in table or document structure: The server represents "template information" as a collection of records including fields such as template identifier, template name, purpose description, recommended component list, and default component order. For example, the server can define a "travel guide" template as: a template used to generate travel-related prompts, which is recommended to include "role setting component information," "task description component information," and "translation instruction component information," etc.

[0157] The server represents the "component information" as a block structure. Each component record contains at least: a block identifier, a block type (such as "role setting," "task description," "translation instructions," "output format," "length control," etc.), a block text template (a natural language pattern sentence containing placeholders), parameter definitions (parameter name, data type, and value range), and display attributes (title, descriptive text, etc. in the editing interface). The server loads this data into memory during initialization, forming a cache structure that can be accessed quickly.

[0158] As an external device, the terminal can be a smartphone, tablet, or personal computer, running a mobile or desktop operating system, and communicating with the server through a browser or local applications. The terminal runs a front-end program in local storage, which can implement a graphical user interface based on a scripting language framework (such as a general component-based front-end framework).

[0159] The terminal displays an editing interface, which includes a template list area, a component selection area, and a preview area for prompts. After receiving template information from the server, the terminal displays the name and purpose of each template in a list or card format in the template list area. The user selects a template via the terminal's touchscreen or pointing device, and the terminal then requests the definition of the corresponding component information from the server.

[0160] After receiving a template details request from the terminal, the server retrieves the set of constituent element information corresponding to the template identifier from the database, and packages fields such as block type, block text template, and parameter definitions into structured response data and returns it to the terminal. During this process, the server uses a data serialization library to convert the internal data structure into a standardized transmission format.

[0161] After receiving the component element information definition, the terminal displays this information as blocks in the component element selection area. Each block is represented as a draggable rectangle or puzzle shape in the graphical interface. The terminal creates an object instance for each block in its local memory, which includes the block type, default text template, current parameter values, and the block's position index in the sequence. Users can perform operations such as dragging and dropping, clicking to add, and clicking to delete these blocks through the terminal's graphical interface.

[0162] After each user action, the terminal updates and sorts the block array locally using a scripting language. For each component information, the terminal performs string substitution operations based on its block text template and the parameter value entered by the user, replacing placeholders with specific parameters. For example, when the user enters "travel expert" in the "Character Setting Component Information," the terminal converts the block template "You (AI) are an expert in {domain}." into "You (AI) are a travel expert."; when the user specifies "Create a detailed sightseeing guide for first-time visitors to Tokyo" in the "Task Description Component Information," the terminal converts the relevant template into the corresponding natural language sentence.

[0163] After generating the text for each component element, the terminal performs string concatenation according to the order of the block array, combining multiple text fragments into a complete prompt statement in a predetermined order. The terminal then displays this complete prompt statement in real-time in natural language in the prompt statement preview area for user confirmation.

[0164] In a specific example, after a user selects the "Travel Guide" template on the terminal and configures the following components, the terminal will generate the following prompt: "You (AI) are a travel expert. Based on the information below, please create a detailed sightseeing guide for first-time visitors to Tokyo. Please include a three-day itinerary, recommended attractions, transportation options, and important notes. Please output the results in Simplified Chinese." In another example, a user can construct the following prompt statement: "You (AI) are a travel expert. Please create a detailed travel guide for first-time visitors to Tokyo based on the information below."

[0165] Please include: 1. Suggested itinerary for a three-day trip; 2. A brief introduction to each attraction and information on transportation; 3. Recommended local food and restaurants; 4. Etiquette and safety precautions to be aware of.

[0166] Please answer in Simplified Chinese and use clear subheadings for paragraph breaks. After the user confirms the prompt, the terminal sends the complete prompt as input data to the server via the communication module using a secure transmission protocol. Upon receiving the prompt, the server creates a corresponding data buffer in memory and invokes the compression module to perform compression processing. The server can use compression algorithms implemented in a general data compression library (such as general byte stream-based compression algorithms) to encode the byte sequence of the prompt, compressing the long text into shorter compressed data.

[0167] Within the compression module, the server can dynamically adjust compression parameters based on the original length of the prompt statement, the amount of information in its constituent elements, and historical call statistics, such as selecting different compression levels. The server generates a compressed byte stream from the input byte stream using an algorithm, recording the lengths before and after compression for subsequent estimation of symbol unit usage.

[0168] When interacting with generative AI models, the server encapsulates prompts in a manner suitable for the model service. If the model interface supports compressed input, the server can encode the compressed byte stream (e.g., using common binary-to-text encoding methods) before transmission; if the model interface does not support compressed input, the server can internally cache the compressed results to optimize storage and network usage, and then decompress the data before sending the resulting raw prompts to the model service.

[0169] In one implementation, the generative AI model is deployed on an inference server with graphics processing unit acceleration, employing a multi-layer self-attention network structure. The model uses an embedded tokenizer to map input prompts into sequences of symbol units; the tokenizer can segment natural language text into fixed- or variable-length tokens based on a sub-word unit encoding algorithm. The model uses an embedding layer to convert each symbol unit into a high-dimensional vector, and performs matrix multiplication, linear transformations, non-linear activations, and attention weight calculations on these vectors through a multi-layer transformer structure.

[0170] During forward inference, the generative AI model calculates the probability distribution of the next symbol unit for each time step based on the previous sequence of symbol units. When calling the model service, the server can specify control parameters such as temperature parameters, maximum generation length, and sampling strategies (e.g., greedy search, random sampling, or bundle search) to balance content diversity and determinism.

[0171] During model training, a combination of supervised and self-supervised learning can be employed. In the pre-training phase, the model uses large-scale corpus data, calculating the difference between predicted and true symbolic units based on a language modeling objective function, and employing cross-entropy as the loss function. The model updates the weight parameters in the network using a series of stochastic gradient descent optimization algorithms (such as an adaptive learning rate optimization algorithm). To improve the model's adaptability to diverse tasks and domains, fine-tuning can be performed after pre-training, continuing training using a dataset containing instruction-response pairs, enabling the model to learn to perform complex tasks based on prompts.

[0172] In the system of this invention, the server achieves fine-grained control over the input size of the generative artificial intelligence model by compressing and monitoring the length of the prompt statements. The server can measure the number of symbol units or byte length of the prompt statement before and after compression each time a request is processed. Based on this measurement result, it determines whether certain component information needs to be pruned, block order adjusted, or the user prompted to reduce the content. Through this mechanism, the server can prevent the prompt statements from exceeding the model's input length limit, avoiding truncation and performance degradation caused by excessively long input.

[0173] Unlike methods that rely solely on manually writing long text prompts, the system of this invention employs a combination of block-based component information and template-driven generation on the terminal side. This makes the structure of the prompts more standardized and parameterized, allowing the server to adopt more refined strategies in compression and resource allocation. Because the component information is stored and combined in a structured form, the server can internally selectively retain or simplify certain segments based on block type and priority. For example, when symbol unit budget is insufficient, priority can be given to retaining "task description component information" and "role setting component information," while reducing secondary decorative information. This block-type-based, unconventional pruning method differs from simple string truncation, thus maintaining the semantic integrity and accuracy of the generated content while reducing input length.

[0174] This invention improves the data flow and resource utilization within a computer system by introducing structured templates, block-based component information, and compression and symbol unit control mechanisms between the server and terminal. The server reduces the amount of data transmitted over the network and the size of the model input data by compressing prompts, enabling generative AI models to handle more requests with the same hardware resources, thereby increasing system throughput. The terminal reduces prompts for incorrect and redundant constructions through a visual editing interface, indirectly reducing invalid computations. By monitoring the length and number of symbol units before and after compression in real time, the server can dynamically balance the quality of individual requests with overall resource utilization efficiency at the system level, thereby improving processing speed and optimizing resource utilization.

[0175] In another implementation, the server can employ different compression strategies based on the category of the prompt statement content. For example, when multiple similar parameter descriptions are detected to appear repeatedly in the prompt statement, the server can internally use dedicated rules to perform pattern encoding on the redundant parts, rather than simply using a general compression algorithm. Through this rule-based compression method that combines template structure and block type characteristics, the server can utilize the semantic structure information of the prompt statement when compressing it, achieving a higher symbol unit saving rate than traditional unstructured compression.

[0176] In another implementation, the terminal can further provide auxiliary information for evaluating the quality of prompt statements. The terminal locally evaluates the current block combination based on preset rules (such as checking whether it contains key blocks like character setting elements, task description elements, and output language settings), and displays the information to the user on the interface indicating which parts are not yet configured. The server can use this structured information to infer the clarity and completeness of the prompt statements, providing a foundation for the controllability of subsequent model output. This method of coupling the block-level structure with the model call flow allows the system to improve the overall generation quality from the external data structure level without increasing the internal complexity of the model.

[0177] In summary, this invention achieves technical improvements in the calling process of generative artificial intelligence models from three levels: data structure, algorithm flow, and resource allocation. These improvements are manifested in reduced symbol usage, increased processing speed, reduced network load, and enhanced stability and consistency of generated results. This is achieved through the server's structured management of template information and constituent element information, the compression of prompt statements and control of the number of symbol units, and the implementation of a visual editing interface on the terminal.

[0178] use Figure 12 The processing flow is explained.

[0179] Step 1: Users request and select template information on the terminal.

[0180] Users open the template list screen in the terminal's graphical interface and select a template by clicking or touching.

[0181] Input: User's selection action (e.g., template identifier).

[0182] Output: Template details request data sent by the terminal to the server.

[0183] Based on the user's selection, the terminal reads the selected template identifier from the local interface state, encapsulates the identifier as a request parameter, constructs a request message through the communication module, and sends it to the server.

[0184] Step 2: The server returns the component information corresponding to the selected template based on the request.

[0185] After receiving the template details request sent by the terminal, the server retrieves the template record and constituent element record corresponding to the template identifier from the database.

[0186] Input: Request data containing template identifiers.

[0187] Output: Response data containing information on multiple constituent elements.

[0188] The server uses query statements to retrieve template names, usage descriptions, a list of recommended block types, and text templates and parameter definitions for each block from the database table. It then organizes the search results into structured data and returns them to the terminal via a network interface.

[0189] Step 3: The terminal builds a visual list of blocks locally and displays an editing interface.

[0190] After the terminal receives the component information definition returned by the server, it creates an object for each component in memory, sets the block type, text template and default parameters, and draws the corresponding draggable block on the interface.

[0191] Input: Definition data of constituent element information returned by the server.

[0192] Output: An editing interface for the template selection area and the component block area presented on the display device.

[0193] The terminal generates interface elements based on the block's display attributes (title, description), calls the graphics rendering engine to render these elements onto the screen, and registers event listeners for drag, click, etc.

[0194] Step 4: Users select and arrange constituent element blocks by dragging and clicking.

[0195] In the terminal's editing interface, users can drag the required component blocks from the candidate area to the editing area, or click the "Add" button to insert blocks, and adjust the order of the blocks by dragging.

[0196] Input: User's drag-and-drop, click, and other interactive events.

[0197] Output: The sequence of constituent element blocks and their position indices maintained internally by the terminal.

[0198] The terminal parses each interaction event, updates the block object array according to the event type, moves the index of the dragged block from its original position to a new position, and rearranges the display position of the block on the interface.

[0199] Step 5: Users input parameter values ​​for each component block on the terminal.

[0200] Users select a section and enter text content or select options in the pop-up parameter input area, such as inputting the field "Travel Expert", destination "Tokyo", and outputting the language "Simplified Chinese".

[0201] Input: The characters or options that the user types in each input box.

[0202] Output: The set of updated parameter values ​​in the terminal block object.

[0203] When the terminal receives an input event, it encodes the string or option entered by the user into an internal format, writes it into the parameter field of the corresponding block object, and synchronously displays the currently set parameter value on the interface.

[0204] Step 6: The terminal generates prompt message fragments based on block templates and parameters.

[0205] The terminal iterates through the block array in the current editing area, reads the text template and parameter value for each block, and replaces the placeholders in the template with user input through string replacement.

[0206] Input: An array of block objects containing the block type, text template, and parameter values.

[0207] Output: A list of prompt message fragments generated for each constituent element block.

[0208] The terminal performs a text processing algorithm on each block, replacing "{domain}" in a template like "You (AI) are an expert in {domain}." with "expert in travel", generating specific natural language clauses, and storing these clauses in a new fragment array.

[0209] Step 7: The terminal concatenates all the fragments in block order to form a complete prompt statement.

[0210] The terminal performs a concatenation operation on the generated prompt statement fragments according to the current order of the block array, inserting appropriate spaces or line breaks between the fragments to form a logically coherent long text.

[0211] Input: A list of sequentially ordered prompt statement fragments.

[0212] Output: The complete text string of the prompt message.

[0213] The terminal uses a string concatenation function to merge multiple fragments, such as "You (AI) are a travel expert." and "Please create a detailed sightseeing guide for first-time visitors to Tokyo based on the information below." into a single long string, which is then displayed in real time in the prompt preview area.

[0214] Step 8: The user confirms the prompt statement on the terminal and sends a generate request.

[0215] Users read the full prompt displayed in the preview area. If the content meets their needs, they can click "Generate Content" or a similar button to instruct the system to start calling the generative artificial intelligence model.

[0216] Input: User's confirmation event.

[0217] Output: The generation request data sent from the terminal to the server, which includes the complete prompt statement and optional control parameters.

[0218] The terminal captures click events, reads the current prompt text from memory, packages it together with parameters such as model name and maximum output length into a request payload, and sends it to the server through the communication module.

[0219] Step 9: The server receives the prompt and performs compression processing.

[0220] The server extracts the prompt text from the request payload, converts it into a byte sequence, and then calls a compression library to execute a compression algorithm.

[0221] Input: The complete prompt statement in text form.

[0222] Output: The compressed byte sequence of the prompt statement and the length information before and after compression.

[0223] The server creates a buffer in memory to store the raw byte stream, calls a compression function to generate a compressed byte stream, and records the original length and compressed length for subsequent symbol unit usage evaluation or log statistics.

[0224] Step 10: The server controls the symbol units based on the length of the prompt statement and the compression result.

[0225] The server compares the length before and after compression with a preset threshold to determine if there is a risk of exceeding the input limits of the generative artificial intelligence model.

[0226] Input: the length of the data before compression, the length of the data after compression, and the system's preset length threshold.

[0227] Output: Control decisions regarding whether the prompt statement needs to be trimmed or simplified.

[0228] Based on the comparison results, if the server finds that the length is close to or exceeds the threshold, it can call the rules module to selectively mark certain non-critical parts for pruning according to the block type priority table, or generate warning messages; if the length is within the safe range, it can directly proceed to the subsequent model call steps.

[0229] Step 11: The server sends the prompt (in compressed or decompressed form) to the generative artificial intelligence model service.

[0230] The server performs necessary encoding or decompression of the compressed byte stream according to the requirements of the model service interface, and prepares the model call parameters.

[0231] Input: Compressed prompt statement, length information, and model call configuration parameters.

[0232] Output: The request message sent to the generative artificial intelligence model service.

[0233] The server sends the prompt data and control parameters (such as temperature and maximum number of generated symbol units) to the model server endpoint via network protocol, triggering the model's forward computation on the inference device.

[0234] Step 12: Generative artificial intelligence models generate content based on prompts and return the results to the server.

[0235] The generative artificial intelligence model runs on an inference server. It converts prompts into symbol unit sequences through a word segmenter. After passing through an embedding layer and a multi-layer self-attention network, matrix operations and probability predictions are performed on the symbol unit sequences to gradually generate the output symbol unit sequences.

[0236] Input: Prompt text indicating the task and model control parameters.

[0237] Output: The generated content text string or structured text.

[0238] The model service converts the generated symbol units back into natural language text, encapsulates them into response data, and sends them back to the server over the network.

[0239] Step 13: The server receives the generated content information and performs necessary post-processing.

[0240] The server extracts the generated content text from the response of the model service, and can optionally perform formatting and basic filtering.

[0241] Input: The generated content text returned by the model service.

[0242] Output: Organized text suitable for display on the terminal.

[0243] The server can remove redundant blank lines, standardize line breaks, split content into paragraphs, perform sensitive word replacement or length truncation operations when necessary, and then load the processed text into the response payload, ready to be sent to the terminal.

[0244] Step 14: The server generates content information and sends it to the terminal.

[0245] The server, through the application server and network server, sends a response message containing the generated content to the requesting terminal.

[0246] Input: The processed generated content text and related metadata (such as generation time and length statistics).

[0247] Output: Response data transmitted to the terminal.

[0248] The server can log data before sending data for subsequent performance analysis and quality assessment.

[0249] Step 15: The terminal receives the generated content and displays it to the user on the interface.

[0250] The terminal receives the server's response in the communication module, parses the generated content field, and displays the text in the results display area.

[0251] Input: The response message returned by the server.

[0252] Output: A view of the generated content that can be scrolled on a display device.

[0253] The terminal uses a typesetting engine to format and display text based on its structure (headings, paragraphs, lists), allowing users to scroll through, copy text, or switch between different generated results.

[0254] Step 16: Users can determine whether they need to edit the prompt statement again based on the displayed results.

[0255] If a user reads the generated content displayed on the terminal and feels that the result needs adjustment, such as wanting a more formal style or adding certain information, they can return to the editing interface to readjust the constituent element blocks and parameters.

[0256] Input: The generated content text and user subjective evaluations.

[0257] Output: New editing operation instructions (such as adding a new block or modifying parameters).

[0258] Users can switch back to the editing interface by clicking buttons such as "Return to Edit" or "Modify Conditions" to continue making modifications at the block level.

[0259] Step 17: The terminal updates the block configuration and regenerates the prompt statement based on the user's new editing operation.

[0260] After receiving the user's modification operation, the terminal updates the internal block object array and parameter values, and repeats the process of generating and splicing the prompt statement fragments to form a new complete prompt statement.

[0261] Input: User-adjusted block order and parameter values.

[0262] Output: The updated complete message text.

[0263] The terminal then displays a new prompt in the preview area for user confirmation. Step 8 and subsequent steps can be repeated to achieve multiple rounds of iterative generation, thereby continuously optimizing the prompt and generation results under a structured and controllable framework.

[0264] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0265] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0266] With the widespread application of generative AI models in scenarios such as dialogue, question answering, translation, and professional consultation, systems typically need to submit prompts to these models containing information such as role settings, task constraints, and output format requirements. In existing technologies, prompts are often directly written or simply assembled manually, leading to the following computer technology-level problems: First, to express complex role settings and multi-step task constraints within a single prompt, the prompts are often excessively long, requiring the model to process a large number of information units (such as characters, subwords, or tags) during inference, thus increasing computational resource consumption, extending response time, and reducing system throughput. Second, existing systems generally lack structured prompt construction mechanisms; the generation of prompts highly relies on human experience, making it difficult to modularly combine and systematically optimize prompt content on the server side through programmatic methods, and thus unable to automatically generate high-quality prompts for different tasks. Furthermore, the compression of user-friendly prompts is problematic. Third, traditional data compression methods primarily optimize the encoding of general files or transmitted data, without integrating them with the semantic structure of prompts or the input characteristics of generative artificial intelligence models. They often compress data only at the transmission link level, without coordinating "text simplification" and "compression encoding strategies" at the application level. Therefore, it is difficult to simultaneously ensure semantic fidelity and reduce the utilization rate of information units. Fourth, in the absence of a unified information structure and a language for describing constituent elements, servers struggle to reuse the same prompt generation mechanism across different application scenarios, resulting in complex software structures, poor scalability, and difficulties in system-level performance tuning and resource management.

[0267] The technical problem to be solved by this invention is to provide a system technical solution within a computer server, by introducing information structure, component assembly language, text simplification processing, and a data compression mechanism adapted to the characteristics of prompt statements, which can significantly reduce the amount of information used, reduce the model inference load, shorten the end-to-end response time, and improve the modularity, automation, and scalability of the prompt statement generation process, while maintaining the semantic clarity of the input of generative artificial intelligence models and the integrity of task constraints. This will achieve a substantial improvement to the computer technology itself (including the server-side prompt generation module, compression module, and model calling process).

[0268] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0269] In this invention, the server includes means for distributing an information structure to a user and providing a general framework, which serves as the information structure, for defining the roles and processing object domains of a generative artificial intelligence model; means for receiving multiple constituent elements described in a constituent element assembly language based on the information structure, obtaining natural language fragments corresponding to each constituent element, and integrating the natural language fragments into the information structure to generate a prompt statement for instructing the generative artificial intelligence model to perform a specific action; means for performing text simplification processing on the prompt statement to reduce redundant expressions and simplify sentence structure while maintaining semantics to obtain a text-simplified prompt statement; means for compressing the text-simplified prompt statement using a general compression algorithm or a special encoding rule for the prompt statement to generate a compressed prompt statement with reduced information unit usage; and means for sending the compressed prompt statement to a terminal or providing the generative artificial intelligence model, obtaining response information obtained by decompressing and processing the compressed prompt statement through the generative artificial intelligence model, and providing the response information to the user. This allows for the automatic generation of high-quality prompts within the server in a structured manner. By combining text simplification and compression strategies tailored to the characteristics of prompts, the number of information units input to generative AI models can be significantly reduced. This reduces the computational resources and time required for model inference, improves the overall system processing efficiency and throughput, and provides a reusable and easily extensible computer implementation framework for role setting and task combination in different application scenarios.

[0270] A "system" refers to a collection of devices consisting of one or more computing devices and programs running on them, used to perform functions such as information structure distribution, prompt generation, text simplification, data compression, and interaction with generative artificial intelligence models.

[0271] A "server" refers to a computing device or cluster of computing devices that provides information structure management, component element parsing, prompt statement generation and compression in a network environment, and communicates with terminals or generative artificial intelligence model providers.

[0272] "Terminal" refers to a computing device operated by a user and interacting with a server or generative artificial intelligence model, including but not limited to mobile terminals, personal computers or other network-connected devices.

[0273] "User" refers to a user who interacts with the system through a terminal, selects information structure, configures constituent elements, and receives response information from generative artificial intelligence models.

[0274] "Information structure" refers to an abstract data structure used to organize and describe the basic framework of prompt statements, and a general framework used to define the roles, processing domains, and task categories of generative artificial intelligence models.

[0275] "Component Assembly Language" refers to a formal representation language used to describe prompts in a modular way. It expresses prompt conditions such as role settings, task requirements, and output formats through multiple combinable components, which facilitates automatic parsing and combination by the server to generate natural language prompts.

[0276] "Construmental elements" refer to the smallest functional units used to represent specific functions or semantic segments in a constituent element assembly language, including but not limited to role designation elements, translation elements, format elements, style elements, etc.

[0277] "Role-assigned elements" refer to constituent elements in constituent element assembly language used to assign specific professional domain roles or functional positioning to generative artificial intelligence models, which are used to constrain the model to perform reasoning and generation in a predetermined role.

[0278] "Translation elements" refer to the constituent elements in a constituent element assembly language that are used to instruct generative artificial intelligence models to translate content between different natural languages, including parameters such as source language, target language, and translation style.

[0279] "Natural language fragments" refer to text fragments written in natural language that correspond one-to-one with the constituent elements. They are used to transform abstract constituent elements into concrete descriptions that can be understood by generative artificial intelligence models when generating prompt statements.

[0280] "Prompt statements" refer to natural language text input generated by the server based on information structure and a combination of multiple natural language fragments, used to instruct generative artificial intelligence models on roles, tasks, and output requirements.

[0281] "Text simplification" refers to the text optimization of prompt statements while preserving their semantics. This includes removing redundant expressions, merging sentences with repetitive meanings, simplifying sentence structure, and standardizing terminology to reduce the number of characters or tags.

[0282] "Data compression" refers to the technical process of encoding a simplified prompt statement into a shorter data representation by using general compression algorithms or specific rules for the prompt statement, in order to reduce the amount of information used per unit.

[0283] "General compression algorithm" refers to a compression algorithm that is widely used in the computer field and is independent of specific application content, including but not limited to lossless compression algorithms based on byte sequences.

[0284] "Encoding rules for prompt statements" refers to compression rules pre-designed based on the language characteristics and structural patterns of prompt statements, including high-frequency phrase replacement, symbol mapping, and templated encoding, which are used to further shorten the encoding length while preserving the semantics.

[0285] "Information unit" refers to the basic unit used to measure the size of text in the input or internal processing of a generative artificial intelligence model, including characters, bytes, subwords, tags, or other measurable language units.

[0286] "Compressed prompt statements" refers to an encoded form that, after text simplification and data compression, contains fewer information units but still retains the semantics and constraint information of the original prompt statement.

[0287] "Generative artificial intelligence models" refer to artificial intelligence models built on machine learning or deep learning that can automatically generate text or other data outputs based on input prompts, including language models, dialogue models, or general generative models.

[0288] "Response information" refers to the result data generated and output by the generative artificial intelligence model based on its internal reasoning process after receiving the decompressed prompts and related inputs. This includes text responses, translation results, explanations, or other forms of output content.

[0289] In a typical implementation, servers are deployed as central computing nodes within a data center. The hardware used by these servers may include a multi-core CPU based on the x86 architecture (e.g., a general-purpose processor with multi-level cache), at least 32GB of memory, solid-state drive storage, and a network interface with gigabit Ethernet or higher bandwidth. On the software side, servers may run a general-purpose operating system (e.g., a Linux distribution) and install application server middleware (e.g., the Python-based FastAPI framework or a JavaScript-based Node.js runtime environment), a database management system (e.g., a relational database management system), and text processing and compression libraries (e.g., the zlib library providing lossless compression).

[0290] In one implementation, the terminal can be a smartphone, tablet, or personal computer. The terminal's hardware may include an ARM- or x86-based processor, a display screen, a touch or keyboard input device, memory, and a wireless or wired network module. On the software side, the terminal may run a mobile operating system or a desktop operating system, and run a web browser or native applications. The terminal establishes an encrypted connection with the server via a secure version of Hypertext Transfer Protocol (HTTP) to send user commands and receive prompts and response information from generative artificial intelligence models generated by the server.

[0291] In this invention, users interact with the server through a graphical user interface on a terminal. Users select information structures, configure constituent elements, input natural language descriptions, and view the results output by the generative artificial intelligence model based on the prompts. Users do not need to directly understand the internal assembly language of constituent elements and the details of compression algorithms; the system automatically performs relevant data processing and calculations in the background.

[0292] In terms of information structure management, the server abstracts each application scenario into an information structure. The server stores these information structures in the database as structured records, including fields describing the generative AI model's role attributes, the domain of the object being processed (e.g., medical, legal, technical fields), task categories, and the set of available constituent elements. The server loads these information structures into memory during initialization for quick retrieval upon receiving user requests. The server uses unified identifiers within the information structures to indicate roles, task types, and output format requirements, thus forming a reusable, general framework.

[0293] In parsing component-based language, the server takes the component representations submitted by the user through the terminal as input data. The server maintains a mapping table in memory, associating each component identifier with its corresponding natural language fragment template. During parsing, the server replaces placeholders in the natural language fragment template according to the component type (e.g., role-specific elements, translation elements, formatting elements) and parameter values ​​to generate specific natural language fragments. The server uses a string processing library to perform operations such as concatenation, replacement, and case normalization, thereby forming a list of composable natural language fragments.

[0294] In generating prompt statements, the server combines the frame text in the basic information structure with the parsed natural language fragments. During this combination process, a fixed order rule can be set, such as adding role settings first, then task descriptions, and finally output format requirements. The server uses string concatenation and placeholder replacement operations to link multiple fragments into a complete prompt statement. The server can automatically remove duplicate expressions or conflicting instructions during the combination process according to preset rules, ensuring the logical consistency of the prompt statements.

[0295] In terms of text simplification, the server automatically optimizes the generated prompts using text analysis algorithms. The server can employ rule-based syntactic simplification strategies, such as deleting consecutive synonymous phrases, rewriting lengthy relative clauses into shorter phrases, and compressing multiple similar requirements into a single summary statement. The server can also use statistical language models or pre-trained language models to help determine which words are low-information-density components in the current context and prioritize their deletion. During simplification, the server preserves three core information categories: "role setting," "task objective," and "output format," thus maintaining the semantic integrity of the prompts while reducing the number of characters or tags.

[0296] In terms of data compression, the server performs two layers of compression on the condensed prompt text. First, the server executes application-layer prompt-specific encoding rules, mapping high-frequency natural language expressions to short tags. For example, the server can replace "You are a medical expert" with the tag "...".<ROLE_MED> Replace "translate the patient's medical condition description from Chinese to English" with the mark "".<TASK_ZH2EN_SYM> At this stage, the server utilizes a mapping table to achieve reversible replacement, which remains consistent on both the server and the generative AI model providing device. The server then treats the replaced text as a byte sequence input to a general compression algorithm (such as a lossless compression algorithm based on dictionary encoding and entropy encoding), calculates the compressed byte stream, and further converts it into a transport-friendly encoded form.

[0297] In terms of compression strategy optimization, the server performs length measurement and information unit statistics on prompt statements or simplified prompt statements. The server can compare the current prompt statement with a preset threshold based on the number of characters, bytes, or tags. When the length or number of information units exceeds the threshold, the server selects a compression method with a higher compression ratio or enables more prompt statement-specific encoding rules; when the length is short, the server can choose a compression method with lower computational overhead or reduce the compression ratio to reduce the time cost of compression and decompression. The server thus dynamically balances compression overhead with the overall overhead of transmission and inference under different input scales, thereby achieving a better end-to-end response time.

[0298] In interactions with generative AI models, the server can act as a proxy node or be configured to communicate directly with external model services. In one implementation, the server invokes a generative AI model deployed on a remote computing cluster. This model can be a neural network architecture based on a multi-layered self-attention mechanism, such as a language generation model with a multi-layered encoder-decoder structure, dozens to hundreds of attention heads, and pre-trained on a large-scale text corpus. Before sending a request to the model, the server organizes compressed prompts and user input (e.g., a description of a patient's condition) into a model input structure, including system-level prompts, user messages, and necessary metadata.

[0299] In terms of model inference, the server decompresses the compressed prompts back into natural language text at the model level, which is then processed by the tag embedding module within the generative AI model. Internally, the model performs tag partitioning, splitting the text into word-level or byte-pair encoding units and mapping each unit to a fixed-dimensional vector. The model then performs matrix multiplication, linear transformation, non-linear activation, and normalization operations on these vectors in a multi-layer self-attention network, calculating the correlation between positions in the sequence through attention weights. At the output, the model transforms the hidden state vectors into a probability distribution over a vocabulary using a parameterized linear layer and a normalized exponential function, progressively sampling or selecting the highest-probability tags as the output sequence. During training, the model uses cross-entropy loss as the error function, calculates gradients through backpropagation, and updates weight parameters using a momentum-based optimization algorithm. The model can be fine-tuned for specific tasks based on pre-training to improve the accuracy of responses to structured constraints in the prompts.

[0300] In this invention, the server does not simply "forward" user input, but introduces explicit rules and algorithms during the construction, simplification, and compression stages of prompt statements. The server transforms task requirements into a rule-based internal representation using an information structure and component assembly language, enabling automated composition, conflict resolution, and resource optimization within the computer. Through text simplification and specialized encoding rules, the server transforms verbose, human-readable natural language into a more compact representation that is more friendly to generative AI models, reducing the number of information units while decreasing the computational complexity of the model in sequence processing. For example, when the prompt statement length is reduced by 30%, the matrix operation size in the multi-head self-attention layer decreases quadratically, significantly reducing inference time and energy consumption.

[0301] In terms of user interface, the terminal displays the information structure provided by the server in the form of lists, trees, or cards. When the terminal obtains an example of a prompt statement returned by the server, it displays the example as a text area to help users understand how the system will constrain the generative artificial intelligence model. The terminal collects the user's text to be processed through input controls, such as descriptions of illnesses and consultation questions, and uses transport layer security protocols to protect data security when sending this text, along with compressed prompt statements, to the model providing device or server agent module.

[0302] Users may see the following example prompt in the terminal during actual use. The prompt generated by the server in one implementation may be: "You are a medical expert. Please translate the patient's medical condition description from Chinese to English, and provide the English translation first, followed by a brief medical explanation." or: "You are a medical expert. You are about to receive a patient's medical condition description. Please first accurately translate this description from Chinese into English, and then attach a brief explanation using professional medical terminology after the English description to help the patient understand their condition." Examples of patient descriptions entered by the user in the terminal include: "I've been coughing for the past two weeks, with a small amount of yellow phlegm. It's worse at night, and sometimes I feel a little tightness in my chest." After combining the above prompts and the patient's description, the server sends an inference request to the generative artificial intelligence model. The model, after processing by its internal neural network, outputs the following result: English translation: I have been coughing for the past two weeks, with a small amount ofyellow phlegm. The cough gets worse at night, and sometimes I feel a bit ofchest tightness. Medical explanation: The persistent cough with yellow phlegm may suggest a respiratoryinfection, such as bronchitis..." The terminal displays the results in sections, allowing users to clearly distinguish between the translation and medical explanation sections.

[0303] In terms of technical effectiveness, the server, through the aforementioned structured prompt generation, text simplification, and data compression processes, effectively reduces the number of tags that the generative AI model's input needs to process. Because the server eliminates redundant information at the application layer before optimizing and compressing at the encoding layer, compared to traditional methods that only compress at the transport layer, it significantly reduces the number of self-attention operations and embedding looks within the model. This reduction directly translates to a smaller matrix operation scale, thereby shortening inference latency and increasing the number of requests that can be processed per unit of time. Simultaneously, because the information structure and component-based assembly language ensure the consistency of the prompts in form and semantics, the generative AI model can more stably understand role and task constraints, thus improving the consistency and stability of the output results.

[0304] In terms of scalability, the server can quickly support new application domains by adding new information structures and new component definitions to the database without modifying the core prompt generation engine. The server reuses the same text simplification and data compression modules across different domains, thus achieving a unified performance optimization strategy at the system level. This modular design makes this invention not merely an automation of human-written prompts, but a systematic improvement to the computer's internal data structures, algorithmic processes, and resource allocation methods—an optimization of computer technology itself.

[0305] In another implementation, the server can collaborate with multiple generative AI models. In this implementation, the server maintains an independent set of compression strategies and encoding rules for each model to accommodate different models' maximum input lengths and labeling methods. Upon receiving the model type specified by the terminal or user, the server selects the corresponding prompt encoding scheme, thus maintaining optimal compression and inference efficiency even in a multi-model environment.

[0306] When used under various network conditions, the terminal can receive compressed prompts from the server, which helps reduce the amount of data transmitted over the network and lowers waiting time in low-bandwidth or high-latency network environments. When receiving compressed prompts, the terminal does not need to perform complex calculations; it only needs to store and forward data. Most computationally intensive tasks are concentrated on the server and model-providing device, thus achieving load balancing in edge-cloud collaboration.

[0307] In other application scenarios, such as legal consultation, programming assistance, and educational tutoring, users can generate prompts suitable for that specific field by selecting different information structures and combinations of constituent elements. The server also performs text simplification and data compression in these scenarios, thus consistently providing advantages in computational efficiency and generation quality across various applications.

[0308] Through the above-mentioned multiple implementation forms, this invention demonstrates how to utilize the collaboration of servers, terminals, and users to implement information structure, component assembly language, text simplification processing, and dedicated compression of prompt statements in the actual operating environment of generative artificial intelligence models, thereby achieving simultaneous improvement in the quality of prompt statement generation and overall system performance.

[0309] use Figure 13 The processing flow is explained.

[0310] Step 1: Users select information structure on the terminal The user views multiple information structure options provided by the server in the terminal interface and selects a target information structure (e.g., a medical expert information structure). The input is the list of information structures previously sent by the server and the user's click action; the output is data containing the identifier of the selected information structure. The terminal encapsulates the user's selection into a request message and sends it to the server over the network. Specific data processing performed by the terminal in this step includes: converting the user's click event into an internal data structure (such as a struct or key-value pair), writing the selected information structure ID, and serializing it into a format that can be transmitted over the network.

[0311] Step 2: The server generates a basic prompting framework based on the information structure. The server receives a request from the terminal containing an information structure identifier, with the information structure ID as input. The server retrieves the corresponding record from the database, reads the pre-stored role description, task domain description, and basic prompt text from the information structure, and outputs the basic prompt frame text. Specific data operations performed by the server in this step include: indexing the information structure table, concatenating a natural language frame from multiple fields, and assembling this frame along with relevant metadata into response data, which is then returned to the terminal.

[0312] Step 3: Users configure constituent elements on the terminal. Based on the basic prompt framework returned by the server, the user selects and configures constituent elements (such as role refinement, translation tasks, output format requirements, etc.) in the terminal interface. The input is the basic prompt framework and a list of constituent elements; the output is a set of user-selected constituent elements and their parameters. In this step, the terminal records the selection status and parameter values ​​of each constituent element as structured data. Specific data processing includes: generating a unique identifier for each constituent element, saving its type, name, and language parameters, and packaging all constituent elements into a request data structure.

[0313] Step 4: The terminal sends the component assembly language description to the server. After the user confirms the configuration, the terminal sends the recorded set of constituent elements as a constituent element assembly language description to the server. The input is a set of constituent elements in an internal data structure format; the output is a serialized request message. The specific operations performed by the terminal in this step include: converting the constituent element objects into a standard transmission format, encoding the data (such as UTF-8 encoding), and sending it to the specified interface of the server via a communication protocol.

[0314] Step 5: Server-side parsing of constituent elements assembly language The server receives a description of the constituent elements from the terminal. The input is structured data containing multiple constituent elements. The server iterates through each constituent element, queries its internal mapping table to obtain the corresponding natural language fragment template, and replaces the placeholders in the template with the parameters of the constituent element. The output is a set of instantiated natural language fragments. The data operations performed by the server in this step include: performing placeholder replacement on the string, concatenating parameter values, and adding category labels according to the constituent element type to form a list of fragments that can be used for subsequent combination.

[0315] Step 6: The server combines information structures and natural language fragments to generate an initial prompt statement. The server combines the basic prompt framework obtained in step 2 with the natural language fragments generated in step 5 in a predefined order. The input is the basic prompt framework text and a list of fragments; the output is the initial prompt statement. Specific data processing performed by the server in this step includes: determining the insertion position of each fragment according to rules (e.g., before or after the role description, within the task description paragraph), performing string concatenation and sentence boundary processing (e.g., adding punctuation and spaces), and forming a coherent natural language prompt statement.

[0316] Step 7: The server performs text simplification on the initial prompt statement. The server takes the initial prompt as input and simplifies the text using preset rules or models. The input is the unsimplified prompt; the output is the simplified prompt. Specific data operations performed by the server in this step include: using pattern matching to remove duplicate phrases, merging multiple similar statements into a single summary sentence, replacing lengthy expressions with phrases, and counting the number of characters or tags before and after simplification, thereby reducing text length while ensuring that information such as "role," "task," and "output format" is not lost.

[0317] Step 8: Server-executed prompt statement special encoding The server takes the condensed prompt text as input and applies a special encoding rule for prompt texts to replace high-frequency natural language expressions with short tags. The input is condensed natural language text; the output is an intermediate text representation containing tags. Specific data processing performed by the server in this step includes: searching a predefined phrase-to-tag mapping table, scanning the text line by line and replacing matching phrases, and generating text with placeholder tags through string replacement operations, thereby compressing text redundancy while preserving semantics.

[0318] Step 9: The server performs data compression on the encoded prompt text. The server encodes the tagged text generated in step 8 into a byte sequence, with the intermediate text representation as input. The server calls a general compression algorithm library to perform compression operations on the byte sequence, outputting a compressed byte stream, which is then re-encoded (e.g., converted into a transmittable text format). The specific data operations performed by the server in this step include: calculating the original byte length, calling the compression function to generate a shorter byte array, calculating the compression ratio, and converting the result into a text representation that can be embedded in the message body.

[0319] Step 10: The server selects a compression strategy based on length statistics. The server performs statistical analysis on the length and number of information units of the prompt statement before and after compression. The inputs are the length of the simplified text, the length of the encoded text, and the length after compression; the output is the selected compression method or compression ratio parameter. The specific data operations performed by the server in this step include: comparing the length at different stages with a preset threshold, and selecting whether to enable a stronger compression algorithm or increase or decrease the use of dedicated encoding rules based on the judgment results, thereby dynamically adjusting the compression strategy to achieve a better balance between time and space.

[0320] Step 11: The server sends compressed prompts to the terminal or model providing device. The server organizes the compressed prompt text along with necessary metadata (such as mapping table version and information structure ID) into a message body. The input is the compression result and metadata; the output is a network message sent to the terminal or generative artificial intelligence model providing device. The operations performed by the server in this step include: encapsulating the message header and message body, setting the target address and authentication information, and sending the message through the network transmission module.

[0321] Step 12: The terminal receives compression prompts and collects user task content. The terminal receives a compression prompt message from the server, with the input being the compressed prompt text sent by the server; the output consists of the locally stored compressed prompt text and the task content to be sent. The specific data processing performed by the terminal in this step includes: parsing the message body, extracting the compressed prompt text, prompting the user to input specific content (such as a description of the illness or a question) on the interface, and saving the user-inputted task text along with the compressed prompt text as input data for the generative artificial intelligence model to be invoked.

[0322] Step 13: The user enters the text to be processed in the terminal. Users input natural language text, such as descriptions of symptoms, problems, or needs, into the input area provided by the terminal. The input is the user's typing; the output is a piece of natural language content to be processed. In this step, the terminal converts the user's input character stream into an internal text object, performs basic encoding (such as UTF-8), length verification, and illegal character filtering to ensure that the data to be sent complies with communication and model interface requirements.

[0323] Step 14: The terminal assembles the model request and sends it to the model providing device or server. The terminal combines the compressed prompt text with the user input text to form a request to invoke the generative artificial intelligence model. The input consists of the compressed prompt text and the user task text; the output is a request message containing both of these parts. Specific operations performed by the terminal in this step include: constructing request fields (such as system prompts and user content), adding necessary control parameters (such as maximum output length and temperature coefficient), serializing it into a network message, and sending it to the model providing device or a server acting as a proxy.

[0324] Step 15: The server or model provider will decompress and restore the prompt message. The server or model provider receives a request from the terminal, with compressed prompt text and user task text as input. The server or model provider first performs reverse encoding and decompression operations on the compressed prompt text, restoring it to a simplified natural language prompt statement with its original semantics. The output is the restored prompt statement and the original user task text. The specific data operations performed by the server or model provider in this step include: decoding the compressed data in text form to restore the byte stream, calling a general decompression function to restore the intermediate text, and using a dedicated prompt statement mapping table to reverse-replace the markers with corresponding natural language phrases to reconstruct the complete prompt statement.

[0325] Step 16: The server or model provider encodes the prompts and task text into model input. The server or model provider combines the restored prompts with the user task text to form the contextual input for the generative artificial intelligence model. The input consists of natural language prompts and user content; the output is a model input vector in the form of a labeled sequence. Specific data processing performed by the server or model provider in this step includes: dividing the text into a labeled sequence using a tokenization algorithm (such as sub-word encoding), obtaining the embedding vector corresponding to each labeled element from a lookup table, and constructing an information tensor containing positional encoding and segmented labels as input for neural network inference.

[0326] Step 17: Servers or model-providing devices perform inference within generative artificial intelligence models. The server or model provider feeds the model input tensor into a multi-layer neural network for computation. The input is an encoded vector sequence; the output is a generated sequence of labels (i.e., response content). The specific data operations performed by the server or model provider in this step include: performing matrix multiplication, linear transformation, self-attention weight calculation, non-linear activation, and normalization operations at each layer; selecting the next label based on a probability distribution at the output layer, iterating until a termination condition is met. The server or model provider utilizes the weight parameters obtained during the pre-training and fine-tuning phases to ensure that the output content conforms to the role settings and task requirements in the prompt statement.

[0327] Step 18: The server or model provider will convert the generated tag sequence into response text. The server or model provider performs post-processing on the labeled sequence output by the generative artificial intelligence model. The input is the labeled sequence; the output is natural language response text. The specific data processing performed by the server or model provider in this step includes: mapping the labeled sequence back to words or characters, merging sub-words, processing special tags (such as line breaks and paragraph marks), and performing basic formatting to obtain structured response text, such as translation results and professional explanations.

[0328] Step 19: The terminal receives and displays the response text. The terminal receives response text from the server or model providing device, with the input being a natural language response; the output is the results area displayed on the interface. Specific operations performed by the terminal in this step include: parsing the response message, extracting the translation and explanation portions, and displaying them in sections according to a preset layout, such as presenting them under the titles "English Translation" and "Medical Explanation," allowing users to intuitively understand the model output.

[0329] Step 20: Users can then interact further or adjust configurations based on the results. Users view the response from the generative AI model on the terminal. The input is the text of the result displayed on the terminal; the output is the user's new operation command, such as continuing to ask follow-up questions, correcting the information structure, or adding or deleting components. After receiving a new user operation, the terminal sends an update request to the server through the above steps again, causing the server to regenerate the prompt statement and optimize the compression strategy, thereby further improving processing efficiency and response quality in subsequent calls.

[0330] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0331] With the widespread application of generative artificial intelligence models in scenarios such as medical data analysis, information retrieval, and recommendation services, users typically issue commands to these models using natural language prompts. However, existing technologies suffer from the following problems: (1) The prompts are mostly written by users in free text, lacking structured design. The differences in expression between different users are huge, which makes the generative artificial intelligence model's understanding of the task unstable, and the controllability and consistency of the output results poor, thus limiting the overall processing effect of the system.

[0332] (2) Prompt statements usually carry a large amount of contextual data. When interacting with large-scale generative artificial intelligence models, the input text often exceeds the reasonable token (symbol unit) limit or leads to excessive resource consumption. Existing systems mostly control the length by simply truncating the prompt statements without performing fine compression based on natural language processing and statistical processing. This not only affects semantic integrity but also fails to effectively reduce token usage while ensuring task constraints, resulting in low computational resource utilization efficiency.

[0333] (3) The existing prompt statement generation and compression process is usually unrelated to the user's emotional state. When the user is in different emotions such as tension, anger, and anxiety, the system still outputs the results with a uniform amount of information, style and interaction rhythm, which can easily cause cognitive load mismatch or inappropriate communication tone, which is not conducive to improving the human-computer interaction experience, and may also indirectly affect the user's understanding and adoption of the generated results.

[0334] (4) Existing emotion recognition technologies are mostly used only for simple emotion tag display or content recommendation, and are not deeply coupled with the prompt statement generation pipeline. They lack an integrated technical solution from "information structure distribution, block instruction construction, prompt statement compression control" to "dynamic emotion adaptive adjustment", making it difficult to achieve fine management and intelligent optimization of the input of generative artificial intelligence models at the system level.

[0335] Therefore, it is necessary to provide a new system that distributes information structures on the server side, assembles block instructions, performs multi-level compression and constraint preservation processing on prompt statements, and combines sentiment estimation results based on multimodal data to dynamically adjust the information content, style, and compression intensity of prompt statements. This can improve the instruction understanding quality of generative artificial intelligence models and the user interaction experience while reducing token usage and improving the efficiency of computing resource utilization, thereby achieving an improvement in computer technology itself.

[0336] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0337] In this invention, the server includes: means for distributing multiple information structures to a user terminal based on structured information from a user terminal; means for generating instruction information for constructing prompt statements for a generative artificial intelligence model by using a component assembly description of multiple components corresponding to the information structures and combining them through visual operations; means for generating prompt statements for causing the generative artificial intelligence model to perform specific processing based on the instruction information and target data obtained from the user terminal; means for applying compression processing based on natural language processing and statistical processing to the generated prompt statements and target data to reduce the number of symbol units contained in the prompt statements and control the symbol utilization rate, and removing redundant information through syntax analysis and summarization processing during the compression process while maintaining constraint information regarding role designation, processing content designation, and output format designation in the generative artificial intelligence model; means for performing emotion estimation processing on the user's emotional state based on image information, voice information, and text information obtained from the user terminal; and means for dynamically adjusting the information content, style, and compression processing intensity contained in the prompt statements according to the estimated emotional state, sending the adjusted prompt statements to the generative artificial intelligence model, and outputting response information obtained from the generative artificial intelligence model to the user terminal. This allows for the formation of an integrated processing chain on the server side, encompassing information structure distribution, block instruction construction, prompt statement compression control, and emotion adaptive adjustment. This makes the input of generative artificial intelligence models more standardized in structure, finely controlled in length, and adaptively matched with the user's emotional state in terms of content and style. Consequently, it reduces token consumption during model invocation, improves the efficiency of computing resource utilization, stabilizes the quality of generated results, and significantly improves the human-computer interaction experience, achieving a comprehensive improvement in computer information processing technology and human-computer interface technology.

[0338] "Information structure" refers to an abstract structure or template set that is predefined by the server and can be distributed to user terminals to guide users in constructing instructions and organizing data. It includes field layout, hierarchical relationships, and placeholder information corresponding to the constituent elements.

[0339] "Constituent elements" refer to the basic units that can be combined and reused in information structures or prompt statements, including modular fragments used to describe roles, tasks, data segments, output formats, tone constraints, etc.

[0340] "Component assembly description" refers to a description method that combines multiple components in a visual or structured way to form complete instructions or prompts. It supports operations such as block dragging, sequential connection, and parameter filling.

[0341] "User terminal" refers to an electronic device used to interact with a server, including but not limited to mobile terminals, wearable terminals, and fixed terminals, which can collect user input, display, or output information returned by the server.

[0342] "Generative artificial intelligence models" refer to data processing models that automatically generate text, code, images, or other content based on input prompts. They are usually based on deep learning structures and learn from large-scale data distributions to produce semantically and contextually appropriate outputs.

[0343] "Prompt statements" refer to textual information used to specify the role, task, input data range, output format, and tone of a generative artificial intelligence model. As one of the main inputs of a generative artificial intelligence model, they are used to guide the model to perform specific processing.

[0344] "Target data" refers to the set of information obtained from the user terminal and provided to the server and generative artificial intelligence model along with the prompt statement for analysis, processing or as context, including structured data, semi-structured data and unstructured data.

[0345] "Structured information" refers to input information with predefined fields or formats that can be directly parsed by the server and mapped to information structures or constituent elements, including form data, configuration parameters, and label data.

[0346] "Compression processing" refers to the process of reducing the length and eliminating redundancy of prompt statements and related text data. It reduces the number of symbol units through natural language processing, statistical analysis, and summary generation, while trying to maintain the semantic and constraint information as much as possible.

[0347] A “symbol unit” refers to the smallest unit of text used to measure length and resource usage in a prompt statement or text. It can be a character, subword, word, or token defined by the model segmenter.

[0348] "Symbol utilization" refers to the ratio between the number of symbol units occupied by prompts and related data when interacting with a generative artificial intelligence model and the upper limit of available symbol resources. It is used to measure the efficiency of input resource utilization.

[0349] Natural Language Processing (NLP) is a general term for technologies that automatically process text, such as word segmentation, part-of-speech tagging, syntactic analysis, named entity recognition, and summary generation. It is used to support the analysis, compression, and reconstruction of prompt statements.

[0350] "Statistical processing" refers to the process of evaluating the importance of each part of a text based on statistical characteristics such as frequency, weight, and importance score, and then deciding whether to retain or delete the content accordingly.

[0351] "Syntactic analysis processing" refers to the process of parsing the syntactic structure of prompt statements, identifying the dependency relationships between words and the sentence structure, thereby discovering redundant modifiers and removable components.

[0352] "Summarization" refers to the process of compressing and rewriting long texts to generate shorter texts while preserving the main semantics and key information. This process is usually completed automatically through rules or models.

[0353] "Emotion estimation processing" refers to the process of identifying and determining a user's current emotional state based on image, voice, or text information, including identifying the emotion category and its intensity.

[0354] "Emotional state" refers to a user's psychological tendency or emotional category at a specific point in time, including but not limited to happiness, anger, sadness, tension, relaxation, etc., and may be accompanied by the degree of intensity.

[0355] "External emotion analysis and processing device" refers to a computing device or service that operates independently of this system and provides emotion recognition functions through a network. It is used to perform emotion analysis on facial expression data, voice data, etc., and return the analysis results.

[0356] "Internal sentiment analysis and processing function" refers to the sentiment recognition module or software function implemented within the system server or user terminal, which can perform sentiment analysis on input data without calling external services.

[0357] "Main sentiment category" refers to the dominant sentiment type determined by a combination of multiple sentiment scores in the sentiment estimation processing results, which is used to drive the adaptive adjustment of subsequent prompts.

[0358] "Emotional intensity" refers to a quantitative indicator corresponding to the main emotional category, used to represent the significance or confidence level of the emotion.

[0359] "Information content" refers to the richness and level of detail of the effective content contained in the prompt statement, including the level of detail in the task description, the amount of background information, and the number of constraints.

[0360] "Style" refers to the stylistic features of a prompt, including formality, politeness, emotional tone, and sentence complexity.

[0361] "Compression intensity" refers to the degree to which the length of the prompt statement is reduced and content is deleted during the compression process, including the compression ratio, target length, and the degree of simplification of the information retained.

[0362] "Response information" refers to the output content of the generative artificial intelligence model after receiving prompts and target data, including analysis conclusions, suggestions, explanations, generated text, etc., which is forwarded by the server to the user terminal.

[0363] The embodiments of this invention will be described in terms of the division of labor among the server, client, and user. In conjunction with the specific hardware and software configuration, data structure, algorithm flow, and internal structure of the generative artificial intelligence model, it will be explained how this system achieves structured generation, compression control, and adaptive emotional adjustment of prompt statements, as well as the resulting improvements in computer technology.

[0364] I. Overall System Composition The server is configured as one or more information processing devices. The server can be a computer device with a multi-core central processing unit, main memory, and persistent storage, such as a general-purpose server equipped with a multi-core general-purpose processor and at least several GB of memory. The server runs an operating system, such as a Unix-like operating system. The server's software architecture includes: The server runs a network application framework to provide application interfaces based on the Hypertext Transfer Protocol. The server runs a database management system, which is used to store information structure, component configuration, user configuration files and log data. The server runs natural language processing libraries, such as text analysis components based on statistical and neural network models, for syntactic analysis and lexical importance calculation; The server runs a model inference library for text summarization and rewriting, and is used to compress prompt statements; The server runs a communication module used to call the interface of generative artificial intelligence models; The server runs a sentiment analysis module, which is used to estimate the sentiment of image, speech, and text information.

[0365] The terminal is configured as a network-connected electronic device, such as a smartphone, head-mounted display, tablet, or desktop terminal. It is equipped with a display unit, input unit, camera unit, and voice acquisition unit, and runs a client application or web application interface for data interaction with the server.

[0366] Users interact with the server through the terminal's display interface, performing operations such as template selection, element combination, data input, and result confirmation. Simultaneously, users provide the server with image and audio data for emotion recognition via the camera and voice acquisition units.

[0367] II. Management of Information Structure and Components The server stores various "information structures" in persistent storage in the form of a relational or key-value database. The server records the identifier, name, applicable scenarios, and associated set of constituent elements for each information structure through database tables or key-value sets. The server uses a set of fields to represent the logical layout within the information structure, such as role description fields, task description fields, data description fields, output format fields, and tone constraint fields.

[0368] At the component level, the server defines corresponding component metadata for each type of instruction function (such as "role assignment," "task description," "data injection," "output format setting," and "tone constraint setting"). Each component has a type field, default text, editable parameter list, and connectable position constraints in the database. Based on this metadata, the server generates a set of structured descriptions when distributing to the client, which the client renders as visual building blocks.

[0369] After receiving the component configuration from the server, the client displays each component as a block-shaped graphic on the display unit, allowing users to combine them by dragging and clicking. Once the user completes the combination, the client sends the result to the server as structured data (e.g., a tree or graph structure). Upon receiving the data, the server maps this structure back to a combination of component identifiers and parameter values, thus reconstructing the structured instruction information on the server side.

[0370] By employing an information structure and component-based assembly description, the server breaks down traditional free-text prompts into a controllable set of modules. This structured management allows the server to avoid logical inconsistencies during the composition phase based on type and location constraints, and to selectively retain or prune modules based on their importance during the compression phase, thus improving the way instructions are represented from the perspective of the computer's internal structure.

[0371] III. Prompt Statement Generation and Data Fusion When generating prompts, the server takes the combined structure of the constituent elements as input. The server first generates the system role description based on the role's constituent elements, for example: "You are a medical data security expert." The server reads the user-defined task objective from the task components, for example: "Please analyze the following medical system configuration and access logs, focusing on assessing whether there is a risk of leakage of patient personal information." The server inserts a summary of the target data received from the user terminal into the data components. The server uses a text processing module to preprocess the raw target data, including field filtering, sensitive information removal, paragraph segmentation, and key sentence extraction. The server appends the processed data as part of the prompt statement, displaying "The following is the data:...".

[0372] The server incorporates constraints on the output format of generative artificial intelligence models into the output format components, for example: "Please list the main risk points and provide actionable improvement suggestions for each risk." The server can add specific style requirements to the elements constituting tone constraints, for example: Please answer users' questions in a calm and reassuring tone. The server assembles these parts in a predetermined order to form a preliminary prompt statement, for example: "You are a healthcare data security expert. Please analyze the following healthcare system configuration and access logs, focusing on assessing the risk of leakage of patient personal information. The data is as follows: ... Please list the main risk points and provide actionable improvement recommendations for each risk." The server can generate different prompts depending on the application scenario. For example, in a product recommendation scenario, the server can generate: "You are an e-commerce product recommendation expert. The user is currently in a relaxed mood and hopes to get some novel and interesting product suggestions. Based on the user's interest tags and recent browsing history below, please recommend 5 products and explain the reasons for each recommendation in 2-3 sentences using light and interesting language." In user complaint handling scenarios, the server can generate: "You are a customer service communication expert. The user is currently quite angry, so please pay special attention to using respectful and understanding language. Below is the complaint the user just entered. Please first express your understanding and apology in a paragraph, and then explain the possible solutions in bullet points." By using this component-driven prompt generation method, the server internally replaces traditional indivisible text with a parsable data structure, thus providing a granular foundation for subsequent compression control and emotion adjustment.

[0373] IV. Compression of Prompt Statements and Control of Symbol Usage After receiving the initial prompt, the server performs multi-level compression. The server then calls a natural language processing library to perform word segmentation, part-of-speech tagging, and dependency parsing on the prompt. Based on part-of-speech and dependency relations, the server identifies modifiers (such as adjectives, adverbs, and modifiers) and optional clauses. The server assigns an importance score to each word or phrase, based on word frequency, part-of-speech, position in the syntactic structure, and type of constituent element (e.g., keywords for role designation and output format have higher weight than modifiers).

[0374] At the statistical processing level, the server removes words below a threshold based on importance scores, thereby reducing the number of symbolic units. For example, "Please analyze the following complex medical data content carefully, comprehensively, and in great detail" is reduced to "Please analyze the following medical data." This step is algorithmically equivalent to a feature selection process, reducing the burden on subsequent model inputs by discarding unnecessary redundant features for generative artificial intelligence model decision-making.

[0375] The server further invokes a summarization model based on a transformer architecture to perform semantic compression on the longer data description portion. This summarization model can employ a multi-layer self-attention network structure, taking serialized prompts or target data text as input and outputting a shorter summary. By controlling the maximum input length, maximum output length, and decoding strategies (such as bundle search width and length penalty coefficient) of the summarization model, the server reduces the total number of symbol units while preserving key meanings.

[0376] The server encodes the compressed prompt statement using a tokenizer, calculates the actual number of symbol units, and compares it with a preset threshold. If the threshold is exceeded, the server repeats the summarization and trimming process, focusing on further blurring and reducing the background and exception descriptions without affecting role assignments, task descriptions, and output format constraints. Because the server possesses information at the component level, it can prioritize retaining the corresponding text of role components and output format components, structurally ensuring that task constraints are not violated.

[0377] Through the above processing, the server internally exercises fine-grained control over the text input to the generative AI model, enabling it to handle larger datasets or more concurrent requests with the same computing resources, thereby improving the system's computational efficiency and response speed. This compression process based on syntactic and statistical features is a specialized optimization strategy for large-scale language model inputs, rather than simple string truncation.

[0378] V. Sentiment Estimation and Adaptive Regulation During user interaction, the terminal periodically captures facial images of the user through the camera unit and records user speech segments through the voice acquisition unit. The terminal reduces the size and formats the images, extracts feature parameters (such as Mel-frequency cepstral coefficients) from the speech, and then sends them to the server.

[0379] In the sentiment analysis module, the server inputs image features into a convolutional neural network or a transformer-based visual coding network, outputting the probability distribution of each sentiment category; it inputs speech features into a temporal neural network or a one-dimensional convolutional network, outputting sentiment-related parameters; and it inputs text into a language model, outputting sentiment polarity and intensity. The server then generates a main sentiment category and its corresponding intensity value using a weighted fusion algorithm or a rule-based fusion strategy.

[0380] The server adaptively adjusts the prompts based on the primary emotion category and intensity. At the information content level, it adds explanatory paragraphs and optional analysis points for happy or curious states; for anxious or angry states, it reduces non-critical explanations, retaining concluding statements and the three most important points. At the stylistic level, it replaces or adds reassuring expressions, such as short sentences like "Please don't be overly nervous; these risks can be effectively controlled through the following measures." At the compression intensity level, it uses the output length of the emotion regulation summarization model and deletion thresholds to provide shorter and easier-to-understand instructions when the user is anxious.

[0381] This emotion-driven modulation is a dynamic adjustment process of control parameters within the server. By reducing semantic redundancy and adjusting tone and word choice, it optimizes the user's understanding of the results and reduces psychological burden. Unlike traditional systems that recommend information solely based on content themes, this invention directly modulates the signal at the input layer of the generative artificial intelligence model, thereby influencing the content structure and style of the model's output.

[0382] VI. Description of Generative Artificial Intelligence Model Invocation and Internal Structure After completing structured generation, compression, and sentiment modulation, the server sends the final prompt along with the necessary target data to the generative AI model. This generative AI model can be an autoregressive language model based on a multi-layer transformer architecture, which consists of multiple stacked self-attention sublayers and feedforward network sublayers, and includes a word embedding layer, a positional encoding layer, and an output layer.

[0383] When invoked, the server encodes the prompts as a sequence of symbolic units and transforms them into a vector sequence within the model using an embedding matrix. At each layer, the model calculates dependencies between different positions using a multi-head self-attention mechanism, then combines the outputs through residual connections and layer normalization. The model is optimized during training using a cross-entropy loss function based on a large corpus; the server does not train it again during the usage phase, but only performs forward propagation in inference mode.

[0384] The compressed prompts provided by the server reduce the time complexity of the model's self-attention computation, as the computational cost of attention is proportional to the square of the input length. The compression step directly reduces the length of the sequences involved in the attention computation, thereby reducing inference latency and energy consumption under the same hardware conditions. Simultaneously, by preserving key constraint information, the model can still produce task-appropriate outputs with shorter inputs, ensuring accuracy.

[0385] VII. Technical Effects and Causal Relationships Through the aforementioned structured information management, grammatical and statistical compression, and sentiment adaptive adjustment, the server has achieved the following improvements at the computer technology level: The server uses a component-based assembly description to represent the prompt statement as a parsable structure, which allows the server to execute compression strategies at the module level rather than blindly truncate at the character level. This allows more task-critical information to be preserved at the same compression ratio, improving the relevance and accuracy of the output of the generative artificial intelligence model. The server identifies redundant components before compression by performing syntactic analysis and importance calculation, transforming the compression process into a feature selection-like process. This reduces computational load while retaining the features that contribute most to the model's predictions, thereby reducing errors. The server dynamically adjusts the amount of information and style based on the sentiment estimation results, making the output more match the user's cognitive state, thereby reducing repeated queries and clarification requests on the user side. This reduces the total number of requests and network communication load at the overall system level. By setting different compression intensity curves for different emotional states, the server prioritizes providing more symbol budget for scenarios that require more detail, and provides short and concise answers for tense scenarios that require rapid response, thus establishing an adjustable balance between time complexity and user experience, given limited hardware resources.

[0386] VIII. Optional Implementation Forms and Variations In one implementation, the server can run the sentiment analysis model locally without calling external sentiment analysis devices, thus enabling adaptive sentiment processing even in unstable network environments. In another implementation, the server can use different types of summarization models or self-trained models to adapt to different languages ​​or professional domains.

[0387] In one implementation, the endpoint can provide only text input and perform sentiment estimation solely based on the text content; in this case, the server only uses the text sentiment analysis module. In another implementation, the endpoint can be a head-mounted display terminal that continuously captures the user's gaze and subtle facial expressions to update the emotional state more frequently, allowing the server to adjust the content and length of prompts more in real-time.

[0388] In a healthcare security scenario, users input access log summaries and configuration descriptions from the hospital information system on their devices. The server generates and compresses prompts, sends them to a generative AI model, and the model's output is transformed by the server into a structured risk list and improvement suggestions, which are then returned to the device. Users then perform system hardening operations based on the structured information output by the server. Because the compression strategy reduces unnecessary context, the generative AI model can quickly locate key risk points within a shorter context, improving diagnostic efficiency.

[0389] In a product recommendation scenario, a user expresses on their device, "I'm feeling a bit down today, I want to buy something to relax." The server analyzes the text and finds the sentiment to be slightly negative, adjusting the prompts to a gentle, soothing style and controlling the compression intensity, allocating more symbolic resources to comforting content describing the recommendation reasons. The recommendation list output by the generative AI model is displayed on the device, allowing users to understand the recommendation reasons with minimal cognitive burden.

[0390] Through the above embodiments, this invention integrates information structure distribution, block instruction construction, compression control, sentiment estimation, and generative artificial intelligence model invocation into a unified technical solution. It improves computer technology at the levels of input representation, algorithm processing, and resource utilization, and achieves comprehensive optimization of processing speed, resource utilization efficiency, output accuracy, and user experience.

[0391] use Figure 14 The processing flow is explained.

[0392] Step 1: The server initializes system configuration and model resources.

[0393] Input: configuration file, environment variables, default model path and database connection information.

[0394] Output: A serviceable application environment, a loaded natural language processing model, a summary model, and a database connection pool.

[0395] The server reads the generative artificial intelligence model interface address, authentication key, database host information, and sentiment analysis service parameters from the configuration file and environment variables; the server calls the natural language processing library to load the word segmentation, part-of-speech tagging, and syntactic analysis models, and calls the summary model loading function to load the pre-trained transformer model into memory; the server establishes a long connection or connection pool to the database to provide a foundation for subsequent queries of information structure and constituent elements.

[0396] Step 2: Each end requests the information structure and configuration of its constituent elements from the server.

[0397] Input: User-selected application scenario identifier (e.g., medical safety diagnosis, product recommendation), user terminal identifier.

[0398] Output: A list of information structures and a configuration interface for constituent elements displayed on the client.

[0399] The client sends a request containing scene identifiers to the server via the network; after receiving the JSON data returned by the server, the client displays the information structure name, description and corresponding constituent elements in a graphical way on the interface, providing a basic interface for the user to combine prompts in subsequent steps.

[0400] Step 3: The server retrieves information structure and constituent element metadata from the database based on the scene identifier.

[0401] Input: Scene identifier from the end.

[0402] Output: Structured data containing information structure definitions, component types, and default text.

[0403] The server executes a query in the database, retrieves the corresponding information structure record and associated constituent element record based on the scene identifier; the server organizes the query results, encapsulates the type, editable parameters, default text and connection rules of each constituent element into structured data, and returns it to the client; this data provides input for the client to render the block editor.

[0404] Step 4: Users can construct the prompt statement structure by dragging and editing the constituent elements on the device.

[0405] Input: The information structure and constituent element blocks displayed on the terminal.

[0406] Output: Structured instruction data representing the order of combination of constituent elements and the content of parameters.

[0407] Users drag and drop "role blocks," "task blocks," "data blocks," and "output format blocks" on a touchscreen or pointing device, connecting them in a logical order. Users click on the editable area in each block to enter specific content, such as "medical data security expert" or "assess privacy leakage risks." After the user's operation, the terminal converts the block structure into a tree or ordered list structure, caches it locally, and sends the structure as structured data to the server when the user confirms.

[0408] Step 5: Users input or upload target data on the client side.

[0409] Input: Input forms and file upload interfaces on the client side.

[0410] Output: Target data after end-to-end preprocessing (such as text logs, configuration instructions).

[0411] Users can input medical system configurations, access log summaries, or product browsing records into forms using a keyboard or touchscreen; users can also select files to upload; the client performs format recognition and preliminary conversion on the files, converting structured files into embeddable text or summary data; the client then sends the user-input text and the converted data, along with the constituent elements, to the server.

[0412] Step 6: The server receives the constituent structure and target data, and generates an initial prompt statement.

[0413] Input: The combination structure of constituent elements from the end and the target data text.

[0414] Output: Uncompressed initial prompt statement.

[0415] The server parses the constituent elements structure, determines the order of the blocks and the parameters of each block; the server generates role instruction text based on the "role block", generates task description text based on the "task block", inserts the target data into the corresponding position of the "data block", and converts the contents of the "output format block" and "tone constraint block" into output requirement text; the server combines the above texts in sequence through string concatenation or template filling to obtain a complete but uncompressed prompt statement.

[0416] Step 7: The server preprocesses and summarizes the target data to control the size of the data portions.

[0417] Input: Raw target data text from the end.

[0418] Output: Preprocessed and summarized target data text.

[0419] The server uses a natural language processing library to perform sentence segmentation, remove meaningless symbols, and delete significantly redundant paragraphs on the target data. The server calls a summarization model to automatically summarize excessively long data segments, compressing lengthy details and retaining only key configuration items or key log events. Through this data processing, the server generates data text suitable for embedding prompt statements, providing a foundation for subsequent overall compression.

[0420] Step 8: The server performs syntax analysis and importance calculation on the initial prompt statement.

[0421] Input: Initial prompt text.

[0422] Output: An intermediate representation of each word or phrase, labeled with its part of speech and importance.

[0423] The server calls a natural language processing library to segment and tag the prompt statements, and constructs a dependency syntax tree. The server calculates the importance of each word or phrase according to preset rules, such as part of speech, position in the syntax tree, type of constituent element, and distance from key task words. The server attaches these importance tags to the internal data structure as the basis for subsequent compression decisions.

[0424] Step 9: The server removes redundant components based on importance and syntactic structure, performing the first stage of compression.

[0425] Input: A prompt statement with importance ratings in the middle.

[0426] Output: Concise suggestion statements with low-importance words and phrases removed.

[0427] The server identifies adjectives, adverbs, and verbose modifiers with importance below a threshold and removes or merges them while maintaining syntactic coherence; for example, it simplifies "analyze carefully, comprehensively, and in great detail" to "analyze"; the server reconstructs sentences to ensure grammatical correctness; through this data calculation, the server reduces the number of symbol units without significantly altering the semantics.

[0428] Step 10: The server calls a digest model to perform a second stage of compression on the prompt statement.

[0429] Input: The prompt statement after the first stage of compression.

[0430] Output: A further shortened but semantically complete prompt.

[0431] The server feeds the prompt statements as input sequences into the summarization model. The summarization model calculates the internal representation based on multi-head self-attention and generates short text during the decoding stage. The server sets the target length and decoding strategy to ensure that the output text covers the role specification, task requirements, and output format constraints as much as possible. The server checks the summarization results. If key sentences are found to be missing, the key sentences are re-inserted from the original text to generate the final compressed version.

[0432] Step 11: The server calculates the number of symbol units in the prompt statement and determines whether the resource constraints are met.

[0433] Input: The compressed prompt text.

[0434] Output: The symbol cell count and the result of whether the threshold is exceeded.

[0435] The server calls a word segmenter to encode the prompt statement, obtaining a sequence of symbol units; the server counts the sequence length and compares it with a pre-set maximum acceptable length; if the length exceeds the threshold, the server marks it as needing further compression, otherwise it is marked as usable directly; this counting process provides a basis for decision-making on whether to continue compression.

[0436] Step 12: The server may further summarize or trim portions of the data as needed.

[0437] Input: The data portion of the prompt statement and the length exceeding the limit marker.

[0438] Output: The final data portion of the text with a controlled number of symbols.

[0439] The server identifies fragments of the target data in the prompt statement and deletes or merges lower-priority fields and minor instances. The server can then call the summary model again to summarize the fragments containing only data, ensuring that key parameters or representative sample records are preserved. Through this targeted pruning, the server keeps the overall number of symbols within the resource-allowed range while maintaining the integrity of the task description.

[0440] Step 13: The system collects users' image, voice, and text information on each device for sentiment estimation.

[0441] Input: The user's current facial expression, voice clip, and text input on the interface.

[0442] Output: Raw sentiment-related data or feature data transmitted to the server.

[0443] The device captures the user's facial image through the camera unit, collects the user's speech through the microphone, and records the words the user uses in the text input box. The device compresses the image, extracts feature parameters from the speech, and performs simple preprocessing on the text. The device packages this data and sends it to the server to provide input for the server to perform sentiment estimation.

[0444] Step 14: The server uses a sentiment analysis module to estimate the user's emotional state.

[0445] Input: Image data, speech features, and text content from the end user.

[0446] Output: Main sentiment category and sentiment intensity value.

[0447] The server inputs image data into a visual emotion classification network to obtain probabilities of multiple emotion dimensions; it inputs speech features into a speech emotion recognition network to obtain emotion tendencies; and it inputs text into an emotion analysis model to obtain emotion polarity and confidence. The server then integrates the multi-source results into a main emotion category (such as "anxiety", "happiness", "anger", "relaxation") and corresponding intensity scores using a weighted fusion algorithm, which serves as the basis for subsequent adaptive adjustment.

[0448] Step 15: The server dynamically adjusts the amount of information and style of the prompts based on the emotional state.

[0449] Input: Compressed prompts and sentiment estimation results.

[0450] Output: A prompt statement that has been adjusted in terms of information content and style.

[0451] Based on the primary emotion category and intensity, the server adds extra explanatory statements and background information for happy or curious states, and removes secondary details while retaining the core points for anxious or angry states. In terms of style, the server adds soothing phrases or gentle expressions for negative emotions and replaces words that may trigger tension. Through this targeted text modification, the server achieves emotional adaptation in both content and style.

[0452] Step 16: The server adjusts the compression intensity parameters based on the emotional state.

[0453] Input: Current compression strategy parameters, sentiment estimation results.

[0454] Output: Compression length target and deletion threshold adapted to emotional state.

[0455] When the server is in a happy or relaxed state, it reduces the compression ratio and relaxes the limit on the number of symbols to make the prompts more detailed; when the server is in an anxious or time-sensitive state, it increases the compression ratio and tightens the upper limit on the number of symbols to make the prompts shorter and more focused. The server converts emotional information into specific algorithm parameters by updating the target output length and importance deletion threshold of the summary model, thereby optimizing the efficiency of subsequent model calculations and user understanding.

[0456] Step 17: The server sends the final prompt and necessary target data to the generative artificial intelligence model.

[0457] Input: Compressed and sentiment-adaptive prompts and simplified target data.

[0458] Output: The response text returned by the generative artificial intelligence model.

[0459] The server constructs a request body containing system messages and user messages, embedding prompts and target data into the message content; the server sends the request to the generative AI model server via a network interface; the generative AI model performs self-attention calculation and decoding internally and then returns the generated text; the server receives the response text, preparing for subsequent structured processing.

[0460] Step 18: The server parses and structures the response text.

[0461] Input: Natural language response text returned by the generative artificial intelligence model.

[0462] Output: Structured risk lists, suggestion lists, or recommendation lists, etc.

[0463] The server uses keyword detection and pattern matching methods to find chapter titles and serial number markers in the response text, and splits the text into multiple entries. The server maps the content of the entries into structured fields, such as "title", "detailed description", "suggested measures", etc. The server combines these fields into a list or dictionary form to provide an organized data structure for the end-to-end visualization.

[0464] Step 19: The system receives structured results from each end and generates a display interface that matches the mood.

[0465] Input: The structured results returned by the server and the main sentiment category.

[0466] Output: The result interface displayed or played on the device.

[0467] The app selects the display mode based on the main emotional category: in an anxious state, it prioritizes displaying the most critical items and folds and hides the others; in a relaxed state, it directly expands all the details; in terms of visual design, the app can bold or highlight key risks or recommended items to help users understand them quickly; the app can also read the results to the user through the speech synthesis module.

[0468] Step 20: Users make decisions or make new requests based on the displayed results.

[0469] Input: Risk assessment, improvement suggestions, or recommended content displayed on the device.

[0470] Output: New operation instructions or feedback information.

[0471] After reading or listening to the results on the client side, users decide whether to adjust the medical system configuration, take certain security measures, or select certain products. If users have further questions, they can enter new questions in the input box or ask questions via voice. The client side submits the new questions along with contextual information to the server, triggering a new round of processes from information structure selection to prompt statement generation, compression, and emotion adaptation.

[0472] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0473] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0474] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0475] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0476] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0477] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0478] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0479] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0480] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0481] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0482] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0483] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0484] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0485] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0486] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0487] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0488] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0489] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0490] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0491] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0492] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0493] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0494] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0495] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0496] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0497] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0498] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0499] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0500] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0501] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0502] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0503] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0504] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0505] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0506] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0507] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0508] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0509] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0510] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0511] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0512] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0513] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0514] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0515] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0516] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0517] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0518] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0519] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0520] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0521] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0522] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0523] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0524] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0525] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0526] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0527] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0528] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0529] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0530] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0531] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0532] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0533] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0534] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0535] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0536] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0537] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0538] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0539] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0540] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0541] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0542] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0543] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0544] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0545] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0546] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0547] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0548] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0549] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0550] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0551] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0552] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0553] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0554] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0555] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0556] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0557] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0558] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0559] In addition, the following notes are provided in response to the above explanation.

[0560] Example 1 (Note 1) An information processing system, characterized in that it comprises: A unit that reads various template data stored in a hierarchical structure from an information storage device by a processing unit of an information processing device, and sends the template data to an information display device via a network; The information display device displays multiple block elements contained in the template data as interface elements for visual operation, and generates block sequence data based on the user's selection, arrangement and parameter input of the block elements, and sends the block sequence data to the unit of the information processing device. The information processing device parses the identifiers of each block element contained in the block sequence data and their corresponding parameter values, and expands the block sequence data into natural language text according to a predetermined conversion rule, thereby generating units of raw data for prompt statements for generative artificial intelligence models. The information processing device performs text conversion processing on the original data of the prompt statement, including merging synonyms, deleting redundant expressions, and shortening sentence structure, to generate compressed candidate data. The device also calculates and compares the number of tags in the original data of the prompt statement and the compressed candidate data, and selects the unit of the final data of the prompt statement to be used according to predetermined conditions. The information processing device sends the final data of the prompt statement to the inference processing device of the generative artificial intelligence model, and after obtaining the response data from the inference processing device, sends the response data to the unit of the information display device; The information display device presents the response data to the user in a visual manner, and modifies the composition of the block sequence data according to the user's re-editing instructions, so as to regenerate the control information of the original data of the prompt statement and send the control information to the unit of the information processing device.

[0561] (Note 2) According to the information processing system described in Appendix 1, the processing unit is configured to include block elements in the template data that include role setting block elements for assigning domain-specific expert roles to generative artificial intelligence models, and translation block elements for converting the original data or final data of the prompt statements into different languages.

[0562] (Note 3) According to the information processing system described in Appendix 1, the processing unit is configured to: in the process of converting the original data of the prompt statement into the final data of the prompt statement, use the reduction in the number of markers and the generation time of the response data obtained from the generative artificial intelligence model as processing efficiency evaluation indicators, and dynamically adjust the conversion rules or compression rate used for the text conversion processing according to the evaluation results.

[0563] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for structuring user-oriented information into various types of template information in an information processing device and sending the template information to an external device; A device for obtaining a selection result of the template information from the external device and information of multiple constituent elements arranged in a visual manner, and constructing a prompt statement for representing the instruction content of the generative artificial intelligence model according to the order of the constituent element information and their respective input values. A means for performing compression processing on the prompt statement to reduce the length of the symbol sequence, and controlling the compression method or compression rate in the compression processing according to the structure or usage of the prompt statement, so as to reduce the number of symbol units used in the generative artificial intelligence model; A device for sending the compressed prompt statement or its decompression result to the generative artificial intelligence model, and for sending the generated content information obtained from the generative artificial intelligence model to the external device; A device for enabling the external device to display the template information and the constituent element information in the form of a visually configurable editing interface, and for receiving user interface processing of adding, deleting or changing the order of the constituent element information according to the user's operation.

[0564] (Note 2) The information processing system according to Appendix 1 is characterized in that, The constituent element information includes role setting information for defining the role of artificial intelligence according to professional fields, and translation instruction information for instructing the semantic content of the prompt statement to be converted into different language systems. The professional field information and language type information are embedded into the prompt statement by selecting and setting at least one of the role setting information and the translation instruction information on the editing interface.

[0565] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to measure the number of symbol units or the data length of the prompt statement before and after compression, and control the length of the prompt statement or the number of constituent element information used for input to the generative artificial intelligence model based on the measurement result, so as to improve the overall processing efficiency using the generative artificial intelligence model.

[0566] Example 2 (Note 1) An information processing system, characterized in that it comprises: A means for distributing an information structure to a user and providing a general framework, as said information structure, for defining the roles assigned to generative artificial intelligence models and the domains of objects they process; An apparatus for receiving multiple constituent elements described in a constituent element assembly language based on the information structure, acquiring natural language fragments corresponding to each constituent element, and integrating the natural language fragments into the information structure to generate prompt statements for instructing specific actions to the generative artificial intelligence model. Apparatus for performing text simplification on the prompt statement to reduce redundant expressions and simplify sentence structure while maintaining semantics, thereby obtaining a text-simplified prompt statement; An apparatus for compressing the simplified text prompt statement using a general compression algorithm or a special encoding rule for the prompt statement, so as to generate a compressed prompt statement with reduced information unit usage. A device for sending the compressed prompt statement to a terminal or providing the generative artificial intelligence model, and a device for obtaining response information obtained by decompressing and processing the compressed prompt statement through the generative artificial intelligence model, and a device for providing the response information to the user.

[0567] (Note 2) The information processing system according to Appendix 1 is characterized in that, The constituent elements of the component-assembly-type language description include role-assignment elements for assigning domain-related roles to the generative artificial intelligence model and translation elements for converting response content generated based on the prompt statement between different natural languages. The system is configured to integrate natural language fragments corresponding to the role-assignment elements and the translation elements into the prompt statement.

[0568] (Note 3) The information processing system according to Appendix 1 is characterized in that, The apparatus for performing the data compression is configured to measure the length of the encoded sequence or the number of information units of the prompt statement or the text-simplified prompt statement, and select a compression method or compression rate based on the measurement results, so as to reduce the processing load of the generative artificial intelligence model while shortening the response time.

[0569] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for distributing multiple information structures to a user terminal based on structured information from the user terminal; An apparatus for generating instruction information for constructing prompt statements for generative artificial intelligence models by combining multiple constituent elements corresponding to the information structure through visual operations in a constituent element assembly description. A device for generating prompt statements for causing the generative artificial intelligence model to perform specific processing based on the indicated information and target data obtained from the user terminal; An apparatus for applying compression processing based on natural language processing and statistical processing to the generated prompt statement and the target data, thereby reducing the number of symbol units contained in the prompt statement and controlling the symbol utilization rate; A means for removing redundant information through parsing and summarizing processes during the compression process, while maintaining constraint information regarding the role specification, processing content specification, and output format specification in the generative artificial intelligence model; A device for performing emotion estimation processing on the emotional state of a user based on image information, voice information and text information obtained from the user terminal; A device for dynamically adjusting the amount of information, style, and compression intensity of the prompt statement based on the estimated emotional state. A device for sending the adjusted prompt statement to the generative artificial intelligence model and outputting the response information obtained from the generative artificial intelligence model to the user terminal.

[0570] (Note 2) The information processing system according to Appendix 1 is characterized in that, The device for performing emotion estimation processing on the user's emotional state is configured to input facial expression information and voice information obtained from the user terminal to an external emotion analysis processing device or an internal emotion analysis processing function, determine the user's main emotion category and emotion intensity based on the analysis results, and change the type of constituent elements used to generate the prompt statement, the length of the prompt statement, and the style of the prompt statement based on the main emotion category and the emotion intensity.

[0571] (Note 3) The information processing system according to Appendix 1 is characterized in that, The apparatus for applying compression processing is configured to perform word segmentation, word class labeling, and importance calculation on the prompt statement, reduce the number of symbol units by deleting statements with importance below a predetermined threshold, and limit the length of the prompt statement to a predetermined range through summary generation processing, thereby improving the efficiency of input resource utilization of the generative artificial intelligence model.

Claims

1. An information processing system, characterized in that, include: processor; The processor is configured to: distribute a specific template to a user; generate prompt text for instructing a generative artificial intelligence model to perform a specific operation based on the user's assembly result of the selected template using a block-based language; compress the generated prompt text to reduce token usage; identify the user's emotions and adjust the generation method and compression rate of the prompt text according to the identified emotions; and send the prompt text to the generative artificial intelligence model and provide the user with the response returned by the generative artificial intelligence model.

2. The information processing system according to claim 1, characterized in that, The processor is configured to provide blocks for assigning specific expert roles and translation blocks for translating prompt text content into different languages ​​via the block-based language.

3. The information processing system according to claim 1, characterized in that, The processor is configured to improve the system's processing efficiency by reducing the number of tokens in the generated prompt text.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A