Information processing system

CN122797604APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610286150.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-10
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]现有的生成式人工智能模型在实际应用中通常通过大规模通用语料进行训练,虽然能够生成自然语言响应,但在以下方面存在不足:其一,难以灵活、系统地赋予模型明确的性格或思想取向,导致生成内容的立场和风格不稳定,用户无法预期模型在特定价值观或性格设定下的一贯表现;其二,缺乏一种将模型所具有的性格或思想以用户易于理解和感知的方式进行可视化展示的机制,用户难以把握模型隐含的立场倾向,从而不利于对生成结果进行正确解读和使用;其三,传统信息获取过程中,人类发言者的知名度、社会地位以及发言文本的修辞方式等因素容易对受众形成显性或隐性的偏见影响,而现有生成式人工智能系统并未充分利用“非人类主体”的特点,缺乏相应的技术方案来降低此类偏见在信息呈现过程中的影响

Benefits of technology

[0004]为了解决上述技术问题,本发明提出了一种信息处理系统,该系统包括处理器,其中,所述处理器被配置为:从互联网或数据库收集与特定性格或思想相关的学习数据,并对所述学习数据进行预处理;利用经预处理的所述学习数据对生成式人工智能模型进行训练,以使所述生成式人工智能模型具有所述特定性格或思想;向所述生成式人工智能模型输入用于指示生成特定响应的提示信息,从而使所述生成式人工智能模型生成相应的响应。通过上述配置,系统能够在训练阶段引入面向特定性格或思想的学习数据,使模型在输出内容时体现出预设的性格特征或思想立场,从而实现性格化、思想化的生成行为。进一步地,所述处理器还被配置为,以用户可理解的形式对所述生成式人工智能模型的性格或思想进行可视化展示,例如通过图形、标签、维度坐标或文字说明等方式向用户呈现模型在多维性格或思想空间中的位置与特征,使用户能够直观理解模型的立场和风格。此外,所述处理器还被配置为,在生成和输出响应的过程中,利用生成式人工智能模型并非人类主体的特性,降低由发言者知名度或发言文本内容所引起的偏见影响,例如在输出界面中弱化或不使用与人类身份、名望相关的外在标识,使用户将注意力更多集中在观点本身,从而有效减轻传统人类信息源带来的主观偏见问题。通过上述技术手段的综合应用,本发明能够实现对生成式人工智能模型性格或思想的可控赋予、透明展示及偏见影响的抑制,从而提高模型输出的可解释性与公正性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797604A_ABST
    Figure CN122797604A_ABST
Patent Text Reader

Abstract

The application provides an information processing system. An information processing system, characterized by comprising: a processor; wherein the processor is configured to: collect learning data related to a specific personality or thought from the Internet or a database, and pre-process the learning data; train a generative artificial intelligence model using the pre-processed learning data, so that the generative artificial intelligence model has the specific personality or thought; input prompt information indicating generating a specific response to the generative artificial intelligence model, so that the generative artificial intelligence model generates a corresponding response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] Existing generative AI models are typically trained on large-scale general corpora in practical applications. While they can generate natural language responses, they suffer from several shortcomings: First, they struggle to flexibly and systematically imbue models with clear personalities or ideological orientations, leading to instability in the stance and style of generated content. Users cannot predict the model's consistent performance under specific value or personality settings. Second, they lack a mechanism to visualize the model's personality or ideology in a way that is easily understood and perceived by users. This makes it difficult for users to grasp the model's implicit biases, hindering the correct interpretation and use of the generated results. Third, in traditional information acquisition, factors such as the speaker's fame, social status, and rhetoric can easily create explicit or implicit biases in the audience. Existing generative AI systems do not fully utilize the characteristics of "non-human subjects" and lack corresponding technical solutions to reduce the impact of such biases in information presentation. Therefore, it is necessary to provide a new system and method that can train generative AI models based on specific learning data to imbue them with specific personalities or ideologies, visualize these personalities or ideologies, and minimize the influence of biases related to human speakers during information output. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes an information processing system. The system includes a processor configured to: collect learning data related to a specific personality or ideology from the internet or a database, and preprocess the learning data; train a generative artificial intelligence model using the preprocessed learning data to imbue the generative artificial intelligence model with the specific personality or ideology; and input prompts to the generative artificial intelligence model to instruct it to generate a specific response, thereby enabling the generative artificial intelligence model to generate the corresponding response. Through this configuration, the system can introduce learning data oriented towards a specific personality or ideology during the training phase, allowing the model to reflect preset personality traits or ideological stances in its output content, thus achieving personalized and ideological generation behavior. Furthermore, the processor is also configured to visualize the personality or ideology of the generative artificial intelligence model in a user-understandable form, such as presenting the model's position and characteristics in a multidimensional personality or ideology space through graphics, labels, dimensional coordinates, or text descriptions, enabling users to intuitively understand the model's stance and style. Furthermore, the processor is configured to, during the generation and output of responses, leverage the non-human nature of generative AI models to reduce bias caused by the speaker's fame or the content of their text. For example, it can weaken or eliminate external identifiers related to human identity or reputation in the output interface, allowing users to focus more on the viewpoint itself, thereby effectively mitigating the subjective bias problems associated with traditional human information sources. Through the comprehensive application of the above techniques, this invention enables the controllable attribution and transparent display of the generative AI model's personality or thoughts, as well as the suppression of bias, thereby improving the interpretability and fairness of the model's output.

[0005] A "system" refers to an overall device or platform consisting of one or more hardware components and / or software modules, used to perform a series of processing procedures related to generative artificial intelligence, such as data acquisition, data processing, model training, model inference, and result output.

[0006] A “processor” refers to an electronic component or computing unit that can execute program instructions, perform data operations and logic control, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and a processing cluster composed of multiple of the above units.

[0007] The "Internet" refers to a global information and communication network consisting of multiple interconnected computer networks used to transmit data and services between different terminals and servers, including the World Wide Web, various online service platforms and their related infrastructure.

[0008] A database is a collection of data that is organized, stored and managed according to a specific data model. It supports data insertion, querying, updating and deletion operations and can be accessed by applications or processors through a database management system.

[0009] "Learning data" refers to various data sets used to train or fine-tune generative artificial intelligence models, including but not limited to text, speech, images and their labels, in which at least some of the data is associated with specific personality traits or ideas.

[0010] "Preprocessing" refers to one or more processing operations performed on the learning data before it is used for model training, such as cleaning, filtering, format conversion, annotation, word segmentation, deduplication, and normalization, in order to improve data quality and adapt it to the needs of model training.

[0011] "Generative AI models" refer to AI models that can automatically generate new content based on input information, including but not limited to generative pre-trained language models, text generation models, and multimodal generation models. These models can be formed by training or fine-tuning learning data.

[0012] "Personality" refers to the stable style characteristics manifested in the output behavior of generative artificial intelligence models, including but not limited to optimism, pessimism, caution, radicalism, humor, seriousness, etc., used to simulate or abstract the emotional tendencies and attitudinal characteristics of humans when expressing themselves.

[0013] "Thoughts" refer to the value orientations or viewpoints reflected in the output of generative artificial intelligence models, including but not limited to environmentalism, economic priority, humanistic care, and safety first, which are used to indicate the model's tendency on specific issues.

[0014] "Training" refers to the process of optimizing or adjusting the parameters of a generative artificial intelligence model by inputting learning data into it, so that the model can learn the statistical regularities, language patterns, and characteristics related to specific personalities or thoughts in the data, thereby producing expected outputs in the inference stage.

[0015] "Prompt information" refers to text or other forms of instruction information input by the processor to the generative artificial intelligence model to guide or limit the content generated by the model, including but not limited to prompt words, system instructions, character setting instructions, or contextual examples.

[0016] "Specific response" refers to the response content output by a generative artificial intelligence model under the combined effect of prompt information and internal model parameters, and has a pre-determined personality traits or ideological stance, including but not limited to natural language text such as comments, answers, suggestions or explanations.

[0017] "Visual presentation" refers to presenting the personality or thought characteristics of a generative artificial intelligence model to users in a form that is easy for users to understand, such as graphics, charts, labels, coordinate systems, diagrams, or text descriptions, so that users can intuitively identify and understand the model's stance and style.

[0018] "Bias influence" refers to the subjective bias or unfair judgment that a speaker makes on the recipient of information due to non-content factors such as the speaker's fame, social status, reputation, or the external form of the text. This includes over-trusting, underestimating, or misunderstanding of viewpoints. Attached Figure Description

[0019] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0020] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0021] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0022] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0023] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0024] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.

[0025] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0026] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0027] Figure 9 This represents an emotion map that maps multiple emotions.

[0028] Figure 10 This represents an emotion map that maps multiple emotions.

[0029] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0030] Figure 12This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0031] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0032] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0033] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0034] First, let me explain the terminology used in the following instructions.

[0035] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0036] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0037] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0038] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0039] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0040] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0041] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0042] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0043] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0044] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0045] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0046] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0047] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0048] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0049] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0050] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0051] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0052] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0053] With the widespread use of generative AI models in computer applications such as dialogue systems and content generation, existing technologies suffer from the following limitations at the computer technology level: First, existing generative AI models are typically trained only on large-scale general corpora, making it difficult to stably and controllably assign specific personality or thought attributes to the model at the system level. This results in unpredictable output styles under the same input, lacking consistency and interpretability of dialogue personalities, which limits the model's usability in specific business scenarios. Second, when constructing training data, most existing systems only perform simple data cleaning and random sampling, failing to explicitly introduce a structured mapping relationship of "personality / thought attribute information—retrieval conditions—text samples—training input" into the data processing pipeline. This prevents the computer from effectively distinguishing language patterns corresponding to different attributes during the training phase, reducing the efficiency of parameter optimization and increasing model convergence time and resource consumption. Third, during the inference phase, existing systems typically only call the model based on the user's input natural language, lacking control signal injection and input encoding mechanisms based on attribute information. This makes it impossible for the computational graph to dynamically adjust the generation strategy according to the target personality or thought during inference, making it difficult to stably generate multiple outputs with significant style differences on the same computing platform. Fourth, existing generative AI systems often fail to systematically feed user interaction logs back into the training data construction and retraining process. They lack mechanisms at the system architecture level for structured storage, reprocessing, and recycling of dialogue history, resulting in an inability to continuously optimize the model's personality / thought performance in real-world deployment environments and low computational resource utilization efficiency. Furthermore, some existing systems directly retain identification information related to human subjects and media in the training data, making the model susceptible to non-content factors such as the speaker's fame and dissemination channels during the learning process. This leads to the generation of responses with undesirable biases during the inference stage, causing a systematic shift in the algorithm's output at the computer system level, affecting the model's fairness and robustness.

[0054] Therefore, it is necessary to provide a technical solution that improves the computer system architecture and data processing flow. By introducing attribute-driven data retrieval and annotation mechanisms, structured preprocessing and encoding mechanisms, attribute-controlled training and inference mechanisms, and closed-loop retraining mechanisms based on dialogue history, generative artificial intelligence models can efficiently, controllably, and sustainably acquire and maintain specific personality or thought attributes on general computing platforms, while suppressing biases introduced by human subjects and media information, thereby substantially improving the computer technology performance related to generative artificial intelligence.

[0055] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0056] In this invention, the server includes: a processing program for generating search conditions based on attribute information corresponding to the personality or thought to be assigned to the generative artificial intelligence model, and obtaining text information from an information providing device and an information storage device according to the search conditions, appending identification information corresponding to the attribute information to the obtained text information, and storing it; a processing program for performing preprocessing operations such as string normalization, useless information deletion, and unit division on the text information, structuring the text information into learning data while maintaining the identification information, and converting the structured learning data into the input format of the generative artificial intelligence model; and a processing program for controlling the information processing device to call the learning device, inputting an input sequence formed by appending control symbols representing attribute information to the learning data into the generative artificial intelligence model, and utilizing error-based... The backpropagation optimization operation adjusts the parameters of the generative AI model to obtain a processing program that produces generative AI models exhibiting different generative behaviors under different personalities or thoughts; a processing program is used to deploy the generative AI model on an inference processing device, encode prompts based on externally received prompts and attribute information to form a model input sequence, and decode the generated output sequence to obtain and output text-based response statements; and a processing program is used to store prompts, attribute information, and response statements as conversation history, update learning data based on the conversation history to form a relearning dataset, and reduce or disable recognition information related to human subjects and speaking mediums during retraining to suppress biases caused by the speaker's fame and expression style. Thus, on a general computing platform, through an attribute-driven data acquisition and preprocessing pipeline, an encoding and training mechanism with control symbols, and a closed-loop dialogue history backflow and retraining mechanism, precise and controllable attribution and continuous optimization of personality or thought attributes of the generative AI model can be achieved, improving the model's stability, consistency, and fairness in multi-style generation tasks, thereby improving the overall technical performance of computer systems related to generative AI.

[0057] "Generative artificial intelligence model" refers to an artificial intelligence model that is based on machine learning algorithms, receives an input sequence, and automatically generates a text sequence through parameterized probability distribution. It can generate natural language responses under given prompts and control information.

[0058] "Personality" refers to the stable expressive style and attitude tendencies reflected in the output of generative artificial intelligence models, including but not limited to linguistic style characteristics such as optimism, pessimism, and neutrality.

[0059] "Thoughts" refer to specific value orientations or viewpoints reflected in the output of generative artificial intelligence models, including but not limited to semantic and thematic preferences such as environmentalism and efficiency priority.

[0060] "Attribute information" refers to abstract identifying data used to represent personality or thought categories, including labels, category numbers, or control codes used to distinguish different personality types or thoughts.

[0061] "Search criteria" refers to the set of conditions derived from attribute information used to query text information in information providing devices and information storage devices, including parameters such as keywords, filtering rules, time range, and source restrictions.

[0062] "Information providing device" refers to a data source device that can output text information through a network or interface, including servers, online service platforms, or other computing devices that provide data to the outside world.

[0063] "Information storage device" refers to a data storage device used to store and manage text information, including database systems, file storage systems, or other non-volatile storage media.

[0064] “Text information” refers to language data represented in the form of character sequences, including sentences, paragraphs, documents, or other natural language content that can be used by generative artificial intelligence models for learning.

[0065] "Identification information" refers to additional data attached to text information to characterize the corresponding attribute information of the text information, and is used to distinguish samples of different personality or thought categories in subsequent processing and learning.

[0066] "Preprocessing" refers to the process of formatting and improving the quality of raw text information before training a generative artificial intelligence model. This includes operations such as string normalization, removal of useless information, and unit division.

[0067] "String normalization" refers to the process of converting the character representation in text information into a unified format, including case uniformity, standardization of special characters, and uniform encoding format.

[0068] "Useless information removal" refers to the process of removing content from text information that does not substantially contribute to the training objective, including deleting advertising text, noisy symbols, meaningless markers, and duplicate content.

[0069] "Unit partitioning" refers to the process of dividing continuous text into smaller processing units, including sentence segmentation, word segmentation, or sub-word segmentation, for subsequent encoding and model input.

[0070] "Learning data" refers to a preprocessed and structured collection of samples used to train generative artificial intelligence models, which includes text content, its corresponding identification information, and necessary metadata.

[0071] "Structured" refers to the process of transforming raw or semi-structured text information into data with fixed fields and a defined format, making it suitable for programs to read and for model training.

[0072] "Input format" refers to the data representation that generative artificial intelligence models can receive and process, including encoded token sequences, mask matrices, and corresponding label information.

[0073] "Control symbols" are special markers added to explicitly introduce attribute information into the model input, used to guide generative artificial intelligence models to produce outputs consistent with specific personalities or thoughts.

[0074] "Input sequence" refers to an ordered sequence of data that serves as input to a generative artificial intelligence model. It is usually formed by a combination of control symbols and encoded learning data or prompt statements.

[0075] "Error backpropagation" refers to the optimization algorithm process in which the model parameters are updated by taking the derivative of the loss function with respect to the model parameters and propagating the gradient back along the computation graph during model training.

[0076] "Optimization computation" refers to the numerical computation process of adjusting model parameters based on loss function and gradient information to minimize training error, including gradient descent and its variants.

[0077] "Inference processing" refers to the process of calculating the output sequence based on the input sequence when the model parameters are fixed, without including updating the model parameters.

[0078] "Information processing device" refers to a computing device that performs data processing, model training, and inference calculations, including a computing platform composed of processors, memory, and related hardware resources.

[0079] "Prompt statements" refer to text content input by users or external systems that triggers generative artificial intelligence models to generate responses; they are usually natural language questions or instructions.

[0080] "Output sequence" refers to the token sequence generated by the generative artificial intelligence model during inference processing, which is then decoded to form a response statement.

[0081] "Response statement" refers to natural language text generated by a generative artificial intelligence model based on prompts and attribute information, which is used as the system's response content to external systems.

[0082] "Conversation history" refers to the collection of prompts, attribute information, and response statements generated and stored during the interaction process, used to represent a record of one or more conversations.

[0083] "Retraining dataset" refers to a dataset used to retrain generative artificial intelligence models after the original learning data has been expanded, filtered, or updated based on the session history.

[0084] "Retraining" refers to the process of updating parameters based on an existing trained model using a relearning dataset to improve or adjust the model's generative behavior.

[0085] "Identification information related to human subjects" refers to identity information that can point to a specific natural person or group, including name, identity markers, social status markers, etc.

[0086] "Identification information related to the medium of speech" refers to the characteristic information used to identify the channel or carrier of information dissemination, including platform name, channel mark, format mark, etc.

[0087] "Bias" refers to the phenomenon that the output of a generative artificial intelligence model systematically deviates from the objective or expected distribution in a statistical sense due to factors such as the human subject's popularity, the characteristics of the medium of speech, or other non-content factors.

[0088] In this invention, the server acts as the central control node, responsible for generating and distributing programs to implement the system's functions, and coordinating data acquisition, data preprocessing, model training control, and retraining data construction on the computing platform. The terminal acts as the execution node for model training and inference, responsible for running the training and inference computation of the generative artificial intelligence model on local computing resources. Users input prompts and receive responses through the terminal's interactive interface, thereby driving the data and control flow between the server and the terminal.

[0089] The server is preferably a computer system with a multi-core central processing unit, a graphics processing unit, large-capacity storage, and a network interface. The operating system can be a general-purpose server operating system. The applications running on the server can use scripting languages ​​to implement data fetching and preprocessing logic, and a database management system to store text information and session history. The server can run natural language processing libraries for text cleaning, word segmentation, and structuring. The terminal can be a desktop computer, a graphics-accelerated workstation, an edge computing device, or a virtual machine in the cloud. The operating system can be a general-purpose desktop or server operating system. The terminal installs a machine learning framework, such as a deep learning framework, to train and infer generative artificial intelligence models. Users can access the interactive interface provided by the terminal or server through a browser or mobile application.

[0090] In one embodiment of the present invention, the server pre-stores a set of attribute information describing personality and ideological attributes, such as attributes like "optimistic" and "environmentalist." The server maps this attribute information to internal identifiers and control symbols in a vocabulary; for example, it maps "optimistic" to the special marker "...". <optimistic>This maps "environmental protectionism" to...<env_protect> The server constructs search criteria based on attribute information. These criteria include a set of keywords (such as "optimistic," "positive," "encouragement," "positive energy," "environmental protection," "low carbon," and "sustainable development"), source site filtering rules, and time range parameters. The server uses these search criteria to drive the network request module to retrieve text information in batches from information providing devices (such as public information service interfaces) and information storage devices (such as internal document libraries).

[0091] After acquiring the text information, the server performs string normalization processing through the text cleaning module. This includes unifying multiple encodings into a unified encoding format, converting uppercase letters to lowercase, standardizing punctuation, and removing invisible control characters. The server then removes useless information, using regular expressions and rule tables to remove advertising paragraphs, menu navigation, script code, and duplicate lines to reduce noise. Next, the server performs unit partitioning, dividing long texts into sentence-level units based on punctuation marks such as periods and question marks, and then further segmenting them into words or subwords according to lexical rules or sub-word algorithms, thereby constructing data units with clear structure and controlled length.

[0092] During the structuring process, the server organizes each data unit into a structured record containing text fields, attribute identifier fields, source fields, and time fields, and stores it in an information storage device. When forming the training dataset, the server retains the attribute identifier information, allowing subsequent training phases to group or label data by attribute. The server can also filter the training data according to preset rules, such as retaining only sentences within a certain length range and removing samples with ambiguous attribute labels, thereby improving the quality and consistency of the training data.

[0093] After receiving the task configuration and data access address from the server, the terminal downloads structured learning data from the information storage device and uses the locally installed word segmentation and encoding module to convert the text content into an input format acceptable to the model. In one implementation, the terminal uses a sub-word encoding algorithm to map character sequences to integer token sequences and generate corresponding attention masks. The terminal adds control symbols before each input sequence, such as inserting "..." at the beginning of the sequence. <optimistic>"or"<env_protect> This explicitly encodes attribute information into the model's input space, enabling the same model to distinguish and learn language generation patterns under different attribute conditions.

[0094] In terms of model structure, the terminal can adopt an autoregressive language model architecture based on a multi-layer self-attention network. The model built by the terminal includes an embedding layer, multi-layer self-attention encoding blocks, and an output projection layer. The embedding layer maps token IDs to fixed-dimensional vectors, and the self-attention blocks calculate the correlation between different positions through multi-head attention and perform non-linear transformation through a feedforward network. The terminal shares or partially shares the embedding space with control symbols and ordinary tokens, and through training, the vectors corresponding to the control symbols encode the preference direction of specific personality traits or thoughts in a high-dimensional feature space.

[0095] During training, the terminal reads the input sequences and corresponding target token sequences in batches from local storage and constructs a computational graph using a machine learning framework. In the forward propagation, the terminal calculates the output logits at each position through the embedding layer, self-attention layer, and output layer, and converts them into a vocabulary probability distribution using softmax. The terminal calculates the error between the predicted probability distribution and the true next token using the cross-entropy loss function, and averages the results across the batch dimension to obtain the batch loss. The terminal uses an optimization algorithm (such as adaptive gradient descent) to calculate the gradient of the model parameters based on the loss, and performs error backpropagation, propagating the gradient back along the weights and biases of each layer to the embedding layer to update the parameters. Through multiple iterations, the terminal continuously reduces the loss value, enabling the model to gradually learn to generate outputs consistent with corresponding personalities or ideas under different control sign conditions.

[0096] Compared to models trained using only general corpora, the terminal explicitly segments the training data in the input space according to attribute information by combining control symbols and attribute labels. This allows backpropagation of errors to form gradient trajectories in the parameter space corresponding to the attribute directions. This structured input and attribute-based target output work together to enable the model to intrinsically form multiple "sub-behavioral patterns" within the same parameter set. As a result, the output style can be switched simply by changing the control symbols during the inference phase, without reloading multiple different models. This reduces storage usage and model management overhead, and improves computational and deployment efficiency.

[0097] After training, the terminal stores the parameter weights for the selected rounds into a model file and loads the model into a local or remote inference service. The terminal implements an input encoding module, a decoding generation module, and a response output module within the inference service. Upon receiving a user's prompt, the terminal determines the control symbol based on the user's selected personality or thought pattern. For example, if the user selects "Optimistic Assistant," the terminal adds "..." before the encoded prompt. <optimistic>When the user selects "Environmental Protection Assistant", the terminal adds "<env_protect> The terminal performs the same word segmentation and encoding steps on the prompt statements as during training, combines the prompt statements with control symbols into an input sequence, and then calls the model in inference mode for autoregressive generation. The terminal controls the generation speed and diversity by setting the maximum generation length, sampling strategy (such as top-k or top-p sampling), and temperature parameters.

[0098] For example, when a user enters the prompt "I'm in a bad mood today, nothing seems to be going well, what should I do?" in the input box on the terminal interface and selects the "Optimistic Assistant" mode, the terminal will internally generate a message like "..." <optimistic>Given the input sequence "I'm in a bad mood today...", the model might generate a response like, "Everyone has their ups and downs, but these experiences make you stronger. Give yourself some time to rest and then gradually adjust your pace." Similarly, if a user inputs, "I want to do something good for the environment, but I don't know where to start. Do you have any suggestions?" and selects the "Environmental Protection Assistant" mode, the terminal might generate a response like, "You can start by reducing disposable items, such as bringing your own shopping bags and water bottles, and trying to use public transportation whenever possible. These small changes, if consistently maintained, will have a significant positive impact."

[0099] In one embodiment of the invention, the server is also responsible for collecting and reprocessing the session history uploaded by the terminal. After each dialogue ends or at a certain time interval, the terminal uploads the prompts used in the interaction, the attribute information corresponding to the control symbols, the generated response statements, and necessary metadata (such as timestamps and session identifiers) to the server. The server performs de-identification processing on these session histories, removing identification information such as names and account tags related to specific individuals, and using a rule-based dialogue filtering algorithm to remove samples containing sensitive information or abnormal output. The server restructures the remaining high-quality dialogue samples and adds them to the learning dataset, forming a relearning dataset. By periodically triggering retraining on the terminal, the server and terminal jointly achieve continuous performance improvement of the generative artificial intelligence model after deployment.

[0100] This invention introduces a mechanism to reduce the identification information related to human subjects and media in data processing and model training. During data preprocessing, the server uses regular expression matching and named entity recognition algorithms to detect personal names, organization names, platform names, and channel markers in the text, replacing or deleting these entities with generic placeholders. This reduces overfitting of the model to specific subjects and platforms during training. The terminal uses this debiased data during retraining, which helps the model generate neutral expressions independent of specific subject authority or media influence during the inference phase, reducing systematic bias. This process not only contributes to output fairness but also prevents excessive occupation of noisy features by internal model parameters, improving the model's ability to express true content features, thereby enhancing generation quality and generalization ability.

[0101] From a computer technology perspective, this invention replaces the traditional approach of manually partitioning datasets and manually setting styles with an automated, structured process jointly driven by the server and terminal, through attribute-driven retrieval condition generation and control symbol encoding. This enables the computer to discriminate between various personalities and thoughts within the same training framework. Because the training data is partitioned and labeled with attribute information at the structural level, the terminal can use more efficient mini-batch sampling strategies during training, such as sampling by attribute grouping. This ensures that gradient updates maintain attribute consistency within the same batch, thereby reducing gradient oscillations and accelerating convergence. During retraining, the server automatically constructs new training samples based on the dialogue history, allowing the model to continuously correct its internal parameter distribution in real-world scenarios. This closed-loop optimization is significantly different from traditional static models trained in a single cycle, and is beneficial for maintaining performance over long periods of operation.

[0102] Furthermore, this invention integrates multi-attribute control into a single model, allowing the terminal to support multiple personality types and thought patterns by calling only one model instance during inference. This avoids the resource consumption and context switching overhead associated with loading multiple large-scale models in memory. The addition of control symbols enables the model to form an internal condition generation mechanism, allowing the terminal to dynamically adjust the control symbols for different task scenarios. This allows for the reuse of most parameters without altering the network structure, achieving parameter sharing and computational reuse. In high-concurrency applications, this structure significantly reduces model loading time and GPU memory usage, thereby improving inference throughput.

[0103] In alternative implementations, the terminal can also employ other neural network structures, such as an encoder-decoder architecture that combines convolutional and attention layers, or different optimization algorithms (e.g., stochastic gradient descent and momentum methods), or different loss functions (e.g., label-smoothed cross-entropy loss), to adapt to different hardware environments or application requirements. The server can also adjust the batch size and compression format of data transmission based on available network conditions to reduce communication load. For example, the server can package preprocessed learning data into compressed files and send them to the terminal, which can then decompress the files before encoding and training, thereby reducing transmission time in bandwidth-limited environments.

[0104] In summary, through close collaboration between the server and terminal in attribute-driven data acquisition, structured preprocessing, control symbolic condition training, inference service deployment, and dialogue history retraining, this invention not only achieves precise and controllable attribution of personality and thought to generative artificial intelligence models, but also improves the efficiency, stability, and fairness of computer systems in generation tasks from multiple levels, including data structure, model structure, and training process. This improvement is not merely a simple automation of manual operations, but rather the implementation of new data organization methods and parameter optimization paths within the computer, resulting in verifiable technical effects in processing speed, generation quality, resource utilization, and bias control.

[0105] use Figure 11 The processing flow is explained.

[0106] Step 1: The server generates attribute information based on a pre-stored list of personalities or thoughts, and generates search criteria based on the attribute information.

[0107] The server takes pre-defined personality or ideology configuration data as input, such as abstract labels like "optimistic" or "environmentalist." The server parses the configuration data, assigns an internal attribute ID to each label, and generates corresponding control tags in the attribute dictionary; for example, it generates a control tag for "optimistic." <optimistic>This has generated "environmental protectionism"<env_protect> The server then constructs a keyword set and filtering rules based on the attribute tags. For example, it maps "optimistic" to keywords such as "optimistic," "positive," "encouragement," and "positive energy," and maps "environmentalism" to keywords such as "environmental protection," "low carbon," "sustainable development," and "reducing plastic." It also sets the data source type, time range, and language filtering conditions. The server packages the generated attribute IDs, control tags, and search conditions into a task configuration record and writes it to the configuration storage or caching system. The output of this step is a structured task configuration for each type of attribute information, which serves as input for subsequent data acquisition and processing.

[0108] Step 2: The server retrieves the original text information from the information providing device and the information storage device according to the search criteria, and stores it with attribute identifiers.

[0109] The server's input is the task configuration generated in step 1, which includes attribute IDs and corresponding search criteria. The server invokes the network request module, constructs a query URL or request parameters based on keywords, sends a request to the external information provider, and receives webpage content, article text, or document data. Simultaneously, the server sends a search request to the internal information storage device, filtering text records matching the search criteria from the existing document library. The server performs basic parsing on each piece of raw text, extracting the title, body, and metadata from the HTML structure, and discarding scripts, styles, and navigation. The server binds the parsed text to its corresponding attribute ID, generating a record containing the fields "text content," "attribute ID," "source type," and "retrieval time," and stores it in the raw corpus. The output of this step is a set of raw text records with attribute identifiers, providing input for the preprocessing steps.

[0110] Step 3: The server performs preprocessing on the raw text information, including string normalization, removal of useless information, and unit partitioning, and structures the results into learning data.

[0111] The server's input is the raw text record set obtained in step 2. The server first calls the string normalization module to standardize the text to a specified encoding format, unifying capitalization, standardizing punctuation, and removing invalid control characters and redundant whitespace. Subsequently, the server uses regular expressions and a noise dictionary to remove irrelevant information such as advertisements, copyright notices, navigation menus, and URL links. Next, the server performs unit partitioning, splitting the normalized text into multiple sentences based on sentence boundaries, and can further utilize word segmentation or sub-word algorithms to divide sentences into words or sub-word units. The server retains the original attribute ID and source metadata for each sentence unit, storing it as a structured record with fields such as "sentence text," "attribute ID," "source," and "length." The output of this step is a cleaned and unit-partitioned structured learning dataset, serving as the basis for constructing the model input.

[0112] Step 4: The server converts structured learning data into data files that can be downloaded to the terminal and distributes data access information to the terminal.

[0113] The server's input is the structured learning dataset generated in step 3. The server groups the data according to different attribute IDs, packaging sets of sentences belonging to the same attribute into data files; for example, the attribute "optimistic" corresponds to one set of files, and the attribute "environmentalism" corresponds to another. The server serializes each record into a standard format, such as each line containing fields like "text content," "attribute ID," and "source," forming a text data file. The server stores these data files in a file storage system or object storage service, generating corresponding access paths or download URLs. Finally, the server writes the task ID, attribute information, and data file access addresses into the task description and returns it to the terminal via an interface. The output of this step is the structured learning data file available to the terminal, along with corresponding access information.

[0114] Step 5: The terminal downloads the learning data file from the server and performs local encoding preparation, converting the text into an input format adapted to the generative artificial intelligence model.

[0115] The terminal's input consists of the task description and data access address provided by the server. The terminal downloads the corresponding data file from the server via a network request and stores it in the local file system. The terminal reads the data file, extracts the "text content" and "attribute ID" fields, and uses a local word segmentation or sub-word encoding tool to convert each piece of text into a token ID sequence, while simultaneously generating an attention mask and sequence length information. The terminal inserts the token ID corresponding to the control symbol determined by the attribute ID at the beginning of each token sequence; for example, the attribute "optimistic" inserts the control token "...". <optimistic>The terminal divides the encoded data into training, validation, and test sets based on samples, and packages these tensor data into a binary format suitable for efficient reading by machine learning frameworks. The output of this step is an encoded dataset that can be directly input into generative artificial intelligence models, including the input sequence, target sequence, and auxiliary mask.

[0116] Step 6: The terminal constructs a generative artificial intelligence model structure and loads pre-trained parameters or initialization parameters as needed.

[0117] The terminal's input consists of model configuration parameters (e.g., number of layers, hidden dimension, number of attention heads) and the dimensionality information of the encoded data features obtained in step 5. The terminal uses a deep learning framework to define a generative language model based on a multi-layer self-attention structure, including embedding layers, multi-layer self-attention sub-layers, and an output projection layer. The terminal initializes the embedding matrix according to the vocabulary size and reserves embedding vector slots for control symbols. The terminal can choose to load some or all parameters from a pre-trained model weight file to accelerate convergence; if no pre-training resources are available, a random initialization strategy is used. The terminal sets the optimizer, learning rate scheduler, and loss function according to the configuration. The output of this step is a trainable generative AI model instance and its initial parameters.

[0118] Step 7: The terminal uses the encoded learning data to train the generative artificial intelligence model and uses backpropagation of errors to adjust the model parameters.

[0119] The terminal's input consists of the training dataset generated in step 5 and the model instance constructed in step 6. The terminal divides the training data into several batches, reading an input sequence and its corresponding target token sequence in each batch. The terminal feeds the input sequence into the model, passing it sequentially through the embedding layer and the self-attention layer, calculating the output logits at each position. The terminal uses the softmax function to calculate the probability distribution of the next token and compares it with the target token, calculating the batch loss using the cross-entropy loss function. The terminal invokes the automatic differentiation function of the deep learning framework, calculates the gradient of the model parameters based on the loss, executes the error backpropagation algorithm, and transmits the gradient from the output layer back to each network layer. The terminal uses an optimizer (such as an adaptive gradient method) to update the weights and biases of each layer, thereby reducing the prediction error in the next training round. The terminal records the training loss and validation set loss during training to monitor training convergence. The output of this step is the trained generative artificial intelligence model parameters, obtained after several training rounds, capable of exhibiting different generative behaviors based on control symbols.

[0120] Step 8: The terminal evaluates the trained generative AI model, selects the best-performing model version, and deploys it to the inference environment.

[0121] The terminal's input consists of multiple completed model checkpoints and the validation and test sets from step 5. During the evaluation phase, parameter updates are disabled; the terminal uses only the model for inference. The terminal inputs the input sequence from the test set (including control symbols and cue fragments) into the model, calculates the output, and evaluates the model's performance using perplexity, accuracy, or other quality metrics. The terminal can further use heuristics to detect whether the output aligns with the target personality or ideology, such as by statistically analyzing the proportion of positive sentiment words or the frequency of environmentally related words. The terminal compares the model weights across different training epochs, selects the version that performs best on the metrics, and archives the parameters of this version into a deployment file. The terminal loads these model parameters during the inference service process, preparing to accept external requests. The output of this step is a deployed generative AI model service instance, ready for online invocation.

[0122] Step 9: Users input prompts through the terminal interface and select their desired personality or thought pattern to trigger the model to generate a response.

[0123] The user's input consists of prompts in natural language and a selection of personality or thought patterns. For example, a user might enter the prompt: "I'm in a bad mood today, nothing seems to be going well, what should I do?" and select "Optimistic Assistant" in the interface; or they might enter the prompt: "I want to do something good for the environment, but I don't know where to start, do you have any suggestions?" and select "Environmentalist Assistant". The terminal receives the user input, packages the prompts and selected patterns into a request message, and uses it as input for subsequent inference calls. The output of this step is a structured request containing the prompts and attribute selection information.

[0124] Step 10: The terminal constructs an inference input sequence based on the user's request, calls a generative artificial intelligence model to perform inference, and generates a response statement.

[0125] The terminal's input is the user request from step 9, which includes a prompt and attribute selection. The terminal determines the corresponding control symbol token ID based on the attribute selection, and concatenates the control symbol and user prompt into an input sequence after word segmentation encoding. During inference, the terminal sets generation parameters (such as maximum generation length, sampling strategy, and temperature) and inputs the input sequence into the model's inference interface. The model progressively generates output tokens, calculating the probability distribution of the next token at each step based on the current context and the embedding vector of the control symbol, and selecting the most likely token or the token that meets the diversity requirements according to the set sampling strategy. The terminal repeats this process until a termination condition is met (e.g., generating an end marker or reaching the maximum length), and then decodes the generated token sequence into natural language text. The output of this step is the response statement corresponding to the user prompt and attribute pattern.

[0126] Step 11: The terminal displays the generated response statement to the user, records this session as the session history, and uploads it to the server for further learning.

[0127] The terminal's input consists of the response statement generated in step 10, the original prompt statement, and attribute selection information. The terminal displays the response text on the user interface, allowing the user to read and continue interacting. Simultaneously, the terminal encapsulates the prompt statement, attribute information, response statement, and metadata such as timestamps and session identifiers into a session record, storing it in a local log or cache. The terminal uploads the session record to the server at appropriate times (e.g., when the session ends or a timed trigger). The output of this step is one or more complete dialogue history records, providing input for subsequent data updates and retraining of the server.

[0128] Step 12: The server receives the session history uploaded by the terminal, cleans and filters it, and updates the learning dataset to support model retraining.

[0129] The server's input is the set of conversation history records uploaded by the terminal in step 11. The server first performs de-identification processing on the conversation history, using entity recognition and rule filtering algorithms to identify sensitive information such as names, accounts, and platform names, replacing or deleting them with placeholders. Next, the server filters dialogue samples according to preset quality standards, such as removing samples that are too short or have abnormal output, retaining dialogues with fluent language and clear attribute characteristics. The server categorizes the retained conversation samples by attribute ID, extracting prompts and responses as new text units and attaching attribute identification information. The server merges these new samples into the existing structured learning dataset, updating the corresponding data files and task configurations. The output of this step is a relearning dataset containing the new conversation samples, providing a foundation for subsequent terminal-initiated retraining.

[0130] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0131] In existing computer-based advertising recommendation technologies, fixed template advertising content is typically pushed to users based solely on simple user interest tags or historical click data, using rule engines or general recommendation algorithms. Such technical solutions suffer from the following problems: First, the server side often uses user behavior data as statistical counts or similarity features, without modeling deep attributes such as user personality and ideological tendencies into iteratively updatable structured information at the system level. This makes it difficult for advertising content to truly align with user values ​​and expression preferences. Second, servers often use static templates or a small number of pre-set texts, lacking a universal mechanism that can dynamically generate prompts based on changes in user behavior and drive generative AI models to produce diverse advertising texts, making it difficult to achieve high-granularity personalized expression. Third, the application of existing generative AI models in advertising scenarios is mostly limited to single-round content generation. There is a lack of a systematic design that allows for the closed-loop feedback of user terminal display, click behavior, and comment content to flow back into the profile model and prompt generation logic, thus failing to form a sustainable adaptive optimization generation control process in the computer system. Fourth, existing systems often only output recommendation results from a single perspective, lacking a computational architecture that can simultaneously schedule multiple role settings and present multi-perspective advertising or response texts in a distinguishable form within the same user interface, making it difficult to fully leverage the advantages of generative AI models in terms of viewpoint and expression diversity.

[0132] Therefore, a new computer implementation method and system architecture are needed to enable the server to: transform user behavior information into updatable personality and thought attribute information within a unified data and model framework; automatically construct prompt statements for generative artificial intelligence models based on this attribute information; generate multi-perspective advertising messages and response texts using multi-role settings; and feed back response information generated by user terminals to dynamically adjust the above processes, thereby substantially improving the computer technology performance related to advertising generation and display while reducing human rule intervention.

[0133] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0134] In this invention, the server includes functional components for: acquiring behavioral information including user browsing history, purchase history, and operation history from external storage devices and performing aggregation, feature extraction, and inference processing to generate attribute information representing user personality and ideological tendencies; selecting at least one role setting from role information based on the attribute information and product information and automatically generating prompt statements with limited output word count, expression style, and tone, containing role settings, attribute information, and product information; inputting the prompt statements into a generative artificial intelligence model and performing inference processing on computing resources to obtain personalized advertising messages, and performing content review and format adjustment to form display data that can be sent to the user terminal; and accumulating data from the user terminal. The system includes functional components for: display status, click status, and comment information, which are then correlated with behavioral information to update attribute and role information, dynamically changing the conditions for generating prompt statements and the input content to the generative artificial intelligence model; dialogue history including comment information, role settings, and attribute information to generate prompt statements for responding to user comments, which are then input into the generative artificial intelligence model to generate response text and send it to the user terminal; and the generation of prompt statements for multiple different role settings, which are then input into the generative artificial intelligence model in parallel or sequentially to obtain advertising messages or response texts from multiple perspectives, and to generate display control data for presenting these multiple contents in a distinguishable form on the user terminal. This allows for the formation of an advertising generation control process within the server, based on deep user attribute modeling, with automatic prompt statement construction at its core, driven by generative artificial intelligence model reasoning, and updated through a closed loop of user response information. This improves the automation and flexibility of personalized advertising generation at the computer technology level, reduces the burden of manual rule configuration, increases the matching degree between advertising content and user personality and thoughts, and improves the overall response efficiency and scalability of the recommendation system.

[0135] "Behavioral information" refers to the recorded data generated by users in the information processing system that is related to operations such as browsing, purchasing, querying, clicking, staying, scrolling, and inputting. Its sources can include terminal logs, server logs, and historical records in external storage devices, and are used to reflect the user's usage behavior over a period of time.

[0136] "Attribute information" refers to abstract characteristic data that represents a user's personality tendencies, ideological tendencies, interests, and preferences, obtained by summarizing, extracting features, and inferring user behavior information. It is usually stored in a structured form and can be used by subsequent processing units to generate prompts and control content generation.

[0137] "Personality traits" refer to relatively stable categories or numerical indicators of personality traits inferred from user behavior information, used to characterize users' emotional expression, risk preferences, communication styles, etc., such as optimistic, rational, conservative, extroverted, etc., which are used to influence the tone and expression of generated content.

[0138] "Thought orientation" refers to the characteristics inferred from user behavior information and their long-term interest distribution, which reflect users' preferences in terms of values, issues of concern, and decision-making orientation. Examples include environmental orientation, price sensitivity, hedonism, and minimalism, which are used to guide the theme and stance of generated content.

[0139] "Product information" refers to a collection of various data related to the promoted service or product, including but not limited to name, category, function description, target audience, price range, usage scenarios, and associated multimedia resources, which constitute the content entity when generating advertising messages.

[0140] "Role information" refers to the data set used in the system to manage multiple role settings. Each role setting includes role name, personality description, value description, language style parameters, etc., which are used to give the generative artificial intelligence model specific "personalized" output characteristics when calling the model.

[0141] "Role setting" refers to the specific role configuration content selected from the role information and applied to the generative artificial intelligence model. It is used to specify the personality traits, ideological stance, tone of expression and communication methods that the model should adopt when generating advertising messages or response text.

[0142] "Prompt statements" refer to instruction information automatically constructed by the server based on attribute information, product information, and role settings, expressed in natural language or structured text, and used as input to generative artificial intelligence models to clarify generation goals, output format, tone requirements, and content restrictions.

[0143] "Generative artificial intelligence model" refers to a data processing model that is trained on machine learning algorithms and large-scale data and can automatically generate text, speech or other content based on input prompts. In this invention, it is used to generate advertising messages and response text based on role settings and prompts.

[0144] "Advertising messages" refer to text content generated by generative artificial intelligence models after receiving prompts and displayed to users to promote a product or service. This content is personalized based on the user's attribute information and may include elements such as reasons for recommendation, scenario descriptions, and action guidance.

[0145] "Response text" refers to the reply content generated by generative artificial intelligence models based on user comments, dialogue history, role settings, and attribute information. It is used to form multi-round interactions between users and the system, enabling the system to answer user questions or feedback in a way that has specific role characteristics.

[0146] "Display data" refers to a structured data set that can be directly interpreted and presented by the user terminal after the server performs content review and format adjustment on the advertising message or response text. It usually includes text content, layout parameters, style information, and role-related identification information.

[0147] "Display control data" refers to the control information generated to guide the terminal in interface layout, grouping display, and label rendering in order to present multiple advertising messages or response texts in a distinguishable form on the user terminal. It includes view structure, sorting method, label identification, and style parameters.

[0148] "Response information" refers to various feedback data generated by user terminals during the display and interaction of advertisements and sent back to the server, including advertisement exposure, clicks, dwell time, scrolling behavior, and user comments, which are used to evaluate content effectiveness and update attribute information and generation strategies.

[0149] "Conversation history" refers to the time-series record of comments, replies, and system-generated content accumulated during multiple rounds of interaction between users and the system around the same topic or advertising content. It is used to provide context for generative artificial intelligence models to generate coherent response text.

[0150] "External storage devices" refer to storage resources located outside the system or connected via a network, including database systems, file storage systems, or distributed storage systems, used to store user behavior information, attribute information, role information, and generated results for extended periods.

[0151] "Computing resources" refers to the hardware and system resources used to run generative artificial intelligence model inference processing and related data processing programs, including processors, graphics processing units, memory, storage devices, and the operating systems and middleware running on them.

[0152] In one embodiment of the present invention, the server comprises a processor, main memory, external storage device, network interface, and optional graphics processing unit. The server runs an application framework and machine learning framework on top of an operating system to perform data acquisition, feature calculation, model inference, and result distribution. The terminal comprises a display device, input device, wireless communication module, and local storage, used to present advertising messages and response text, and to send user behavior information to the server. Users generate actions through the terminal's input device, thereby forming behavioral information and comment information.

[0153] In one implementation, the server uses a general-purpose processor as the central processing unit and a graphics processing unit (GPU) as the acceleration unit for generative AI model inference and training. On the software side, the server can use a scripting language-based backend framework as the network service framework, a relational database management system or document-oriented database as the storage system, a data analysis library for feature computation, and machine learning libraries or deep learning frameworks (such as those supporting tensor operations and automatic differentiation) and model libraries supporting the Transformer architecture as the foundation for implementing generative AI models. In its specific implementation, the server can divide different functions into multiple modules, including a data acquisition and preprocessing module, a profile generation module, a role management module, a prompt statement generation module, a generative AI model inference module, a content review and formatting module, a feedback processing and profile update module, and a multi-role display control module, etc.

[0154] The server can process user behavior information in the following ways: The server reads the behavior information from external storage, which is stored as log records. Each record includes fields such as user ID, timestamp, resource ID, operation type (browsing, clicking, purchasing, etc.), dwell time, and device type. The server uses a data analysis library to construct a tabular data structure for these records in main memory, accelerating queries by user ID through hash indexes or B-tree indexes. The server performs cleaning and aggregation operations on the behavior information, including deleting records with missing necessary fields, standardizing the time format, normalizing URLs or resource IDs, and mapping product categories to integer or one-hot encodings. Based on this, the server calculates statistical features (such as the number of visits to a certain category of content, average dwell time, and visit time distribution) and constructs a sparse or dense vector representing the distribution of users' long-term interests.

[0155] In the profile generation module, the server uses machine learning algorithms to map the aforementioned features into attribute information of personality and ideological tendencies. The server can employ supervised learning models such as feedforward neural networks, gradient boosting decision trees, and logistic regression, taking user feature vectors as input and pre-labeled or clustered personality and ideological tendency labels as output. When using feedforward neural networks, the server stores the network structure configuration in storage, including the input layer dimension, the number of hidden layers and nodes per layer, activation function type (e.g., ReLU, tanh), output layer form (multi-label or multi-class), loss function (e.g., cross-entropy loss), and optimization algorithm (e.g., stochastic gradient descent, Adam, etc.). During the training phase, the server iteratively updates the weights and bias parameters using batch data, adjusting the parameters based on the gradient of the loss function through backpropagation, so that the predicted personality and ideological tendencies output by the model approximate the labels of the training samples.

[0156] During the inference phase, the server inputs the preprocessed user feature vector into the trained profile model to obtain a low-dimensional vector or probability distribution representing personality and ideological tendencies. The server then discretizes this vector into several labels based on the maximum probability or a preset threshold, such as "optimistic," "rational," "environmentally conscious," and "price-sensitive," and stores this attribute information in a profile table in the database. The structure of this profile table can include fields such as user identifier, personality label set, ideological label set, and update time, thereby supporting rapid retrieval and updates.

[0157] The server maintains a set of role information in the role management module. Each role information record includes a role identifier, role name, role personality description, role ideology description, language style parameters (e.g., formality, humor level), and a set of applicable product categories. The server loads the role information into a dictionary or key-value mapping structure in main memory to facilitate role selection based on attribute information. Based on the user's personality and ideology, the server selects at least one suitable role setting from the role information. For example, if the attribute information indicates that the user has a high reading rate on environmental topics and an overall optimistic mood, the server can select the role setting "Optimistic Environmental Advocate," which would include descriptions such as "caring about the environment," "inspiring tone," and "emphasizing long-term values."

[0158] In the prompt generation module, the server constructs prompts for the generative AI model based on attribute information, product information, and role settings. The server can organize the prompts into natural language templates and use string interpolation or a template engine to replace variables with specific content. When constructing the prompts, the server explicitly limits the output word count, expression style, and tone to control the output distribution of the generative AI model, thereby reducing unnecessary post-processing overhead and improving inference efficiency. In a specific example, the server can generate prompts in the following text format: "You are an optimistic and passionate environmentalist. Please create a Chinese advertising copy for an eco-friendly reusable water bottle, written in the first person. Please emphasize the following points:" - Safe and recyclable materials - Reduce single-use plastics - Suitable for commuting and outdoor activities The copy should be no longer than 150 characters, with a positive and inspiring tone, suitable for display in mobile ad slots. The server can also generate different prompts for multi-role and multi-perspective output, for example: "You are an optimistic environmentalist. Please write a short Chinese recommendation for an eco-friendly laundry detergent in a friendly and sincere tone. Highlight its environmentally friendly ingredients, water-saving formula, and suitability for families with children and pets." "You are a rational and calm price analyst. Please evaluate the same environmentally friendly laundry detergent from the perspective of cost-effectiveness. Please compare its price with that of ordinary laundry detergent, explain the potential cost savings from long-term use, and write an advertising description of no more than 120 words in concise and objective language." "You are a minimalist lifestyle advocate. Please explain why choosing this eco-friendly laundry detergent helps simplify your life, reduce the number of items you own, and doesn't sacrifice cleaning effectiveness. Please write an advertising copy of no more than 120 words in a gentle and quiet tone." Once the server receives a user comment, it can construct a prompt message for the conversation, for example: "System character setting: You are a sincere and professional environmental product consultant with an optimistic but not exaggerated personality."

[0159] User review: 'Is this eco-friendly water bottle really that durable? I'm worried it will leak after a while.' Please reply to this user in Chinese, explaining the product's durability and after-sales service in a friendly and specific manner. Avoid excessive marketing and keep the response under 120 characters. The logic for constructing the aforementioned prompt statements is implemented programmatically within the server, avoiding manual writing of each statement, thus forming an extensible prompt statement generation subsystem within the computer.

[0160] In the generative AI model inference module, the server uses a generative model that supports the Transformer architecture. During model deployment, the server stores model parameters in external storage and loads them into the GPU's memory upon system startup. This generative AI model typically consists of a multi-layer self-attention encoder and decoder structure, or an autoregressive structure containing only a decoder. Each layer includes a multi-head attention sublayer and a feedforward sublayer, and layer normalization and residual connections are used to stabilize training. During the training phase, the server pre-trains the model using a large corpus, utilizing language modeling loss (e.g., minimizing the cross-entropy of the next token) and performing gradient descent updates in the parameter space via backpropagation. The server can also fine-tune the model for specific advertising scenarios, using labeled advertising text and comment replies as training samples. By minimizing the difference between the generated results and the target text, the model is optimized to achieve higher generation quality and relevance in the advertising domain.

[0161] During the inference phase, the server converts the prompt statements into a sequence of segmented or sub-word tokens, maps them to vectors through a word embedding layer, and then inputs them sequentially into the Transformer layer. The server can adjust sampling parameters, such as temperature, top-k, or top-p, to balance the diversity and stability of the generated content. The server generates the output token sequence through beam search or sampling strategies and decodes it into text. The server performs regular expression matching, sensitive word filtering, and length checks on the generated text in memory. Based on business strategies, the server can adjust generation parameters or regenerate when it detects non-compliant content or deviations from the prompt statement requirements, thus achieving closed-loop control of the generation process. This closed-loop control is executed automatically by the algorithm, reducing the burden of manual review and improving the overall system processing efficiency.

[0162] After receiving display data from the server, the terminal parses the structured data locally and presents the advertising messages and response text in the user interface. Based on the layout information in the display control data, the terminal can display multiple advertising messages generated by different roles in a card-style arrangement or tab-switching format, allowing users to distinguish content from different perspectives. When the user scrolls and clicks, the terminal uses local scripts to record exposure time, scroll depth, and click events, and sends them back to the server via the network interface at appropriate times. In high-latency network environments, the terminal can employ a local caching mechanism to batch upload behavioral information generated within a short period, thereby reducing the number of communications and lowering network load.

[0163] When users read advertising messages on their devices, they can click, save, or jump to relevant content based on their personal interests, and they can also enter natural language text in the comment input area. User comments are transmitted from the device to the server and written into the dialogue history table. The server constructs corresponding prompts based on this dialogue history and user roles, calls a generative artificial intelligence model to generate response text, and then returns it to the device for display. Users can gradually obtain more tailored information through multiple rounds of interaction. This multi-round dialogue process is automatically managed by the server, and the corresponding data stream includes comment text, prompts, generated results, and subsequent user behavior information, thus forming a long-term, learnable data loop within the computer system.

[0164] In one embodiment of the invention, the server performs fine-grained statistical analysis on the response information from the terminal. The server can correlate metrics such as exposure count, click-through rate, dwell time, and comment sentiment with corresponding attribute information and role settings, and update the prompt generation strategy using regression models or reinforcement learning algorithms. For example, when the server detects that a certain role setting has a good conversion effect among a specific user group, it can increase the selection probability of that role during the role selection stage. The server can also adjust prompt parameters such as generated text length, tone intensity, and emphasis based on the response information, making the output of the generative artificial intelligence model more suitable for that user group. In this way, the server not only achieves personalization at the content level but also adaptive optimization at the generation control strategy level, thereby improving model invocation efficiency and generation quality within the computer.

[0165] Compared to traditional systems that only use simple rule engines, this invention maps user behavior information into structured attribute information and tightly integrates it with role information and prompt generation modules. This allows the server to introduce fine-grained control at the front end of the generative AI model, reducing irrelevant generation and invalid requests, and lowering the average computational load of model inference. Furthermore, by systematically feeding user response information back into the attribute and role information updates, the server can improve overall prediction accuracy and content matching through iterative upstream control logic without altering the core structure of the generative AI model. This achieves the technical effects of increasing processing speed, reducing the number of repeated generation attempts, and lowering the error rate.

[0166] In another alternative implementation, the server can employ different generative AI model structures and training methods. For example, the server can use an encoder-decoder sequence-to-sequence model, taking prompts as the source sequence and advertising copy or response text as the target sequence, aligning and generating the sequence through an attention mechanism. In this implementation, the server can improve the model's generalization ability using techniques such as teacher-forced training, label smoothing, and data augmentation (such as synonym replacement and sentence rewriting). The server can also employ multi-task learning, jointly training the advertising generation task and the comment reply task in the same model parameter space, thereby improving parameter utilization efficiency and reducing model storage resource consumption based on shared representations.

[0167] In another implementation, the server can map multiple roles to different "virtual channels" within the same generative AI model. By embedding special markers or role embedding vectors in prompts, the model is guided to select different subspaces for generation. With this design, the server does not need to deploy independent model instances for each role. Instead, by injecting role information into the input, multiple roles can share a single set of model parameters, thereby reducing memory usage and inference latency, and improving the system's concurrency performance in multi-role, multi-user scenarios.

[0168] As can be seen from the above embodiments, this invention does not simply automate the human advertising writing process, but introduces a prompt statement generation mechanism based on attribute information and role settings within the server. Combined with the deep learning structure and multi-round feedback update strategy of the generative artificial intelligence model, it improves the technical performance of the computer system at multiple levels such as data structure, algorithm process and resource scheduling, and realizes an overall improvement in the accuracy, speed and scalability of the advertising content generation and display process.

[0169] use Figure 12 The processing flow is explained.

[0170] Step 1: The server reads behavioral information from external storage devices.

[0171] Input: Raw log records including user ID, timestamp, resource ID, operation type (browse, click, purchase, etc.), dwell time, and device type.

[0172] The server uses database queries to extract records from the log table by user ID and time range, and loads the results into an in-memory tabular data structure. The server standardizes the timestamp field format, normalizes URLs or resource identifiers, and removes records with missing key fields. Output: A cleaned collection of user behavior records.

[0173] Step 2: The server performs statistical aggregation and feature extraction on behavioral information.

[0174] Input: A collection of cleaned behavior records.

[0175] The server uses a data analysis library to perform group-by operations on records by user identifier and content category, calculating statistics such as pageview count, average dwell time, and click-through rate for each content category. The server encodes discrete fields such as product category and keywords into integer or one-hot vectors, and concatenates the statistical features with the encoded features to form a fixed-length feature vector. Output: A set of user-generated behavioral feature vectors.

[0176] Step 3: The server infers a user's personality and ideological tendencies based on behavioral feature vectors.

[0177] Input: A set of behavioral feature vectors.

[0178] The server inputs each user's feature vector into a pre-trained profiling model (e.g., a feedforward neural network or gradient boosting tree), performs forward computation on the processor, and obtains the probability distribution or score of personality and ideology tendencies. Based on a maximum probability or threshold rule, the server converts the continuous output into discrete labels (e.g., "optimistic," "environmentally oriented," etc.) and constructs attribute information records. Output: Attribute information containing user personality and ideology tendency labels.

[0179] Step 4: The server selects the appropriate character settings based on the attribute information.

[0180] Input: Attribute information and character information database.

[0181] The server iterates through or retrieves the role information database in memory, comparing each role's applicability conditions (such as preferred themes and tone of voice) with the user's attribute information. It then calculates the match degree for each role using matching rules (such as weighted scores or rule tables). The server selects one or more role settings with the highest match degree and constructs a role setting structure containing role identifiers, role descriptions, and language style parameters. Output: One or more role settings associated with the target user.

[0182] Step 5: The server generates prompts based on attribute information, role settings, and product information.

[0183] Input: User attribute information, selected role settings, and product information.

[0184] The server invokes the template generation module, filling predefined natural language templates with parameters such as character descriptions, personality tags, thought tags, product names, and selling points. The server inserts constraints into the template, including output word count, tone requirements, and applicable scenarios, forming complete prompt text. The server can generate different prompts for different characters. Output: One or more prompts used to drive the generative AI model.

[0185] Step 6: The server inputs the prompts into the generative artificial intelligence model to generate advertising messages.

[0186] Input: One or more prompt statements.

[0187] The server uses a word segmentation or sub-word encoder to convert the prompt statements into token sequences, and then invokes a generative AI model based on a Transformer architecture to perform inference on the graphics processing unit. The server sets generation parameters (maximum length, temperature, top-k, top-p, etc.) to control the generation process, progressively generating the output token sequence through autoregressive decoding. The server decodes the token sequence into natural language text to obtain the original advertising message. Output: A set of advertising message texts corresponding to each prompt statement.

[0188] Step 7: The server performs content review and format adjustment on advertising messages.

[0189] Input: The original collection of advertising message text.

[0190] The server uses a sensitive word dictionary and regular expressions to scan the advertising text, deleting or replacing words that do not conform to the specifications, checking whether the text length is within the range specified in the prompt statement, and pruning or triggering regeneration for excessively long text. The server also adds line breaks, short titles, or highlights according to the terminal's display requirements, constructing a display data structure containing text content and style information. Output: Advertising display data that can be directly sent to the terminal.

[0191] Step 8: The server sends advertising display data to the terminal.

[0192] Input: Ad display data and user identifier.

[0193] The server encapsulates the display data into a structured response message via the network interface and returns it to the requesting terminal using a communication protocol. The server logs the sending time, ad identifier, and user identifier for later statistical analysis. Output: A successful ad display response sent to the terminal.

[0194] Step 9: The terminal parses and displays the advertisement locally.

[0195] Input: Ad display data sent by the server.

[0196] The terminal uses a parsing module to read structured data and extract advertising text, style parameters, and role identification information. The terminal calls graphical interface components to create text and label views, draws the advertising content on a designated area of ​​the screen, and sets the font, color, and layout according to the style parameters. Simultaneously, the terminal starts a local timer to record the advertisement's exposure duration and prepares to collect interaction data such as scrolling and clicks. Output: The advertising interface presented on the display device and the initial display state to be reported.

[0197] Step 10: Users read and interact with advertisements on their devices.

[0198] Input: The advertising interface displayed on the terminal.

[0199] Users interact with advertisements through touch, click, or scrolling, such as clicking the "Buy Now" button, opening a details link, or entering text in the comment box. Each user click and input is recorded by the terminal's event listening logic, forming a local interaction event list containing timestamps, control identifiers, and input content. Output: User-generated interaction events and comment text.

[0200] Step 11: The terminal reports interaction events and comment information to the server.

[0201] Input: A list of interactive events and comment text.

[0202] When certain conditions are met (such as the number of events reaching a threshold or the required interval), the terminal packages the collected display information (exposure duration, scroll depth), click information (click target and number of clicks), and comment content into a behavior reporting request and sends it to the server via the network interface. After receiving confirmation from the server, the terminal can clear its local cache of reported events. Output: A network request message containing response information.

[0203] Step 12: The server receives the response information and updates the behavior and attribute information.

[0204] Input: Display status, click status, and comment information reported by the terminal.

[0205] The server writes the response information to a new behavior table or appends it to the existing log table, marking it as "feedback source". The server merges this feedback with the original behavioral features, recalculates several statistical indicators for the user (such as recent ad click-through rate and comment sentiment score), and inputs the updated feature vector into the user profile model to perform incremental inference. The server adjusts personality and ideology labels based on the new output, updating user attribute information records as needed. Output: Updated behavior database records and attribute information.

[0206] Step 13: The server generates prompts for replying based on comments and conversation history.

[0207] Input: The user's latest comment text, historical comments and replies, current character settings and attribute information.

[0208] The server retrieves historical conversations related to the user and the advertisement from the dialogue history table and concatenates them chronologically to form a contextual summary. The server constructs new prompts that include system role descriptions, the latest user comment, a personality description and ideological stance of the selected role, and the response task requirements (e.g., "explain product features," "reassure," etc.). The server sets character limits, tone requirements, and prohibited content in the prompts. Output: Response prompts for generative AI models.

[0209] Step 14: The server invokes a generative artificial intelligence model to generate response text.

[0210] Input: Response prompt statement.

[0211] The server repeats the encoding and reasoning process, transforming the prompt statement into a token sequence input model. Using control parameters, it generates a response text that is semantically relevant to the user's comment and conforms to the user's role. The server performs the same sensitive word filtering and length checks on the output text, correcting or regenerating any content that does not meet the requirements. Output: A response text that meets the constraints.

[0212] Step 15: The server encapsulates the response text into display data and sends it to the terminal.

[0213] Input: Response text and dialog identifier.

[0214] The server constructs a display data structure containing the response text, role name, timestamp, and dialogue round number, and sends it to the terminal via the network interface. The server records the response content for that round in the dialogue log table for subsequent dialogue history queries. Output: Response display data that the terminal can directly view.

[0215] Step 16: The terminal displays the response text and continues to collect interaction data.

[0216] Input: The response data sent by the server.

[0217] The terminal parses the response content, adds a new chat bubble in the comment area, and displays the user's name and reply text. The terminal continues to monitor subsequent user input and clicks, adding new comments and feedback events to the local event list. Output: The updated dialogue interface and subsequent interaction events, providing data for the next round of reporting.

[0218] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0219] Example 2 The process flow of a specific process in Example 2 will be described below. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. In addition, the data processing device 12 is referred to as the "server" and the smart device 14 is referred to as the "terminal".

[0220] With the widespread application of generative artificial intelligence models in information retrieval, content creation, and decision support, the following technical problems are prevalent in existing technologies. First, existing systems typically directly invoke generative AI models to generate responses based on a single prompt, lacking structured modeling and automatic decomposition mechanisms for "viewpoint information." This results in generated results often remaining at a single or implicit standpoint, making it difficult for users to obtain multi-perspective, structured responses in a timely manner. Second, in the model invocation chain, most existing systems simply pass user input directly to the model without performing multi-perspective programmatic expansion and task orchestration of the prompts on the server side. This fails to effectively reduce biases introduced by the speaker's attributes or external evaluation information at the system architecture level, resulting in insufficient bias control during the computation process. Third, existing systems lack visual support for the aptitude information or judgment tendencies of generative AI models. Servers typically only return plain text results, making it difficult for users to understand the model's inherent preferences and judgment logic under different viewpoints from the system output. This hinders users from conducting reliance assessments and risk control of the model results.

[0221] More specifically, in traditional computer systems, server-side processing is mostly "request-forwarding": after receiving a request from a terminal, the server almost never restructures the prompt statements, nor does it automatically generate sub-prompt statements based on a preset viewpoint and perform multi-path parallel inference; the server lacks dedicated task management, viewpoint association, and structured integration mechanisms for generative AI inference tasks. Consequently, on the one hand, it cannot fully utilize computing resources to uniformly schedule and optimize batch inference for multi-view tasks on the server side; on the other hand, it cannot provide a unified technical solution for "multi-view structured output" and "bias reduction" at the system level, limiting the interpretability and fairness capabilities of the computing architecture.

[0222] Therefore, an improved computer technology solution is needed. This solution introduces task generation, symbol sequence conversion, inference scheduling, and structured integration mechanisms on the server side, based on prompt statements and viewpoint information. This would enable the server to automatically transform a single prompt statement into multi-viewpoint sub-prompt statements, call generative artificial intelligence models to generate responses for each sub-prompt statement, and return them in a multi-viewpoint structured format. Simultaneously, the server-side needs to visualize the model's orientation information and judgment tendencies. This would improve the diversity of information generation, reduce the impact of biases, and enhance the interpretability and controllability of generative artificial intelligence services at the computer system level, thereby improving the overall technical performance of the computer system and the quality of user interaction.

[0223] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0224] In this invention, the server includes: a receiving and assembling unit for responding to user operations in an information processing device, receiving input prompt statements and related prompt parameters from a terminal, and generating request information containing the prompt statements and prompt parameters; a receiving and assembling unit for automatically generating multiple sub-prompt statements based on the prompt statements and viewpoint information in the request information, and constructing a viewpoint splitting and task generation unit for each sub-prompt statement containing model input information for a generative artificial intelligence model; and a receiving and assembling unit for converting the sub-prompt statements in the model input information into symbol sequences adapted to the symbolic processing of generative artificial intelligence models, and executing the generation using numerical computing resources. The system includes: a reasoning processing unit for generative AI models, which encodes and infers the generated response information for each sub-prompt statement; a multi-view integration and output unit, which associates and integrates multiple generated response information according to viewpoint information, organizes the integrated generated response information into structured response information representing a multi-viewpoint structure, and sends the structured response information to the terminal via a network communication device; and a personality visualization unit, which performs correlation analysis and visualization processing on the aptitude information or judgment tendency reflected in the generated response information based on user-understandable explanatory information and the multi-viewpoint structure, thereby generating model aptitude explanatory data and providing it along with the structured response information. This allows for multi-viewpoint splitting and task scheduling of a single prompt statement on the server side in a programmatic manner, enabling multi-path inference calls and structured integration output of the generative AI model. From the system architecture level, this reduces the bias caused by the speaker's attributes and external evaluation information, improves the balance and diversity of the generated results across multiple perspectives, and enhances the interpretability and controllability of the generative AI model, thereby improving the overall technical performance of the computer system in information generation, resource scheduling, and human-computer interaction.

[0225] "Information processing device" refers to an electronic device that executes programs through one or more processing units, processes input data, and outputs processing results, including but not limited to server devices, computing platforms, or computing nodes.

[0226] "Terminal" refers to an electronic device operated by a user to send prompts to an information processing device and receive returned results, including but not limited to mobile terminals, fixed terminals, or computing devices with communication and display functions.

[0227] "Prompt statements" refer to text data or equivalent data expressions used to instruct generative artificial intelligence models to perform information generation and processing. Their content is used to define the topic, scope, or requirements of the generation task.

[0228] "Prompt parameters" refer to additional information associated with prompt statements that controls the output characteristics or processing conditions of generative artificial intelligence models, including but not limited to language type, output style, target viewpoint, length limit, or other control parameters.

[0229] "Request information" refers to a structured data set generated by an information processing device based on prompt statements and prompt parameters received from a terminal, which is used for subsequent model calls. It includes at least the prompt statement and control information related to the prompt statement.

[0230] "Viewpoint information" refers to the control information used to instruct generative artificial intelligence models from which angles, positions, or dimensions to generate response content, including but not limited to technical, economic, social, or other analytical dimensions.

[0231] "Sub-prompt statements" refer to refined prompt texts automatically generated by the information processing device based on the original prompt statements and viewpoint information, which are targeted at specific viewpoints or sub-tasks and are used to drive generative artificial intelligence models to generate information under the corresponding viewpoints.

[0232] "Model input information" refers to the set of input data that enables a generative artificial intelligence model to perform inference processing, and includes at least sub-prompt statements and control parameters or context information related to those sub-prompt statements.

[0233] "Generative artificial intelligence models" refer to information processing models that use training data to learn parameters based on artificial intelligence technologies such as statistical learning or deep learning, and can automatically generate output information such as text based on input prompts.

[0234] "Symbol sequence" refers to a discrete representation sequence obtained by encoding a prompt or sub-prompt statement in natural language form, including but not limited to tag sequences, index sequences, or other forms of discrete encoded data.

[0235] "Numerical computing resources" refers to hardware or software computing resources used to perform generative artificial intelligence model inference processing, including but not limited to central processing units, graphics processing units, accelerators, computing cores, or corresponding virtualized computing resources.

[0236] "Inference processing" refers to the process of performing forward computation based on the input information of the generative artificial intelligence model, after the model has been trained or deployed, in order to generate corresponding output information.

[0237] "Generated response information" refers to the response content or generated result, expressed in natural language or equivalent form, output by the generative artificial intelligence model after receiving model input information such as sub-prompt statements.

[0238] "Structured response information" refers to output data with a clear hierarchical or labeled structure formed by the information processing device after associating and integrating multiple generated response information according to viewpoint information, which is used to present the data in a multi-viewpoint structure on the terminal.

[0239] "Network communication device" refers to a communication function module used to transmit request information and response structured information between an information processing device and a terminal via a network, including but not limited to network interfaces, telecommunications modules, or communication protocol stacks.

[0240] "Preference information" refers to attribute information that reflects the preferences, tendencies, or characteristics of a generative artificial intelligence model in a specific topic, viewpoint, or context, including but not limited to stance preference, expression style preference, or judgment standard preference.

[0241] "Judgment bias" refers to the systematic judgment pattern exhibited by generative artificial intelligence models when analyzing, comparing, or evaluating input information, including but not limited to a preference for a certain type of viewpoint, sensitivity to risk, or emphasis on the strength of evidence.

[0242] "Personality visualization unit" refers to a functional module in an information processing device that generates representational data based on generated response information and a multi-view structure, which allows users to understand the personality information or judge tendencies of generative artificial intelligence models.

[0243] "Display data" refers to formatted data generated by an information processing device and sent to a terminal for presenting structured and model-oriented information of responses on the terminal interface, including but not limited to data structures with labels, paragraphs, or graphic markers.

[0244] "Information generating entity" refers to the logical entity that undertakes the task of information generation. In this invention, it specifically refers to a generative artificial intelligence model rather than a human individual, in order to distinguish it from content sources that are directly written or published by humans.

[0245] "Attribute information" refers to identifying or descriptive information associated with the speaker, including but not limited to identity category, social status, occupation, background tags, or other characteristic data that can be used to identify or evaluate the speaker.

[0246] "External evaluation information" refers to evaluation data that originates from outside the speaker and is used to measure or describe the speaker's influence or reputation, including but not limited to popularity, ranking, historical evaluation records, or social feedback information.

[0247] The embodiments of the present invention will be described in conjunction with the system structure defined in the appendix. In the following description, the server, as an information processing device, performs the main data processing and calculations, and the terminal, as a user interface device, performs input and display processing. The user sends prompts to the server through the terminal and utilizes the generated results. The present invention is not limited to the specific embodiments described below, and those skilled in the art can make various modifications and substitutions without departing from the spirit of the invention.

[0248] Servers can employ computing platforms based on general-purpose processors and accelerators in their hardware architecture. Servers may include multi-core central processing units (CPUs), such as processors based on general-purpose instruction set architectures; servers may also include one or more graphics processing units (GPUs) for performing large-scale matrix operations and parallel vector computations. Servers can run server operating systems, such as a server operating system based on an open-source kernel; servers can run middleware and application frameworks, such as a web server software and a backend application framework (such as a web framework in a Python environment or a server framework in a JavaScript environment). Servers can also install deep learning frameworks, such as PyTorch or TensorFlow, for deploying and executing generative artificial intelligence models.

[0249] In terms of hardware, the terminal can be a smart mobile device or a personal computing device, equipped with a processor, memory, display screen, and network interface. The terminal can run a mobile operating system or a desktop operating system and communicate with the server through a browser or dedicated application. At the software level, the terminal can use a front-end framework to implement the user interface, used to display input areas for prompts, multi-view result areas, and model orientation visualization areas.

[0250] Users input prompts through the terminal's graphical interface. The prompts are in natural language text, for example: "Please analyze the possible future development trends and impacts of AI technology from three perspectives: technology, economy, and society." "Please explain the main advantages of the new technology from three aspects: efficiency improvement, cost saving, and user experience." Please explain the advantages and disadvantages of large-scale use of artificial intelligence in education from the perspectives of both supporters and opponents. "Assuming you are an economics researcher, please analyze the short-term and long-term impacts of the large-scale application of artificial intelligence on the job market." The terminal combines the user-input prompts with parameters selected on the interface (such as target language, output style, viewpoint list, etc.) to form a structured data object, which is then sent to the server via a network protocol. The server receives this structured data, parses and stores the prompts and parameters. The server manages request information in memory as key-value pairs, including at least the following fields: original prompt field, viewpoint information field, output control field, and session identifier field. The server can temporarily store request information in a cache or in-memory database for later retrieval during the inference process.

[0251] The server constructs a mapping from viewpoint information to sub-prompt statements at the application layer. Based on the configured viewpoint template, the server generates more specific sub-prompt statements for each viewpoint. For example, when the original prompt statement is "Please analyze the possible future development trends and impacts of AI technology from three perspectives: technology, economy, and society," the server can generate the following sub-prompt statements: "Explain in detail the possible future development trends and technical bottlenecks of AI technology from a technical perspective." "Explain in detail the potential economic impact of AI technology on industrial structure and the macroeconomy in the future." "Explain in detail, from a social perspective, the potential future impacts of AI technology on employment structure, education models, and social relations." The server employs a set of rules that go beyond simple string concatenation when generating sub-prompt statements. It can maintain a viewpoint template library containing placeholders and fixed phrases. By replacing placeholders with keywords or time ranges from the original prompt statements, the sub-prompt statements achieve a consistent expressive structure. This rule-based generation ensures that different requests from the same viewpoint exhibit similar feature distributions, contributing to the consistency and stability of generative AI models at specific viewpoints and making response results more semantically comparable.

[0252] Before invoking the generative AI model for each sub-prompt statement, the server performs symbolic processing on the natural language text using a tokenizer or sub-word encoder. The server can use the tokenization tools provided by deep learning frameworks to map the sub-prompt statements into a token sequence, i.e., an integer index sequence. The server uses a pre-trained vocabulary during symbolic processing to ensure consistency with the parameters of the embedding layers within the generative AI model. The server stores the resulting token sequence in a tensor data structure and constructs auxiliary tensors such as attention masks and positional encoding indices as needed to control the scope of the model's attention calculation and the sequence length.

[0253] For generative AI models, the server can employ a large-scale language model based on the Transformer architecture. This model typically includes multi-layered self-attention encoding and decoding units, with each layer containing a multi-head self-attention module, a feedforward neural network module, and a layer normalization module. During the inference phase, the server uses fixed model weight parameters, which have been learned from large-scale corpus data during pre-training. During pre-training, the server or another training platform uses a loss function (e.g., cross-entropy loss) to measure the difference between predicted and true tokens, updating the model weights through backpropagation and an optimizer (e.g., stochastic gradient descent or adaptive learning rate optimizer); the weight update process is achieved through gradient accumulation and batch training. These training steps are completed before deployment; this implementation focuses on the specific application during the inference phase.

[0254] During inference, the server inputs token tensors into the forward computation graph of the generative AI model. At each layer, the model performs operations such as matrix multiplication, vector addition, and non-linear activations (e.g., ReLU or GELU), utilizes multi-head attention to calculate the relevance weights between tokens at each location, and aggregates contextual information based on these weights. The server can use sampling parameters (e.g., top-k or top-p, temperature coefficient) to control the randomness and concentration of the output token distribution, thereby balancing the diversity and controllability of answers across multiple viewpoints.

[0255] To improve computational efficiency, the server can batch process multiple sub-prompt statements from the same or multiple users, merging multiple token sequences into a single batch tensor and performing forward inference uniformly within a single GPU call. This reduces model loading and memory switching overhead, increases throughput, and lowers average response latency. The server can also employ tensor parallelism or pipelined parallelism, depending on the model size and hardware configuration, distributing different layers or tensors across multiple compute nodes to further enhance inference performance.

[0256] After the generative AI model outputs a token sequence, the server decodes the token sequence, using a tokenizer to restore the integer indices into natural language text. The server then performs post-processing on the decoded results, including removing special markers, correcting encoded abnormal characters, and dividing the text into paragraphs by periods or paragraph marks. The server can also perform simple filtering on the text according to preset rules, such as removing duplicate paragraphs and merging excessively short sentences, to improve readability.

[0257] The server binds the generated response information for each viewpoint to the corresponding viewpoint tag, forming a key-value mapping structure. For example, the server can construct the following logical structure: "Viewpoint-Technology" corresponds to the technical viewpoint response text, "Viewpoint-Economy" corresponds to the economic viewpoint response text, and "Viewpoint-Society" corresponds to the social viewpoint response text. The server packages this mapping structure into structured response information and adds metadata to the structure, such as the generation timestamp, the model version used, and important parameters (e.g., temperature, maximum length). This structured organization method facilitates the terminal's partitioned display on the interface and makes it easier for subsequent analysis modules to compare and statistically analyze the results from different viewpoints.

[0258] To visualize model aptitude information or judgment biases, a aptitude analysis module can be added to the application layer. The server can extract features from the response text for each viewpoint, such as keyword distribution, sentiment polarity score, and stance bias score. The server can use traditional natural language processing algorithms (such as dictionary-based sentiment analysis) or smaller-scale classification models to score the generated text according to preset dimensions. Based on these scores, the server generates aptitude descriptions, such as "more inclined to emphasize risk control from a technology viewpoint" or "more inclined to emphasize growth potential from an economic viewpoint." The server sends these aptitude descriptions along with the corresponding viewpoint results as display data to the terminal, allowing users to understand the inherent preference structure of the model's output under different viewpoints.

[0259] The server reduces bias introduced by external evaluation information by explicitly separating speaker identification and reputation information during the prompt processing chain. Specifically, when prompts contain names or expressions identifying specific entities, the server can use named entity recognition or pattern matching algorithms to replace these identifications with generic titles or neutral appellations. This prevents generative AI models from over-relying on external reputation to generate biased content. By reconstructing the input at the viewpoint level, rather than directly copying the original sentence at the surface level, the server technically weakens the source of bias structurally.

[0260] Through viewpoint splitting, batch inference, and structured output, the server not only automates the execution of user tasks but also improves performance and quality at the computer technology level. On one hand, by batch processing multiple sub-hint statements, the server improves hardware utilization and reduces the average time for single-request inference. On the other hand, the structured and templated nature of the viewpoint sub-hint statements leads to a more stable input distribution, helping generative AI models maintain consistent output across different viewpoints, thus improving logical coherence and content contrast. Furthermore, the server's multi-viewpoint structured display and aptitude visualization at the output end allow users to quickly locate information differences, reducing repetitive manual queries and comparisons, thereby minimizing unnecessary communication rounds and data transmission volume at the interaction level.

[0261] After receiving the structured response information from the server, the terminal parses the structured data, identifying viewpoint labels and corresponding text fields. The terminal allocates an independent display area for each viewpoint in the display interface, such as using tabs, collapsible panels, or a column layout, allowing users to visually compare the generated results from different viewpoints. The terminal can also utilize the aptitude description data provided by the server, visually annotating the degree of preference under different viewpoints using icons, colors, or graphical bar charts. The terminal interface provides an area for re-entering prompts, enabling users to continue asking follow-up questions or request more granular analysis for a specific viewpoint.

[0262] In typical use cases, users can leverage the technical advantages of this invention's system in the following ways. For example, when a user wants to evaluate the "role of new technologies in corporate decision-making" from multiple perspectives, the user can input a prompt on the terminal: "Please explain the role of generative artificial intelligence in corporate decision-making from three perspectives: efficiency improvement, cost control, and risk management." The server then generates three sub-prompt statements corresponding to these perspectives and calls the generative artificial intelligence model to produce detailed explanations for each. The user can simultaneously see the sub-perspective results for "efficiency improvement," "cost control," and "risk management" on the terminal, and understand whether the model leans more optimistically or cautiously at each perspective through aptitude descriptions, thus obtaining structured and comparable information in a short time.

[0263] Through the aforementioned modular structure and data flow, the server not only achieves routine processing of prompt statements but also technically transforms the traditional "single input—single path—single output" generation model. Internally, the server introduces mechanisms such as viewpoint-oriented subtask orchestration, symbol sequence batch processing, parallel inference scheduling, and structured response integration, resulting in an architectural-level change in how generative AI models are invoked within the system. The resulting technical effects include: reducing average computational overhead through batch inference; improving output stability through templated sub-prompt statements; improving data management and access efficiency through multi-view structured data; and enhancing overall fairness of results through bias-weakening rules.

[0264] This invention is not limited to a specific generative artificial intelligence model or a specific hardware implementation. The server can replace language models of different sizes and structures according to application requirements. For example, it can use a model with fewer parameters to deploy in a resource-constrained environment, or use a model with more parameters to deploy in a high-performance computing environment. The server can flexibly choose local inference mode or remote inference mode (e.g., through remote model service interface calls) according to network conditions and load. However, regardless of the method used, the server achieves the above-mentioned technical effects through a unified mechanism of viewpoint splitting, batch processing, and structured output.

[0265] Through the above implementation, the server, terminal, and user form a computing system oriented towards multi-viewpoint generation and bias control, centered around a generative artificial intelligence model. The modules in this system do not simply correspond to traditional manual operation steps, but rather, through specially designed data structures and control flows, fully utilize the parallel capabilities of computing hardware and the characteristics of the model structure to achieve comprehensive technical improvements in processing speed, response quality, structured management, and bias suppression.

[0266] use Figure 13 The processing flow is explained.

[0267] Step 1: The user enters a prompt statement on the terminal. Users can input prompts in natural language via keyboard or touchscreen on the terminal's input interface, and select parameters such as viewpoint, language, and style as needed.

[0268] Input: The user's natural language text (e.g., "Please analyze the possible future development trends and impacts of AI technology from three perspectives: technology, economy, and society.") and viewpoint information checked or selected by the user in the interface (e.g., "technology, economy, society") and other control parameters.

[0269] Output: A structured input data object within the terminal, containing prompt strings, viewpoint parameters, output style parameters, etc. The terminal temporarily stores this data in memory, preparing to send it to the server.

[0270] Step 2: The terminal encapsulates the user input and sends it to the server. The terminal reads the prompts and parameters from the interface components, assembles them into a request data structure, and sends it to the server via a network protocol.

[0271] Input: The structured input data object generated in step 1.

[0272] Data processing and computation: The terminal encodes the prompt statements and parameters, usually converting the strings into UTF-8 byte sequences; the terminal organizes these fields into key-value pairs and serializes them into JSON format or equivalent format; the terminal uses network libraries to construct HTTP / HTTPS request messages and attaches necessary header information (such as Content-Type).

[0273] Output: A request message sent to the server over the network, carrying parameters such as the prompt text and viewpoint.

[0274] Step 3: The server receives and parses the terminal request. The server receives request messages from the terminal at the network interface and parses them into internally usable data structures.

[0275] Input: An HTTP / HTTPS request message sent by the terminal (including encoded prompts and parameters).

[0276] Data processing and computation: The server uses a web server or application framework to read the request body, deserializes the JSON string to obtain an object containing fields such as "prompt statement", "viewpoint information", and "other parameters"; the server performs basic cleaning of the prompt statement, such as removing leading and trailing spaces and standardizing line breaks; the server performs simple validation, such as checking whether the prompt statement is empty or exceeds the length limit.

[0277] Output: A request information object in the server's memory, which contains the cleaned prompts and parsed parameters in a structured form.

[0278] Step 4: The server generates multiple sub-prompt statements based on viewpoint information. Based on the viewpoint information in the request information, the server performs rule-based splitting of the original prompt statement and generates multiple sub-prompt statements corresponding to different viewpoints.

[0279] Input: The request information object obtained in step 3 (containing the original prompt statement and a list of viewpoints, such as [technology, economy, society]).

[0280] Data processing and computation: The server reads the template string corresponding to each viewpoint from the viewpoint template library and inserts the topic content or key phrases from the original prompt statement into the template; the server performs string replacement, concatenation and formatting operations on each viewpoint to generate more specific and targeted sub-prompt statements, such as "detailed explanation from a technical perspective..." or "detailed explanation from an economic perspective..."; the server establishes a mapping relationship between each sub-prompt statement and its corresponding viewpoint and stores it in a list or dictionary structure.

[0281] Output: A collection of sub-tooltip statements, each element containing a paired data structure of "viewpoint label + sub-tooltip statement text".

[0282] Step 5: The server performs symbolic encoding on each sub-prompt statement. The server converts the sub-prompt statements into a sequence of symbols (token sequences) suitable for processing by generative artificial intelligence models.

[0283] Input: The set of sub-prompt statements generated in step 4.

[0284] Data processing and computation: The server calls the word segmenter or sub-word encoder to perform word segmentation or sub-word segmentation on each sub-prompt statement, decomposing the text into words or sub-word units; the server maps these units to integer IDs and generates a token ID sequence based on the pre-trained vocabulary; the server constructs the corresponding attention mask and length information, and packages the token sequences of multiple sub-prompt statements into a batch processing tensor to improve the computational efficiency of subsequent inference.

[0285] Output: Encoded input data used for model inference, including the token ID sequence corresponding to each sub-prompt statement and its auxiliary tensors (such as mask, length), usually in tensor or array form.

[0286] Step 6: The server invokes a generative artificial intelligence model to perform inference. The server inputs the encoded sub-prompt statement into the generative artificial intelligence model, performs forward computation, and generates a response token sequence for the corresponding viewpoint.

[0287] Input: the token ID sequence batch and auxiliary tensor obtained in step 5, and inference control parameters (such as maximum generation length, temperature, top-k / top-p, etc.).

[0288] Data processing and computation: The server uses a deep learning framework to perform forward computation of the generative artificial intelligence model on the CPU or GPU; inside the model, matrix multiplication, vector addition, self-attention calculation and non-linear activation are performed layer by layer to extract features and conditionally model the input token sequence; according to the sampling strategy, the server selects the next token ID from the probability distribution of the output at each step, and repeats the iteration until a termination flag is generated or the maximum length is reached; the server obtains a response token sequence for each sub-prompt statement in the batch.

[0289] Output: A set of generated token sequences corresponding to each sub-prompt statement. Each element of the set is a list of integer IDs, representing the response content generated under that viewpoint.

[0290] Step 7: The server will generate a token sequence and decode it into natural language response text. The server converts the token sequence output by the model into readable natural language text.

[0291] Input: The set of generated token sequences obtained in step 6.

[0292] Data processing and computation: The server calls the decoding function of the word segmenter to convert each token ID list into a string; the server performs post-processing on the string, including deleting special control characters, correcting encoding error symbols, and removing extra spaces or line breaks; the server performs simple sentence segmentation based on periods, punctuation, etc., to enhance readability.

[0293] Output: A collection of response texts for each viewpoint, where each element contains a viewpoint label and generated response information in natural language.

[0294] Step 8: The server generates structured information for multi-view responses and performs aptitude analysis. The server integrates the response texts from various viewpoints into a unified structure and performs orientation and judgment bias analysis on them.

[0295] Input: The set of response texts for each viewpoint obtained in step 7.

[0296] Data Processing and Calculation: The server constructs a mapping structure (such as a dictionary or object) based on viewpoint tags, binding "viewpoint → response text"; the server calls pre-configured analysis algorithms (such as sentiment analysis, keyword statistics, and stance classification) to calculate feature values ​​for the response text under each viewpoint, such as sentiment score, risk preference coefficient, and conservatism index; the server generates concise aptitude description text based on these feature values, such as "technical viewpoints emphasize security and controllability" and "economic viewpoints emphasize growth and returns"; the server combines the response text and aptitude description into structured response information, including the main prompt statement, viewpoint list, response content for each viewpoint, and corresponding aptitude description.

[0297] Output: A structured response information object containing multi-view response text and orientation descriptions for display on the terminal side.

[0298] Step 9: The server sends the structured response information to the terminal via the network. The server encapsulates the structured results into a response message and returns it to the terminal that initiated the request.

[0299] Input: The structured information object of the response generated in step 8.

[0300] Data processing and computation: The server serializes structured information into JSON or an equivalent format string and encodes it using UTF-8; the server sets HTTP response headers and status codes, and returns the encoded data as the response body via network protocols; the server may log information before returning, including request identifier, viewpoint configuration, generation length, and time consumption, for subsequent system performance analysis.

[0301] Output: A response message containing structured information is transmitted to the terminal via the network.

[0302] Step 10: The terminal parses the structured response information and displays it from multiple viewpoints. After receiving the server's response, the terminal parses the data and displays it on the interface according to the viewpoint.

[0303] Input: The response message returned by the server in step 9.

[0304] Data processing and computation: The terminal uses a JSON parser to extract the main prompt statement, viewpoint list, response text for each viewpoint, and orientation description from the response body; the terminal constructs a display data structure in memory to classify the results for different viewpoints; the terminal creates a display area (such as a tab, collapsible panel, or list item) for each viewpoint according to the preset UI layout, and fills the corresponding area with text content; the terminal sets colors, icons, or labels based on the orientation description data to visualize the judgment tendency of different viewpoints.

[0305] Output: A multi-view response interface is displayed on the terminal screen, allowing users to read and compare the generated results and their orientation descriptions from different viewpoints.

[0306] Step 11: Users can read the generated results and enter new prompts as needed. Users view the multi-view information returned by the server on their terminals and decide whether to initiate a new round of interaction.

[0307] Input: The multi-view response text and orientation description displayed on the terminal in step 10.

[0308] Data processing and computation: Users understand and compare text from different viewpoints, analyzing information at a cognitive level; users input new prompts on the terminal as needed, such as asking follow-up questions about a particular viewpoint: "Please further explain the specific security risks and protective measures of this solution from a technical perspective."; users pass the new prompts to the terminal application logic by clicking the send button.

[0309] Output: The new natural language prompt and updated parameter selection are repackaged into request data by the terminal and enter a new round of processing starting from step 1.

[0310] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0311] The following technical problems exist in existing computer systems that utilize generative artificial intelligence models: First, servers typically drive generative AI models based solely on fixed prompts manually written by developers, lacking the ability to automatically construct prompts based on target content attribute information. This results in insufficient relevance between the generated results and specific application scenarios (such as video content, text content, etc.), inefficient consumption of computing resources, and difficulty in fully leveraging overall information processing performance.

[0312] Second, servers generally do not dynamically adjust the attributes or prompts of generative AI models based on user state or emotional information. The same model outputs a single style of response information to different users at different emotions and interaction stages, which can easily lead to a decline in user experience. At the same time, multiple rounds of requests are required to obtain appropriate output, thereby increasing network transmission and computing load and reducing system throughput.

[0313] Third, existing systems often only call a single generative artificial intelligence model to generate a single-path response result. They lack a mechanism for multiple information processing units with different attributes to generate multiple types of response information in parallel under the same input information. Users cannot obtain multi-perspective results in a single interaction, which leads to the need to call the model multiple times or manually switch model configurations, increasing the number of requests and latency, and affecting the overall processing efficiency.

[0314] Fourth, when generating evaluation or comment information and sending it to the information distribution platform, the server usually separates the "content selection" and "user context matching". There is no unified management of response information for different emotional states within the system, nor can it automatically select the result that best matches the current user state from multiple response information. This results in a discrepancy between the generated content and the user's real-time needs. The algorithm also lacks fine control over the utilization of network load and platform storage resources.

[0315] Fifth, in the above-mentioned processing, the control of training data and prompts for generative artificial intelligence models is relatively crude, which is not conducive to suppressing biases based on human factors at the system level. It is difficult to build a general information processing infrastructure that can continuously output small deviations, controllable styles, and adaptable to user emotions in various scenarios.

[0316] Therefore, how can we build a system on the server side that can: (1) Automatically generate prompt statements from content attribute information; (2) Dynamically adjust model attributes and prompt statements based on user status or emotional information; (3) Manage multiple information processing units with different attributes and generate multiple types of response information under the same input; (4) Select and output the appropriate response information based on the user's state or emotion from a variety of responses; (5) Simultaneously suppress human bias at the level of training data and prompt statements. Improving the overall processing performance, response efficiency, and output quality of generative artificial intelligence models in network environments has become a pressing computer technology issue that needs to be addressed in this field.

[0317] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0318] In this invention, the server includes a unit for acquiring learning information from an information source to assign specific attributes to a generative artificial intelligence model and preprocessing the learning information; a unit for training the generative artificial intelligence model using the preprocessed learning information; a unit for inputting a predetermined prompt statement into the generative artificial intelligence model and causing the generative artificial intelligence model to generate response information; a unit for automatically generating the prompt statement based on attribute information of target content acquired by a user terminal; a unit for automatically generating evaluation information for the target content based on the prompt statement using the generative artificial intelligence model and sending the evaluation information to an information distribution platform; a unit for dynamically changing at least one of the attributes and the prompt statement assigned to the generative artificial intelligence model based on user state information or emotional information acquired by the user terminal; a unit for setting multiple information processing units assigned different attributes as the generative artificial intelligence model, so that each information processing unit generates multiple types of response information for the same input information and outputs the response information in list form; and a unit for selecting response information from the multiple types of response information that is compatible with the user's state information or emotional information and prompting the user terminal. This allows for the automated processing of content attribute-driven prompt generation, emotion-aware model attribute adjustment, multi-model parallel response generation, and user state-based response selection within the server. This reduces network round trips and unnecessary computations while improving the relevance, response speed, and output quality of generative AI models across multiple scenarios, thereby enhancing the overall technical performance of computer systems in information generation and distribution tasks.

[0319] "System" refers to an overall device or combination of devices consisting of one or more computing devices and software programs running on them, used to perform functions such as information acquisition, data processing, model reasoning, and result output.

[0320] A "server" is a computing device that provides data processing, model running, and service response functions in a network environment. It typically includes a processor, memory, and a network interface for communicating with user terminals.

[0321] "User terminal" refers to an information processing device operated by a user and communicating with a server via a network, including but not limited to computer equipment, mobile terminal equipment, or other electronic devices with input / output functions.

[0322] "Information source" refers to the storage medium or service system that provides learning information or content information, including network resources, data storage systems or other data providing devices.

[0323] "Learning information" refers to textual data, structured data, or other forms of data used to train or adjust the parameters of generative artificial intelligence models.

[0324] "Preprocessing" refers to the process of cleaning, formatting, filtering, normalizing, or otherwise processing the learning information before it is input into a generative artificial intelligence model for training.

[0325] "Generative artificial intelligence models" refer to artificial intelligence models built based on machine learning algorithms that can generate text, speech, or other content information based on input information.

[0326] "Attributes" refer to the parameter settings given to generative artificial intelligence models to represent their behavioral characteristics or output style, including but not limited to personality tendencies, opinion tendencies, language style, or evaluation criteria.

[0327] "Training" refers to the process of using learned information to adjust the internal parameters of a generative artificial intelligence model through optimization algorithms, so that it can generate output information that meets expected characteristics based on input information.

[0328] "Prompt statements" refer to textual information provided as input to generative artificial intelligence models to instruct or constrain the models to generate specific types of response information or content information.

[0329] "Response information" refers to text information or other forms of output information automatically generated by generative artificial intelligence models based on prompts and input information, used to answer questions or provide opinions.

[0330] "Target content" refers to object content that users browse or interact with, including video content, audio content, document content, web page content, or other digital content.

[0331] "Attribute information" refers to metadata or content feature information used to describe the characteristics of target content, including but not limited to title information, category information, tag information, duration information, or other relevant information.

[0332] "Evaluation information" refers to comments, opinions, ratings, or other forms of evaluative output information generated by generative artificial intelligence models for target content.

[0333] An "information distribution platform" refers to a service system used to receive, store, and provide content or evaluation information to multiple users, including content sharing platforms, social platforms, or other online information provision platforms.

[0334] "User status information" refers to relevant information used to characterize the current user interaction status, including but not limited to browsing behavior information, interaction history information, or physiological status information.

[0335] "Emotional information" refers to relevant information used to characterize a user's emotional state, which is obtained by analyzing user input, facial expressions, voice, or behavior to obtain data on the category or intensity of emotion.

[0336] "Information processing unit" refers to a functional module or virtual instance that is logically divided in the system and used to independently perform information processing or response generation tasks based on specific attributes.

[0337] "Multiple types of response information" refers to multiple response results that differ from each other in viewpoint, tone, style, or content structure, generated by information processing units with different attributes under the same input information conditions.

[0338] "List output" refers to the output method that arranges and displays multiple response messages in an ordered or grouped list structure, allowing users to view or compare the response messages side by side.

[0339] "Human bias" refers to the bias or lack of objectivity in response information introduced by human subjectivity, such as the source of training data, the writing of prompt statements, or the way the model is used.

[0340] In one embodiment of the present invention, the server, the terminal, and the application program running on the user terminal collaboratively constitute the basic structure for implementing the system of the present invention. The server includes a processor, a memory, and a network interface, while the terminal includes a processor, a display device, an input device, a camera, a microphone, and a network communication module. The server stores software components such as a generative artificial intelligence model, a prompt statement generation module, an attribute management module, a sentiment analysis module, and a response selection module in its memory.

[0341] In one implementation, the server employs a deep learning-based generative artificial intelligence model, which can be a multi-layer sequence-to-sequence neural network based on a self-attention mechanism. The server's model structure can include a word embedding layer, a multi-layer encoder, a decoder, and a multi-head attention layer. During the training phase, the server uses the cross-entropy loss function as the error function and employs a gradient descent-based optimization algorithm (such as adaptive moment estimation) to iteratively update the model parameters. The server can also introduce data augmentation techniques during training, such as randomly deleting, replacing, or rearranging words in some training samples, to improve the model's robustness to noise and diverse expressions.

[0342] In this invention, the server trains a generative artificial intelligence model using learning information provided by an information source. In its implementation, the server can crawl large amounts of text information from network resources or data storage systems, and use data processing software to clean and structure the crawled data, removing invalid tags, duplicate content, and abnormal characters. During the preprocessing stage, the server divides the text into sentences or paragraphs, performs word segmentation, part-of-speech tagging, and stop word removal on the sentences, and organizes the processing results into a unified data structure stored in a data table.

[0343] In further preprocessing, the server uses a feature extraction module to map sentences into vector representations, which include semantic, sentiment polarity, and topic label dimensions. During training, the server uses these vectors as input to a neural network, learning internal parameter configurations that can represent specific attributes (such as optimism, pessimism, neutrality, rationality, and emotionality) through multi-layer nonlinear transformations and attention mechanisms. This allows the server to select different parameter subspaces for generation based on the attributes during the inference stage.

[0344] In one implementation, the server configures different parameter vector clusters or masks for different attributes. During inference, the server selects the corresponding parameter subset based on pre-defined attributes (e.g., "optimistic perspective," "pessimistic perspective," "neutral scientific perspective," etc.), enabling the same base model to generate response information with different styles and viewpoints under different attribute settings. Through this attribute control mechanism, the server transforms the traditional generation method that relies on fixed model weights into a non-habitual inference process based on dynamically selecting weights according to attributes. This allows for multi-style output without increasing large-scale model copying, reducing storage resource consumption and inference latency.

[0345] In this invention, the terminal is responsible for acquiring content information and user status information related to user interaction. When a user browses multimedia content, the terminal reads the attribute information of the target content from the content platform interface or the local player, including metadata such as title, category, tags, duration, and current playback position. The terminal organizes this metadata into a structured record and sends it to the server via the network interface. The terminal can also collect the user's facial expressions and voice signals, and use a local or remote sentiment analysis module to convert these signals into sentiment tags and sentiment intensity values, which are then sent to the server as user sentiment information.

[0346] In a typical scenario, a user watches video content or reads text content through a terminal, and allows the terminal to collect their interactive behavior and emotional responses within the interface. The user can enter a question or draft comment in a text input box, and the terminal then sends this input to the server. This allows the server to consider both the user's language content and emotional state when constructing prompts and generating responses.

[0347] After receiving the target content attribute information, the server automatically generates prompts to drive the generative artificial intelligence model based on the attribute information using a prompt generation module. When generating prompts, the server no longer relies on pre-hard-coded fixed text. Instead, it uses a set of rules and a machine learning model to parse the content attribute information, combining elements such as category, theme, plot keywords, and user scenario into prompts with clear role settings and output constraints. For example, when the server identifies the target content as a movie review video categorized as science fiction, it can generate the following prompt: "You are an ordinary viewer with an objective yet slightly sentimental personality, currently watching a video reviewing a science fiction movie. Based on the theme and content of this video, please generate a short comment suitable for posting in the comment section, no more than 40 characters, in a natural tone." In another implementation, when faced with an abstract question from a user, the server generates multiple prompts based on a pre-defined multi-role configuration. For example, when a user enters "Please tell me your views on climate change.", the server can generate: "You are an optimistic environmental scholar. Please answer the following question in a positive and hopeful tone, within 150 words. Question: 'Please tell me your views on climate change.' Requirements: Emphasize the opportunities for humanity to address climate change and the positive changes brought about by technology." "You are a pessimistic social commentator. Please answer the following question from the perspective of risks and negative impacts, within 150 words. Question: 'Please tell me your views on climate change.' Requirement: Accurately point out the potentially serious consequences." By automatically constructing prompts using target content attributes and user input, the server transforms model invocation from a single static instruction-driven approach to a multi-dimensional context-driven approach. This reduces the workload of manually designing prompts for different scenarios and significantly improves the relevance and consistency between the generated results and the target content, thereby reducing the storage and transmission resource consumption of invalid generated results.

[0348] When combining user state or emotional information, the server uses a sentiment analysis module to parse the sentiment tags from the terminal, mapping different emotions to different attribute configurations and prompt message adjustment strategies. For example, the server can increase the weight of mapping "joy" to the "optimistic attribute," increase the weight of mapping "sadness" to the "empathy attribute," and increase the weight of mapping "anger" to the "soothing attribute." When generating prompt messages, the server explicitly specifies the character's attitude and tone in the text, for example: "You are a calm and empathetic customer service assistant. The user is very angry. Please respond to the following statement in a soothing and understanding tone. Do not argue with the user, blame the user, or use a commanding tone. User comment: 'This product is completely useless!' Please generate a reply of no more than 80 characters." The server achieves fine-grained control over the generation style by encoding sentiment information into constraints in prompt statements and adjusting the model's internal parameter subspace through the attribute management module. Compared to traditional methods that rely solely on post-output filtering, the server constrains the model search space during the generation phase, suppressing irrelevant or inappropriate outputs in the probability space. This reduces the overhead of multiple regenerations, improves overall inference speed, and increases the proportion of effective outputs.

[0349] In this invention, the server manages multiple information processing units with different attributes as virtual model instances. During implementation, the server can define different parameter masks or adaptation layers for different attributes on the same basic neural network, obtaining attribute-related outputs by selecting different masks during forward propagation. When receiving the same input information, the server can schedule multiple information processing units in parallel to reason about different variations of the same prompt or the same question under different attribute settings, generating various types of response information. The server then organizes these response information into a unified list structure, attaching metadata such as attribute identifiers, confidence scores, and sentiment tendencies, and returns it to the terminal.

[0350] When displaying multiple types of response information, the terminal can present them on the interface in a list or multi-card format, allowing users to view results from different perspectives side-by-side in a single interaction. Simultaneously, the terminal feeds back user selection preferences to the server. For example, if a user frequently selects "rational analysis" type responses, the server can increase the weight of the rationality attribute in subsequent calls. By recording selection feedback, the server updates the weight parameters in the attribute selection module, enabling the sorting and selection of response information to gradually adapt to user habits, achieving adaptive alignment between model output and human preferences.

[0351] In applications that need to send evaluation or comment information to an information distribution platform, the server first generates an evaluation-type prompt statement based on the target content attributes, for example: "You are an objective yet slightly sentimental film critic. Please generate a comment of no more than 40 characters based on your overall impression of the following film and post it in the video comment section. Film Genre: Science Fiction; Film Theme: Exploring Time and Family." After the server outputs a short review using a generative artificial intelligence model, it packages the review into a data structure that conforms to the platform's interface and sends the review information via a network interface call to the platform's application programming interface. When selecting a response, the server automatically selects the most suitable review from multiple candidate reviews based on the user's current state or sentiment information, thereby reducing the user's manual selection steps and the number of repeated calls, and reducing network round-trip latency.

[0352] Through the structured data flow and algorithmic processes described above, the server enables the use of generative AI models in the system to go beyond simply automating the process of manually writing text. Instead, it improves the internal processing mechanisms of the computer by jointly optimizing training data, attribute settings, and prompt statements. The server introduces unconventional combination strategies in attribute control, emotion adaptation, and multi-model parallelism, allowing multiple outputs adapted to different attributes and emotional states to be obtained in a single forward propagation during each inference call. This significantly reduces the resource consumption caused by traditional multi-round calls to different models or multiple regenerations.

[0353] During training and inference, the server improves overall processing speed by reducing redundant computations, compressing attribute-related parameter space, and utilizing parallel execution. By explicitly encoding scene information, sentiment information, and output constraints in the prompt statements, it enhances the semantic relevance and stylistic consistency of the generated results, thereby reducing the error output ratio and the burden of manual post-processing under the same hardware conditions. At the storage layer, the server improves data management structure by uniformly managing training data and prompt statement templates, facilitating subsequent expansion with new scenes and attribute settings.

[0354] In other implementations, the server can also employ different neural network architectures, such as introducing an overlay mechanism or pointer network in the decoding stage to more accurately reference key details in the target content when generating evaluation information; the server can also use a multi-task learning method to enable the same generative artificial intelligence model to learn sentiment classification and text generation tasks simultaneously, extract more general feature representations in the parameter sharing layer, and then fine-tune them through a task-specific layer to further improve the accuracy of sentiment-adaptive generation.

[0355] In other implementations, the terminal can also undertake some lightweight inference tasks. For example, it can run a simplified sentiment classification model locally to perform preliminary analysis of the user's voice or facial expressions, and send the results as high-level sentiment tags to the server, thereby reducing bandwidth consumption and privacy risks associated with uploading raw audio and video data. In these implementations, users still perform natural reading, viewing, and input operations through the terminal, while the server performs complex calculations and coordination related to the generative artificial intelligence model in the background.

[0356] Through the above embodiments, the present invention realizes the technical solutions of automatically generating prompt statements based on content attributes, dynamically adjusting generation attributes based on user emotions, and generating and selecting response information in parallel by multiple information processing units with different attributes under the same input conditions. This transforms the calling method of generative artificial intelligence models in network systems from static, single-channel, and manual-dependent to dynamic, multi-channel, and context-adaptive, thus achieving improvements over traditional technologies in terms of processing speed, output quality, communication load, and resource utilization.

[0357] use Figure 14 The processing flow is explained.

[0358] Step 1: The terminal acquires content attribute information and user input information. Taking the target content as input, the terminal reads attribute information such as title, category, tags, duration, and current playback position from the local player or content platform interface, and outputs it as a structured record. Simultaneously, the terminal takes user actions as input, collecting draft questions or comments typed by the user in the input box, and encapsulates this text along with the attribute information into request data as output. During the encapsulation process, the terminal performs character encoding and length truncation on the text to ensure the stability of subsequent transmission and parsing.

[0359] Step 2: The terminal acquires user status and emotional information. Taking user facial expressions, voice signals, or text input as input, the terminal performs data calculations such as facial feature extraction, voice pitch statistics, or emotional vocabulary matching through a local emotion recognition module or a remote emotion recognition interface. This converts the raw signal into emotion tags and emotion intensity values, and the emotion tags, along with the user identifier, are written into structured data as output. When necessary, the terminal downsamples and compresses the raw multimedia data to reduce the amount of data to be uploaded.

[0360] Step 3: The terminal sends comprehensive request data to the server. Taking the content attribute information output from step 1, the user input information, and the sentiment information output from step 2 as input, the terminal merges this information into a single request message. This message is then serialized and encrypted via the network communication module and sent to the server as a network data packet, with the output being a "request sent" status. Before sending, the terminal generates a timestamp and a request ID for subsequent result matching and timeout handling.

[0361] Step 4: The server parses the request and constructs its internal data structure. Taking the request data sent by the terminal as input, the server receives it through the network interface, deserializes it, and stores the content attribute information, user input text, and sentiment tags into data objects in memory, which are then output as the parsing results. During the parsing process, the server performs data validation and field normalization, such as mapping category strings to standardized category codes and sentiment tags to internal enumeration values, for unified processing by subsequent modules.

[0362] Step 5: The server automatically generates basic prompt statements based on content attribute information. Taking the parsed content attribute information as input, the server performs keyword extraction and rule matching operations through the prompt statement generation module to determine the content type, theme, and scenario. It then combines this information into basic prompt statement text according to a preset template, which is the output. During generation, the server selects appropriate role descriptions and output constraints based on the attribute information, such as limiting the number of characters or specifying the tone, thereby generating prompt statements that are scenario-appropriate.

[0363] Step 6: The server adjusts the prompt statements and model attributes based on user sentiment information. Taking the basic prompt statements output in step 5 and the sentiment tags from step 4 as input, the server performs sentiment-to-attribute mapping operations through the attribute management module to determine the currently applicable combination of model attributes. It then inserts or modifies descriptions about tone, attitude, and response purpose into the basic prompt statements, forming sentiment-adapted prompt statements as output. Simultaneously, the server selects parameter subspaces or attribute masks within the generative AI model based on the sentiment tags, preparing attribute control parameters for subsequent inference stages.

[0364] Step 7: The server constructs multi-role prompts to generate various types of response information. Taking the user's original input text and the sentiment adaptation information generated in step 6 as input, the server constructs multiple prompts based on preset multi-role configurations (e.g., optimistic, pessimistic, neutral, reassuring, etc.). Each prompt contains specific role settings and output requirements, and a list of these prompts is output. During the construction process, the server sets different viewpoint constraints and language style descriptions for different roles, thereby creating multi-perspective generation conditions for the same input.

[0365] Step 8: The server invokes a generative artificial intelligence model to generate response information. Taking the prompt statements (single or multiple) output from step 6 or 7 and their corresponding model attribute parameters as input, the server encodes the prompt statements into a model input sequence. The generative artificial intelligence model then performs forward inference operations to obtain one or more response information sequences as output. During the inference process, the server performs probability sampling or beam search algorithms on the output sequences to balance generation quality and diversity, and performs length pruning and illegal character filtering on the generated results.

[0366] Step 9: The server organizes and scores various types of response information. Taking one or more response messages generated in step 8 as input, the server performs semantic integrity checks, repetition rate statistics, and sentiment consistency verification through a quality assessment module. It also calculates a quality score and a matching score for each response message, generating a list of responses with scores and attribute tags as output. During the organization process, the server stores each response along with its corresponding role attributes and sentiment adaptation strategy for terminal display and subsequent learning.

[0367] Step 10: The server selects target responses based on user status or sentiment information. Taking the response list output from step 9 and the user status information and sentiment tags from step 4 as input, the server performs sorting and filtering operations through the response selection module, selecting one or more responses that best match the current user context as the target responses, and generating a result structure containing a "primary recommended response" and a "list of alternative responses" as output. The server can weigh quality scores, sentiment matching, and user historical preferences during the selection process, ensuring that the selection process conforms to a predetermined technical strategy.

[0368] Step 11: The server generates evaluation information for the target content and sends it to the information distribution platform when needed. Taking the target response information and content attribute information as input, the server encapsulates the response information as evaluation information or comment text into the message structure required by the platform, sends it to the information distribution platform via an external interface protocol, and outputs a "sending result status" and an identifier returned by the platform. During the encapsulation process, the server performs field mapping and format checks to avoid interface failures due to format errors.

[0369] Step 12: The server returns the final response and related metadata to the terminal. Taking the target response selected in step 10, the alternative response list from step 9, and possible sending result statuses as input, the server constructs a response message containing information such as the main response text, a list of multi-perspective responses, role tags, and sentiment tags. This message is sent to the terminal via the network interface, with the output being a "response sent" status. The server may re-encode and compress the text before returning the message to reduce transmission load.

[0370] Step 13: The terminal receives and displays the response information returned by the server. Taking the server's response message as input, the terminal extracts the main response text, multi-perspective responses, and relevant tags through a parsing module. It then displays the main response on the interface and simultaneously shows multi-perspective responses in a list or multi-card format, outputting the user-visible interface content. During display, the terminal adds different identifiers or colors based on role tags or emotional tags, allowing users to quickly distinguish between different types of responses.

[0371] Step 14: Users interact with and select or edit responses displayed on the terminal. The user takes the main response and multi-view responses displayed on the terminal interface as input, confirms by clicking a response, modifies the response text using an edit box, or provides feedback on their "preference type" by selecting a button. These interactive operations are then used as output. The user's selections or editing results are recorded by the terminal as feedback data for subsequent transmission to the server.

[0372] Step 15: The terminal sends user selection results and preference information back to the server. Taking the user's interaction results from step 14 as input, the terminal encapsulates the selected response identifier, edited text, and preference type (e.g., rational preference analysis, brief preference reply, etc.) into a feedback message, sends it to the server via the network interface, and outputs a "feedback sent" status. The terminal may also include a timestamp and session ID when necessary, so that the server can update the user profile and model attribute weights.

[0373] Step 16: The server updates its attribute management strategy and prompt message templates based on user feedback. Taking feedback messages from the terminal as input, the server uses a learning module to statistically analyze user acceptance levels of different attributes and prompt message styles, adjusts attribute weight parameters and prompt message generation rules, and writes the updated strategy parameters into memory as output. In subsequent sessions, the server uses these updated parameters to generate prompt messages and responses that better match user preferences, thus achieving adaptive optimization and improved computational resource allocation at the technical level.

[0374] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0375] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0376] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0377] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0378] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0379] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0380] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0381] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0382] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0383] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0384] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0385] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0386] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0387] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0388] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0389] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0390] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0391] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0392] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0393] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0394] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0395] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0396] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0397] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0398] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0399] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0400] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0401] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0402] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0403] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0404] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0405] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0406] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0407] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0408] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0409] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0410] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0411] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0412] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0413] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0414] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0415] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0416] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0417] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0418] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0419] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0420] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0421] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0422] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0423] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0424] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0425] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0426] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0427] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0428] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0429] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0430] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0431] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0432] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0433] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0434] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0435] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0436] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0437] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0438] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0439] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0440] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0441] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0442] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0443] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0444] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0445] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0446] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0447] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0448] In the emotion mapping, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This is when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This is when there are positive feelings such as "wanting more" or "wanting to know more."

[0449] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0450] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0451] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0452] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0453] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0454] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0455] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0456] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0457] As an example of a single processor, there are two main approaches: First, a processor is composed of a combination of one or more CPUs and software, functioning as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0458] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0459] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0460] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0461] In addition, the following notes are provided in response to the above explanation.

[0462] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device that generates retrieval conditions corresponding to the attribute information corresponding to the personality or thought to be assigned to the generative artificial intelligence model, obtains text information from the information providing device and the information storage device according to the retrieval conditions, attaches the obtained text information to the identification information corresponding to the attribute information and stores it. An apparatus for performing preprocessing on the text information, including string normalization, deletion of useless information, and unit division, structuring the text information into learning data while maintaining the identification information, and converting the structured learning data into the input format of a generative artificial intelligence model; An apparatus for generating generative artificial intelligence models that exhibit different generative behaviors under different personalities or thoughts by inputting an input sequence formed by appending control symbols representing attribute information to the learning data into the generative artificial intelligence model and adjusting the parameters of the generative artificial intelligence model using optimization operations based on error backpropagation. The device deploys the generative artificial intelligence model on an information processing device for inference processing, encodes the prompts based on externally received prompts and attribute information to form the input sequence of the generative artificial intelligence model, and decodes the generated output sequence to obtain and output a text-based response statement. An apparatus for storing the prompt statements, attribute information, and response statements as conversation history, and updating the learning data based on the conversation history, thereby forming a data set for relearning the generative artificial intelligence model.

[0463] (Note 2) The information processing system according to Appendix 1 is characterized in that, It also includes: a device for extracting feature quantities corresponding to the personality or thought assigned to the generative artificial intelligence model based on the conversation history and the output of the generative artificial intelligence model, and displaying the feature quantities in the form of statistical information or visualization information, thereby presenting the personality or thought in a user-recognizable form.

[0464] (Note 3) The information processing system according to Appendix 1 is characterized in that, In the process of generating the response statement by the generative artificial intelligence model, the device generates the response statement by reducing or invalidating the identification information related to the human subject and the identification information related to the speaking medium from the learning data, thereby suppressing the influence caused by the speaker's fame and expression.

[0465] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for obtaining behavioral information, including a user's past browsing history, purchase history, and operation history, from an external storage device, and for summarizing, extracting features, and inferring the behavioral information to generate attribute information representing a user's personality and ideological tendencies. An apparatus for selecting at least one role setting from role information used to manage multiple role settings assigned to generative artificial intelligence models based on the attribute information and product information, and for automatically generating a prompt statement containing the role setting, the attribute information, and the product information, and limiting the number of words, expression style, and tone of the output. An apparatus for inputting the prompt statement into a generative artificial intelligence model, performing inference processing of the generative artificial intelligence model on computing resources, thereby generating personalized advertising messages based on user personality and ideological tendencies, and performing content review and format adjustment on the generated advertising messages to generate display data that can be sent to the user terminal. An apparatus for accumulating response information obtained from the user terminal, including the display status, click status, and comment information of the advertising message, and associating the response information with the behavioral information to update the attribute information and the role information, thereby dynamically changing the generation conditions of the prompt statement and the content input to the generative artificial intelligence model; An apparatus for constructing a prompt statement for generating response text for user comments based on dialogue history including comment information obtained from the user terminal, the role setting, and the attribute information, inputting the prompt statement into the generative artificial intelligence model to generate the response text, and sending the response text to the user terminal; A device for generating prompt statements for multiple different roles, and inputting the prompt statements into the generative artificial intelligence model in parallel or sequential manner to obtain multiple advertising messages or response texts from different perspectives, and generating display control data for presenting the multiple advertising messages or response texts in a distinguishable form on the user terminal.

[0466] (Note 2) According to the information processing system described in Appendix 1, the display data includes attribute information corresponding to the role settings of the generative artificial intelligence model as graphic information or text information, thereby displaying the personality and ideological tendencies of the generative artificial intelligence model in a user-recognizable form.

[0467] (Note 3) According to the information processing system described in Appendix 1, the generative artificial intelligence model is used as the generator of the advertising message and the response text, instead of using human-provided speaker information, thereby reducing the impact of judgment bias caused by the speaker's fame and the content of the speech text when presenting the advertising message and the response text.

[0468] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for responding to user operations in an information processing device, receiving input prompt statements and prompt parameters from a terminal, and generating request information containing the prompt statements and prompt parameters; A device for automatically generating multiple sub-prompt statements based on viewpoint information according to the prompt statements contained in the request information, and for generating model input information containing a generation request issued to a generative artificial intelligence model for each of the multiple sub-prompt statements; An apparatus for converting sub-prompt statements contained in the model input information into symbol sequences adapted to the symbolic processing of the generative artificial intelligence model, and for performing the inference processing of the generative artificial intelligence model using numerical computing resources, thereby obtaining generated response information for each sub-prompt statement; A device for associating and integrating the generated response information according to viewpoint information, organizing the integrated generated response information into a structured response information representing a multi-viewpoint structure, and sending the structured response information to the terminal through a network communication device. A device for visualizing the orientation information or judgment tendency contained in the generated response information based on user-understandable explanatory information and explanatory information of the multi-view structure for the generative artificial intelligence model. An apparatus for enabling the acquisition of generation response information from multiple viewpoints by including the viewpoint information in the generation request, thereby reducing the influence of bias caused by the attribute information of the speaker or external evaluation information.

[0469] (Note 2) According to the information processing system described in Appendix 1, the information processing device generates display data, which associates the generated response information of each viewpoint with the orientation information or judgment tendency of the generative artificial intelligence model in the structured response information, and uses the display data to prompt the personality or thoughts of the generative artificial intelligence model in a user-understandable form.

[0470] (Note 3) According to the information processing system described in Appendix 1, the information processing device is characterized in that, based on the generative artificial intelligence model as a non-human information generating subject, it removes the speaker's identification information and fame information from the prompting statements, and obtains generated response information through multiple sub-prompting statements based on the viewpoint information, thereby providing information that reduces the deviation caused by speaker attributes and text content.

[0471] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A unit for acquiring learning information from an information source for assigning specific attributes to a generative artificial intelligence model and for preprocessing the learning information; A unit for training the generative artificial intelligence model using the preprocessed learning information; A unit for inputting predetermined prompt statements into the generative artificial intelligence model and causing the generative artificial intelligence model to generate response information; A unit for automatically generating the prompt statement based on attribute information of the target content obtained from the user terminal; A unit for automatically generating evaluation information for the target content based on the prompt statement using the generative artificial intelligence model and sending the evaluation information to the information distribution platform; A unit for dynamically changing at least one of the attributes and the prompt statements assigned to the generative artificial intelligence model based on user status information or emotional information obtained from the user terminal; A unit for setting up multiple information processing units with different attributes as the generative artificial intelligence model, so that each information processing unit generates multiple types of response information for the same input information, and outputs the response information in list form; A unit for selecting response information from the various types of response information that is appropriate to the user's status information or emotional information and prompting the user terminal.

[0472] (Note 2) The information processing system according to Appendix 1 is characterized in that, It also includes a unit for visually displaying the attributes and response information of the generative artificial intelligence model generated by the unit, enabling the user to identify them.

[0473] (Note 3) The information processing system according to Appendix 1 is characterized in that, It also includes a unit for controlling at least one of the learning information and the prompting statements, so that the generative artificial intelligence model generates response information that reduces human bias.< / optimistic> < / optimistic> < / optimistic> < / optimistic> < / optimistic> < / optimistic>

Claims

1. An information processing system, characterized in that, include: processor; The processor is configured as follows: Collect learning data related to specific personality traits or ideas from the Internet or databases, and preprocess the learning data; The generative artificial intelligence model is trained using the preprocessed learning data so that the generative artificial intelligence model possesses the specific personality or thought; The generative artificial intelligence model is input with prompts instructing it to generate a specific response, thereby enabling the generative artificial intelligence model to generate the corresponding response.

2. The information processing system according to claim 1, characterized in that, The processor is further configured to visualize the personality or thoughts of the generative artificial intelligence model in a user-understandable form.

3. The information processing system according to claim 1, characterized in that, The processor is further configured to reduce the bias caused by the speaker's fame or the content of the speech text when the generative artificial intelligence model generates a response, since it is not a human subject.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A