Information processing system

CN122797486APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610326716.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-17
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]现有的财务知识讲解系统通常仅基于预先编写的固定文本或规则库向用户提供财务概念解释,难以及时反映财务知识体系的更新,且在语言风格、专业深度和表达方式上缺乏柔性与个性化,无法根据不同用户的理解能力和需求动态生成合适的解说内容,从而导致用户学习效率低、理解不透彻的问题

Benefits of technology

[0007]为解决上述问题,本发明提出了一种信息处理系统,该系统包括处理器,其中,所述处理器被配置为向用户提供用于接收用户问题的接口,对通过所述接口接收到的用户问题采用自然语言处理技术进行解析,从数据库中获取与所述用户问题相关的信息;并基于所获取的信息生成用于指示生成特定财务概念相关解说内容的提示信息,将所述提示信息输入至生成式人工智能模型,以使所述生成式人工智能模型生成所述特定财务概念的解说内容,从而实现对多种财务概念的自动化、灵活化、可扩展的讲解。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797486A_ABST
    Figure CN122797486A_ABST
Patent Text Reader

Abstract

The application provides an information processing system. An information processing system comprises a processor configured to: provide an interface for receiving a user question to a user; parse the user question received through the interface using a natural language processing technique, obtain information related to the user question from a database; generate prompt information for indicating generation of specific financial concept related explanation content based on the obtained information, and input the prompt information to a generative artificial intelligence model to make the generative artificial intelligence model generate the explanation content of the specific financial concept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] Existing financial knowledge explanation systems typically provide explanations of financial concepts to users based solely on pre-written fixed texts or rule bases. This makes it difficult to reflect timely updates to the financial knowledge system. Furthermore, they lack flexibility and personalization in terms of language style, professional depth, and expression, failing to dynamically generate appropriate explanations based on different users' comprehension abilities and needs. Consequently, this results in low learning efficiency and incomplete understanding for users.

[0004] In addition, traditional systems mostly provide static explanations of financial concepts themselves, lacking the ability to recognize and respond to users' emotional states. They cannot automatically adjust the tone and content of the explanation based on the user's emotional changes (such as confusion, anxiety, boredom, etc.) during the explanation process, making the interactive experience rather stiff, making it difficult to continuously attract users' attention, and also not conducive to relieving psychological pressure during the learning process.

[0005] Furthermore, when providing financial services or education to specific users, the existing technologies have relatively weak automated analysis and personalized risk warning functions for users' own financial data. Usually, professionals are required to manually interpret the reports and make risk judgments, which is inefficient, highly subjective, and difficult to provide real-time and detailed financial risk warnings to a large number of individual or corporate users on a large scale.

[0006] Therefore, how to provide an information processing system that can generate personalized and dynamic explanations of specific financial concepts based on generative artificial intelligence models, automatically adjust the explanation style according to the user's emotional state, and further combine user financial data to generate warning information for specific financial risks has become the technical problem that this invention urgently needs to solve. Summary of the Invention

[0007] To address the aforementioned problems, this invention proposes an information processing system. The system includes a processor configured to provide an interface for receiving user questions, parse the user questions received through the interface using natural language processing technology, retrieve information related to the user questions from a database, and generate prompts based on the retrieved information to instruct the generation of explanatory content related to specific financial concepts. These prompts are then input into a generative artificial intelligence model, enabling the model to generate explanatory content for the specific financial concept, thereby achieving automated, flexible, and scalable explanations of various financial concepts.

[0008] Furthermore, to improve the user's interactive experience and enhance the knowledge transfer effect, the processor is also configured to present the generated explanation content to the user in the form of text, voice, or graphical interface, and to collect and analyze the user's tone of voice and facial expressions to identify the user's current emotional state. When the processor identifies the user's emotions as confusion, anxiety, low interest, etc., it dynamically adjusts the tone and / or content of the explanation content based on the emotion recognition results, such as simplifying technical terms, adding examples, slowing down the speaking speed, or using more encouraging wording, thereby adaptively optimizing the explanation method based on the user's real-time feedback.

[0009] Furthermore, to achieve intelligent analysis and risk warning of users' financial situation, the processor is also configured to analyze the financial data provided by users (including but not limited to assets, liabilities, income, expenses, cash flow, etc.), calculate or extract key indicators related to solvency, profitability, cash flow status, and leverage level; based on this, the processor generates a prompt message to indicate the generation of warning information related to specific financial risks according to the analysis results, and inputs the prompt message into the generative artificial intelligence model, so that the generative artificial intelligence model generates natural language warning information corresponding to the specific financial risk, thereby providing users with personalized and interpretable financial risk prompts.

[0010] Through the above-mentioned technical means, the information processing system of the present invention can automatically combine relevant information in the database with generative artificial intelligence models to generate high-quality financial concept explanations after receiving users' natural language questions. It can also perceive the user's emotional state in real time during the interaction process and dynamically adjust the explanation style. At the same time, it can generate targeted risk warning information based on the user's own financial data, thereby effectively solving the technical problems of rigid explanation content, lack of emotional perception and insufficient personalized risk warning capabilities in the prior art.

[0011] "Information processing system" refers to a computer system used to receive user questions, obtain relevant financial information, and generate specific financial concept explanations and / or financial risk warning information through generative artificial intelligence models.

[0012] A "processor" refers to an electronic hardware device or its logical unit used to execute program instructions, control the operation of various functional modules of the system, and complete processing tasks such as problem analysis, data acquisition, prompt information generation, and interaction with generative artificial intelligence models.

[0013] An "interface" refers to a human-computer interaction channel provided by a processor for receiving user input questions and outputting explanations or warnings to the user, including graphical user interfaces, text input boxes, voice input interfaces, etc.

[0014] Natural Language Processing (NLP) technology refers to the techniques used to process user input in natural language, such as word segmentation, part-of-speech tagging, syntactic analysis, intent recognition, and entity extraction, in order to obtain a structured representation of the question.

[0015] A "database" is an organized collection of data that stores information related to financial concepts, financial knowledge, and / or user financial data, and can be retrieved and accessed by a processor.

[0016] "Relevant information" refers to definitions, formulas, explanations, examples, or other knowledge-based content retrieved from the database based on the user's question analysis results, which correspond to the financial concepts or topics involved in the question.

[0017] "Prompt information" refers to instructional or guiding text or data generated by a processor and input into a generative artificial intelligence model, used to instruct the model to generate explanations of specific financial concepts or warnings of specific financial risks.

[0018] "Generative AI models" refer to AI models that can automatically generate natural language text or other content based on input prompts, including but not limited to large language models.

[0019] "Specific financial concepts" refer to specific financial professional concepts that are pointed to by user questions and / or identified by the system, such as free cash flow, debt ratio, discounted cash flow, etc.

[0020] "Explanatory content" refers to natural language text or its audio or graphical presentation form that is output by a generative artificial intelligence model to explain, interpret, or analyze specific financial concepts or financial risks.

[0021] "Voice tone" refers to the changes in pitch, volume, speech rate, and pauses in a user's voice when they speak, which are used to reflect the user's emotions or attitudes.

[0022] "Face" refers to the visual characteristics displayed by users through changes in facial muscles during interaction, used to reflect the user's emotional state, such as confusion, focus, or pleasure.

[0023] "Emotions" refer to the psychological state that users exhibit during interaction with the system, including but not limited to calmness, confusion, anxiety, and varying levels of interest, which can be identified by analyzing tone of voice and facial expressions.

[0024] "The tone of the narration" refers to the stylistic characteristics of the narration in terms of its expression, including the formality of the wording, the seriousness or lightness of the tone, and whether it is encouraging or neutral.

[0025] "User's financial data" refers to structured or semi-structured data related to a user's financial situation, including but not limited to information such as assets, liabilities, income, expenses, cash flow, investment holdings, and loan records.

[0026] "Specific financial risks" refer to specific types of risks identified based on user financial data analysis that are related to debt repayment ability, liquidity, leverage level, cash flow status, and return volatility.

[0027] "Warning messages" refer to natural language text or its presentation form generated by a generative artificial intelligence model to alert users to potential adverse situations and their possible impacts in response to specific financial risks. Attached Figure Description

[0028] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0029] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0030] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0031] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0032] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0033] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.

[0034] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0035] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0036] Figure 9 This represents an emotion map that maps multiple emotions.

[0037] Figure 10 This represents an emotion map that maps multiple emotions.

[0038] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0039] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0040] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0041] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0042] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0043] First, let me explain the terminology used in the following instructions.

[0044] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0045] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0046] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0047] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0048] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0049] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0050] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0051] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0052] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0053] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0054] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0055] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0056] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0057] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0058] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0059] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0060] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0061] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0062] In traditional financial reporting systems or conversational question-and-answer systems, it is known that generative artificial intelligence models are used to generate natural language responses. However, most of these systems merely rely on static text generation without depending on external financial data or the user's environment, and do not actively improve the computer's processing methods or data structures. Therefore, the following computer technology problems exist.

[0063] First, when a server collaborates with a generative artificial intelligence model to use financial data obtained from an external information provider, it typically combines structured data with generated text through simple text retrieval or insertion of fixed templates. In this configuration, the data acquisition, data update, and text generation processes within the server are loosely coupled. Because each process executes independently, this leads to redundant memory accesses, wasted network bandwidth, and increased processing latency, resulting in low efficiency in the use of computing resources.

[0064] Second, in traditional systems, because the user's questioning and explanation history is not fully considered when constructing the input to the generative AI model, the context is easily interrupted in multi-turn dialogues, making it difficult to maintain consistency in the explanation of the same financial concept. As a result, the server either sends a large amount of redundant context with each question or invokes the model when the necessary context is missing. Either way, this leads to increased network traffic and model inference load, and causes inefficiencies in computer technology, such as unstable response quality.

[0065] Third, in existing financial analysis support systems, server-side time-series data processing and visualization generation are often performed independently of the output of generative artificial intelligence models, and the presentation of which financial indicators should be in which graphical form depends on pre-defined rules. Therefore, the lack of correspondence between the items of interest pointed out by the model in the explanatory text and the graphical information generated by the server leads to the display of charts with low relevance to the user. Simultaneously, the server repeatedly performs unnecessary graphical generation processes, causing meaningless execution of graphical generation algorithms or database accesses, thus hindering the efficient use of computing and storage resources.

[0066] Fourth, when feeding back user feedback or additional input to the system, traditionally, this relies mainly on manual debugging or offline model relearning, lacking a sufficiently robust mechanism to automatically update the generation conditions of prompts or the style of explanation output online. Therefore, the server cannot dynamically optimize prompt content based on different users' levels of understanding or preferences, resulting in operation with fixed prompt word designs. Even repeating similar processes fails to improve the effectiveness of responses or the efficiency of computational resource utilization.

[0067] This invention aims to solve the aforementioned computer technology problems by integrating the acquisition, structuring, and updating of external financial data; the generation of prompts for generative artificial intelligence models; the generation of explanations by combining model output with structured data; the generation of graphical information based on time-series data; the context management of session history; and the online adaptation based on user feedback into a server-side collaborative processing flow. This integrated optimization of various processes such as network communication, memory access, model inference, and graphics generation improves the processing efficiency, response performance, and resource utilization efficiency of the computer system.

[0068] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0069] In this invention, the server includes: an information acquisition and storage means for acquiring financial information related to a user's natural language question from an external information providing device via a communication interface, and organizing the financial information into structured information and storing it in a predetermined storage device; a prompt statement generation means for presenting an inquiry interface on a user device and receiving a natural language question input by the user, taking the natural language question and the structured information as input, using natural language processing to parse the question intent and financial concepts, and generating, based on the parsing results and the structured information, a prompt statement for instructing a generative artificial intelligence model to output an explanation containing financial concept definitions, formulas, numerical examples, and time series change descriptions, and inputting the prompt statement into the generative artificial intelligence model; and a prompt statement generation means for receiving the generative artificial intelligence model. The output explanatory text, based on the structured information, inserts specific numerical examples into the explanatory text, and automatically generates graphical information representing changes in financial indicators based on the time series data contained in the structured information. The processed explanatory text and graphical information are then sent to the user's device. This includes an explanatory generation and sending mechanism; a conversation management mechanism for managing the user's natural language question history and explanatory presentation history by conversation unit, extracting summary information related to the current question and incorporating it into new prompt statements to maintain consistency in explanations across multiple rounds of interaction; and an adaptive update mechanism for receiving user evaluation information and additional input from the user's device, accumulating it in a storage device, and dynamically updating the generation conditions of prompt statements and the output style of explanatory content based on the accumulated results to adaptively adjust subsequent generation behaviors. This allows for the formation of a collaborative processing pipeline within the server, encompassing data acquisition, structured storage, context management, prompt generation, model inference result processing, and graphics generation. While ensuring the relevance and coherence of financial explanations, it reduces redundant data transmission and invalid graphics generation operations, lowers the overhead of memory access and model calls, and thus significantly improves the overall processing efficiency and resource utilization of the computer system in financial question-and-answer and visualization tasks.

[0070] A "system" refers to a collection of technical devices consisting of one or more computing devices, storage devices, and communication devices, which work together through hardware and software to perform data processing, communication, display, and control processes.

[0071] A "server" refers to a computing device that undertakes core processing functions such as data acquisition, data processing, generative artificial intelligence model invocation, and result distribution in a system. It typically includes a processor, memory, and communication interface, and runs a program to implement the method steps of this invention.

[0072] "User device" refers to a terminal device operated by a user to input natural language questions, receive and display explanatory text and graphic information, including but not limited to mobile terminals, desktop terminals or web clients.

[0073] A "communication interface" refers to the hardware and software modules used for data transmission and reception between a server and external information providing devices or user devices, supporting the transmission of request and response signals via network protocols.

[0074] "External information providing device" refers to a data source device or service system that provides financial data or related information to a server, including public data service platforms, internal enterprise database systems, or other information providing systems.

[0075] "Natural language questions" refer to queries entered by users through their devices, expressed in human natural language (such as Chinese), to request explanations or analysis of financial concepts or information.

[0076] "Financial information" refers to data related to the economic activities of an enterprise or entity, including but not limited to financial statement data, market price data, and various financial indicator data.

[0077] "Structured information" refers to a data set in which the server formats, organizes, and stores the raw financial information it has acquired in a storage device according to a predetermined data structure. This set has field definitions and a structure that can be accessed by programs.

[0078] "Storage device" refers to physical or logical storage resources used to store structured information, session history, evaluation information, and program code, including main memory, secondary memory, and database systems.

[0079] Natural Language Processing (NLP) refers to the computer technology process of processing text data such as natural language problems through word segmentation, part-of-speech tagging, syntactic analysis, and semantic understanding in order to extract intent and key concepts.

[0080] "Financial concepts" refer to abstract concepts or indicators used in financial analysis or accounting practice, such as rate of return, cash flow, profitability indicators, etc., to explain or evaluate financial conditions.

[0081] "Generative artificial intelligence models" refer to artificial intelligence models built on machine learning that can automatically generate natural language text or other output content based on input prompts.

[0082] "Prompt statements" refer to input text constructed to instruct generative artificial intelligence models on the target content and style of output, which includes at least a natural language question and related contextual information.

[0083] "Explanatory text" refers to the natural language text output by the generative artificial intelligence model based on prompts, used to explain financial concepts or financial information, including definitions, formula explanations, example analyses, and trend descriptions.

[0084] "Time series data" refers to financial indicator data recorded in chronological order at multiple points in time or over a period of time, which can reflect the changes in the indicators over time.

[0085] "Graphical information" refers to graphical data generated based on time series data or other structured information, used to visualize changes or distributions of financial indicators, including line charts, bar charts, and other statistical graphs.

[0086] A “question interface” refers to a graphical user interface displayed on a user device for receiving natural language questions and feedback from the user, including a text input area, a send button, and related controls.

[0087] "Information acquisition and storage means" refers to the hardware and software functional modules in a server that enable the acquisition of financial information from external information providing devices through communication interfaces, the organization of this information into structured information, and its storage in a storage device.

[0088] "Prompt statement generation method" refers to the functional module in the server that receives natural language questions and structured information, performs natural language processing, and constructs and outputs prompt statements to drive generative artificial intelligence models to generate explanations.

[0089] "Commentary generation and transmission method" refers to the functional module in the server that receives the narration text from the generative artificial intelligence model, processes it, combines it with graphic information, and then sends it to the user's device.

[0090] "User interface means" refers to the functional modules in a user device that present explanatory text and graphic information, receive user input and evaluation information, and send them to the server.

[0091] "Session management methods" refers to a functional module in the server that records, organizes, and summarizes natural language questions and explanatory texts by session unit, and utilizes relevant historical information when generating new prompts.

[0092] "Adaptive update methods" refer to functional modules in the server that dynamically adjust the conditions for generating prompts and the output style of explanations based on accumulated user feedback and additional input, in order to improve subsequent generation behavior.

[0093] In one embodiment of the present invention, the server comprises a computer device equipped with a multi-core processor, main memory, persistent storage, and a network interface. The server may employ a general-purpose server hardware platform, such as a rack-mount server equipped with a multi-core processor, and install a server operating system, such as a Unix-like operating system. The server runs application server software, such as a web server, middleware, and backend business logic programs, on this operating system.

[0094] The server deploys relational database management software, such as a relational database system, on the storage device to store structured information, session history, evaluation information, etc. The server also deploys data processing libraries, such as data analysis libraries, on the storage device to parse and calculate externally acquired financial information. The server calls network communication libraries, such as HTTP client libraries, within the application to access the network interface of external information provision devices via HTTP or HTTPS protocols.

[0095] In this embodiment, the server employs a generative artificial intelligence model as its text generation engine. The server can invoke a deep learning-based sequence-to-sequence model, such as a multi-layered self-attention transformer network. This transformer network includes an encoder and a decoder, each consisting of several stacked self-attention layers, feedforward fully connected layers, and normalization layers. During the model invocation phase, the server uses language model parameters trained by an external service provider, or stores pre-trained and fine-tuned model weights locally. When invoking the model, the server specifies hyperparameters, such as temperature parameters, maximum generation length, and probability truncation parameters, to control the diversity and length of the output content.

[0096] The server uses multiple training samples during model training or fine-tuning. Each sample contains a natural language question, the corresponding financial concept, a structured financial data summary, and the target explanatory text. During training, the server uses the cross-entropy loss function as the error function, calculating the loss value based on the difference between the word distribution of the model's output and the word distribution of the target text. The server uses gradient descent-type optimization algorithms, such as adaptive moment estimation, to update the model weights. During training, the server can implement data augmentation strategies, such as using different expressions and numerical examples for the same financial concept, to improve the model's robustness to diverse inputs.

[0097] In terms of structured information management, the server organizes the financial information returned by external information providers into multiple data tables. The server can set up data structures such as "entity tables," "time series tables," and "indicator tables" in the database. For example, the server records basic attributes such as company identifiers and names in the entity table, basic values ​​such as stock price and cash flow corresponding to the company in the time series table by date or reporting period, and various financial indicators derived from the basic values, such as earnings per share and yield. The server links these tables through primary keys and foreign keys, enabling rapid retrieval of required data by company and time range.

[0098] During the generation of structured information, the server cleans and transforms the raw data received from external information providers. The server checks the completeness of fields and the reasonableness of values ​​for each record; when the server detects missing or outlier values, it can interpolate or discard them using preset rules. The server uses a data analysis library to perform arithmetic operations on the raw numerical fields, such as calculating earnings per share based on net profit and the number of issued shares, and calculating a valuation metric based on stock price and earnings per share. After completing the calculations, the server writes the results to an indicator table and marks each record with a timestamp and source information to ensure data traceability in subsequent processing.

[0099] In this embodiment, the terminal can be a mobile computing device, a desktop computing device, or a browser-based terminal. The terminal runs a user interface program locally and calls a graphical user interface library to present a query interface on a display device. The terminal provides a text input area and send controls in the query interface, allowing the user to input natural language questions via a software keyboard. The terminal can also use the speech recognition service provided by the operating system to convert the user's voice input into a text-based natural language question.

[0100] Users input natural language questions as prompts via the terminal. Examples of questions in the following text format can be entered by the user: "Please use a generative artificial intelligence model to explain to me from scratch what the price-to-earnings ratio (PER) is, including its definition, calculation formula, and the general meaning of high and low values, and provide an example using data from a specific company." "Assuming a company's stock price is 500 yen and its earnings per share (EPS) is 50 yen, please calculate the PER using a generative artificial intelligence model and explain where this PER is roughly at." "Based on the cash flow statement data of the most recent three years, please use a generative artificial intelligence model to analyze the changing trend of the company's operating cash flow, and use charts and text to explain the possible operating problems." "If a company's profits have been growing for three consecutive years, but its operating cash flow has been negative for three consecutive years, please use a generative artificial intelligence model to list three to five common reasons and explain what risks investors should pay attention to." After the user confirms the question, the terminal packages the natural language question along with metadata such as user identifier and session identifier into a message and sends it to the server via the network interface. Upon receiving the explanatory text and graphic information returned by the server, the terminal displays the text and graphics in a unified interface through the user interface module. When displaying graphics, the terminal uses vector graphics libraries or image controls to present line charts or bar charts on the screen, supporting user scrolling, zooming, and other operations to facilitate viewing details.

[0101] After receiving a natural language question from the terminal, the server associates the question with the current session. The server maintains a historical question and explanation summary for each session in its session management module. When generating a new prompt, the server extracts several rounds of question-and-answer summaries related to the current question's topic from the session history and inserts them as context into the preceding text of the prompt. During this process, the server uses algorithms such as keyword matching and semantic similarity calculation to select the most relevant content from the historical records. Through this session management strategy, the server can provide sufficient context for the generative AI model without sending all historical content, reducing network traffic and model input length, and improving inference efficiency.

[0102] In the prompt generation module, the server concatenates natural language questions, structured information summaries, and conversation summaries into a single input text according to a predetermined format. The server can include multiple parts in the prompt, such as a system description section, a user question section, and a data summary section. In the system description section, the server specifies the model's behavior, for example, instructing the model to explain financial concepts in concise Chinese and prioritizing the use of structured numerical data as examples. In the data summary section, the server introduces key figures relevant to the current company or indicator, such as the latest year's earnings per share, stock price, and cash flow. In the user question section, the server retains the user's original sentence to ensure the integrity of the user's intent. This structured prompt construction method allows the generative AI model to simultaneously utilize pre-trained language knowledge and current structured data when generating text, thereby improving the consistency between the explanation content and real-time data.

[0103] After receiving the explanatory text output by the generative AI model, the server performs post-processing. The server can use rule-based text processing algorithms to segment the generated text, remove redundant prefixes and suffixes, and standardize terminology. During post-processing, the server also verifies the generated content against structured information. If the server detects inconsistencies between the model's output numbers and the structured data, it can correct the inconsistencies using rules or re-insert the correct values ​​using structured data. Through this rule-based data verification process, the server can reduce numerical biases caused by the model's generation and improve the numerical accuracy of the explanatory text.

[0104] In the graph generation module, the server automatically selects the corresponding time series data for the indicators and time ranges mentioned in the explanatory text by the generative AI model. The server parses the explanatory text in this module, identifying the indicator names and time intervals, such as the description "operating cash flow for the last three years." The server then extracts time series data matching this description from the structured information. The server uses a graph library to construct graphical objects, such as line charts or bar charts, and generates coordinate values ​​for each data point. The server exports the generated graphs as image files or encodes them as graphical data streams, recording the graph's metadata (such as coordinate scales and data labels) for interpretation by the terminal during display. Because the server only generates graphs for the indicators actually involved in the explanatory text, it avoids graphical processing of irrelevant data, thereby reducing the number of graph generation attempts and the computational load.

[0105] In the adaptive update module, the server receives user feedback and additional input uploaded by the terminal. The server records data such as satisfaction metrics and types of user follow-up questions for each explanation in its storage device. Based on this data, the server analyzes the effectiveness of different prompt templates and output styles. For example, the server can analyze whether user satisfaction increases when using prompt templates with more numerical examples; and whether users ask fewer follow-up questions when using more detailed definitions. The server adjusts the prompt generation rules based on the statistical results, such as increasing or decreasing the proportion of definitions, formulas, or case studies in the model's generated content. When calling the generative AI model, the server can also dynamically set the output length and level of detail based on historical user feedback to adapt to the needs of different user groups.

[0106] Through a series of specific technical means, including structured information management, session context control, prompt statement construction, model output correction, graphics generation control, and online adaptive updates, the server achieves multiple technical benefits at the computer technology level. By reducing the transmission of irrelevant historical content and redundant data, the server lowers network communication load. By embedding structured data summaries in prompt statements, the server enables the model to generate explanations consistent with the current data with shorter inputs, thereby reducing model inference complexity and response time. Furthermore, by validating structured information and triggering graphics generation on demand, the server reduces unnecessary database access and graphics rendering operations, improving storage access efficiency and graphics generation efficiency. Because the above processing flows are implemented by the server using specific data structures, algorithmic sequences, and model invocation methods, rather than simply automating manual operations, this system represents a substantial technical improvement in internal computer resource management and data flow control.

[0107] Users in this system receive not only textual explanations of financial concepts through the terminal, but also numerical examples and intuitive graphical displays consistent with current financial data. When displaying this content, the terminal can utilize hardware-accelerated graphics rendering capabilities to enhance interactive smoothness. Since the server has already filtered and compressed the transmitted data, the terminal only needs to process text and graphics relevant to the current question, thereby reducing the computational load and energy consumption on the terminal side.

[0108] In other implementations, the server can employ different types of generative AI models. For example, the server can choose a language model containing only a decoder structure and encode structured information into feature vectors through an additional feature concatenation layer, then concatenate this vector with the prompt text to input the model. The server can also adopt a hybrid architecture, combining traditional rule engines with neural network models to enforce rule validation for certain high-risk values. Regarding learning methods, the server can employ an incremental learning strategy, periodically fine-tuning the model using new dialogue data to further improve its adaptability to user group characteristics.

[0109] The terminal can also adopt different layouts in terms of interface design. For example, the terminal can use a two-column layout, with one column displaying text explanations and the other displaying relevant charts; or it can use a tabbed format, allowing users to switch between "definition descriptions," "numerical examples," and "trend graphs." These changes do not change the core idea of ​​this invention: the server jointly processes structured information and generative artificial intelligence models.

[0110] In the above implementation, the server, terminal, and user each assume different responsibilities. The server is responsible for data management, algorithm execution, and model invocation, serving as the main computing platform for realizing the technical solution of this invention. The terminal is responsible for interacting with the user and is the carrier for presenting the technical effects of this invention. The user provides questions and evaluation data to the server through natural language input and feedback, enabling the server to continuously optimize its internal processing strategies at the technical level. Through the above configuration, this invention achieves collaborative optimization of data structure, algorithm flow, and model invocation mode within the computer system, thereby achieving technical effects such as improved processing speed, enhanced data consistency, and reduced communication load in financial explanation tasks.

[0111] use Figure 11 The processing procedure is explained.

[0112] Step 1: The server initializes the data structure and loads the model configuration.

[0113] Inputs: Database connection configuration in the storage device, interface configuration of the external information provider, and access key and parameter configuration of the generative artificial intelligence model.

[0114] The server reads the database address, username, password, and table structure definition from the configuration file and establishes a connection to the relational database; the server constructs the basic request headers for the HTTP client based on the URL of the external information provider, authentication token, and other information; the server initializes the model invocation client based on the interface address, model name, temperature parameters, maximum generation length, and other information of the generative artificial intelligence model service.

[0115] The server creates a mapping structure in memory for caching structured information (such as a dictionary with enterprise identifiers as keys and lists of time-series data as values), and a hash table for recording session history (with session IDs as keys and lists of questions and answers as values).

[0116] Output: The established database session object, the external API client object, the generative artificial intelligence model client object, and the data cache structure residing in memory.

[0117] Step 2: The server obtains raw financial data from external information providers and performs structured processing.

[0118] Input: API interface description of the external information provider (including query parameters such as enterprise identifier and date range), and the HTTP client object initialized in step 1.

[0119] The server constructs an HTTP request, setting query parameters for each target company and time range; the server sends the request through a network interface and receives the returned data stream in JSON or CSV format. The server parses the received text stream, mapping each record to an internal temporary data object.

[0120] The server uses a data processing library to clean the parsed results: removing null records, checking whether numeric fields are within a reasonable range (e.g., the rate of return is not non-numeric), and uniformly formatting date fields. The server performs arithmetic operations on fields such as net profit and number of shares to calculate derived indicators such as earnings per share, and writes these calculation results into internal structured records.

[0121] Output: A cleaned and computed set of structured financial data in an in-memory data structure indexed by company identifier, indicator name, and timestamp.

[0122] Step 3: The server writes structured financial data to the database and updates the cache.

[0123] Input: The structured financial data set generated in step 2 and the database session object.

[0124] For each structured record, the server generates an SQL statement (e.g., INSERT or UPDATE) that writes the company identifier, reporting period, raw numeric fields, and derived metric fields to the corresponding database table. After performing the write operation, the server updates the time series list corresponding to that company in the memory cache, inserting or overwriting the new record in chronological order.

[0125] If the server detects a database constraint conflict (such as a duplicate primary key) during the write process, it executes update logic to replace the old value with the latest data.

[0126] Output: The latest structured financial data table in the database, along with its synchronized in-memory cache content.

[0127] Step 4: The terminal displays a prompt interface and receives natural language questions input by the user.

[0128] Input: Local interface templates, input methods and display components provided by the operating system.

[0129] The terminal renders a text input box and a send button on the screen, and calls the system input method service to receive the content typed by the user. The user enters a natural language question in the input box of the terminal, such as "Please use a generative artificial intelligence model to explain to me from scratch what the price-to-earnings ratio (PER) is, including the definition, calculation formula, general meaning of high and low values, and give an example with a specific company's data." When the user clicks the send button, the terminal combines the text in the input box with the locally stored user identifier and session identifier into a message object.

[0130] Output: A request message containing a natural language question text, a user identifier, and a session identifier.

[0131] Step 5: The terminal sends the user's question as a prompt to the server.

[0132] Input: The request message generated in step 4 and the terminal network stack.

[0133] The terminal encapsulates the natural language question text into fields in the HTTP request body, puts the user identifier and session identifier into the header information or request parameters, and sends a POST request to the server's interface address using the HTTPS protocol.

[0134] During the transmission process, the terminal calls the system network library to establish an encrypted channel, writes the serialized message content into the transmission buffer, and waits for the server's response.

[0135] Output: The HTTP request that arrives at the server, which includes the user's natural language question and metadata.

[0136] Step 6: The server receives user questions, preprocesses them, and associates them with the session.

[0137] Input: The HTTP request sent by the terminal in step 5.

[0138] The server receives the request within the web framework, parses the request body, and extracts the natural language question text, user identifier, and session identifier. The server removes leading and trailing whitespace characters from the question text and checks if the text length is within a preset range; if it is too long, the server performs a truncation operation, retaining only the most important initial content.

[0139] The server searches for the record with the corresponding session ID in the session hash table. If it does not exist, a new session entry is created; if it exists, the current question is appended to the session's history list. The server extracts summaries of the most relevant rounds of questions and answers from the session history and constructs a short contextual fragment.

[0140] Output: Cleaned natural language question text, context summary related to the session, and information on the association between the user and the session.

[0141] Step 7: The server extracts structured financial data related to the problem from the database and cache.

[0142] Input: The problem text and context summary obtained in step 6, and the database and cache structure built in step 3.

[0143] The server identifies company names or codes, time ranges, and key metrics (such as "operating cash flow for the last three years" and "current PER") in the query text, and determines query conditions through string matching or simple semantic parsing. The server then uses these conditions to issue queries to the cache or database, retrieving numerical records for the corresponding company and time period from the entity table, time series table, and metric table.

[0144] The server organizes the query results into a structured data summary, such as extracting the annual operating cash flow amount for the most recent three years, the stock price and earnings per share at the most recent point in time, forming a list of key-value pairs or a compact text summary.

[0145] Output: A structured data summary related to the current problem, including indicator names, time points, and corresponding values.

[0146] Step 8: The server provides prompts for constructing generative artificial intelligence models.

[0147] Inputs: Natural language question text and context summary from step 6, structured data summary from step 7, and model invocation configuration.

[0148] The server concatenates multiple text segments in memory: First, it constructs a system description segment, instructing the model to "explain financial concepts in professional yet accessible Chinese, and prioritize the use of input data for examples"; then, it inserts a structured data summary segment, describing in concise sentences such as "The company's operating cash flows for the most recent three years were..."; finally, it appends the user's original question and necessary conversational context question-and-answer summaries.

[0149] During the server's construction process, the length of the prompt statements is controlled to ensure that the total number of characters does not exceed the model's input limit, and different parts are separated by line breaks and labels.

[0150] Output: A complete text of prompts used to drive the generative artificial intelligence model.

[0151] Step 9: The server invokes a generative artificial intelligence model to generate explanatory text.

[0152] Input: The prompt statement and model client object constructed in step 8.

[0153] The server passes the prompt statement as input parameters to the model service interface, sets parameters such as temperature and maximum generation length, and then sends a request to the model service over the network. Upon receiving the response, the server receives the text fragments returned by the model in token order and assembles them into a complete narration text.

[0154] The server performs a preliminary check on the generated results, such as confirming whether the text contains definitions, formulas, and at least one numerical example; if there are serious omissions, the server can reconstruct the prompt statement according to the rules and request the model again.

[0155] Output: The original explanatory text output by the generative artificial intelligence model.

[0156] Step 10: The server corrects and enhances the narration text based on structured data.

[0157] Input: The explanatory text obtained in step 9 and the structured data summary from step 7.

[0158] The server scans the numerical values ​​and indicator names appearing in the explanatory text and compares them with the actual values ​​in the structured data. If it finds that the model-generated numbers are inconsistent with the actual data, the server replaces the incorrect numbers in the text with the correct values ​​from the structured data.

[0159] The server inserts explicit calculation examples based on the structured data summary, such as "the current stock price is X, earnings per share are Y, therefore the valuation metric is X / Y," and annotates these examples in the text. The server can also supplement the description with phrases like "showing an upward / downward trend over the last three years" based on time-series trends in the data to enhance the completeness of the explanation.

[0160] Output: Corrected and enhanced explanatory text, with numerical values ​​consistent with structured data, and accompanied by specific calculation examples and trend descriptions.

[0161] Step 11: The server generates graphical information based on time-series data.

[0162] Input: The time series data obtained in step 7 and the enhanced explanatory text obtained in step 10.

[0163] The server identifies the timeframes and metrics mentioned in the explanatory text, such as "operating cash flow over the last three years" or "changes in this metric over the past five years." The server then extracts numerical sequences from the time-series data that match these descriptions.

[0164] The server calls a graphics library to generate a line chart or bar chart for the selected time series, mapping each time point to the x-axis and the corresponding indicator value to the y-axis. The server sets axis labels, titles, and data point markers, generates a complete graphical object, and exports it as an image file or encoded as an image binary stream.

[0165] Output: Graphical information showing how financial indicators change over time, including the generated image data and its metadata (title, coordinate labels, etc.).

[0166] Step 12: The server combines the text and graphical results and sends them to the terminal.

[0167] Input: Enhanced explanatory text from step 10 and graphical information from step 11.

[0168] The server constructs a response message, placing the explanatory text in the text field, the image data or image access path in the image field, and attaching necessary metadata (such as issue ID and timestamp). The server returns this message to the terminal via an HTTP response.

[0169] Before sending, the server can compress the data to reduce the size of network transmission and record a summary of the response in the session history for subsequent context extraction.

[0170] Output: A response message containing explanatory text and graphical information sent to the terminal.

[0171] Step 13: The terminal receives the server's response and displays the narration and graphic information.

[0172] Input: The response message sent by the server in step 12.

[0173] The terminal parses the text and graphic fields in the response, creates a text display area on the screen, renders the explanatory text by paragraph, and highlights formulas or key sentences. The terminal decodes the image data contained in the graphic field and embeds line charts or bar charts near the text, allowing users to see both text descriptions and graphical trends simultaneously.

[0174] During the rendering process, the terminal calls the graphics rendering library to scale the image to adapt to the screen resolution and provides interactive operations such as clicking to zoom in and swiping to browse.

[0175] Output: A comprehensive explanatory interface displayed on the terminal screen, including text descriptions and corresponding financial graphs.

[0176] Step 14: Users read or view the graphics and provide feedback or ask additional questions.

[0177] Input: The explanation interface displayed on the terminal in step 13.

[0178] After reading the text explanation and observing the graphical trends, users can determine whether the content is easy to understand and whether it resolves their current questions. Users can click feedback buttons such as "Helpful" or "Not Helpful" on the terminal, or they can enter additional questions in a new input box, such as "Based on the explanation just given, please compare the level of this indicator with that of other companies in the same industry."

[0179] Output: User-generated rating information (such as satisfaction ratings) and new natural language question text.

[0180] Step 15: The terminal collects user feedback and additional questions and sends them back to the server.

[0181] Input: User feedback and additional question text from step 14.

[0182] The terminal packages the user-selected feedback flag, optional text comments, and new natural language questions into a message, encapsulating the session identifier and the previous question ID. The terminal then sends this message to the server's feedback and Q&A interface via HTTPS.

[0183] After sending the question, the terminal displays the new question in the local chat list and waits for the server's next round of explanation.

[0184] Output: Feedback and additional question requests received from the server.

[0185] Step 16: The server updates session history and adaptive parameters.

[0186] Input: The feedback and append question request sent by the terminal in step 15, as well as the currently saved session history and prompt statement generation rules.

[0187] The server adds new natural language questions to the history of the corresponding session and writes user satisfaction ratings and comments to the evaluation data table. The server updates the statistics for that session or user, such as cumulative satisfaction scores and the frequency of additional questions.

[0188] The server adjusts the parameters for generating prompts based on accumulated statistics. For example, it reduces terminology density and increases the number of numerical examples for users who frequently report "the explanation is too difficult"; and it shortens the generated text for users who prefer concise answers. The server stores the adjusted parameters in the configuration cache for use the next time a prompt is constructed.

[0189] Output: Updated session history and prompt generation configuration, enabling subsequent calls to generative AI models to adapt to user needs and improve technical processing efficiency.

[0190] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0191] Existing personal financial management assistance systems mostly analyze users' income and expenditure data in a fixed format using preset rules or simple statistical functions, and then present the results to users in templated text. These systems suffer from the following technical problems: First, on the server side, the problem parsing and data analysis processes are loosely coupled. The server often simply maps the user's natural language question into a few keywords before executing business logic, failing to integrate the natural language understanding results with the results of large-scale structured financial data analysis within the same technical process. This leads to poor consistency between the dialogue content and the data analysis results, and the generated content lacks specificity. Second, when the server calls generative AI models, most only use the user's original question as input, lacking structured embedding of the user's financial status data. The prompts are crudely constructed, resulting in model outputs that are insensitive to individual user data, making it difficult to provide personalized suggestions based on specific expenditure structures and budget situations. Third, existing systems typically do not technically manage historical interaction data as part of the model input construction. The server cannot automatically update prompts using dialogue history at the technical level, thus failing to reliably support cross-round, context-continuous financial consultations, resulting in a lack of contextual coherence between different rounds of dialogue. Fourth, the relationship between the terminal and the server is mostly a simple request-response model. The terminal only acts as a display terminal and does not have unified control over the dual-channel output of text and voice of the generated results. Therefore, it is impossible to guarantee the consistency and traceability of the same auxiliary speech data in multiple presentation methods at the system level.

[0192] Against this backdrop, an improved computer implementation scheme is needed, enabling the server to: on the one hand, perform fine-grained intent parsing of user natural language questions and conduct statistical and aggregate analysis of user financial data within the same processing flow; on the other hand, embed the parsing and analysis results into the prompts of the generative artificial intelligence model in a structured manner, and dynamically update the prompts in combination with historical interaction data, thereby technically improving the input quality and relevance of model calls, and ultimately realizing a financial assistance service that is continuous, targeted, and multimodal, thereby improving the human-computer interaction effect and computing resource utilization efficiency of the entire platform.

[0193] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0194] In this invention, the server includes: means for acquiring transaction record data and income record data from an external storage device based on user identification information, and generating user financial status data representing expense category expenditure amount, expense category expenditure ratio, and budget information through statistical processing and summary processing; means for executing a natural language processing program through a processor to extract intent information, object information, and time information from question data sent by the terminal to determine the question type and generate summary information associated with the user financial status data; means for constructing prompt statement character data containing user expenditure composition, income information, budget information, and historical interaction information based on the question type and the user financial status data, and controlling the processor to input the prompt statement character data into a generative artificial intelligence model; and means for acquiring answer data from the generative artificial intelligence model, generating suggestion data including financial concept interpretation information or saving suggestion information based on the answer data, and sending it to the terminal through a communication device. This allows for the creation of an integrated processing flow within the server that tightly couples natural language parsing, structured financial data analysis, and prompt statement construction. This significantly improves the relevance and contextual integrity of the data input to the generative artificial intelligence model, thereby generating suggestion information that is more in line with the user's individual financial situation and maintains continuity in multi-turn dialogues. Furthermore, it enables the terminal to uniformly control the text and voice output of the same suggestion data, thus improving the overall technical performance of the computer system in terms of human-computer interaction quality, data processing efficiency, and model inference effectiveness.

[0195] A "terminal" refers to an electronic computing device carried or operated by a user for executing applications, inputting and outputting information, and interacting with a server through a communication network. This includes mobile phones, tablet computers, and personal computing devices.

[0196] A "server" refers to an electronic computing device or set of devices equipped with a processor and storage devices, used to receive data requests from terminals and perform data processing, program execution, and result return in a network environment.

[0197] "Information input / output device" refers to a human-machine interface or hardware / software module used on a terminal or server to receive user input information and output information to the user or other devices, including touch screens, keyboards, displays, network interfaces, microphones, speakers, etc.

[0198] "Communication device" refers to the general term for hardware and software components used to send and receive data between a terminal and a server, or between a server and an external device, including network interface cards, wireless communication modules, wired communication modules, and corresponding communication protocol stacks.

[0199] "External storage devices" refer to data storage resources used for persistent storage of user transaction records, revenue records, and other related data, including database servers, network storage devices, disk arrays, etc.

[0200] "Transaction record data" refers to a set of structured or semi-structured data related to a user's financial expenditures, including information such as expenditure time, expenditure amount, expenditure category, payment method, and counterparty.

[0201] "Income record data" refers to a collection of structured or semi-structured data related to various types of income earned by users, including income time, income amount, income source category, account information, etc.

[0202] "Statistical processing" refers to the process by which the server performs mathematical operations and data analysis on the acquired transaction and revenue data, including summation, counting, mean calculation, variance calculation, frequency statistics, etc.

[0203] "Summary processing" refers to the process by which a server aggregates and categorizes statistically processed data according to predetermined rules, including grouping and summarizing data by time interval, expense category, payment method, and other dimensions.

[0204] "Expense Category Amount" refers to the total expenditure amount corresponding to each expense category obtained by classifying and statistically analyzing user expenditure data according to pre-defined expense categories within a predetermined period.

[0205] "Expense Category Spending Ratio" refers to the proportion of a certain expense category to the user's total expenditure during a predetermined period, used to represent the relative proportion of each expense category in the overall expenditure structure.

[0206] "Budget information" refers to data related to the target spending limits or planned values ​​set by users for different expense categories or overall expenditures, including budget amount, budget period, budget category, etc.

[0207] "User financial status data" refers to a comprehensive data set that reflects a user's income and expenditure structure and financial status during a predetermined period, obtained by the server after statistical processing and summarizing transaction record data, income record data, and budget information.

[0208] A "natural language processing program" refers to a software program or program module that runs on a server's processor and is used to analyze and understand natural language text, including functions such as word segmentation, part-of-speech tagging, syntactic analysis, entity recognition, and intent recognition.

[0209] "Intent information" refers to semantic information extracted from user question data by natural language processing programs to represent the goal or task type that the user expects the system to perform.

[0210] "Object information" refers to semantic information extracted from user question data by natural language processing programs to represent target entities or topics of interest to the user, including expenditures, income, budgets, and specific expense categories.

[0211] "Time information" refers to semantic information extracted from user question data by natural language processing programs to represent the time range or point in time involved in the question, including "this month", "this week", "within a year", etc.

[0212] "Issue type" refers to the issue category identifier obtained by the server after classifying user issues based on intent information, object information, and time information. It is used to distinguish different types of issues such as saving suggestions, concept interpretation, and risk warnings.

[0213] "Summary information" refers to structured or semi-structured data generated by the server based on the user's financial status data and natural language problem analysis results. This data is used to summarize the user's financial characteristics and key points of the problem, and is used to construct subsequent prompt statements.

[0214] "Prompt statements" refer to character data generated by the server and input into the generative artificial intelligence model. Their content includes a comprehensive description of the model's role settings, task descriptions, user questions, user financial status data, and historical interaction information, which are used to guide the generative artificial intelligence model to generate expected outputs.

[0215] "Generative artificial intelligence models" refer to artificial intelligence models trained using machine learning methods that can automatically generate natural language text or other forms of content based on input prompts, including large-scale language models that employ deep neural network structures.

[0216] "Character data" refers to a sequence of information generated by the server and represented in text form, including text, numbers, punctuation marks, etc., used to construct prompts or transmit problem data, suggestion data, etc.

[0217] "Response data" refers to the response content data, mainly in natural language, output by the generative artificial intelligence model after receiving prompts. This data is used to generate financial concept interpretation information or cost-saving suggestions.

[0218] "Suggested data" refers to data generated by the server based on the response data, which includes information on financial concepts, saving suggestions, or other suggestions related to the user's financial status, and is sent to the terminal for presentation.

[0219] "Display device" refers to an output device used on a terminal to present text or graphic data in a visual form, including liquid crystal displays, organic light-emitting displays, etc.

[0220] A "speech synthesis program" refers to a software program or program module that runs on a terminal or server and is used to convert text data into playable audio data.

[0221] "Audio data" refers to data generated by a speech synthesis program that represents speech waveforms and can be played through a sound output device, including compressed or uncompressed digital audio data.

[0222] "Sound output device" refers to a hardware device used to convert audio data into sound signals that can be perceived by the human ear, including speakers, headphones, etc.

[0223] "Historical data" refers to suggestion data, question data, and related information stored on the terminal or server that are related to the user's past interactions. This data is used to support the updating of prompts and the maintenance of context in subsequent rounds of dialogue.

[0224] "Additional question data" refers to new natural language question data that users continue to input based on previous questions and suggestions during multiple rounds of interaction with the system.

[0225] The embodiments of this invention will be described in conjunction with the hardware structure, software modules, data structure, and algorithm flow. The following descriptions are merely exemplary embodiments, and those skilled in the art can make various modifications and substitutions without departing from the spirit of this invention.

[0226] In the descriptions of each implementation, the server, terminal, and user are described as acting entities to highlight the distributed computing characteristics of the system in a network environment.

[0227] I. Overall System Structure The server is equipped with a processor, main memory, external storage devices, and a network interface. The server processor can be a multi-core general-purpose processor, the main memory can be semiconductor memory, and the external storage devices can be relational database servers or distributed storage systems. The server runs multiple software modules on the operating system, including: a network communication module, a natural language processing module, a financial data analysis module, a prompt generation module, a generative artificial intelligence model invocation module, and a suggestion generation module, etc.

[0228] The terminal is equipped with a processor, display device, touch input device, sound output device, local storage device, and wireless communication module. The terminal runs a financial management application that accepts user input through a graphical user interface, exchanges data with a server via a network, and performs text display and speech synthesis output locally.

[0229] Users input questions in natural language through the terminal and authorize the terminal and server to access their financial data. Users do not directly operate the modules on the server, but participate in the system's closed loop through their natural language input and viewing of output results.

[0230] II. Data Structure and Storage Method The server maintains multiple logical data tables or datasets in external storage. The server maintains the following data structure for each user: 1. User basic information record: The server records metadata such as user identifier, authentication information and preference settings for each user.

[0231] 2. Transaction Record Data Structure: The server creates a record for each expenditure, including at least the following fields: record identifier, user identifier, timestamp, amount, expense category code, payment method code, counterparty identifier, and remarks text. The server supports fast retrieval by user and time range through an index structure.

[0232] 3. Revenue record data structure: The server creates a record for each revenue transaction, which includes at least the timestamp, amount, revenue category code, and source identifier.

[0233] 4. Budget Information Data Structure: The server records fields such as budget period, budget category, budget amount, and current consumption amount for each budget item.

[0234] 5. Historical Interaction Data Structure: The server records a summary of the user's question text, the server-generated prompts, the output text of the generative artificial intelligence model, and the final suggestion data for each round of dialogue, which is used for contextual reference in subsequent rounds of dialogue.

[0235] By organizing the data into a matrix or key-value pair structure, the server enables the financial data analysis module to efficiently construct vector or matrix data in memory for aggregation and statistical operations. This structured design allows the server to quickly aggregate multi-source data at the same user level, thus reducing data preparation time.

[0236] III. Natural Language Processing and Feature Extraction The server uses a natural language processing module to perform fine-grained parsing of user questions. In this module, the server uses pre-trained language models or a combination of rule-based and statistical methods to establish a mapping from natural language text to structured features. After converting the user-input question text into an internally unified encoding format, the server performs the following technical processing: 1. The server performs word segmentation and part-of-speech tagging on the text, splitting the sentence into a sequence of lexical units and tagging the part of speech for each lexical unit.

[0237] 2. The server performs dependency parsing to construct a directed graph structure representing the syntactic relations between lexical units, and identifies the relationship between predicates and arguments through this graph.

[0238] 3. The server classifies the entire sentence using an intent classification model. This intent classification model can be a shallow neural network or a classification head based on a pre-trained language model, outputting one of the following categories: "saving suggestion," "concept interpretation," or "risk warning."

[0239] 4. The server identifies object information and time information through entity recognition and rule matching. For example, it extracts key information such as "expenditure", "this month" and "daily necessities" from the question to form a structured feature vector.

[0240] Through the aforementioned natural language processing, the server transforms natural language text that cannot be directly associated with financial data into structured features that can directly correspond to database fields, thereby establishing a semantic-to-field binding relationship within the computer. This binding allows the server to perform matching based on the same feature space when subsequently constructing prompts and analyzing financial data, thus reducing the number of manual rules and improving reasoning speed and accuracy.

[0241] IV. Financial Data Analysis and Aggregation Algorithms The server aggregates transaction and income data in the financial data analysis module. The server constructs a two-dimensional data array or table structure in memory, mapping each record to a row, with columns including numerical or discrete features such as amount, timestamp, and category code. The server generates user financial status data using the following methods: 1. The server filters records based on time information, forming a subset of records within a predetermined period.

[0242] 2. The server groups expenditure records according to expense category codes and performs a summation operation on each group to obtain the expenditure amount for that expense category. The server calculates the expenditure ratio for each category, which is obtained by dividing the amount for that category by the total expenditure amount to obtain a normalized weight value.

[0243] 3. The server performs concatenation and calculation with the above expenditure aggregation results to generate derived features such as "budget amount", "expenditured amount", and "remaining budget percentage" for each expense category.

[0244] 4. The server compresses the above results into an intermediate structure—the user's financial status vector, which contains numerical features such as the proportion of each type of expenditure, budget deviation, and total income and expenditure difference.

[0245] By using this data structure that represents the user's financial status in vector form, the server significantly reduces the number of fields that need to be repeatedly queried from the database when generating subsequent prompts. The server only needs to process the compressed status vector in memory to quickly filter out the features that have the most influence on financial advice, thereby reducing storage access latency and improving overall processing speed.

[0246] V. Construction of Prompt Statements and Invocation of Generative Artificial Intelligence Models In the prompt generation module, the server combines the structured features output by the natural language processing module with the user's financial state vector generated by the financial data analysis module to generate prompts for invoking the generative artificial intelligence model. Compared to using only the user's original question as input, this implementation technically constructs composite text containing various semantic and numerical information, thereby imposing more granular constraints on the model on the input side.

[0247] In one implementation, the server maps the numerical features in the financial state vector into natural language descriptions. For example, the server encodes "This month's food and beverage expenditure accounts for 40%" into a text fragment such as "The user's total expenditure this month is 8000 yuan, of which food and beverage expenditure is 3200 yuan, accounting for 40%", and combines it with the user's question text.

[0248] The server-generated prompt statement is shown below: "You are a professional personal financial advisor. Based on the following user spending data and user questions, provide specific and actionable saving suggestions and answer them in Simplified Chinese."

[0249] User's total expenditure this month: 8000 yuan Of this, food and beverage expenses amounted to 3200 yuan, accounting for 40%. Transportation expenses: 800 yuan, accounting for 10% Entertainment expenses: 1200 yuan, accounting for 15% User question: 'How can I reduce my expenses this month?' Please analyze the user's spending structure, identify the main spending categories that can be reduced, and provide 2 examples. Three specific suggestions. When the server needs to explain financial concepts, it generates another type of prompt, such as: "You are a financial literacy instructor. Please explain the following financial concepts to non-professional users in simple Chinese and give a simple example from everyday life."

[0250] Financial concept: Compound interest User background: Just starting to learn about personal finance, not good at math. Please keep it under 300 words. In this way, the server embeds structured numerical data and natural language instructions into the prompts, enabling the generative AI model to simultaneously reference both the user's specific data and the task description during inference. This "structured-text integration" input method, compared to simply transmitting the original question, can form a more concentrated and separable feature representation in the model's internal vector space, thereby improving the relevance of the output, reducing irrelevant information, and lowering errors.

[0251] VI. Structure and Training of Generative Artificial Intelligence Models In the generative AI model invocation module, the server interacts with the generative AI model deployed on cloud computing resources via a network interface. The generative AI model used by the server can be a large-scale language model based on the Transformer architecture. Specifically: 1. The model invoked by the server consists of a multi-layer self-attention encoder-decoder structure. Each layer includes a multi-head self-attention sublayer and a feedforward fully connected sublayer. The prompts provided by the server are first segmented and mapped into word vectors, and then propagated forward through the multi-layer network.

[0252] 2. During the model training phase, the server adopts an autoregressive language to model the objective, uses the cross-entropy loss function to measure the difference between the predicted sequence and the true sequence, and updates the weight matrix through gradient descent algorithm and parameter update rules (such as Adam optimization algorithm).

[0253] 3. During training, the server performs data augmentation on the input text, such as synonym replacement and sentence structure transformation, to improve the model's robustness to different expressions. The server adjusts hyperparameters such as learning rate, batch size, and training epochs to achieve a balance between generalization ability and fitting ability.

[0254] 4. The server sets temperature parameters and top parameters during inference. k or top A p-sampling strategy is used to control the diversity and stability of the output text. The server explicitly specifies the output language, style, and length in the prompt statement, further constraining the model's search space.

[0255] Through the aforementioned structure and training methods, the generative AI model invoked by the server can technically and effectively encode the numerical and natural language information in the prompts into a high-dimensional vector space, achieving conditional generation. By carefully crafting the prompts, the server enables the model to focus on keywords such as expense categories and budget deviations within its internal attention mechanism, thereby reducing attention allocation to irrelevant words and improving the model's effective computational utilization in financial advice tasks.

[0256] VII. Suggestions for Generation and Terminal Presentation After receiving the response data from the generative artificial intelligence model, the server performs post-processing. The server performs structural checks on the output text, such as dividing it into suggestion items based on periods and line breaks, checking for empty or duplicate content, and adding disclaimers if necessary. The server stores the processed text as suggestion data in the historical interaction data structure and sends it to the terminal via the communication module.

[0257] After receiving the suggestion data, the terminal loads it into local memory and presents it as multiple text segments on the display device. The terminal can emphasize the numerical parts of the suggestion, for example, by using different colors or font sizes, to improve the user's speed of recognizing key information.

[0258] When the user enables the voice playback option, the terminal invokes a speech synthesis program to convert the suggested data text into audio data. The terminal then converts the digital audio data into analog electrical signals via its internal sound card interface and outputs them through a speaker or headphones, allowing the user to receive the suggestions through hearing. The terminal caches the generated audio data locally so that it does not need to be re-synthesized when the user plays it again, thereby reducing communication and computational load.

[0259] VIII. Multi-round dialogue and utilization of historical data In multi-turn dialogue scenarios, the server combines historical interaction data with new question data to generate prompts. The server extracts summaries of recent dialogues from historical data, compresses them into short text blocks, and incorporates them into new prompts, such as: "In the previous dialogue, you suggested that the user reduce the number of times they eat out to lower food expenses." In this way, the server provides the model with the necessary context without transmitting the entire history to it, thus maintaining dialogue coherence under conditions of limited memory and network bandwidth.

[0260] Because the server filters and summarizes historical data instead of simply concatenating it, the length of prompts is controlled, the number of input tokens to the model is reduced, and inference time is shortened. Through this technique, the present invention not only improves the coherence of the dialogue but also brings substantial improvements in computational efficiency and communication load.

[0261] IX. Technical Effects and Improvement Mechanisms The server achieves high-quality input construction for generative artificial intelligence models by uniformly embedding natural language parsing results, financial data analysis results, and historical interaction data into prompt statements in a structured manner. Compared with traditional systems that rely solely on keyword retrieval or templated output, this system achieves improvements at the computer technology level in the following aspects: 1. Improved processing speed: The server first performs aggregation operations and generates compressed financial status vectors at the database layer, avoiding the need to interpret the original detailed data line by line during the model call phase, thereby reducing the amount of data transmission and computation.

[0262] 2. Improved output accuracy: The server embeds quantitative features such as expense category expenditure ratio and budget deviation into the prompt statements, making it easier for the generative artificial intelligence model to capture important financial features when allocating internal attention, reducing interference from irrelevant information, and thus reducing the probability of incorrect suggestions.

[0263] 3. Improved data management: The server manages transaction records, budget information and historical interaction data with a unified data structure, making data flow between multiple modules traceable and reusable, and facilitating the future expansion of other analysis modules or model calling strategies.

[0264] 4. Communication load reduction: The server only transmits the aggregated key features in the prompt statement, rather than the complete history, which significantly reduces the amount of data transmission between the terminal and the server, and between the server and the model service.

[0265] 5. Error control and robustness improvement: During the training and inference phases, the server uses data augmentation, sensitive word filtering, and structure detection algorithms on the prompt statements and output text to reduce the risk of abnormal output by the model and improve the overall stability of the system.

[0266] In this system, users only need to ask questions in natural language, without needing to understand complex operational instructions or financial formulas. Internally, the server drives the generative AI model using a non-traditional "structured-text hybrid" prompting method. This is not simply automating the scripts of a human advisor, but rather technically constructing a new model input encoding method, thereby constraining the model's internal attention distribution and vector representation. It is this new input construction and data flow design that enables this invention to demonstrate superior technical performance compared to traditional systems in terms of processing accuracy, speed, and resource utilization.

[0267] use Figure 12 The processing procedure is explained.

[0268] Step 1: Users launch the financial management application on the terminal and enter a question in natural language.

[0269] Input: The user's natural language question text.

[0270] Output: A data packet generated internally by the terminal, containing the problem text and user identification information.

[0271] Users type questions like "How can I reduce my expenses this month?" into input boxes via the terminal's touchscreen. The terminal combines this text with locally stored user IDs and authentication tokens to construct structured data (containing fields such as "user_id", "question", and "token"), ready to send it to the server.

[0272] Step 2: The terminal sends a data packet containing the problem data to the server.

[0273] Input: Data packet generated internally by the terminal (problem text + user identification information).

[0274] Output: The request message sent to the server via the communication network.

[0275] The terminal invokes the network communication module to encapsulate data packets into a network request, which is then transmitted via a wireless communication module (such as a cellular network or Wi-Fi). Fi) Sends the data to the specified interface address of the server via the HTTPS protocol. The terminal records the time of this transmission and waits for the server's response.

[0276] Step 3: The server receives and parses requests from the terminal.

[0277] Input: A network request message sent by the terminal (containing the question text and user identification information).

[0278] Output: The question text and user identification information structure stored in the server's memory.

[0279] After the server receives the request through the network interface, the web service program parses the message body, deserializes the JSON or similar structure into an internal object, extracts fields such as "user_id", "question", and "token", and builds the corresponding data structure in memory for subsequent processing.

[0280] Step 4: The server verifies the user's identity and decides whether to continue processing.

[0281] Input: User identification information (user_id) and authentication token (token).

[0282] Output: Authentication result flag and the user ID of the authenticated user.

[0283] The server invokes the authentication module to perform signature verification and validity checks on the token, and compares the user information carried in the token with the user_id in the request. If authentication is successful, the server marks the request as "valid" in memory and passes the authenticated user identifier to subsequent modules; if it fails, the server generates an error response and terminates subsequent processing.

[0284] Step 5: The server retrieves the user's transaction and revenue records from external storage devices.

[0285] Input: The authenticated user ID and the scheduled time range (e.g., the current month).

[0286] Output: A set of transaction records and a set of revenue records loaded into memory.

[0287] The server constructs query conditions based on the user identifier, sends a query request to the database, and retrieves all expenditure and income details within a predetermined time range. The server receives the result set returned by the database, maps each record to a data row or object in memory, and forms a list of transaction records and a list of income records for use by the financial analysis module.

[0288] Step 6: The server performs statistical and summary processing on the acquired financial records to generate user financial status data.

[0289] Input: A set of transaction records and a set of revenue records.

[0290] Output: User financial status data including expense amounts by expense category, expense ratio by expense category, total revenue, total expenditure, and budget deviation information.

[0291] The server groups transaction records by expense category field, sums the amounts in each group to obtain the expenditure amount for each expense category. The server then calculates the total expenditure by summing the amounts of all expenditure records and divides the expenditure amount of each category by the total expenditure to obtain the expense category expenditure ratio. The server sums the income records to obtain the total income. The server further reads the budget for each category from the budget table, calculates the difference between the budget and the actual expenditure, and obtains the budget deviation. All these numerical results are organized into a structured object: the user's financial status data.

[0292] Step 7: The server performs natural language processing on the user's question, extracting intent information, object information, and time information.

[0293] Input: User's question text.

[0294] Output: Structured semantic features, including intent information, object information, time information, and question type.

[0295] The server inputs the question text into a natural language processing program for word segmentation, part-of-speech tagging, and syntactic analysis. Then, it uses a classification model to determine whether the question belongs to "saving advice," "concept interpretation," or other types. The server also applies entity recognition and pattern matching algorithms to extract information from the text that represents the user's goal (such as "expenditure") and time range (such as "this month"), and encodes it into structured feature objects, which serve as input for generating subsequent prompt statements.

[0296] Step 8: The server combines semantic features with financial status data to generate the intermediate descriptions needed for the prompt statements.

[0297] Input: Semantic feature objects (intent information, object information, time information, question type) and user financial status data.

[0298] Output: Intermediate text description and data summary representing the user's income and expenditure structure and the background of the problem.

[0299] The server selects the corresponding template logic based on the type of question. For example, for the "saving suggestions" type, the server selects the expense category with the highest expenditure percentage from the financial status data and generates a text fragment such as "The user's total expenditure this month is X yuan, of which food and beverage expenditure is Y yuan, accounting for Z%". The server also combines time information to limit the analysis to the corresponding period. By formatting each numerical feature into natural language phrases, the server obtains a set of intermediate text descriptions and a summary of relevant data points.

[0300] Step 9: The server constructs prompts for input into the generative artificial intelligence model.

[0301] Input: Intermediate text description, data summary, and original user question text.

[0302] Output: The complete prompt text.

[0303] The server combines role settings, task descriptions, user financial data descriptions, user question text, and constraints (such as response language and length limits) into a continuous text according to a pre-defined prompt structure. The server inserts this intermediate text description into the designated position of the prompt, allowing the generative AI model to simultaneously acquire quantitative data and task context during inference. The final generated prompt is a plain text string, which serves as direct input for subsequent model calls.

[0304] Step 10: The server will input the prompts into the generative artificial intelligence model and obtain the response data.

[0305] Input: The complete prompt text.

[0306] Output: The response text output by the generative artificial intelligence model.

[0307] The server calls a generative AI model through a model interface, taking the prompts as the input sequence. Internally, the model encodes the prompts through word vector mapping, multi-layer self-attention computation, and feedforward network operations, gradually generating the output word sequence. The server sets parameters such as temperature and maximum length during the call to control the generation process. After the model completes inference, it returns the generated natural language response text to the server, which receives it as response data and stores it in memory.

[0308] Step 11: The server performs post-processing on the response data to generate suggestion data.

[0309] Input: The response text output by the generative artificial intelligence model.

[0310] Output: Structured suggested data (including text content and optional category labels and summary information).

[0311] The server segments the response text into sentences and paragraphs, separating specific suggestions. Based on keywords or patterns, the server can tag the suggestions with labels such as "saving measures" and "risk warnings," and add fixed disclaimer text when necessary. The server encapsulates the processed text and its metadata into suggestion data objects, which are then sent to the terminal and saved to the history.

[0312] Step 12: The server sends suggested data to the terminal and updates historical interaction data.

[0313] Input: Suggested data object and current session identifier.

[0314] Output: The response message sent to the terminal and the updated history of interactions.

[0315] The server serializes the suggested data into a response message via the communication module and sends it to the requesting terminal through the network interface. Simultaneously, the server creates a new record in the historical interaction data structure, storing the user's questions, prompts, and suggested data from this round of dialogue for reference in subsequent rounds.

[0316] Step 13: The terminal receives and parses the suggestion data from the server.

[0317] Input: The response message sent by the server (containing suggested data).

[0318] Output: Suggested text and related metadata structure within the terminal.

[0319] The terminal receives the server's response via the network communication module, deserializes the message body, extracts the suggested text and tags, and stores them in local memory. The terminal saves the current suggestion record in a local database or cache so that users can view historical suggestions.

[0320] Step 14: The terminal displays suggested text on the display device.

[0321] Input: Suggested text and segmentation information stored internally in the terminal.

[0322] Output: Visual suggestions displayed on the screen.

[0323] The terminal invokes a graphical interface component to lay out the suggested text on the display screen according to paragraph and item formats, highlighting or bolding key numbers or keywords to improve readability. The terminal automatically adjusts the font size and line spacing based on the screen size, allowing users to browse the main suggested content in a single interface.

[0324] Step 15: The terminal performs speech synthesis based on user settings and outputs speech suggestions.

[0325] Input: Suggested text and voice broadcast settings.

[0326] Output: Audio signal played through a sound output device.

[0327] When the terminal detects that the user has enabled the voice broadcast function, it sends the suggested text to the speech synthesis program, which converts the text into an audio data file. The terminal then calls the audio playback module to stream the audio data to the sound output device, where it is played through a speaker or headphones, allowing the user to receive the same suggested content audibly.

[0328] Step 16: Users interact further based on the suggestions, triggering a new processing loop.

[0329] Input: Suggested content to be displayed or played to the user.

[0330] Output: A new round of natural language questions from the user based on the suggested content.

[0331] After reading or listening to the advice, users can enter a new question in the terminal, such as "If I reduce eating out, how much more money can I save in a month?" The terminal binds the new question to the existing session context and sends it to the server again, thereby triggering a new round of processing starting from step 1, forming a multi-round continuous dialogue financial consultation process.

[0332] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0333] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0334] With the widespread application of natural language processing technology and generative artificial intelligence models, traditional information provision systems, while capable of generating answers using general language models, still suffer from several shortcomings at the computer technology level: First, these systems typically input the user's original question directly into the generative AI model, lacking a structured parsing process based on intent recognition and keyword extraction. This results in high noise levels in the model's input information, low efficiency in utilizing computational resources, and increased inference latency. Second, the systems do not utilize backend databases or knowledge bases for precise retrieval of relevant information. The generative model can only rely on the knowledge implicit in its own parameters for reasoning, making it difficult to reflect external data updates in a timely manner, leading to insufficient accuracy and controllability of the generated results. Third, the systems often use fixed templates for constructing prompts, failing to dynamically generate multi-level prompts based on the parsing results of the user's query and the retrieved reference information, thus limiting the generative AI model's... Fourth, in the post-processing stage of model output, the system typically only performs simple formatting, lacking a systematic process for deleting useless information and adjusting expression methods for explanatory information. This results in redundant content and inconsistent style in the generated text, increasing the user's understanding burden. Fifth, existing systems rarely utilize the user's real-time voice and image signals for emotional state estimation during interaction. They cannot dynamically adjust the prompt statement conditions and model output strategies based on emotional state at the computer system level, making it difficult to achieve adaptive human-computer interaction while ensuring system stability. Sixth, traditional systems mostly process user-related numerical and historical information at the simple statistical level, lacking a structured generation mechanism for risk assessment results input to generative artificial intelligence models. This makes it difficult to construct high-quality risk warning prompt statements, limiting the effective performance of the model in risk warning scenarios.

[0335] Therefore, it is necessary to propose a system that optimizes the computer architecture and program flow as a whole. This system can improve the accuracy, relevance and interactivity of the generated results under the same hardware resource conditions, reduce invalid reasoning calculations, and improve the technical performance of the entire human-computer question-answering process.

[0336] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0337] In this invention, the server includes: means for controlling an input / output interface to receive natural language query information from a user; means for performing string validity verification and length limitation processing on the query information and converting it into a predetermined data structure to generate a communication request message; means for performing morphological analysis, intent classification, and keyword extraction processing on the query information to generate parsing result data containing query intent information and related word information; means for retrieving information resources in an information storage device based on the parsing result data, obtaining multiple reference information according to relevance evaluation, and shaping the reference information into a predetermined text format; means for constructing a prompt statement for input to a generative artificial intelligence model based on the query information and the reference information, and encoding the prompt statement into labeled sequence data; and means for inputting the labeled sequence data into the generative artificial intelligence model to generate a response through numerical computation. The server includes a means for performing useless information deletion and expression adjustment processing on the response text to generate explanatory information, and a means for presenting the explanatory information to the user through the input / output interface; the server may also include a means for acquiring the user's voice and image information when presenting the explanatory information, performing emotion estimation processing on them to extract emotional state information, and adjusting the conditions of the prompt statement and the expression form output by the generative artificial intelligence model according to the emotional state information to dynamically adjust the tone or content of the explanatory information; and a means for parsing numerical and historical information related to the user to generate risk assessment information, generating a prompt statement containing warning content of a specific risk event based on the risk assessment information and the parsing result data, and inputting the prompt statement into the generative artificial intelligence model to generate explanatory information containing the warning content. This allows for the simplification and enrichment of model input through structured parsing and relevant information retrieval before the generative AI model is invoked. During the model inference stage, highly relevant prompts can improve the utilization of computing resources and the quality of output. After the model outputs, post-processing based on explanatory information and dynamic control mechanisms based on emotional state and risk assessment can comprehensively improve the computer technology performance of the question-answering system in terms of response accuracy, controllability of generated content, human-computer interaction adaptability, and computational efficiency.

[0338] A "system" refers to an overall technical device consisting of one or more information processing devices, storage devices, and communication devices, which coordinates the execution of user query processing, data retrieval, and generative artificial intelligence model invocation through program instructions.

[0339] "User" refers to a human user or external terminal that submits natural language queries to the system through input / output interfaces and receives explanatory information in return from the system.

[0340] "Input / output interface" refers to the hardware or software interface used to exchange data between the system and the user or external terminal, including graphical user interface, programming interface, network interface, etc.

[0341] "Natural language query information" refers to text data or speech-transcribed data that is entered by the user in natural language and used to request explanations, suggestions or other information from the system.

[0342] "String validity verification processing" refers to the inspection and processing performed on natural language query information to confirm whether the character encoding, character set and format meet the predetermined rules and to remove illegal characters or control characters.

[0343] "Length limit processing" refers to the process of detecting the length of natural language query information and performing truncation or prompting when it exceeds a predetermined limit, in order to prevent excessively long input from causing abnormal consumption of system resources.

[0344] "Predefined data structure" refers to a predefined data organization format for transmission or storage within a system, such as data records or message objects containing field names and field values.

[0345] A "communication request message" refers to a request message or data unit generated based on natural language query information and used for transmission through a communication channel within the system or between the system and external devices.

[0346] "Information processing device" refers to an electronic computing device that executes program instructions to complete data reception, parsing, calculation and transmission, including servers, terminal devices or other computing nodes.

[0347] A "response message" refers to a message or data unit generated and returned by an information processing device in response to a communication request message, which at least includes explanatory information or result information in response to natural language queries.

[0348] "Morphological analysis processing" refers to the text preprocessing process that decomposes natural language query information into words or subwords and optionally performs operations such as part-of-speech tagging and syntactic segmentation.

[0349] "Intent classification processing" refers to the process of using statistical methods or machine learning models to classify natural language query information, in order to determine the purpose or task type of the user's query.

[0350] "Keyword extraction processing" refers to the process of identifying and selecting words or phrases that can represent the main semantic content from natural language query information.

[0351] "Inquiry intent information" refers to data obtained through intent classification processing that represents the category of the user's inquiry purpose or task.

[0352] "Related word information" refers to a set of words or phrases that represent the core semantic elements in natural language query information, obtained through keyword extraction.

[0353] "Analysis result data" refers to the structured data set obtained after performing morphological analysis, intent classification, and keyword extraction on natural language query information.

[0354] "Information storage device" refers to a data storage component or system used to permanently preserve information resources, including database systems, file storage systems, or knowledge bases.

[0355] "Information resources" refer to various types of data content stored in information storage devices that can be retrieved and used by the system, including text data, structured records, or knowledge entries.

[0356] "Relevance assessment" refers to the process of calculating the degree of relevance between information resources and natural language query information based on the parsed data.

[0357] "Reference information" refers to text or data fragments selected from information resources based on relevance assessment results to support the generation of explanatory information.

[0358] "Preset text formatting" refers to the text processing process that adjusts reference information to meet preset rules such as length, structure, punctuation, and paragraphs.

[0359] "Generative artificial intelligence models" refer to automatic generation models that use neural networks or other machine learning structures to generate natural language text by performing numerical operations on input labeled sequences.

[0360] "Prompt statements" refer to input text constructed to guide generative artificial intelligence models to produce the desired output. They typically include role settings, task descriptions, reference information, and answer requirements.

[0361] "Tagged sequence data" refers to a sequence of discrete token identifiers obtained by an encoder or word segmenter from the conversion of prompt statements.

[0362] "Numerical computation processing" refers to the forward inference computation process performed internally by generative artificial intelligence models based on vector and matrix operations, in order to generate response text from labeled sequence data.

[0363] "Response text" refers to the natural language text output by a generative artificial intelligence model based on prompts.

[0364] "Useless information removal processing" refers to the process of filtering and editing response text to remove content that is irrelevant to or repetitive with the user's query.

[0365] "Expression style adjustment processing" refers to the process of modifying the language style, word order, or sentence structure of the response text without changing the main semantics.

[0366] "Explanatory information" refers to the textual information presented to the user as the final answer after the removal of useless information and the adjustment of expression.

[0367] “Voice information” refers to audio data or its characteristic representation data that are obtained from the user during the presentation of explanatory information and reflect the user’s voice characteristics.

[0368] "Image information" refers to image data or video frame data that reflects the user's facial or posture features and is obtained from the user during the presentation of explanatory information.

[0369] "Emotion estimation processing" refers to the process of analyzing voice and image information to infer the user's emotional state or attitude.

[0370] "Emotional state information" refers to structured data that represents the category or intensity of a user's emotions, obtained through emotion estimation processing.

[0371] "Prompt statement conditions" refer to the parameters or constraints used to control the behavior of generative artificial intelligence models when constructing prompt statements, including tone requirements, length limits, or content emphasis.

[0372] "Numerical information" refers to quantitative data that is relevant to users and can be used to analyze risk, including amount, frequency, proportion or other measurable data.

[0373] "Historical information" refers to behavioral data, status data, or interaction records that are related to users and recorded in chronological order.

[0374] "Risk assessment information" refers to structured data that characterizes the degree of risk of a user or event, obtained through analysis of numerical and historical information.

[0375] "Specific risk event warning content" refers to explanatory text or information generated based on risk assessment information to indicate that a certain type of risk event may occur or has already occurred.

[0376] In the following description, the embodiments of the present invention will be described using the terms "server," "terminal," and "user," respectively. This description focuses on the information processing system structure described in Appendix 1 to 3, emphasizing the internal data structure design, processing flow, algorithm structure, and the resulting technical effects of the computer, without listing the order of processing steps.

[0377] In one implementation, a server serves as the core information processing device, deployed in a data center or cloud computing platform. Its hardware includes a multi-core central processing unit, a graphics processing unit, high-speed memory, solid-state storage, and network interface cards. The server can utilize general-purpose server hardware platforms, such as rack-mount servers based on multi-core processors and supporting GPU acceleration. At the software level, the server runs an operating system (e.g., a general-purpose server operating system) and deploys web server middleware (e.g., general-purpose HTTP server software or reverse proxy software), application service frameworks (e.g., Python-based web frameworks, Java-based web frameworks, or JavaScript-based web frameworks), database management systems (e.g., general-purpose relational databases or document databases), and deep learning inference frameworks (e.g., PyTorch, TensorFlow, or ONNX Runtime).

[0378] In one embodiment, the terminal includes a smartphone, tablet, or personal computing device, whose hardware includes a central processing unit, a graphics processing unit, a touchscreen display, a microphone, a camera, a memory, and a wireless communication module. At the software level, the terminal runs a general-purpose mobile or desktop operating system and installs a browser or dedicated application to form an input / output interface, sending natural language queries to the server and receiving explanatory information returned by the server. In this invention, the terminal not only performs user interface presentation functions but also performs basic validity checks and encoding processing on the input data, thereby filtering illegal characters and abnormally long inputs before communication, reducing the overhead of exception handling on the server side.

[0379] In one implementation, users input natural language queries through a text input interface or speech-to-text interface on a terminal. Users can input explanatory questions related to finance, risk, indicators, etc., or queries related to other technical fields. This invention does not limit the knowledge domain of the question; any query that can be expressed in natural language can be processed by this system.

[0380] After receiving a natural language query from the terminal, the server uses its natural language processing module to perform a series of data processing and operations on the query. In one implementation, the server uses a word segmentation and morphological analysis library (such as a general Chinese word segmentation library or a general natural language processing toolkit) to perform word segmentation and morphological analysis on the text, obtaining a structured representation composed of lexical units, part-of-speech tags, and syntactic boundaries. The server further calls an intent classification sub-model based on a pre-trained language model. This sub-model can employ a bidirectional encoding model based on a Transformer structure, mapping the input text into high-dimensional vectors and applying fully connected layers and Softmax to the vectors to obtain the probability distribution of multiple intent labels. In its implementation, the server selects the one with the highest probability as the query intent information, such as "definition explanation," "calculation method description," or "risk warning."

[0381] The server can employ various algorithms in keyword extraction. For example, in one implementation, the server utilizes TensorFlow. The IDF algorithm calculates the importance of text terms and selects words with higher weights as relevant word information. In another implementation, the server uses a keyword extraction method based on semantic vector space, such as using encoders like Sentence-BERT to map terms or phrases into vectors and calculate the cosine similarity with the entire sentence vector, thereby selecting representative keywords. The server then generates parsed result data containing query intent information and relevant word information. This data is organized using a key-value structure, including fields such as intent label, keyword list, and original text. Through such structured parsing, subsequent retrieval and suggestion construction can use structured data as input, thereby narrowing the search scope and reducing meaningless computation.

[0382] During the information retrieval phase, the server accesses the information storage device. In one implementation, the information storage device employs a relational database. The server uses keywords from the parsed result data to construct structured query statements and performs retrieval in tables containing concept entries, descriptive text, and tag fields. The server can enable full-text indexing at the database level or connect to an external search engine system to support keyword-based inverted index retrieval. In another implementation, the server uses a document database or vector database to pre-store the semantic vectors of reference texts and uses nearest neighbor search algorithms (e.g., approximate nearest neighbor search based on inverted document indexes or approximate nearest neighbor search based on graph structures) to obtain reference information documents semantically similar to the question. By calculating and sorting the query results based on relevance scores, the server selects several highly relevant reference information entries and performs predetermined text formatting on them, including removing invalid tags, truncating excessively long content, and standardizing punctuation and paragraph structure, thereby obtaining a set of reference information text fragments that can be directly used to construct prompt statements.

[0383] When constructing prompt statements, the server combines the query intent information, relevant word information, and the aforementioned reference information from the parsed result data into one or more segments of structured natural language instruction text. In a typical implementation, the server defines the role and task of the generative artificial intelligence model, summarizes the reference information, and specifies the format requirements for the generated results. For example, the server can generate the following prompt statement text: You are a senior financial analyst. Please answer users' questions in Simplified Chinese.

[0384] References 1. Free cash flow is the cash flow remaining for a company after deducting capital expenditures necessary to maintain normal operations.

[0385] 2. Free cash flow can be used to pay off debt, pay dividends, or make new investments.

[0386] 3. One of the commonly used formulas for free cash flow is: Free cash flow = Net cash flow from operating activities - Capital expenditures.

[0387] User Issues What is free cash flow? Please explain in simple terms and give a numerical example.

[0388] Answer requirements - Use simple and easy-to-understand expressions.

[0389] - First, give a brief definition, then provide examples.

[0390] - Keep it under 300 characters.

[0391] In another implementation, the server can construct different prompt statements for different risk warning scenarios, for example: You are a risk management expert. Based on the following user transaction history and risk assessment results, please generate a risk warning statement for ordinary users.

[0392] Risk assessment information - The number of high-risk transactions has increased significantly in the past 30 days; - Large single transactions concentrated in late-night hours; Historically, similar patterns have been highly correlated with fraud incidents.

[0393] User Issues Have I made any mistakes with my recent spending? What risks might I be facing? Answer requirements - Clearly identify potential risk points; - Provide 2-3 specific risk control recommendations; - Maintain an objective and cautious tone.

[0394] After constructing the prompt statement, the server calls the word segmenter or encoder of the generative AI model to convert the prompt statement into a tokenized sequence data. The word segmentation process can employ algorithms based on sub-word units, such as Byte Pair Encoding, WordPiece, or SentencePiece. The server represents the encoded tokenized sequence as a tensor and loads it into the graphics processing unit or central processing unit via a deep learning inference framework. In one implementation, the generative AI model employs a Transformer decoder structure, containing multiple layers of self-attention sublayers and feedforward network sublayers. The parameters of each layer are optimized during pre-training by minimizing a language modeling loss function (e.g., cross-entropy loss). Model parameter learning can utilize autoregressive language modeling methods, pre-training on large-scale text corpora, iteratively updating network weights using gradient descent algorithms (e.g., the Adam optimizer), and improving training stability through learning rate decay, gradient clipping, and regularization. During the inference phase, the server no longer updates the model weights but instead performs forward propagation calculations with fixed parameters to predict the conditional probability distribution of the output tokens at each step, and then uses Topology to perform the calculations. k or Top The p-sampling strategy selects the next generated tag. Through this structure, the server fully utilizes the matrix operation capabilities of the graphics processing unit at the hardware level to achieve highly parallel attention computation, thereby significantly improving generation speed and throughput compared to rule-based serial text generation methods.

[0395] After receiving the labeled sequence output by the model, the server converts it into natural language response text using a decoder and performs unnecessary information removal and expression adjustment on the response text. The server can define a set of rules in its implementation, such as removing lengthy background descriptions unrelated to the user's question, removing potentially repetitive prompts from the model, and standardizing terminology. The server can also incorporate a small discriminative sub-model to score text fragments and remove low-relevance sentences. Through this post-processing mechanism, the server significantly reduces the redundancy in the final description information, lowers the user's reading cost, and reduces the length of returned data at the network transmission level, thereby reducing communication load.

[0396] In the implementation described in Appendix 2, the server further integrates an emotion estimation module. While displaying explanatory information, the terminal captures the user's voice and image information and transmits them to the server via a secure communication channel. In this implementation, the server utilizes a voice feature extraction algorithm (e.g., a feature extraction algorithm based on Mel-frequency cepstral coefficients or an acoustic feature extraction network based on convolutional neural networks) and a facial expression recognition network (e.g., an emotion recognition model based on convolutional neural networks and attention mechanisms) to extract high-dimensional feature vectors from the input signal and input them into the emotion classification network. The emotion classification network can employ a multilayer perceptron or a Transformer-based multimodal fusion structure, undergoing supervised training by minimizing cross-entropy loss, and outputting emotion state labels such as "neutral," "confused," "anxious," and "satisfied," along with their probabilities. Based on the obtained emotion state information, the server dynamically adjusts the conditions of subsequent prompts. For example, when the user exhibits confusion, the server adds constraints such as "Please use simpler language and provide more examples" to the newly constructed prompt; when the user exhibits anxiety, the server can control the balance between reassurance and risk warning in the response tone. This dynamic control based on emotional states is not a simple replacement for manual operation, but rather introduces a feedback loop at the computer system level to transform multimodal signals into structured control parameters, thereby changing the input conditions of the generative artificial intelligence model, achieving adaptive output style, and improving the overall technical performance and user interaction quality of the system.

[0397] In the implementation described in Appendix 3, the server performs risk assessments on user-related numerical and historical information. The server can read time-series data from transaction record databases, interaction log databases, or behavioral analysis systems, and analyze it using statistical feature extraction and machine learning models. For example, the server can calculate feature vectors such as the frequency, amount distribution, and time distribution of high-risk operations over a recent period, and input these features into a supervised learning model (e.g., a gradient boosting decision tree model or a deep learning-based sequence model). During the training phase, this model learns the mapping from feature vectors to risk level labels by minimizing binary or multi-class classification losses (e.g., log loss), thereby enabling rapid risk assessment of the current user state during the inference phase. The server encodes the output risk assessment information into structured data, including risk levels, triggered risk characteristics, and historical similar cases, and constructs prompts for generative artificial intelligence models based on this structured data. By explicitly indicating the source of risk, statistical basis, and expected output structure in the prompts, the server enables the generative model to prioritize key risk points in its responses, thereby improving the relevance and consistency of risk prompts while maintaining the readability of natural language. This mechanism differs from manually writing warning messages; instead, it characterizes, structures, and automatically generates risk assessment results at the algorithmic level, improving scalability and consistency when handling large-scale users.

[0398] In this invention, the terminal not only serves as a display device but also undertakes local preprocessing and security control functions. After receiving user input, the terminal performs character filtering, partial masking of sensitive content, and length limitation processing, converting the processing results into a predetermined data structure before sending them to the server. This design can pre-filter data with obvious formatting errors or abnormal lengths on the terminal side, reducing the server's parsing burden and minimizing unnecessary network round trips. In another embodiment, the terminal can cache recent queries and corresponding explanations. When a user repeatedly asks the same question within a short period, the terminal can directly return the results from the local cache, thereby reducing the frequency of calls to the server and generative artificial intelligence models, significantly saving computing and communication resources.

[0399] Users can utilize this system in various ways across different implementation scenarios. Users can continuously ask multiple rounds of questions via text through the terminal interface. When constructing prompts, the server embeds a summary of the previous round of dialogue as contextual information, thus achieving context-aware continuous question-and-answer. Users can also receive proactively pushed risk information from the server in risk-related scenarios. For example, when the server detects an abnormal pattern based on a risk assessment model, it automatically constructs prompts and calls a generative artificial intelligence model to generate explanatory text, which is then presented to the user through the terminal. In this process, the server uses a non-conventional process combining its internal rule set with model output to determine when to trigger risk warnings. These rules can include threshold rules, pattern matching rules, and model confidence threshold rules. Compared to traditional static rule systems, this system combines structured risk assessment information with the natural language generation capabilities of a generative artificial intelligence model, enabling detailed and personalized explanations for complex and ever-changing data patterns while maintaining interpretability.

[0400] As can be seen from the above implementation forms, this system, through a series of technical means such as introducing parsing result data structures, reference information retrieval and shaping, constructing prompt statements for generative AI models, tag sequence encoding, GPU-accelerated Transformer inference, and post-processing for explanatory information, as well as dynamic control driven by sentiment estimation and risk assessment, makes the input of generative AI models more refined and highly relevant, reducing the interference of noise information on model inference. Simultaneously, by adding multimodal sentiment feedback and structured risk assessment control at the technical level, a closed-loop optimization mechanism from input, retrieval, generation to output is established. This achieves comprehensive technical effects such as improved response accuracy, increased inference speed, reduced network load, and enhanced human-computer interaction under the same hardware resource conditions. These effects are not simply the automation of human writing work, but rather a substantial improvement in the overall performance of the computing system through improvements in the computer's internal data representation methods, algorithmic processes, and model control strategies.

[0401] use Figure 13 The processing procedure is explained.

[0402] Step 1: The user enters the query information on the terminal. Users can enter natural language queries through text input boxes or voice transcription interfaces on the terminal, such as "What is free cash flow?".

[0403] Input: Natural language query information (in text form or text obtained by speech transcription).

[0404] Output: The raw text data displayed in the terminal input box.

[0405] The terminal listens for user key or voice input events. When the user finishes inputting and clicks the "send" button, the terminal temporarily stores the current text content in memory, preparing for subsequent preprocessing.

[0406] Step 2: The terminal performs legality verification and length limitation on the query information. The terminal performs a string validity check on the raw text generated in step 1, including verifying whether the character encoding is in a predetermined encoding format, scanning for the presence of control characters or illegal characters, and removing characters that do not conform to the rules. The terminal also checks whether the text length exceeds a preset limit (e.g., maximum number of characters); if it does, it performs truncation or prompts the user on the interface to shorten the text.

[0407] Input: Raw natural language query text.

[0408] Output: Normalized text that passes the validity check and meets the length limit.

[0409] The terminal internally uses string processing functions to traverse the text, determine the category of each character in turn, generate a string containing only the allowed character set, and record the length information as the basis for subsequent data structure construction.

[0410] Step 3: The terminal encapsulates the normalized text into a communication request message and sends it to the server. The terminal, based on a predetermined data structure, organizes standardized text along with metadata such as user identifier, terminal type, and language settings into a request record, and serializes it into JSON or other message formats. The terminal then invokes the network library provided by the operating system to send the communication request message to the specified address on the server via an HTTPS channel.

[0411] Input: Standardized text and locally stored user identifiers and terminal environment information.

[0412] Output: The communication request message data packet sent to the server.

[0413] In specific operations, the terminal constructs data objects in the form of key-value pairs, calls network interface functions to write the HTTP request body, and attaches content type and authentication information to the request header. Then, it transmits the data packets to the server network interface via wireless or wired network.

[0414] Step 4: The server receives the communication request message and parses the message body. The server listens on the network port through a web server component, receives HTTPS requests from the terminal, and obtains the HTTP message content after TLS decryption. The server reads JSON or message body strings from the request body and calls a parsing library to convert them into internal data structures, extracting fields such as query information text, user identifier, and language settings.

[0415] Input: A communication request message data packet sent by the terminal.

[0416] Output: An internal request object containing the query text and related metadata.

[0417] During this process, the server performs a network buffer read operation, verifies the message format and the existence of required fields. If a missing field is detected, an error response is generated; otherwise, the parsing result is stored in the request processing queue for use by the subsequent natural language processing module.

[0418] Step 5: The server performs morphological analysis and basic preprocessing on the query information. The server retrieves the query information text from the internal request object, calls a natural language processing library (such as a Chinese word segmentation tool or morphological analyzer) to perform word segmentation, punctuation normalization, and stop word filtering on the text, and obtains the word sequence and its basic attributes.

[0419] Input: Original query information text.

[0420] Output: Preprocessed results consisting of lexical units, part-of-speech tags, and sentence boundaries.

[0421] In this step, the server first standardizes the text by unifying capitalization and converting full-width and half-width characters, then calls the word segmentation function to split the sentence into an array of words, and optionally adds part-of-speech tags so that subsequent intent recognition and keyword extraction can more accurately utilize part-of-speech information.

[0422] Step 6: Server execution intent classification and keyword extraction processing The server inputs the word sequence or original text obtained in step 5 into a pre-trained intent classification model. This model is based on the Transformer encoding structure, which encodes the text into a vector representation and outputs the probability distribution of different intent categories through a fully connected layer and a Softmax function. The server selects the category with the highest probability as the query intent information.

[0423] The server then calculates TF based on the word sequence. IDF weights, or semantic vector similarity methods, can be used to select several high-weight words or high-similarity phrases as related word information.

[0424] Input: Preprocessed word sequence and original text.

[0425] Output: Parsing results data consisting of query intent information and related word information.

[0426] In practice, the server calculates the word frequency for each lexical unit and combines it with a pre-calculated inverse document frequency table to calculate the TF. The IDF values ​​are sorted in descending order of weight. The top few words are selected to form a keyword list, which is then packaged together with the intent tags into a structured data object.

[0427] Step 7: The server retrieves information resources from the information storage device based on the parsed data. The server uses relevant word information from the parsed data to construct database query conditions, such as generating SQL query statements containing multiple keywords or calling full-text search interfaces. The server retrieves document entries that match these keywords from the information storage device and scores each record according to a predefined relevance calculation method (such as keyword matching degree, vector similarity).

[0428] Input: Parsing result data (intent labels, keyword list).

[0429] Output: A set of candidate information resource records sorted by relevance.

[0430] During execution, the server interacts with the database driver, embedding keywords into WHERE conditions or full-text search queries, receiving the record set returned by the database, calculating a comprehensive relevance score for each record, and generating a sorted list of results for subsequent filtering of reference information.

[0431] Step 8: The server selects reference information and performs pre-defined text formatting. The server selects several records with the highest relevance from the candidate records obtained in step 7, and performs length truncation, tag filtering, and paragraph rearrangement on the content fields of each record to make it a concise, coherent, and clearly structured text fragment.

[0432] Input: Candidate information resource records sorted by relevance.

[0433] Output: Several structured reference information texts that can be directly embedded with prompt statements.

[0434] In this step, the server scans the recorded content, removes HTML tags or special marks, truncates paragraphs exceeding the predetermined maximum length by sentence boundaries, and rearranges them in logical order to make the reference information more suitable for presentation in the prompts, reducing the processing burden on the generative artificial intelligence model for irrelevant content.

[0435] Step 9: The server constructs prompts for generative artificial intelligence models. The server combines the original query text, parsed result data, and formatted reference information into a multi-part natural language instruction text, i.e., a prompt statement. The server clearly defines the model's role, task, available background information, and response requirements in the prompt statement, such as word limits, tone requirements, and whether examples are needed.

[0436] Input: Query information text, parsing result data, and reference information text.

[0437] Output: A semantically clear prompt text that includes background and constraints.

[0438] When generating prompts, the server places reference information in the "References" section, user questions in the "User Questions" section, and answer requirements in the "Answer Requirements" section, thus forming a clearly structured instruction text that facilitates the use of context by generative artificial intelligence models.

[0439] Step 10: The server encodes the prompt statement into a token sequence data. The server invokes the word segmenter corresponding to the generative artificial intelligence model to perform sub-word-level segmentation or encoding on the prompt statement obtained in step 9, mapping the text into a sequence of tag IDs. The server then combines this sequence into a multi-dimensional tensor, containing information such as tag IDs and positional encodings, to conform to the model's input format.

[0440] Input: Prompt text.

[0441] Output: Labeled sequence data that can be processed by generative artificial intelligence models.

[0442] When the server actually executes the code, it calls the word segmenter interface to encode the prompt statement, obtains an array of integer sequences, and truncates or pads them according to the maximum sequence length specified by the model, storing the results in a tensor structure in memory.

[0443] Step 11: The server performs numerical calculations on a generative artificial intelligence model to generate response text. The server transmits the labeled sequence data to the graphics processing unit (GPU) or central processing unit (CPU), invoking a deep learning inference framework to perform forward propagation computation. Internally, the generative AI model uses a multi-layered self-attention mechanism and a feedforward network to perform matrix multiplication, weighted summation, and nonlinear transformations on the input tensor, predicting the probability distribution of the next label for each token. The server employs a Top-level architecture. k or Top The p-sampling strategy selects specific labels from the distribution and gradually generates a complete sequence of response labels.

[0444] Input: A label sequence data tensor.

[0445] Output: The sequence of response labels predicted by the model.

[0446] In this step, the server fully utilizes the parallel matrix computing capabilities of the GPU, reducing the average latency of a single call through batch inference, thereby improving the overall processing throughput while ensuring the quality of the generated data.

[0447] Step 12: The server decodes the response token sequence into natural language text and performs post-processing. The server uses the decoding function of the model's tokenizer to convert the response token sequence into a natural language text string. The server then performs useless information removal, such as removing duplicate prompts from the model and lengthy introductory statements irrelevant to the question, and adjusts the expression, such as standardizing terminology and simplifying complex sentences, to generate the final explanatory information.

[0448] Input: Response tag sequence.

[0449] Output: Cleaned and adjusted explanatory text.

[0450] In its implementation, the server can set a set of compliance rules or call a small discriminative model to detect and delete low-relevance or repetitive sentences, and replace or annotate some terms to make the explanatory information more understandable and concise.

[0451] Step 13: The server encapsulates the description information into a response message and sends it to the terminal. The server encapsulates the explanatory information obtained in step 12, along with the original question, timestamp, etc., into a response data structure, serializes it into JSON or other response formats, and sends it to the terminal that initiated the request via an HTTPS channel.

[0452] Input: Description text and related metadata.

[0453] Output: Terminal-oriented response message data packet.

[0454] Before sending, the server sets the content type and character encoding to ensure that the terminal can parse it correctly, and can attach diagnostic information such as processing time in the response header so that the terminal or upper layer system can perform performance monitoring.

[0455] Step 14: The terminal receives and parses the response message. The terminal receives response data packets from the server through the network library, reads the response body content, and calls the JSON parser to restore it into an internal data object, extracting explanatory information text and related fields.

[0456] Input: The response message data packet sent by the server.

[0457] Output: Explanatory text and associated metadata that can be displayed on the interface.

[0458] During the parsing process, the terminal checks whether the response status code and necessary fields exist. If an error is found, it prompts the user to retry or displays an error message on the interface; otherwise, it passes explanatory information to the interface rendering component.

[0459] Step 15: The terminal displays instructional information to the user. The terminal fills the pre-defined display area with explanatory text, automatically wraps and typesets the text according to the terminal screen size and font settings, and enables a scroll view to accommodate longer content when necessary.

[0460] Input: Description text and layout parameters.

[0461] Output: A readable response displayed on the screen.

[0462] During operation, the terminal triggers an interface refresh, replacing the old content with new instructions. It can also record the question and answer session for later viewing or for local caching. Users can read the instructions generated by the server by browsing the screen.

[0463] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0464] In existing information processing technologies, most systems for handling financial-related natural language queries simply use fixed rules or simple retrieval methods to extract preset text from data storage devices and return it directly to the user. These systems suffer from the following technical problems: (1) In the data processing flow inside the computer, the semantic understanding of natural language queries is separated from the subsequent data operation and processing. The processing device usually cannot dynamically generate appropriate data processing instructions based on the specific intent of the query and the data analysis results, resulting in a rigid processing flow and poor scalability. (2) Although the existing system can use machine learning models to score risks, it lacks a unified control mechanism to efficiently couple numerical analysis results with generative artificial intelligence models. It cannot automatically construct high-quality prompts within the processing device, resulting in the output quality of generative artificial intelligence models being highly dependent on manual design and difficult to be stably reused in large-scale online scenarios. (3) Most conversational financial systems do not deeply integrate the emotion recognition engine with the input control of the generative artificial intelligence model. The processing device cannot finely adjust the tone and expression style of the instructions at the model input level according to the user's emotional state, making it difficult to adaptively generate responses with appropriate tone and level of detail under the same computing architecture. (4) For multi-source heterogeneous data such as corporate financial risk, transaction counterpart risk and personal budget risk, traditional systems usually use decentralized analysis modules. Each module outputs results separately and then they are manually integrated. The lack of a unified data analysis device and model input / output control device in the computer system leads to the inability to unify and abstract the calculation logic under different risk scenarios, which increases the system complexity and reduces the computational efficiency.

[0465] Therefore, at the computer technology level, how to build a system that can automatically analyze and assess financial data and risks by combining natural language processing algorithms, numerical computing libraries and statistical learning algorithms in a single processing architecture, and automatically generate and dynamically adjust prompts for use by generative artificial intelligence models, and adaptively adjust the output content and tone based on the user's emotional state, thereby improving the overall quality of human-computer interaction and computational efficiency, has become a technical issue that urgently needs to be solved in this field.

[0466] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0467] In this invention, the server includes a processing device for executing program instructions, a user information input / output device for interacting with the user, an information acquisition control device for acquiring financial information from an information storage device and an external information providing device based on a natural language processing algorithm, a data analysis device for performing feature calculations and risk assessments on the financial information using a numerical computation processing library and statistical learning algorithms, a model input / output control device for dynamically constructing prompt statements as input instructions for a generative artificial intelligence model based on the natural language query category, the analyzed financial indicators and risk assessment values, and the target concept, and inputting the prompt statements into the generative artificial intelligence model, and an information presentation device for organizing the text output by the generative artificial intelligence model into explanatory or warning information and presenting it to the user through the user information input / output device. The processing device is also configured to perform emotion estimation on the voice signal, image signal, or text information acquired through the user information input / output device to determine the user's emotional state, and adjust the tone indication and expression conditions in the prompt statements based on the user's emotional state. This allows for the formation of an integrated data processing chain within the computer, encompassing natural language query parsing, structured data acquisition and analysis, automatic generation of prompts, and emotion-adaptive output control. This not only improves the automation and computational efficiency of financial data analysis and risk assessment but also significantly enhances the response quality and interactive experience of generative artificial intelligence models in different contexts, achieving technical improvements to the architecture and processing flow of existing human-computer dialogue-based financial analysis systems.

[0468] A “processing unit” refers to an electronic computing unit used to execute program instructions and control, calculate and logically process input data. It may consist of one or more central processing units, graphics processing units or other programmable processing units.

[0469] "User information input / output device" refers to a collection of hardware and software components used to transmit information between humans and machines, including input interfaces for receiving user input and display interfaces for presenting output information to users. Specifically, it may include a display screen, keyboard, touch screen, microphone, speaker or camera, etc.

[0470] A “natural language query” refers to a statement or set of statements expressed by a user in natural language (including text or voice) to obtain information or request analysis and processing.

[0471] "Natural Language Processing Algorithms" refer to programs and models used in computing devices to perform operations such as word segmentation, part-of-speech tagging, syntactic analysis, entity recognition, and intent recognition on natural language text or text obtained through speech conversion, in order to extract structured information from natural language.

[0472] "Information storage device" refers to a data storage system used to store structured or semi-structured data in a searchable form, which may include relational databases, key-value databases, document databases, or file storage systems.

[0473] "External information providing device" refers to a device or service located outside this system that provides data to this system through a communication network, including remote data servers, network application interfaces, or online data service platforms.

[0474] A “structured information set” refers to a set of grouped data that is organized and stored according to a predetermined data pattern or data model and can be directly accessed through query statements or programming interfaces.

[0475] "Financial information" refers to data related to an organization's or individual's assets, liabilities, income, expenses, cash flow, budget, or other economic activities, including accounting data, financial statement data, transaction records, and statistical indicators.

[0476] "Information acquisition control device" refers to a software and hardware module used to generate query conditions based on the results of natural language processing and to control access to information storage devices and external information providing devices, thereby selectively acquiring the required financial information.

[0477] "Data analysis device" refers to the software and hardware modules that use numerical processing libraries and statistical learning algorithms to clean, transform, aggregate, and calculate the acquired data, and generate various indicators and prediction results.

[0478] A "numerical computation processing library" refers to a software component provided in the form of a program library in a computer, used to perform numerical computation operations such as matrix operations, statistical calculations, and time series calculations.

[0479] "Statistical learning algorithms" refer to machine learning algorithms that model and infer data based on statistical theory, including regression models, classification models, clustering models, or other predictive analysis models.

[0480] "Financial indicators" refer to quantitative values ​​that reflect financial condition or operating results, obtained through calculation or statistics of financial information, such as current ratio, debt-to-equity ratio, profit margin, or free cash flow.

[0481] "Risk assessment value" refers to a numerical value or level obtained by quantitatively assessing financial risk through statistical learning algorithms or rule models, used to represent the degree of default risk, liquidity risk, or other financial risks.

[0482] "Model input / output control device" refers to the hardware and software modules used to construct generative artificial intelligence models based on analysis results and contextual information, and to control the sending of inputs to the model and the receiving of its output results.

[0483] "Generative AI models" refer to AI models that are trained on large-scale data and can automatically generate text and other content based on input instructions, including generative language models or multimodal generative models.

[0484] "Prompt statements" refer to instruction information expressed in natural language or structured text that is used as input to a generative artificial intelligence model to instruct the model to generate output of a specific type, style, or content.

[0485] "Information presentation device" refers to the hardware and software components used to present text information or other content output by generative artificial intelligence models to users in the form of graphics, text, audio or video.

[0486] "Emotion estimation algorithms" refer to algorithms that analyze speech signals, image signals, or text information to identify the category and intensity of a user's emotions, including speech emotion recognition algorithms, facial expression recognition algorithms, and text emotion analysis algorithms.

[0487] "User emotional state" refers to the type and intensity of a user's emotions at a specific point in time, inferred through emotion estimation algorithms, such as anxiety, anger, neutrality, or happiness.

[0488] "Tone instructions" refer to textual descriptions or control information in prompt statements used to constrain or guide generative artificial intelligence models to respond with a specific tone or attitude.

[0489] "Expression mode conditions" refer to the constraints or requirements imposed on the form, style, level of detail, or structure of the output content in the prompt statement.

[0490] "Transaction record information" refers to data that reflects various transactions or income and expenditure activities of an entity within a certain period of time, including transaction time, amount, counterparty identification, and transaction category.

[0491] "Income and expenditure information" refers to a collection of data related to an entity's income and expenditure activities, including records of wages, interest income, daily consumption, fixed expenses, etc.

[0492] "Asset and liability information" refers to data used to describe the asset and liability status of an entity at a certain point in time, including various asset items, liability items and their amounts.

[0493] "Financial risk indicators" are quantitative values ​​used to measure the level of credit risk, liquidity risk, debt repayment risk, or other financial risks. They are obtained by analyzing and calculating transaction records, income and expenditure information, and asset and liability information.

[0494] "Future risk forecast value" refers to the quantitative prediction result of the level of financial risk in the future time period inferred by statistical learning algorithms or prediction models based on historical data and the current state.

[0495] The embodiments of this invention will be described primarily using servers, terminals, and users as examples. The following embodiments are merely illustrative, and those skilled in the art can make various modifications and substitutions without departing from the spirit of this invention.

[0496] I. Overall System Composition Servers are deployed in data centers or cloud computing platforms. A server includes hardware such as processing units, storage devices, network interfaces, and graphics processing units. The server runs an operating system (such as a general-purpose server operating system) and multiple software components. The server's software components include at least: (1) Web / application server component, used to receive requests from the terminal; (2) Natural Language Processing Module, based on a natural language processing library (e.g., a Chinese processing model based on spaCy); (3) Data access module, used to access relational database systems (such as relational database management systems) and external data interfaces; (4) Data analysis module, based on numerical computation and statistical learning libraries (e.g., based on NumPy, pandas, scikit-learn). (5) Generative AI model calling module, used to call generative AI models (such as language models based on Transformer structure) deployed in the cloud or locally through network interface. (6) Sentiment analysis module, used to call the text sentiment analysis interface or image / speech sentiment recognition interface (e.g., natural language sentiment analysis service, image emotion recognition service); (7) Response building module, used to organize and format terminal-oriented return data.

[0497] The terminal can be a smart mobile terminal, a personal computing device, or other device equipped with a display device, input device, microphone, and camera. The terminal runs a client application or browser and communicates bidirectionally with the server via a network. The terminal functions as a user information input / output device through a user interface.

[0498] Users can input natural language queries via text or voice through the terminal interface. Users can select the target subject (such as a company or account) or the target concept (such as a financial term). During the interaction with the terminal, users can also choose whether to enable voice emotion recognition or facial expression emotion recognition assistance functions.

[0499] II. Modular Structure and Data Processing of Server-Side Programs In terms of software architecture, the server divides the program into several logical modules. Each module interacts with a clearly defined data structure. This modular division is intended to create an efficient data flow and a clear control flow within the computer, thereby improving processing speed and accuracy at the algorithm level.

[0500] 1. Natural Language Processing Module The server parses the user's natural language query using its natural language processing module. The server uses natural language processing algorithms to perform the following data processing operations on the input text: The server first performs word segmentation and part-of-speech tagging on the text, transforming the original string into a sequence of words and parts of speech. The server then performs dependency parsing and named entity recognition, identifying key elements such as financial-related terms, entity names, and interrogative intents. The server represents the analysis results in structured data format, using internal fields such as "intent category," "target concept," and "target entity."

[0501] The server eliminates the ambiguity of plain text descriptions through this structured representation, allowing subsequent data query and analysis processes to directly select branches based on "intent category" and "target concept," thereby reducing multiple string matching and rule judgments and improving processing efficiency.

[0502] 2. Data Access Module and Storage Structure The server constructs multiple logical data tables in the storage device, such as a terminology definition table, a financial indicator table, and a user ledger table. Each table uses a relational modeling approach, employing primary keys and foreign keys to link financial entries with entity identifiers. The server uses indexes and preprocessed views to reduce query latency.

[0503] When the server detects an intent related to "terminology explanation," it retrieves the corresponding record from the terminology definition table based on the target concept, including a brief definition, detailed explanation, typical application scenarios, and formula description. When the server detects an intent related to "corporate financial risk," it calls an external data interface to obtain the target entity's balance sheet, income statement, and cash flow statement, among other financial data. When the server detects an intent related to "personal budget risk," it summarizes the current period's budget and actual expenditures from the user's ledger.

[0504] During the data access phase, the server uses a batch query and caching strategy to retrieve multiple related records at once and build an intermediate data structure (such as a pandas DataFrame object) in memory. This design reduces the latency caused by multiple network round trips and disk accesses, and facilitates centralized execution of vectorized computations on the server side.

[0505] 3. Data Analysis Module and Feature Calculation The server inputs the raw or preprocessed data returned by the data access module into the data analysis module. In the data analysis module, the server performs matrix operations using a numerical computation library, primarily executing the following data processing steps: The server calculates numerical indicators such as current ratio, quick ratio, debt-to-equity ratio, interest coverage ratio, and free cash flow from corporate financial data. It also calculates budget utilization rate, category expenditure ratio, and time-series volatility indicators from individual expense records. Furthermore, the server uses a pre-trained model with statistical learning algorithms to predict risk. During the training phase, this model takes a multi-dimensional financial feature vector as input and known default or risk events as labels, obtaining weight parameters through supervised learning.

[0506] During runtime, the server assembles various numerical indicators into fixed-dimensional feature vectors. These vectors are then used for forward inference through a loaded classification or regression model to obtain a risk assessment value. The server employs vectorized computation, batch prediction, and pipelined data flow to reduce the overhead of the Python interpretation layer and improve overall computational speed.

[0507] The server unifies all indicators and risk scores into a structured result object, storing them as a mapping table of "indicator name-value" and a risk score field. This allows the model's input / output control devices to directly access these indicators when constructing prompts and selectively embed them as needed.

[0508] III. Generative Artificial Intelligence Models and Prompt Statement Construction 1. Structure and Invocation Methods of Generative Artificial Intelligence Models The generative AI model invoked by the server is a language model based on a multi-layer self-attention network. During the training phase, the server uses a large-scale text corpus and trains the model in an autoregressive manner, updating parameters such as multi-layer attention weights and feedforward network weights by minimizing the cross-entropy loss function for predicting the next word. During the runtime phase, the server encodes the prompt statement as a discrete word sequence. The model performs forward propagation on this sequence to calculate attention weights and implicit representations, and uses a softmax layer to generate the subsequent word distribution, thereby generating a natural language response.

[0509] When invoking a model, the server can choose different sampling strategies, such as greedy search, temperature sampling, or top-k / top-p sampling, to balance output diversity and stability. The server can adjust the temperature parameter and maximum generation length based on the question type; for example, a lower temperature can be used for answers related to financial warnings to avoid excessive arbitrariness.

[0510] 2. Internal construction logic of prompt statements The server generates prompt statements according to a predefined template in the model input / output control device. During the construction process, the server inserts information from different sources in layers: The server first inserts role settings based on the intent category, such as "You are a financial instructor," "You are a financial analyst," or "You are a personal financial advisor." Next, the server inserts the user's original question text, preserving the user's expression. The server then inserts structured financial indicators and risk assessment values, explained in natural language, such as "The current ratio is 1.6, the debt-to-equity ratio is 70%, and free cash flow has been positive for three consecutive years." Finally, the server inserts output requirements constraints, such as answer length, paragraph structure, tone, and whether a number of suggestions should be provided.

[0511] Through this structured splicing strategy, the server enables generative AI models to simultaneously receive three key types of information—"problem context," "numerical basis," and "task constraints"—in a single prompt statement. This reduces the burden on the model to guess the context from massive amounts of parameters, resulting in more stable output quality.

[0512] When the server detects a specific emotional state, it adds tone instructions to the prompts. For example, when the sentiment analysis module gives the result of "anxiety", the server adds instructions such as "the tone should be gentle and reassuring" and "avoid using overly exaggerated language" to the prompts; when the sentiment result is "anger", the server adds instructions such as "the tone should be direct and frank, highlighting factual evidence", thereby guiding the model to adjust its expression style on the same numerical basis.

[0513] 3. Example of a prompt statement The server can use the following prompt statements in terminology explanation scenarios: "You are a financial instructor. User question: 'What is EBITDA?'" Database definition: EBITDA is a company's operating profit plus depreciation and amortization, and is often used to measure a company's operating profitability.

[0514] Require: 1. Explain this concept in plain Chinese; 2. Provide a simple calculation example; 3. Answer in 3–6 sentences. In enterprise risk analysis and user anxiety scenarios, the server can use the following prompt statements: "You are a professional financial analyst."

[0515] User question: 'Is this company's financial situation stable?' User sentiment: Anxious.

[0516] Financial indicators already calculated: - Current ratio: 1.6 - Debt-to-equity ratio: 70% - Free cash flow: Positive for the last three years, but slightly declining. - Machine learning risk score (0–1, higher scores indicate greater risk): 0.78 Please answer in Chinese: 1. First, conduct an overall assessment of the company's financial soundness; 2. Identify 2-3 key risk points; 3. Maintain a gentle and reassuring tone, but be sure to explain the risks truthfully; 4. Provide two specific suggestions to help users make a decision. The server can use the following prompt statement in a personal budget scenario: "You are a personal financial advisor."

[0517] User question: 'Will this payment affect my budget?' User sentiment: Anxious.

[0518] data: - Monthly budget: 5000 yuan - Expenditure to date: 4200 yuan - Amount due: 800 yuan Please explain in 4–7 sentences: 1. The percentage of this payment relative to the monthly budget and remaining budget; 2. Will it bring significant pressure? 3. Offer one or two suggestions in a friendly and encouraging tone. IV. Combining Sentiment Analysis and Generative Control In the sentiment analysis module, the server uses a text sentiment classification model for text input. This model can be trained based on a bidirectional recurrent neural network or a Transformer encoder structure. The server inputs user text through word segmentation and embedding transformation before feeding it into the sentiment model. The final classifier layer outputs sentiment labels and probability distributions. The server can set a threshold; when the probability of a certain sentiment exceeds the threshold, that category is considered the user's sentiment state.

[0519] The server can invoke external emotion recognition services to obtain facial expression categories and voice emotion labels for both images and speech. The server then weights and fuses the emotion results from different modalities to form the final emotional state. The server subsequently writes the emotion label and confidence score into a unified context object for the model's input / output control device to read.

[0520] By incorporating explicit instructions regarding tone and expression into the prompts, the server transforms the discrete labels output by the sentiment analysis module into natural language control parameters that the generative AI model can understand. Compared to simply filtering the output afterward, this approach guides the model to favor specific expressions during the decoding path, even at the generation stage. This improves the match between the response and the user's emotion while maintaining generation efficiency.

[0521] V. Terminal-side display and interaction processing After receiving data from the server, the terminal uses the response text and structured metrics for interface presentation. The terminal can display explanatory information from the generative AI model's output in an explanation area, and use metrics to draw curves or bar charts in the graph area, such as showing debt ratio trends or budget usage.

[0522] The terminal can invoke a speech synthesis service to convert text-based responses into spoken audio, improving accessibility. Simultaneously, the terminal can adjust interface elements based on the emotional state and risk level returned by the server; for example, using prominent colors or icons when the risk is high. When the user evaluates the response or asks further questions, the terminal sends new information back to the server so that the system can continuously optimize its internal models and prompt templates.

[0523] VI. Technical Effects and Improvements in Computer Technology Through the aforementioned structured data flow and modular algorithm design, the server has achieved multiple levels of technical improvements: 1. The server directly drives the data access and feature generation modules through natural language parsing results, making query construction, feature calculation and model calling form a seamless pipeline, thereby reducing intermediate format conversion and redundant rule matching, and improving the overall processing speed.

[0524] 2. The server centrally performs vectorized operations based on the numerical operation processing library through the data analysis module, avoiding cyclic comparison of each record, significantly reducing time complexity, and improving the calculation efficiency of risk assessment values.

[0525] 3. By embedding structured metrics and sentiment control information into the prompt statements, the server transforms the rigid process of traditional "query → retrieval → template filling" into "data-driven generative control," enabling generative AI models to maintain flexibility while having clear numerical basis, thus improving output accuracy and consistency.

[0526] 4. The server uses a unified model input / output control device to construct a framework of unified prompt statements for different application scenarios (terminology explanation, enterprise risk, personal budget), which not only reduces system complexity, but also facilitates reuse through incremental configuration when adding new scenarios, thereby improving system maintainability and scalability.

[0527] 5. By directly applying sentiment analysis results to the tone and expression conditions of prompts, rather than simply filtering them after output, the server achieves feedforward control over the generation process. From an information theory perspective, this reduces invalid generation paths, thereby improving generation efficiency and lowering the proportion of irrelevant content.

[0528] 6. When training generative artificial intelligence models and emotion recognition models, the server adopts an iterative optimization method based on gradient descent. By minimizing cross-entropy loss or mean squared error loss, the network weights are continuously updated. The server can use batch normalization, dropout strategies and data augmentation methods to improve the model's generalization ability, thereby maintaining high accuracy across a wider range of data distributions during runtime.

[0529] Through the above implementation, the server, terminal, and user collaboratively realize a data-driven financial interpretation and risk warning system with generative artificial intelligence models and prompt statements as its core. This system realizes an innovative combination of specific data structures, processing sequences, and control strategies within the computer, which not only improves the speed and accuracy of information processing but also enhances the quality of human-computer interaction, representing an improvement to computer technology itself.

[0530] use Figure 14 The processing procedure is explained.

[0531] Step 1: The user enters an inquiry in the terminal. Users open a client application or browser interface on the terminal, type their natural language question in the input box, or speak their question through the microphone, and can select the target subject or concept from the drop-down list. The terminal temporarily caches the text or voice input in local memory.

[0532] Input: User's keyboard input text and / or raw speech signal.

[0533] Output: Natural language question text (e.g., UTF-8 string) represented internally by the terminal, optional target entity identifier, and raw speech data.

[0534] Step 2: Terminal collects and preprocesses user data The terminal preprocesses user input: when the user inputs speech, the terminal calls a speech recognition service to convert the speech into text and merges the recognition result with the user's manual input to form a complete question text; when emotion recognition is enabled, the terminal uses a camera to capture facial image frames and a microphone to collect short audio clips, encoding the images and audio (e.g., compressing them into JPEG images and compressed audio). The terminal assembles the request data structure, encapsulating the question text, target subject identifier, and optional audio / video data into JSON or form data.

[0535] Input: Natural language question text obtained in step 1, raw speech signal, image frame, and target subject information selected by the user.

[0536] Output: The request data packet assembled by the terminal, which contains the question text, the body identifier, and the encoded audio and video data.

[0537] Step 3: The terminal sends a request to the server. The terminal establishes an encrypted connection with the server using the HTTPS protocol through the network stack, and sends the request data packet constructed in step 2 to the server's predetermined interface address. After sending the data, the terminal enters a waiting state and displays a "Analyzing" message on the interface.

[0538] Input: Terminal internal request data packet.

[0539] Output: The HTTP request message from the terminal to the server, and the waiting status indication on the terminal.

[0540] Step 4: The server receives and parses the request. The server receives HTTP requests from the terminal via a network interface and forwards them to the application through the web server component. The server parses JSON or form fields from the request body, extracting natural language question text, target identifiers, audio / video data, and request flags. The server constructs a session context object in memory for subsequent modules to share this input information.

[0541] Input: An HTTP request message from the terminal.

[0542] Output: A session context object inside the server, containing the question text, subject identifier, and raw or encoded audio and video data.

[0543] Step 5: The server performs natural language parsing and intent recognition. The server uses natural language processing algorithms in its natural language processing module to perform word segmentation, part-of-speech tagging, dependency parsing, and named entity recognition on the question text. Based on the analysis results, the server extracts the target concept (such as a financial term), the target subject (such as an organization name), and the intent category (such as "term explanation", "enterprise risk analysis", "personal budget assessment", etc.).

[0544] Input: Natural language question text from the session context object in step 4.

[0545] Output: Structured semantic information, including intent category, target concept, target subject, and a list of key entities, updated in the session context.

[0546] Step 6: The server selects the data source and constructs the query based on the intent. The server reads the intent category and target concept / subject information. When the intent is "terminology explanation," the server constructs query conditions for the internal database (e.g., the terminology field equals the target concept). When the intent is "enterprise risk analysis," the server constructs request parameters for the external financial data interface based on the target subject. When the intent is "personal budget assessment," the server generates conditions for querying the user's ledger and budget based on the user identifier. The server encapsulates these query conditions into data access commands.

[0547] Input: Structured semantic information such as intent category, target concept, and target subject obtained in step 5.

[0548] Output: Database query statements and / or external API request parameter sets, written to the session context for use by the data access module.

[0549] Step 7: The server retrieves raw data from a database and / or external interfaces. The server executes the queries constructed in step 6 through the data access module: it performs SQL queries on the internal relational database to obtain terminology definitions, historical financial records, or user budget data; and it retrieves JSON data such as financial statements of the target entity through HTTP requests to external interfaces. The server then converts the obtained data into a standardized format, such as a tabular data structure.

[0550] Input: The query statement generated in step 6 and the external interface request parameters.

[0551] Output: The original or pre-cleaned dataset, including terminology definition records, financial statement data, user transaction records, budget setting data, etc., and stored in the data portion of the session context.

[0552] Step 8: The server uses a numerical computation library to perform index calculations. The server processes the data obtained in step 7 using a numerical computation library within the data analysis module. For corporate financial data, the server calculates financial indicators such as current ratio, debt-to-equity ratio, free cash flow, and profit margin; for personal account data, the server calculates budget utilization rate, the proportion of each category of expenditure, time series mean, and variance. The server performs these calculations using vectorized operations to reduce the number of loops and improve computational efficiency.

[0553] Input: The standardized dataset obtained in step 7.

[0554] Output: A feature set consisting of various financial indicators and statistics, written into the session context in key-value pairs.

[0555] Step 9: The server performs risk assessment using statistical learning algorithms. The server reads the feature set calculated in step 8, converts it into a fixed-dimensional feature vector, and inputs it into a pre-trained risk assessment model (e.g., a model based on logistic regression, random forest, or deep neural networks). The server performs forward inference operations on the model and outputs risk assessment values, such as default risk scores, liquidity risk scores, or budget overrun risk scores.

[0556] Input: The feature set obtained in step 8.

[0557] Output: One or more risk assessment values ​​and their corresponding confidence indices, written to the session context.

[0558] Step 10: The server estimates the user's emotional state. When the session context contains audio / video data or text data, the server performs emotion recognition in the sentiment analysis module. For text input, the server uses a sentiment classification model to embed the text and outputs sentiment labels (such as "anxious," "angry," "neutral," etc.) through a classification network; for images or speech, the server calls an external sentiment recognition service to obtain the expression or speech emotion category. The server then fuses the results from different modalities into a single sentiment state.

[0559] Input: Text, audio, and image data obtained in step 4.

[0560] Output: User sentiment state label and its confidence level, written to the session context.

[0561] Step 11: Prompt statements for the server to build generative artificial intelligence models The server reads the intent category, target concept, target subject, financial indicator results, risk assessment value, and user emotional state from the model input / output control device, and generates prompt statements based on predefined templates. The server first adds role settings and task descriptions, then inserts the user's original question text, then embeds financial indicators and risk scores in natural language, and finally adds tone requirements and output format requirements based on the emotional state to form a complete explanatory text that can be understood by the generative artificial intelligence model.

[0562] Input: Structured information from the session context formed in steps 5, 8, 9 and 10.

[0563] Output: A prompt text used to drive the generative artificial intelligence model, cached in the session context.

[0564] Step 12: The server invokes a generative artificial intelligence model to generate the response text. The server, through the generative AI model invocation module, sends the prompt statement generated in step 11 to the generative AI model interface deployed locally or in the cloud. The server includes generation control parameters such as model type, maximum generation length, and temperature parameters in the interface request. The server receives the response text returned by the model, which includes explanations of terminology, descriptions of the company's or individual's financial situation, and corresponding suggestions or warnings.

[0565] Input: The prompt statement constructed in step 11 and the model call parameters.

[0566] Output: Natural language response text generated by the generative artificial intelligence model and added to the conversation context.

[0567] Step 13: The server organizes and constructs the response data. In the response construction module, the server integrates the generated answer text with structured financial indicators, risk assessment values, user sentiment status, and other information. Based on the terminal capabilities and interface requirements, the server uses the answer text as the primary field and organizes the financial indicators and risk scores into chart data structures or explanatory additional fields. The server generates a standardized response object, prepared for return in JSON format.

[0568] Input: The response text obtained in step 12 and the metrics and risk assessment results saved in the conversation context.

[0569] Output: Terminal-oriented response data object, which contains the response text, indicator data, and risk information.

[0570] Step 14: The server sends a response to the terminal. The server encapsulates the response object constructed in step 13 into an HTTP response message and returns it to the requesting terminal via an HTTPS connection. The server sets a success status code and necessary caching control information in the response header.

[0571] Input: The response data object generated in step 13.

[0572] Output: HTTP response messages from the server to the terminal.

[0573] Step 15: The terminal parses the response and renders the interface. The terminal receives the HTTP response from the server, parses the JSON data, and extracts the response text, indicator data, and risk score. The terminal displays the response text in the display area, draws graphs (such as line charts and bar charts) based on the indicator data in the chart area, and highlights the risk level using color or icons when necessary. The terminal fine-tunes the text style or prompts based on emotional state information, for example, highlighting reassuring statements when the user is anxious.

[0574] Input: The HTTP response message from step 14.

[0575] Output: A graphical interface and visual charts rendered by the terminal, as well as prepared voice-over text (if needed).

[0576] Step 16: The terminal can optionally perform voice broadcasting and collect feedback. If the terminal has speech synthesis enabled, it calls a speech synthesis service to convert the answer text into audio and plays it to the user through the speaker. Simultaneously, the terminal displays feedback buttons such as "Helpful / Neutral / Not Helpful" on the interface, inviting the user to rate the answer. After the user clicks the feedback option, the terminal organizes the feedback information into a new request data packet, ready to send it to the server for further optimization.

[0577] Input: The answer text parsed in step 15 and the user's operation intent.

[0578] Output: The audio signal played to the user, and a request data packet containing feedback information (when the user submits feedback).

[0579] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0580] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0581] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0582] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0583] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0584] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0585] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0586] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0587] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0588] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0589] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0590] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0591] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0592] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0593] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0594] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0595] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0596] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0597] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0598] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0599] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0600] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0601] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0602] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0603] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0604] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0605] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0606] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0607] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0608] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0609] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0610] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0611] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0612] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0613] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0614] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0615] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0616] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0617] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0618] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0619] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0620] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0621] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0622] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0623] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0624] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0625] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0626] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0627] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0628] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0629] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0630] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0631] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0632] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0633] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0634] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0635] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0636] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0637] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0638] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0639] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0640] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0641] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0642] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0643] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0644] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0645] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0646] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0647] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0648] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0649] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0650] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0651] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0652] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0653] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0654] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0655] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0656] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0657] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0658] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0659] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0660] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0661] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0662] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0663] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0664] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0665] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0666] In addition, the following notes are provided in response to the above explanation.

[0667] Example 1 (Note 1) An information processing system, characterized in that it comprises: Interface means for displaying an inquiry interface on a user device and receiving natural language questions from the user device; Information acquisition and storage means for sending a query signal to an external information providing device through a communication interface, obtaining financial information related to the natural language question from the external information providing device as an information acquisition response signal, and storing and updating the financial information in a predetermined storage device in the form of structured information; This is a prompting statement generation method that takes the natural language question and the structured information as input, uses natural language processing to parse the intent of the natural language question and the financial concept, and generates prompt statements based on the parsing results and the structured information to instruct the generative artificial intelligence model to generate explanations including the definition, formula, numerical examples and time series change descriptions of the financial concept, and inputs the prompt statements into the prompt statement generation method of the generative artificial intelligence model. The method is used to obtain the explanatory text output from the generative artificial intelligence model, insert specific numerical examples into the explanatory text based on the structured information, generate graphical information representing changes in financial indicators based on the time series data contained in the structured information, and send the explanatory text and the graphical information to the explanatory generation and transmission means of the user device. The means for displaying the received explanatory text on the user device, generating a graphical display screen based on the received graphical information to provide visual prompts for the financial concept, and further receiving additional evaluation information or additional input from the user and sending it to the user interface of the processing device. And adaptive updating means for accumulating the evaluation information and the additional input in the storage device, and updating the output style of the prompts or explanations of the generative artificial intelligence model based on the accumulated results.

[0668] (Note 2) The information processing system according to Appendix 1 is characterized in that the system further includes: This is a conversation management method used to manage the history of natural language questions sent from the user's device and the presentation history of explanatory text by conversation unit, obtain summary information of past questions and explanations in the conversation unit corresponding to the current question, and include the summary information in the prompt statement of the generative artificial intelligence model and input it into the generative artificial intelligence model, so as to maintain the consistency of the financial concept explanation in multi-turn dialogue.

[0669] (Note 3) The information processing system according to Appendix 1 is characterized in that the system further includes: A means for generating and linking graphics data to extract time series data from the structured information, corresponding to the items of interest pointed out in the explanatory text output by the generative artificial intelligence model, from financial indicators of multiple periods, automatically generating line charts or bar charts as graphic information using the time series data and sending them to the user device, and displaying the graphic information and the explanatory text together on the same screen in the user device.

[0670] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: The terminal provides an information input / output device for receiving natural language question data from users and sending the received question data and user identification information to the server via a communication device; In the server, a device is used to obtain the user's transaction record data and income record data from an external storage device based on the user identification information, and to classify the transaction record data and income record data by means of statistical processing and summary processing, and to calculate the amount of expenditure of each expense category and the expenditure ratio of each expense category within a predetermined period, thereby generating the user's financial status data. In a server, a device is used to execute a natural language processing program through a processor to extract intent information, object information, and time information from the question data to determine the type of question and generate summary information associated with the financial status data. In the server, character data is generated as input to the generative artificial intelligence model based on the type of question and the financial status data. The character data includes the user's expenditure composition, income information and question content, and the processor inputs the character data into the generative artificial intelligence model. In a server, a device obtains the output answer data from the generative artificial intelligence model, generates suggestion data including financial concept interpretation information or saving suggestions based on the answer data, and sends the suggestion data to the terminal through a communication device. In a terminal, the received suggestion data is displayed as text on a display device, and then converted into audio data using a speech synthesis program and output through a sound output device.

[0671] (Note 2) According to the information processing system described in Appendix 1, the server, when generating a prompt statement input to the generative artificial intelligence model, embeds the expense category expenditure amount, expense category expenditure ratio, and budget information contained in the user's financial status data into the prompt statement, so that the prompt statement contains the determination results of major expenditure items and explanatory instructions on expenditure reduction candidate items, and controls the operation of the processor accordingly.

[0672] (Note 3) According to the information processing system described in Appendix 1, the terminal stores the suggestion data received from the server as historical data in a recording device, and when receiving additional question data from the user, associates the historical data with the additional question data and sends it to the server, thereby enabling the server to update the prompt statements input to the generative artificial intelligence model based on the historical data and the additional question data to generate financial advice in the form of a continuous dialogue.

[0673] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for controlling input / output interfaces to receive natural language queries from users; An apparatus for performing string validity verification and length limitation processing on the query information, and converting the query information into a predetermined data structure to generate a communication request message; A device for sending a communication request message to an information processing device via a communication channel and receiving a response message from the information processing device; A device for performing morphological analysis, intent classification, and keyword extraction on the query information to generate parsing result data containing query intent information and related word information; An apparatus for retrieving information resources in an information storage device based on the parsing results, obtaining multiple reference information based on relevance evaluation, and shaping the obtained reference information into a predetermined text format. A means for constructing prompt statements for input to a generative artificial intelligence model based on the query information and the reference information, and encoding the prompt statements into labeled sequence data; An apparatus for inputting the labeled sequence data into a generative artificial intelligence model to generate response text through numerical computation, and for performing useless information deletion and expression mode adjustment on the generated response text to generate explanatory information; A means for presenting the explanatory information as a response message to the user through the input / output interface.

[0674] (Note 2) The information processing system according to Appendix 1 is characterized in that, The system further includes: a device for acquiring the user's voice information and image information when presenting the explanatory information, performing emotion estimation processing on the voice information and image information to extract emotional state information, and adjusting the conditions of the prompt statement and the expression form of the response text output by the generative artificial intelligence model according to the emotional state information, thereby dynamically adjusting the tone or content of the explanatory information.

[0675] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system further includes: a device for parsing numerical and historical information related to the user to generate risk assessment information, generating a prompt statement containing warning content about specific risk events based on the risk assessment information and the parsing result data, and inputting the prompt statement into the generative artificial intelligence model to generate explanatory information containing the warning content.

[0676] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: Processing unit, used to execute program instructions; A user information input / output device is used to receive natural language queries from users and output information to users through a display device and an input device; An information acquisition control device is used to apply a natural language processing algorithm to the natural language query, parse the natural language query to extract the target concept and intent therein, and select and acquire financial information from a structured information set in an information storage device and an external information providing device based on the extracted target concept and intent. A data analysis device is used to calculate multiple financial indicators and risk assessment values ​​based on the financial information acquired by the information acquisition and control device, using a numerical processing library and statistical learning algorithms, and to organize the calculation results into a predetermined data structure. The model input / output control device is used to dynamically construct prompt statements as input instructions for the generative artificial intelligence model based on the category of the natural language query, the financial indicators and risk assessment values ​​obtained by the data analysis device, and the target concept, and input the prompt statements into the generative artificial intelligence model. An information presentation device is used to organize the text information output by the generative artificial intelligence model into explanatory or warning information in response to the natural language query, and present it to the user through the user information input / output device.

[0677] (Note 2) The information processing system according to Appendix 1 is characterized in that, The processing device is configured to: apply an emotion estimation algorithm to the voice signal and image signal or text information acquired through the user information input / output device to determine the user's emotional state, and adjust the tone indication and expression conditions in the prompt statement constructed by the model input / output control device according to the determined emotional state, thereby controlling the tone and content of the explanation information or warning information output by the generative artificial intelligence model.

[0678] (Note 3) The information processing system according to Appendix 1 is characterized in that, The processing device is configured to: acquire transaction record information, income and expenditure information, and asset and liability information related to the user or trading partner organization through the information acquisition control device; calculate multiple financial risk indicators and future risk prediction values ​​based on the information through the data analysis device; and input a prompt statement containing the financial risk indicators and the future risk prediction values, and used to instruct the generation of warning information related to a specific financial risk, through the model input-output control device, into the generative artificial intelligence model, so that the generative artificial intelligence model generates warning information for the specific financial risk.

Claims

1. An information processing system, characterized in that, Includes a processor, the processor being configured to: Provide users with an interface for receiving user questions; The user questions received through the interface are parsed using natural language processing techniques, and information related to the user questions is retrieved from the database. Based on the acquired information, a prompt message is generated to instruct the generation of explanatory content related to a specific financial concept, and the prompt message is input into a generative artificial intelligence model so that the generative artificial intelligence model generates explanatory content for the specific financial concept.

2. The information processing system according to claim 1, characterized in that, The processor is also configured to present the generated narration to the user, analyze the user's tone of voice and facial expressions to identify the user's emotions, and adjust the tone and / or content of the narration according to the identified emotions.

3. The information processing system according to claim 1, characterized in that, The processor is also configured to analyze the user's financial data, generate a prompt message based on the analysis results to indicate the generation of warning information related to a specific financial risk, and input the prompt message into the generative artificial intelligence model so that the generative artificial intelligence model generates the warning information for the specific financial risk.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A