Information processing system

CN122797481APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610273977.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-07
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

这种方式存在以下问题:首先,教师在繁忙的教学与行政工作之外,还需投入较多时间和精力处理联络簿回复,造成工作负担过重、业务时间过长

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797481A_ABST
    Figure CN122797481A_ABST
Patent Text Reader

Abstract

The application provides an information processing system. An information processing system, characterized by comprising: a processor configured to: receive and parse contact book content sent by a guardian; input a prompt text to a generative artificial intelligence model according to a parsing result to generate a reply text scheme; and present the generated reply text scheme to a terminal of a teacher to allow the teacher to modify the reply text scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] Current communication between schools and guardians primarily relies on text-based methods such as contact books, requiring teachers to manually read guardians' messages and write responses one by one. This approach has several problems: First, in addition to their busy teaching and administrative work, teachers must dedicate significant time and energy to processing contact book responses, resulting in an excessive workload and extended working hours. Second, due to individual differences in teachers' expression abilities, emotional states, and time pressures, the consistency and quality of responses lack stability, making it difficult to convey care and professional judgment to guardians in a timely and accurate manner, potentially affecting guardians' sense of security and trust in the school and teachers. Third, when dealing with a large volume of repetitive and similar messages, the existing manual method lacks auxiliary tools to help teachers efficiently generate appropriate responses, leading to low communication efficiency and difficulty in maximizing communication effectiveness. Therefore, it is necessary to provide a system that can automatically generate response texts based on contact book content, while simultaneously enhancing guardians' sense of security and trust, reducing teachers' workload, shortening working hours, and improving communication quality and efficiency. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides an information processing system comprising a processor configured to: receive and parse contact book content sent by a guardian, extracting information related to the student's situation, the guardian's concerns, emotional tendencies, and communication needs; input prompt text into a generative artificial intelligence model based on the parsing results, enabling the model to automatically generate a corresponding response text based on an understanding of the guardian's intentions and context; and further present the generated response text to a teacher's terminal, allowing the teacher to review, modify, and confirm the response text based on their own judgment. Through these methods, the system can provide teachers with high-quality, draft-level response suggestions while ensuring the teacher's final approval, significantly reducing the time teachers spend writing contact book responses. Preferably, the processor is configured to control the response text generated by the generative artificial intelligence model to include content aimed at enhancing the guardian's sense of security and trust, such as through polite language, positive feedback, specific explanations of the student's situation, and clear expressions of follow-up plans, enabling the guardian to more fully understand the school and teacher's response measures and attitudes. Furthermore, by utilizing the generated response text scheme, the processor enables teachers to complete high-quality responses in a shorter time, thereby shortening the overall business processing time and maintaining the consistency and professionalism of the response content throughout multiple rounds of communication, maximizing the communication effectiveness between the school and guardians. Through the above structure and processing flow, this invention achieves the technical effects of reducing teachers' workload, improving response quality, and enhancing guardians' peace of mind and trust.

[0005] "System" refers to an overall device or platform consisting of hardware and / or software, used to perform a combination of processing functions as described in this invention, such as receiving contact book content, parsing content, calling generative artificial intelligence models, and providing reply text solutions to teacher terminals.

[0006] A "processor" is a computing unit that can execute program instructions to perform operations such as data reception, parsing, feature extraction, interaction with generative artificial intelligence models, and result output. It can be a single physical processor, a multi-core processor, a processor cluster, or a computing resource composed of multiple logical processing modules working together.

[0007] "Contact Book Contents" refers to text information related to students that is entered by guardians electronically or on paper and then digitized and sent to the system. This includes, but is not limited to, written records about students' learning, life, health status, behavior, requests, feedback, and inquiries.

[0008] "Guardian" refers to a natural person who has a legal responsibility to protect or care for a student, including but not limited to parents, grandparents, legal guardians, and other student caregivers recognized by the school.

[0009] "Prompt text" refers to text instructions or contextual information generated by the processor based on the parsed contact book content and input into the generative artificial intelligence model. It is used to guide the generative artificial intelligence model to generate corresponding response text schemes according to the expected tone, structure, content focus and target effect.

[0010] "Generative AI models" refer to AI models trained using machine learning and deep learning techniques that can automatically generate natural language text output based on input prompts. These include, but are not limited to, large language models, dialogue generation models, and other models with text generation capabilities.

[0011] "Response text scheme" refers to a draft text automatically generated by a generative artificial intelligence model after receiving a prompt text, used to respond to the contents of the guardian's contact book. This draft can be sent directly or after being modified by the teacher as a formal response to the guardian.

[0012] "Teacher's terminal" refers to the electronic device used by teachers to receive and view the system-generated response text schemes, as well as to modify and confirm the schemes, including but not limited to personal computers, tablets, smartphones, or other display and input devices that can access the system.

[0013] "Reassurance" refers to the psychological stability that a guardian experiences after reading a teacher's reply, when they have a full understanding and trust in the student's situation at school and the school and teachers' response measures. It manifests as a sense of reassurance and comfort regarding the school's educational and communication outcomes.

[0014] "Trust level" refers to the degree of trust that guardians have in the school and teachers in terms of their educational abilities, communication attitudes, information transparency, and problem-solving abilities. It is usually reflected in the willingness to continue to cooperate with the school and the tendency to accept the school's suggestions.

[0015] "Business time" refers to the total time teachers spend on various tasks related to replying to the contact book, including the time spent on reading the contact book, thinking about the reply content, writing the reply text, proofreading and sending, etc.

[0016] "Communication effectiveness" refers to the quality and efficiency of information transmission between schools and guardians achieved through the system of this invention, specifically including the comprehensive results such as whether the information is accurate, whether the expression is clear, whether the emotional communication is in place, and whether it helps to enhance understanding and cooperation. Attached Figure Description

[0017] Figure 1This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0018] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0019] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0020] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0021] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0022] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.

[0023] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0024] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0025] Figure 9 This represents an emotion map that maps multiple emotions.

[0026] Figure 10 This represents an emotion map that maps multiple emotions.

[0027] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0028] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0029] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0030] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0031] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0032] First, let me explain the terminology used in the following instructions.

[0033] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0034] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0035] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0036] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0037] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0038] First Implementation Method Figure 1An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0039] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0040] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0041] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0042] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0043] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0044] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0045] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0046] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0047] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0048] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0049] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0050] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0051] With the widespread adoption of online and mobile payments, the volume of transaction data has increased dramatically. Traditional methods for detecting improper transactions, which rely on manual rules and fixed models, have revealed shortcomings in computer technology in the following aspects: First, servers typically only make threshold judgments on a few dimensions such as transaction amount, region, and time, resulting in insufficient feature representation capabilities and an inability to fully utilize the complex correlation between historical transaction information and user behavior patterns, thus reducing accuracy in identifying new improper transaction patterns. Second, existing systems mostly adopt a single-stage, one-time judgment architecture, where the server directly provides a risk level after receiving transaction data, lacking tiered classification. The in-depth analysis process of the segment is prone to false positives and false negatives, and it is difficult to dynamically adjust computing resources according to different risk levels. Third, when the server calls the intelligent model, it often uses fixed input forms for reasoning, lacking explicit control over the model's "role," "output format," and "evaluation benchmark," resulting in unstable model output and difficulty in standardizing and integrating it into the existing risk control engine. Fourth, the existing system lacks a structured recording and feedback mechanism for model input and output, making it difficult for the server to effectively accumulate and reuse historical judgment results and transaction characteristic information, resulting in difficulty in iteratively optimizing the model and risk control logic, and making it difficult to improve the overall processing efficiency and real-time performance of the system.

[0052] Therefore, in large-scale concurrent transaction scenarios, traditional fraud detection systems suffer from bottlenecks in computer technology aspects such as feature extraction, risk assessment process control, model invocation methods, and result feedback mechanisms. They cannot fully leverage the expressive power of generative artificial intelligence models and struggle to balance security with processing performance and scalability. Thus, it is necessary to provide a server-side fraud detection system that utilizes generative artificial intelligence models and prompt statement control mechanisms. By improving the server's feature extraction methods for transaction data, model invocation processes, and result storage and reuse mechanisms, the accuracy, efficiency, and maintainability of fraud detection can be improved at the computer technology level.

[0053] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0054] In this invention, the server includes a device for extracting payment characteristic information based on payment information and historical transaction information obtained from an information processing device and generating a prompt statement containing the payment characteristic information; a device for inputting the prompt statement into a generative artificial intelligence model to obtain a first judgment result of improper transaction risk, updating the prompt statement based on the first judgment result and the payment characteristic information to call the generative artificial intelligence model again to obtain a final judgment result of improper transaction risk; and a device for explicitly adding judgment condition information indicating the role, output format, and evaluation benchmark of the generative artificial intelligence model to the prompt statement and storing the final judgment result and the payment characteristic information as record information. This allows for a phased improper transaction judgment process driven by prompt statements on the server side. On the one hand, the expressive power of input features is improved through the joint extraction of multi-dimensional payment characteristic information and historical transaction information; on the other hand, the standardization and structured output of risk judgment results are achieved through explicit control of the role and output format of the generative artificial intelligence model; and through the continuous accumulation of the final judgment result and characteristic record information, the efficiency of subsequent improper detection processing and the stability of model calls are improved, thereby improving the security review performance and scalability of the online payment system at the computer technology level.

[0055] A "system" refers to a collection of devices consisting of multiple functional components, used to perform a series of processes such as data acquisition, data transmission, judgment of improper transactions, and result output in an electronic trading environment.

[0056] "Information processing device" refers to terminal equipment capable of inputting, processing and outputting data, including but not limited to computing terminals, mobile terminals and other electronic devices with communication functions.

[0057] "Payment information" refers to a set of structured or semi-structured data related to an electronic transaction, including but not limited to account information, transaction amount, transaction time, network address, device identifier, and other parameters related to the transaction environment.

[0058] "Encrypted communication method" refers to a communication protocol or mechanism with encryption function used when transmitting data between a terminal and a server. It is used to encrypt the transmitted content to ensure the confidentiality and integrity of the data during the transmission process.

[0059] A "server" refers to a network-side computing device equipped with a processor and memory, configured to receive data from information processing devices, perform illicit transaction detection and processing, and output processing results.

[0060] "Historical transaction information" refers to a collection of transaction data that is associated with and recorded within a predetermined time frame, including past payment information, records of improper transactions, and statistical information related to user behavior patterns.

[0061] "Payment characteristic information" refers to the characteristic data extracted or generated by the server based on payment information and historical transaction information to describe transaction behavior patterns, including but not limited to amount characteristics, time characteristics, geographical location characteristics, device characteristics, and frequency characteristics.

[0062] "Prompt statements" refer to textual or instructional information that is constructed by the server based on payment characteristic information and input into the generative artificial intelligence model. They are used to specify the model's task objectives, input context, output format, and evaluation benchmark.

[0063] "Generative artificial intelligence models" refer to computational models that use machine learning algorithms to infer from given prompts using pre-learned parameters, thereby generating text output or judgment results, including but not limited to language models based on deep learning.

[0064] "Improper transaction risk" refers to the qualitative or quantitative indicators of the probability or degree of danger that electronic transactions may violate normal use purposes or cause financial losses.

[0065] "First-stage judgment result" refers to the result information output by the generative artificial intelligence model after receiving an initial prompt statement containing payment characteristic information, which is the result of the first-stage judgment on the risk of improper transactions.

[0066] "Final judgment result" refers to the comprehensive final judgment output given by the generative artificial intelligence model after combining the initial judgment result with the updated prompt statement for further inference regarding the risk of improper transactions.

[0067] "Decision condition information" refers to the control information used in the prompt statements to limit the behavior of generative artificial intelligence models, including the setting of model roles, constraints on output format, and explanations of evaluation criteria.

[0068] "Risk differentiation" refers to the level or category formed by classifying or categorizing the risks of improper transactions according to pre-set judgment criteria, such as classifying risks into high risk, medium risk, and low risk.

[0069] "Payment review result" refers to the output information generated by the server based on the final judgment result, which indicates whether the target transaction is considered a secure transaction or whether further manual review is required.

[0070] "Record information" refers to the data set that is formed and stored by the server during the process of detecting improper transactions, including the final judgment result, payment characteristic information, and other metadata related to the judgment process.

[0071] The embodiments of the present invention will combine the collaborative processing of the server, terminal and user, and will provide a detailed description of the system's hardware structure, software modules, data structure and the internal workings of the generative artificial intelligence model, so that those skilled in the art can implement the present invention accordingly.

[0072] I. Overall System Structure The server is configured with a processor, memory, and network interfaces. The processor can be a combination of a multi-core central processing unit and / or a graphics processing unit, such as a computing node based on a general-purpose processor and a graphics processing unit. The memory includes high-speed random access memory and non-volatile storage media for storing the operating system, applications, generative artificial intelligence model parameters, and historical transaction data.

[0073] The server runs a server operating system, such as a Unix-like operating system, at the software level, and also runs network server software, application services, and generative artificial intelligence model inference services. The network server can be general-purpose network server software, the application services can be implemented using a service-oriented architecture, and the generative artificial intelligence model inference services can be deployed through independent processes or containers.

[0074] The terminal is equipped with a processing unit, a storage unit, a display unit, an input unit, and a network communication unit. The terminal can be a mobile terminal or a computing terminal. The terminal runs a terminal operating system and communicates with the server in encrypted form through a browser or a dedicated application.

[0075] Users input payment information into the system through the input unit of the terminal and receive the payment review result returned by the server through the display unit.

[0076] II. Server-side functional modules and data structures A server logically comprises multiple functional modules, which can be implemented through software modules or hardware circuits: 1. Communication Management Module The server receives encrypted data packets sent by the terminal through the communication management module and performs decryption and integrity verification through the encryption library. The communication management module then forwards the decrypted application-layer data to the transaction processing module.

[0077] 2. Transaction Processing Module The server parses the payment information from the terminal through the transaction processing module and encapsulates it into an internally unified data structure. This data structure can be a collection of key-value pairs or a record structure, containing at least an abstract representation of the following fields: - Account identifier field; - Transaction amount field; - Transaction time field; - Network address field; - Device identification field; - Merchant category field; - Geographic location field.

[0078] The transaction processing module can also generate a unique transaction identifier for each transaction, so that it can be associated and recorded in subsequent modules.

[0079] 3. Historical Data Access Module The server retrieves historical transaction information related to the current transaction from the data storage system through the historical data access module. Historical transaction information may include: - Transaction records within the previous scheduled time window; - Records of improper transactions; - Chargeback records; - Behavioral statistics associated with device identifiers and network addresses.

[0080] This module accesses the storage medium through database query statements and converts the query results into internal feature vectors, which are then input to the feature extraction module.

[0081] 4. Feature Extraction and Construction Module The server performs multi-dimensional feature processing on payment information and historical transaction information through a feature extraction and construction module. This module performs specific data computation operations on the processor, such as: - Calculate the ratio of the amount field to the user's historical average amount; - Calculate the degree of deviation between the transaction time field and the user's usual transaction time period; - Calculate the distance or country difference between the geographic location field and the user's frequently used geographic locations; - Calculate the number of transactions within the most recent multiple time windows for transaction frequency; - Calculate whether the combination of device identifier and network address indicates a new or abnormal environment.

[0082] The server organizes the above calculation results into continuous or discrete features, forming payment characteristic information. This payment characteristic information can be represented as a vector or stored in an in-memory structure as a set of fields, for use by the subsequent prompt generation module and model inference module.

[0083] 5. Prompt Statement Generation and Update Module The server uses a prompt generation and update module to convert payment feature information into text format that the generative artificial intelligence model can understand. This module performs text concatenation and template filling operations on the processor to construct prompts that include model roles, task descriptions, input data summaries, and output requirements.

[0084] For example, the prompt statement generated by the server for a single determination could be: "You are a generative artificial intelligence model used to detect the risk of improper credit card payments."

[0085] Based on the following transaction information and common characteristics of improper payments, determine the level of improper risk for this transaction and output only one of 'high risk', 'medium risk', or 'low risk'.

[0086] Transaction information: - Amount: 9800 yuan - Trading time: 2026-01-30 23:58:12 (late night) - Country of origin for IP address: Country A - Device type: Mobile terminal - Cardholder's registered country: Country C - Average transaction amount over the past 6 months: 300 yuan - Number of transactions in the last 24 hours: 10 answer:" After receiving a judgment result, the server generates an update prompt statement for the final judgment using the same module, for example: "You are now acting as the generative AI model for the final review of improper payments. Please make a comprehensive judgment based on the preliminary risk assessment."

[0087] First assessment result: High risk.

[0088] Detailed information: - The current transaction amount differs significantly from the average transaction amount per user; - The country of the current transaction is inconsistent with the country where the user registered; - Trading frequency has increased significantly in the last 24 hours; - There is a chargeback record within the last 30 days; Based on the information above, please give your final judgment: 1. Output only one of the following: 'Safe', 'Possibly improper', or 'Requires manual confirmation'; 2. Explain the main reason in no more than 40 words.

[0089] answer:" The server explicitly adds decision criteria information to these prompts to specify the role, output format, and evaluation benchmark of the generative artificial intelligence model, thereby constraining the model output.

[0090] 6. Model Invocation and Inference Module The server sends prompts to the generative AI model through the model invocation and inference module and receives the inference results. The generative AI model can be a neural network language model based on a self-attention mechanism, whose network structure includes an embedding layer, multiple coding units, and an output layer.

[0091] The server performs the following computation and control operations in this module: - Segment and encode the prompt statements, mapping the text into a vector sequence; - Perform multi-head self-attention computation and feedforward network operations within the encoding unit; - Generate a token sequence by calculating the probability distribution in the output layer; - The generated natural language results are parsed into risk levels and reasons, with specific keywords corresponding to the risk distinctions predefined by the system.

[0092] In one judgment, the server calls the generative artificial intelligence model to obtain an initial risk distinction. In the final judgment, the server calls the generative artificial intelligence model again based on the update prompt statement, thus realizing a two-stage reasoning process.

[0093] 7. Risk Result Analysis and Audit Result Generation Module This module allows the server to parse the natural language output by the generative artificial intelligence model into internally standardized risk labels and descriptive fields. The server can predefine the correspondence between terms such as "high risk," "medium risk," "low risk," "safe," "potential for impropriety," and "requires human confirmation" and their internal codes.

[0094] The server generates a payment approval result object based on the final judgment result. This object includes: - Audit status field; - Brief description of the field; - Related transaction identifier field.

[0095] The server returns the payment approval result to the terminal and simultaneously records it in the data storage system.

[0096] 8. Result Recording and Iterative Utilization Module This module stores the final judgment result, payment characteristic information, and relevant environmental information as records in persistent media. The server can further perform statistical analysis based on these records to adjust the prompt statement template, optimize feature selection strategies, or use them for subsequent offline retraining.

[0097] III. Terminal-side functions and data processing When implementing this invention, the terminal mainly performs the following technical processes: 1. Data Acquisition and Local Verification The terminal collects payment information entered by the user through the input unit and performs local format checks in the processing unit, such as card number length verification, algorithm verification, and amount validity verification. By performing some preprocessing at the terminal, the consumption of server resources by illegal requests can be reduced.

[0098] 2. Encrypted communication processing The terminal communicates with the server via a network communication unit using encrypted protocols. The terminal initiates an encrypted handshake within the operating system's network stack, and serializes and encrypts the payment information after the handshake is complete. The terminal only sends the encrypted message to the server, without retaining the complete plaintext in insecure storage areas, thus cooperating with the server to achieve secure end-to-end data transmission.

[0099] 3. Results Reception and Presentation After receiving the payment review result from the server, the terminal parses the review status and explanatory text in the processing unit and displays the result to the user using clear interface elements in the display unit. The terminal can adopt different interaction strategies based on different review statuses, such as automatically refreshing the interface or guiding the user to contact customer service.

[0100] IV. User Operations and System Integration In this invention, users actively input payment information via a terminal, and the system uses this input as a trigger for a series of subsequent internal computer processes. Users do not directly interact with the generative artificial intelligence model; all processing related to prompt construction, feature extraction, and model invocation is completed on the server side, thus ensuring security and consistency.

[0101] V. Generative Artificial Intelligence Model Structure and Training Methods The generative artificial intelligence model used by the server in this invention has the following typical structure: - Input embedding layer: The server maps each symbol in the prompt statement to a vector and adds sequence information through positional encoding; - Multi-layer coding unit: Each layer includes a multi-head self-attention sub-layer and a feedforward network sub-layer. In each layer, the server performs linear transformations, matrix multiplications and non-linear activation operations on the input vector. - Output generation layer: The server performs a linear transformation on the encoded results at this layer and outputs the sequence through a probability distribution.

[0102] During model training, the server uses a large amount of labeled transaction data and improperly labeled data as training samples to perform supervised fine-tuning of the model. The server employs a loss function, such as cross-entropy loss, to measure the error between the model output and the target label, and uses optimization algorithms to update the model parameters. The server can also employ data augmentation strategies and regularization mechanisms during training to improve the model's generalization ability and reduce overfitting.

[0103] VI. Explanation of Technical Effects and Causal Relationship In this invention, the server controls the generative artificial intelligence model's reasoning through multi-stage prompt statements. Instead of performing a single judgment on transaction data, it executes a two-stage risk assessment process. In the first call, the server obtains a judgment result through prompt statements containing basic features. In the second call, the server adds the judgment result along with more refined features to the update prompt statements, thereby achieving a deeper level of comprehensive judgment within the same model structure.

[0104] This structured, multi-stage prompt statement control method, compared to the traditional approach that relies solely on fixed rules or single model calls, enables the server to automatically adjust the analysis depth for different risk levels, avoiding the use of the most complex rule set from the outset, thereby reducing the average computational burden.

[0105] Meanwhile, the server explicitly specifies the model role, output format, and evaluation benchmark in the prompt statement, making the output of the generative artificial intelligence model more focused and easier to parse, reducing the complexity of post-processing rules, and improving the overall processing speed and stability.

[0106] After each judgment, the server stores the final judgment result and payment characteristic information as a record, and uses this record information to optimize the feature extraction strategy and the logic for constructing prompt statements. Through this closed-loop feedback mechanism, the server gradually improves feature selection and prompt design, increasing the accuracy of illegitimate transaction detection and reducing the false positive rate without changing basic hardware resources.

[0107] Because the server internally organizes feature vectors and judgment results with specific data structures and calls generative artificial intelligence models in a modular manner, the system of this invention can better utilize the resources of computing nodes and graphics processing units in high-concurrency scenarios, thereby improving processing throughput and shortening response time.

[0108] VII. Optional Implementation Forms and Variations In one implementation, the server uses a single generative AI model to complete both the initial and final judgments, distinguishing the stages only through different prompts. In another implementation, the server can configure generative AI models with different parameter scales or be specially fine-tuned for the initial and final judgments respectively, to further balance processing speed and judgment accuracy.

[0109] In one implementation, the server uses only text-based prompts; in another implementation, the server can encode some structured features as specific tags and embed them into the prompts to enhance the model's perception of numerical features.

[0110] In some implementations, the server can also dynamically adjust the length and level of detail of the prompt statements based on factors such as terminal type, network conditions, and merchant category, in order to reduce unnecessary model calculations and improve overall computing efficiency.

[0111] Through the above-mentioned various implementation forms, a set of technical solutions for detecting improper transactions in the electronic payment environment is formed between the server, terminal and user. This invention improves the efficiency of feature utilization and reasoning process control by introducing a generative artificial intelligence model and prompt statement control mechanism inside the server, and achieves the technical effect of improving processing speed and judgment accuracy while ensuring security.

[0112] use Figure 11 The processing flow is explained.

[0113] Step 1: Users input payment information using a terminal. Users enter their account identifier, transaction amount, transaction time, cardholder information, and necessary authentication information on the terminal's payment interface.

[0114] Input: Raw keyboard / touch input from the user.

[0115] The terminal fills the user-input data into the local transaction data structure and forms a transaction record in memory that includes fields such as account identifier, amount, time, network address, and device identifier.

[0116] Output: A structured local payment information object (not encrypted, not transmitted).

[0117] Step 2: The terminal performs local verification of the payment information and prepares for encrypted transmission. The terminal performs format and consistency checks on the payment information object formed in step 1, including account number length verification, amount validity verification, and mandatory field integrity verification.

[0118] Input: Local payment information object.

[0119] Based on preset verification rules, the terminal performs conditional judgment operations and string length calculations on each field. If an error is found, the terminal outputs an error message on the display interface and terminates the subsequent process. If the verification passes, the terminal calls the operating system network stack to establish an encrypted connection session and generates an application layer message to be sent.

[0120] Output: The verified payment information message and the established encrypted communication session context.

[0121] Step 3: The terminal sends payment information to the server via encrypted communication. The terminal uses the established encrypted session to serialize the payment information into a request message, and then encrypts the message using an encryption algorithm before sending it to the server.

[0122] Input: Verified payment information message and encrypted session key.

[0123] The terminal performs serialization operations locally, converting the data structure into text or binary format; then the terminal calls the encryption library to perform symmetric encryption on the serialization result, generating ciphertext; finally, the terminal fragments the ciphertext into network data packets through the communication interface and sends them to the specified address of the server.

[0124] Output: The encrypted payment data packet sent to the server.

[0125] Step 4: The server receives and decrypts the payment information. The server receives encrypted data packets from the terminal at the network interface and decrypts and verifies their integrity using an encryption library.

[0126] Input: Encrypted payment data packet from the terminal.

[0127] The server invokes the decryption algorithm and session key to perform decryption operations on the ciphertext in the data packet, recovering the original application layer message. Subsequently, the server performs verification and validation of the message as well as protocol consistency checks, eliminating abnormal or tampered data. Finally, the server parses the decrypted message into an internal payment information object, which contains structured fields such as account identifier, amount, time, network address, and device identifier.

[0128] Output: Server-side structured payment information object.

[0129] Step 5: The server retrieves historical transaction information and constructs feature inputs. Based on the received payment information object, the server retrieves historical transaction records related to the account, device, network address, etc. from the data storage.

[0130] Input: The current payment information object.

[0131] The server performs query operations in the database, retrieving a set of transaction records within a predetermined time window based on keys such as account identifier, device identifier, and network address. The server then performs aggregation calculations on this set, such as calculating the average transaction amount, transaction frequency, country distribution, and historical chargeback count, and organizes these values ​​and classification information into a vector or set of payment characteristic information fields.

[0132] Output: A set of historical transaction information associated with the current transaction and the corresponding payment characteristic information.

[0133] Step 6: The server generates a prompt statement for a single determination. The server constructs a prompt statement for a decision based on payment information and payment characteristic information, and explicitly specifies the role and output format of the generative artificial intelligence model in it.

[0134] Input: Payment information object, payment characteristic information.

[0135] The server performs string concatenation and template filling operations in the processor, converts numerical features into natural language descriptions, and inserts preset task descriptions and output constraints into a unified template; an example of the prompt statement generated by the server is as follows: "You are a generative artificial intelligence model used to detect the risk of improper credit card payments."

[0136] Based on the following transaction information and common characteristics of improper payments, determine the level of improper risk for this transaction and output only one of 'high risk', 'medium risk', or 'low risk'.

[0137] Transaction information: - Amount: 9800 yuan - Trading time: 2026-01-30 23:58:12 (late night) - Country of origin for IP address: Country A - Device type: Mobile terminal - Cardholder's registered country: Country C - Average transaction amount over the past 6 months: 300 yuan - Number of transactions in the last 24 hours: 10 answer:" Output: A text message containing the transaction details, characteristic information, and output constraints as a decision prompt.

[0138] Step 7: The server calls a generative artificial intelligence model to make a judgment. The server inputs the prompt generated in step 6 into the generative artificial intelligence model to obtain a judgment result on the risk of improper transactions.

[0139] Input: The text of the judgment prompt statement.

[0140] The server first performs word segmentation or sub-word segmentation on the prompt statement, and then maps the symbols into vector embeddings. The server performs matrix multiplication and nonlinear transformation in the multi-layer self-attention network of the model, and calculates a high-dimensional representation based on attention weights. Then, the server selects the most likely output word sequence through probability distribution in the output layer, and parses the generated result into one of the predefined labels such as "high risk", "medium risk" or "low risk".

[0141] Output: A label indicating the risk of improper transactions and its corresponding original text output.

[0142] Step 8: The server generates a prompt statement for the final judgment based on the initial judgment result and characteristic information. The server uses the initial judgment result along with more detailed payment characteristic information to construct an update prompt statement for the final judgment.

[0143] Input: a judgment result label, payment characteristic information, and payment information object.

[0144] The server performs conditional checks and text construction operations in the processor. Based on the result of a check, it selects different prompt templates and embeds the risk level and abnormal characteristics (such as amount deviation, geographical location change, abnormal frequency, chargeback records, etc.) into the prompt statement in natural language. An example of the final prompt statement generated by the server is as follows: "You are now acting as the generative AI model for the final review of improper payments. Please make a comprehensive judgment based on the preliminary risk assessment."

[0145] First assessment result: High risk.

[0146] Detailed information: - The current transaction amount differs significantly from the average transaction amount per user; - The country of the current transaction is inconsistent with the country where the user registered; - Trading frequency has increased significantly in the last 24 hours; - There is a chargeback record within the last 30 days; Based on the information above, please give your final judgment: 1. Output only one of the following: 'Safe', 'Possibly improper', or 'Requires manual confirmation'; 2. Explain the main reason in no more than 40 words.

[0147] answer:" Output: The text of the update prompt statement used for the final judgment.

[0148] Step 9: The server calls a generative artificial intelligence model to make the final judgment. The server inputs the prompt generated in step 8 back into the generative artificial intelligence model to obtain the final judgment result and explanation of the reasons for the risk of improper transactions.

[0149] Input: The text of the final judgment prompt.

[0150] The server performs the same word segmentation, embedding, self-attention calculation, and output generation operations as the first judgment, but because the input contains the first judgment result and richer feature descriptions, the model internally performs a deeper synthesis of risk patterns; the server parses the model output text, maps the natural language result to internal states such as "safe", "possibly improper" or "requires human confirmation", and extracts brief reasoning sentences.

[0151] Output: Structured final judgment result (risk status label) and brief explanation of the reasons.

[0152] Step 10: The server generates and records the payment approval result. The server generates a payment review result object based on the final judgment result and the explanation of the reason, and stores the characteristic information related to the result as record information.

[0153] Input: Final judgment result label, reason explanation text, payment characteristic information, transaction identifier.

[0154] The server constructs the audit result record structure in the processor, encodes the risk status as internal code, adds the reason explanation as a text field, and stores the associated characteristic vectors or fields in sequence form; the server calls the database interface to perform insertion operations and persist the audit results; at the same time, the server can create indexes on some fields for subsequent retrieval and statistical analysis.

[0155] Output: Audit records stored in the data system, and audit result objects that can be returned to the terminal.

[0156] Step 11: The server returns the payment verification result to the terminal. The server converts the audit result object into a response message and sends it to the terminal through an encrypted communication channel.

[0157] Input: Payment audit result object.

[0158] The server performs serialization operations, encoding the audit status and reason explanation into a response structure; the server encrypts the response content using an encryption library and sends the encrypted data packet to the terminal address via the network interface.

[0159] Output: An encrypted audit result data packet sent to the terminal.

[0160] Step 12: The terminal receives and decodes the review results, and the user can view the review status. The terminal receives the encrypted audit result data packet sent by the server, decrypts and parses it locally, and then displays it to the user.

[0161] Input: An encrypted audit result data packet from the server.

[0162] The terminal reassembles the data packets through the network stack and calls the encryption library to perform decryption operations, recovering the plaintext response message. The terminal parses the review status and reason explanation fields in the message and generates corresponding prompts on the display interface, such as "Transaction is secure, you can continue processing" or "Improperty has been detected, please contact the card issuer." The user views the status through the terminal display unit and makes subsequent action choices based on the prompts.

[0163] Output: Visualized review results presented on the terminal interface, as well as the user's awareness and response to the review results.

[0164] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0165] With the widespread adoption of electronic trading and online settlement technologies, detecting fraudulent transactions using computing devices has become a fundamental function of financial systems. However, existing fraudulent transaction detection technologies typically rely on fixed threshold rules or traditional machine learning models, which suffer from significant computer technology limitations in the following aspects: First, existing systems often employ pre-defined rule engines or single numerical feature models when conducting risk assessments on large-scale, real-time transaction data at the server side. This approach struggles to fully leverage the complex semantic relationships of users' historical behavior when dealing with cross-regional, highly volatile transaction patterns. This results in limited accuracy in server-side risk scoring calculations, leading to high false positive or false negative rates. Consequently, significant computational resources are wasted on ineffective manual review, reducing overall processing performance.

[0166] Second, traditional irregularity detection systems typically only input structured numerical features into the model at the server level, failing to effectively utilize contextual information in natural language at the system level. For complex factors such as "changes in transaction location," "deviations in amount relative to historical data," and "equipment changes," the expressive power of existing models is limited, causing the server to be unable to generate high-quality risk assessment information in a single determination. This necessitates multiple rule comparisons or manual intervention, increasing the complexity and latency of the server-side processing flow.

[0167] Third, existing systems, after receiving the model output, mostly only provide a binary judgment (pass / reject), lacking intermediate information structures that can be used to optimize subsequent calculation processes. For example, the lack of multi-level quantification of risk differentiation and the lack of linkage mechanisms with recommended actions (notification only or direct interception) cause the server to require additional logical branches when selecting decision paths, increasing the complexity of program implementation and maintenance costs, and is not conducive to maintaining stable response performance under high concurrency conditions.

[0168] Fourth, regarding user confirmation feedback, existing technologies often simply treat user feedback as a business record without systematically mapping it to model inputs (such as prompts) and corresponding transaction data. This prevents the construction of a closed-loop data chain on the server: "transaction data—natural language prompts—model output—final user confirmation." This deficiency makes it difficult for generative AI models to be incrementally trained and updated with labeled data from real-world environments in a timely manner. Consequently, the system's detection capabilities are insufficient to adapt to new irregularity patterns, impacting the overall robustness and scalability of the computing system.

[0169] Fifth, in terms of real-time performance, traditional risk control systems often distribute feature construction, model invocation, threshold judgment, and notification triggering across multiple heterogeneous subsystems, achieving coordination through message queues or batch processing. This introduces multiple network overheads and intermediate storage, resulting in significant end-to-end latency. Consequently, the server struggles to complete the entire process from "transaction occurrence" to "user receiving risk notification" in a timely manner, reducing the system's real-time response capability to large-scale concurrent transactions.

[0170] Therefore, a technical solution is needed to optimize the transaction data processing path within the server: in one or a group of servers, transaction reception, statistical feature calculation, prompt statement generation, generative artificial intelligence model invocation, graded risk judgment, notification generation, and closed-loop training of user feedback are organically integrated in a unified program flow to improve the server's computational efficiency, expressive power, scalability, and online learning ability for detecting improper transactions, thereby substantially improving the technical performance of the computer system in this type of application scenario.

[0171] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0172] In this invention, the server includes a processing component for receiving transaction information from a terminal and storing it along with corresponding user information; a processing component for calculating statistical information from historical transaction records based on the transaction information and the user information and standardizing the transaction information; a processing component for generating prompt statements based on the standardized transaction information and statistical information and inputting them into a generative artificial intelligence model; a processing component for inputting the prompt statements into the generative artificial intelligence model and obtaining a primary judgment result regarding the likelihood of the transaction information being incorrect; and a processing component for comparing the primary judgment result with a preset threshold to determine the risk level of the transaction information and, when the risk level is within a predetermined range, generating a detailed prompt statement containing multiple related transaction records and inputting it again into the generative artificial intelligence model. The model includes a processing component for obtaining a final judgment result and recommended processing content regarding whether the transaction information is incorrect; a processing component for associating and recording the initial judgment result and the final judgment result with the transaction information and determining the transaction information as a suspicious transaction based on the risk level and the final judgment result; a processing component for generating notification information containing suspicious transaction details and user confirmation options and sending it to the terminal; a processing component for receiving the user's confirmation result of the suspicious transaction from the terminal and updating the transaction information status accordingly, and sending a control request to an external settlement processing device regarding the continuation or termination of the transaction; and a processing component for associating the user confirmation result with the corresponding transaction information and prompt statements and storing it as learning data for online learning or periodic updates of the generative artificial intelligence model. This allows for automated processing within the server, encompassing everything from transaction reception, natural language prompt generation, generative AI model inference, tiered risk decision-making, to user interaction and feedback closed-loop training, all within a unified program flow. By combining structured feature computation with natural language expression capabilities, the accuracy and robustness of fraudulent transaction detection are improved. Furthermore, by reducing data round trips between heterogeneous subsystems and human intervention, end-to-end processing latency is significantly reduced, thereby improving the overall processing performance and resource utilization efficiency of the computer system in high-concurrency electronic trading scenarios.

[0173] A "system" refers to an overall technical solution consisting of one or more computing devices and programs running on them, used to perform processing flows such as receiving, storing, analyzing, reasoning, notifying, and controlling transaction information.

[0174] "Terminal" refers to an information processing device used by a user and communicating with a server, including but not limited to mobile communication terminals, portable computing devices, and fixed computing devices.

[0175] A "server" is a computing device equipped with a processor and memory and connected to terminals and external devices through a communication network, used to perform transaction information processing, irregularity detection and control processing.

[0176] "Transaction information" refers to various types of data related to an electronic transaction, including but not limited to user identification, transaction amount, currency type, transaction location, transaction time, merchant identification, and terminal identification.

[0177] "User information" refers to attribute and behavioral data related to the entity conducting the transaction, including but not limited to user identification, historical transaction records, behavioral characteristics, and preference information.

[0178] "Historical transaction records" refer to a collection of one or more transaction information that are stored in the server in chronological order and are associated with a user's past transactions.

[0179] "Statistical information" refers to aggregated data obtained through mathematical operations based on historical transaction records to characterize user transaction behavior, including but not limited to average transaction amount, maximum transaction amount, transaction frequency, and frequently used transaction locations.

[0180] "Standardization processing" refers to the preprocessing of raw transaction information by the server, such as format standardization, time and location normalization, numerical scale adjustment, and feature extraction, to facilitate subsequent model input and comparison calculations.

[0181] "Generative artificial intelligence models" refer to models that are based on machine learning and deep learning technologies, trained on large-scale data, and can automatically generate text or other forms of output based on input information, and be used to reason about the possibility of transactions being unfair.

[0182] "Prompt statements" refer to natural language text constructed by the server and used as input to a generative artificial intelligence model. The text contains at least a portion of transaction information, user information, and statistical information, and is used to guide the model in performing incorrect detection reasoning.

[0183] "Initial judgment result" refers to the result output by the generative artificial intelligence model based on the prompt statement for the first time, which is used to indicate the possibility of incorrect transaction information, including at least the probability of incorrectness or risk score expressed in numerical form.

[0184] "Preset threshold" refers to one or more numerical boundaries set in advance during system design or operation, which are used to compare the results of a judgment to distinguish different risk levels.

[0185] "Risk level" refers to the transaction risk category determined by the server based on a comparison between a single judgment result and a preset threshold. It is used to indicate whether a transaction is in a low-risk, medium-risk, or high-risk range.

[0186] "Detailed prompts" refer to natural language text generated by the server when the risk level is within a predetermined range. This text contains richer contextual information, such as multiple historical transaction records related to the current transaction and user information, and is used for final judgment.

[0187] "Final judgment result" refers to the conclusion information output by the generative artificial intelligence model based on detailed prompts, which characterizes whether a transaction is irregular, including label information that classifies the transaction as irregular or normal.

[0188] "Recommended actions" refers to the system action suggestions that the generative artificial intelligence model should take for the current transaction, based on the prompt statements output by the model. These suggestions include, but are not limited to, simply sending a notification to the user or directly terminating the transaction.

[0189] "Suspicious transactions" refer to transactions that the server identifies and marks from all transactions that have a certain probability of being suspicious, based on risk level and final judgment.

[0190] "Notification message" refers to message data generated by the server and sent to the terminal in response to suspicious transactions. It is used to alert the user to the existence of suspicious transactions and guide the user to confirm them. The message includes at least the details of the suspicious transaction and the user's confirmation option.

[0191] "User confirmation result" refers to the user's confirmation opinion on the terminal regarding a suspicious transaction, including result information such as whether the transaction is classified as a normal transaction or an improper transaction.

[0192] "Transaction information status" refers to the status identifier associated with a specific transaction record in the server, used to indicate that the transaction is in different processing stages such as pending review, confirmed as normal, confirmed as incorrect, or terminated.

[0193] "External settlement processing device" refers to an external information processing device that is connected to the server through a communication network and is responsible for actually performing payment settlement, transaction authorization, transaction cancellation or freezing, etc.

[0194] A "control request" is an instruction sent by the server to an external settlement processing device based on the user's confirmation and risk assessment, used to request the execution of continued processing, termination processing, or other control operations on the target transaction.

[0195] "Learning data" refers to one or more sets of data stored on a server for training or updating generative artificial intelligence models, containing relationships between transaction information, prompts, model outputs, and user confirmations.

[0196] "Online learning" refers to a generative artificial intelligence model that updates its parameters based on new learning data during system operation to improve its ability to recognize the latest incorrect patterns.

[0197] "Periodic updates" refers to a model update method in which generative artificial intelligence models are retrained or fine-tuned in batches based on accumulated learning data at predetermined time intervals in order to improve overall detection performance.

[0198] In various embodiments of the present invention, the server, terminal, and user each undertake different technical functions. The following description focuses on the server's program structure, data structure, the composition of the generative artificial intelligence model, the learning method, and the linkage with external settlement processing devices to describe the embodiments of the present invention. Those skilled in the art can make appropriate substitutions to the hardware and software environment without departing from the spirit of the present invention.

[0199] I. System Hardware and Software Composition In one embodiment, the server is implemented using general-purpose server hardware. The server may include a multi-core central processing unit (e.g., an x86_64-based processor), main memory (e.g., 64GB or more of random access memory), solid-state storage (e.g., an SSD array), and a network interface controller (e.g., a gigabit or 10-gigabit Ethernet adapter). The server connects to multiple terminals and external settlement processing devices via a communication network.

[0200] On the software side, the server can run a general-purpose operating system, such as a server operating system based on the Linux kernel. Application server software is then deployed on this operating system, such as a web server and application runtime environment based on a reverse proxy component (e.g., Nginx and Python runtime environment, Node.js runtime environment, etc.). The server further deploys database management software, such as a relational database management system, to store transaction information, user information, statistical information, model input / output logs, etc.

[0201] The server also deploys inference and training services for generative AI models. In one example, the server can use a model runtime environment based on a deep learning framework, such as an inference server based on PyTorch or TensorFlow. The server loads a pre-trained generative AI model into this inference server, inputs prompts to the model via an internal API call, and receives outputs regarding the likelihood of fraudulent transactions.

[0202] In one embodiment, the terminal can be a mobile communication terminal, a portable computing device, or a fixed computing device. The terminal runs a mobile operating system or a desktop operating system and carries applications with payment and notification display functions. The terminal interacts with the server via a secure communication protocol (such as HTTPS).

[0203] In an embodiment of the present invention, the user initiates an electronic transaction through a terminal and, upon receiving a suspicious transaction notification from the server, manually confirms or denies the suspicious transaction.

[0204] II. Server-side program modules and data structures In one implementation, the server divides its functionality into multiple processing components. Each component can be physically implemented by one or more software modules, or logically divided into different program units.

[0205] 1. Server transaction receiving and storage components The server has a transaction receiving component. It receives network requests from terminals and parses the transaction information sent by the terminals. After parsing, the server writes the transaction information and corresponding user information into the transaction record table and user information table of a relational database.

[0206] The server can establish a unique identifier field for each transaction in the transaction record table, such as a transaction identifier. The server can also set fields for amount, currency type, origin location, origin time, terminal identifier, merchant identifier, risk score, risk level, model's initial judgment result, model's final judgment result, recommended processing content, user confirmation result, and transaction status in this record table.

[0207] By employing a unified data formatting and indexing structure design for the aforementioned fields, the server reduces query latency and improves batch scanning efficiency. This data structure design enables the server to quickly retrieve historical data related to the current transaction in high-concurrency transaction scenarios, thus providing a foundation for subsequently generating statistical information and prompt statements.

[0208] 2. Server statistical information calculation and standardization components In one implementation, the server includes a statistical calculation component. Within this component, the server reads historical transaction records over a specific period from a transaction log table, each associated with a specific user identifier. Through database queries and in-memory calculations, the server obtains statistical information such as average transaction amount, maximum transaction amount, number of transactions within a given time window, and a list of frequently used transaction countries or regions.

[0209] During standardization, the server can convert transaction amounts into relative values ​​at the individual user scale, for example, by dividing the current amount by the average amount over the past N days to obtain the amount ratio. The server can also map the transaction location to a standardized country code and corresponding geographic area code through a lookup table, and uniformly convert the local time to Coordinated Universal Time (UTC).

[0210] Through this standardization process, the server transforms raw transaction information in heterogeneous formats into a unified numerical space that is easy to compare and model. This process reduces the model's sensitivity to input scale, thereby improving the stability and generalization ability of the model's inference.

[0211] 3. Server prompt statement generation component In one implementation, the server includes a prompt message generation component. By reading current transaction information and corresponding user statistics, the server embeds this information into a natural language template, automatically generating prompt messages suitable for processing by a generative artificial intelligence model.

[0212] The server can generate a relatively short prompt during a single determination phase, for example: User ID: 12345, Current transaction amount: $1,000, Transaction location: France, Transaction time: 2026-01-30 10:00 UTC.

[0213] In the past 30 days, all of this user's transactions have taken place in Japan, with an average transaction amount of $80 and a maximum transaction amount of $150.

[0214] Based on the information above, estimate the probability that this transaction is fraudulent, ranging from 0 to 1.

[0215] Please output only a single number; do not provide any other explanation. When the risk level is medium to high, the server can generate more detailed alerts, which include multiple historical transaction records, for example: "You are an assistant for detecting irregular settlements in banks."

[0216] Current transaction User ID: 12345 - Amount: $1000 Location: France - Time: 2026-01-30 10:00 UTC The last 10 transactions 1. Location: Japan, Amount: US$60, Time: 2026-01-28 09:00 UTC 2. Location: Japan, Amount: US$90, Time: 2026-01-27 14:30 UTC ... Statistical information - Frequently traded with: Japan - Average amount over the past 30 days: $80 - Largest amount in the past 30 days: $150 - History of irregular transactions: 0 Please determine whether the current transaction is an invalid transaction and provide a JSON output containing three items: 1) "label": "fraud" or "legit" 2) "action": "notify_only" or "block_transaction" 3) "reason": Briefly explain the reason. The server uses natural language prompts to transform structured numerical information scattered across multiple tables into a unified semantic context, thereby stimulating the sequence modeling and cross-feature combination reasoning capabilities of generative AI models. This prompt generation method does not simply display data, but rather uses specific template arrangements and explicit comparisons of distribution differences to enable the model to notice non-linear relationships such as "location changes," "amount jumps," and "frequency anomalies," thus improving the accuracy of error detection.

[0217] III. Structure and Learning Methods of Generative Artificial Intelligence Models In one implementation, the server deploys a generative AI model in a local inference environment. The model can employ a transformer architecture based on a self-attention mechanism. The server can configure a multi-layer encoder structure in this model, with each layer including a multi-head self-attention sublayer and a feedforward network sublayer, which are combined through residual connections and normalization operations.

[0218] On the model input side, the server decomposes the text in the prompt statements into sub-word units and maps them into fixed-dimensional vectors through an embedding layer. The server can also add positional and segmental encodings to the embedding layer to provide sequence order and positional information for different semantic segments. On the model output side, the server generates the corresponding text through a decoding structure, such as risk score values ​​or text containing tags and recommendation processing content.

[0219] The server can employ supervised learning methods during training. It constructs a training sample set where each sample includes: a prompt statement, its corresponding ground truth label (e.g., incorrect / normal), a numerical rating, and a recommended action. During training, the server defines an error function, such as a weighted combination of classification and regression errors. The classification error can use the cross-entropy loss function, and the regression error can use mean squared error or absolute error.

[0220] During training, the server updates model parameters using backpropagation and stochastic gradient descent optimization algorithms (such as the Adam optimization algorithm). The server can employ batch training, iteratively updating model weights with mini-batch samples, and uses techniques such as learning rate decay and weight regularization to reduce the risk of overfitting.

[0221] In one implementation, the server associates the user confirmation result with the prompt statement and transaction information, forming learning data containing a complete data chain. The server can extract newly labeled samples from the learning data table at predetermined time intervals to periodically retrain or fine-tune the model. This online or periodic update mechanism enables the model to dynamically adapt to newly emerging error patterns, thereby improving detection accuracy and robustness over long-term operation.

[0222] IV. Server Judgment Logic and Decision Rules In one implementation, the server not only uses the output of a generative artificial intelligence model as a reference, but also defines a set of unconventional hierarchical thresholds and decision rules. After obtaining a numerical score in a judgment phase, the server compares the score with multiple preset thresholds to classify it into low-risk, medium-risk, and high-risk levels.

[0223] When the server is in a low-risk zone, it can directly mark the transaction as normal, avoiding unnecessary model calls and notifications, thus reducing computational and communication load. When the server is in a medium-risk zone, it generates detailed prompts and calls the model again for final judgment. When the server is in a high-risk zone, it can directly trigger strong alerts or pre-freeze requests according to the strategy.

[0224] Through this multi-level decision-making structure, the server combines the output of the generative artificial intelligence model with rule-based judgments, enabling the system to reduce resource consumption and improve overall processing speed while ensuring accurate identification of high-risk transactions. This hierarchical judgment is not a simple manual rule, but a dynamic control logic designed around the numerical characteristics of the model's output, thereby achieving adaptive balancing of computational load in high-concurrency environments.

[0225] V. Interaction Patterns Between Terminals and Users In one implementation, the terminal receives notification information sent by the server via a push service. Upon receiving a notification about a suspicious transaction, the terminal displays the notification in a structured format on the interface, including the transaction amount, location, time, and necessary risk explanations. The terminal also provides the user with interactive options, such as "Was this your operation?" and "Was this not your operation?".

[0226] The user makes a selection on the terminal based on the actual situation. The terminal packages this selection along with a transaction identifier into a request and sends it to the server.

[0227] After receiving the user's confirmation, the server writes the result to the transaction record table and updates the transaction status and control requests to the external settlement processing device based on the result. The server also stores the confirmation result along with the prompt statement that triggered it, for use in subsequent model training.

[0228] VI. Technical Integration with External Settlement Processing Devices In one implementation, the server communicates with an external settlement processing device via an application programming interface (API). When the server determines that a transaction is illegitimate or high-risk, it can send a control request to instruct the external device to freeze, cancel, or further review the transaction.

[0229] When orchestrating control requests, the server can utilize transaction status and risk information fields to generate request messages containing necessary instructions and parameters. After processing, external devices can return status information, which the server stores to form a complete audit chain.

[0230] Through this linkage with external devices, the system of the present invention not only logically marks transaction risks, but also plays a direct control role in the actual flow of funds, demonstrating a close coupling with the physical settlement process in the real world.

[0231] VII. Technical Effects and Improvements in Computer Technology In the above implementation, the server improves computer technology itself in the following ways: 1. By introducing standardized processing and statistical feature calculation, the server compresses high-dimensional, heterogeneous transaction data into a unified numerical space and conveys these features in a structured manner through natural language prompts. This enables generative artificial intelligence models to learn more fully the nonlinear relationships across fields, thereby improving the accuracy of error detection and reducing false positives and false negatives.

[0232] 2. The server combines model inference results with rule judgments through hierarchical threshold decision logic, quickly releasing low-risk transactions and reducing redundant calculations, while performing in-depth inference on medium- and high-risk transactions. This optimizes the overall allocation of computing resources and improves the system's throughput and response speed in high-concurrency scenarios.

[0233] 3. By forming a closed-loop learning data system with user confirmation results, prompts, and transaction information, the server constructs a complete supervision chain from input text and model output to real labels. This allows the generative AI model to be updated online or periodically, continuously adapting to new misconduct patterns. This learning method differs from the traditional approach of only manually reviewing records; it is a technical mechanism for optimizing the model's internal parameters, directly impacting the quality of model inference.

[0234] 4. By specially designing the data structure and module division, the server enables the processing of transaction reception, statistical calculation, prompt statement generation, model inference, result recording, notification sending and control requests to form an ordered data flow within the same computing system. This reduces the communication and format conversion overhead between heterogeneous systems, thereby reducing latency and improving stability from a system architecture perspective.

[0235] 5. In terms of model training, the server employs explicit error function design and parameter update strategies, enabling generative AI models to converge quickly with smaller training batches while maintaining sensitivity to a small number of outlier samples. This training configuration directly impacts the model's inference latency and memory usage in a production environment, thus providing technical support for large-scale real-time trading scenarios.

[0236] VIII. Alternative Implementation Forms and Expansion In other implementations, the server can employ different types of generative AI models. For example, the server can use a hybrid structure consisting of an encoder and a classification head to encode the prompts into latent vectors and directly output the probability of incorrectness and the recommended action; it can also use a model that supports multimodal input, simultaneously receiving numerical feature vectors and text prompts, and fusing the two types of information through a multi-head cross-attention mechanism.

[0237] The server can adopt an adaptive adjustment strategy for threshold settings, such as dynamically adjusting low-risk and high-risk thresholds based on recent overall transaction volume and historical false alarm rates, thereby further optimizing the balance between computing resource utilization and detection effectiveness.

[0238] In terms of data storage, the server can distribute transaction records and learning data across multiple nodes to support larger data scales, while shortening batch processing time through distributed queries and parallel statistical computing.

[0239] Through the above-mentioned various implementation forms and alternative methods, the system of the present invention is not limited to a specific hardware and software environment, but rather substantially improves the processing speed, accuracy and resource utilization efficiency of computer technology in the field of unfair transaction detection through specific data structures, prompt statement generation methods, generative artificial intelligence model structures and closed-loop learning mechanisms.

[0240] use Figure 12 The processing flow is explained.

[0241] Step 1: The terminal receives user input and generates transaction information. The terminal receives the transaction amount, selected payment method, and transaction location information obtained by the location module from the user's input in the application interface, and generates a transaction information object when the user clicks the payment button. The input consists of the amount entered by the user on the interface, the payment method, and the time and location information provided by the operating system. The output is structured transaction information containing fields such as user identifier, transaction amount, currency type, transaction location, transaction time, terminal identifier, and merchant identifier. The terminal serializes these fields into JSON data in preparation for sending it to the server.

[0242] Step 2: The terminal sends transaction information to the server. The terminal uses the HTTPS protocol via a secure communication channel to send the JSON-formatted transaction information generated in step 1 as a request body to the interface address provided by the server. The input is a structured transaction information object, and the output is an HTTP request message transmitted over the network. After sending, the terminal waits for the server to return a synchronous response and displays a "Processing" status on the interface.

[0243] Step 3: The server receives the transaction information and parses its format. The server receives HTTP requests from the terminal and forwards them to the application via a web server component. The input is an HTTP request message containing a JSON string, and the output is a data structure representing transaction information (e.g., a collection of key-value pairs) in memory. The server parses the JSON string, validating the existence of required fields and the validity of their data types and value ranges; if the validation passes, the server writes each field into the internal data structure. Data processing includes string parsing, type conversion, and preliminary validity checks.

[0244] Step 4: The server stores the original transaction records. The server writes the transaction information parsed in step 3 into the transaction record table of the database. The input is a transaction information data structure in memory, and the output is a newly added transaction record in the database, generating a unique transaction identifier. The data processing performed by the server involves mapping each field to a column in the database table, storing the time and location fields in their raw form initially, and initializing the transaction status to pending review.

[0245] Step 5: Standardized processing of server execution time and location The server reads the newly inserted transaction record from the database and standardizes its time and location. The input is a transaction record containing the original time and location; the output is an updated transaction record supplemented with standardized time fields (unified to UTC time) and standardized location fields (e.g., country code, region code). The server converts local time to a unified time using a timezone conversion algorithm and converts free-text locations to standard codes using a pre-built location mapping table, thus achieving consistent processing of time and location data.

[0246] Step 6: The server reads the user's historical transaction records and calculates statistical information. The server retrieves a user's historical transaction records from the transaction record table within a predetermined time window (e.g., the past 30 days) based on the user's identifier in the transaction records. The input consists of the user identifier and the time window parameter. The output is a set of historical transaction records and the statistical information calculated from them, including average transaction amount, maximum transaction amount, number of transactions, and a list of frequently used transaction countries. The server obtains these statistical results by summing, finding the maximum value, and counting operations on the amount field of the historical records, and by performing frequency statistics on the location field.

[0247] Step 7: Server-generated structured features and standardized numerical features The server takes current transaction records and statistical information as input to construct a numerical feature vector for risk analysis. Inputs include fields such as current transaction amount, standardized location, standardized time, average amount, maximum amount, and number of transactions. Outputs are a set of numerical features, such as the ratio of the amount to the average amount, whether it represents a frequently used country, and the time interval since the last transaction. The server generates these features through arithmetic operations (division and subtraction) and logical judgments (comparing whether a transaction is listed in the frequently used country list) to reflect the degree of deviation between the current transaction and historical patterns.

[0248] Step 8: The server generates a prompt statement for a single determination. The server generates natural language prompts based on current transaction and statistical information, which are then input into the generative AI model. The input consists of structured transaction and statistical information, and the output is a text prompt. The server embeds the user ID, current amount, transaction location, transaction time, and the average and maximum amounts over the past 30 days into a predefined natural language template, generating content in the following format: User ID: 12345, Current transaction amount: $1,000, Transaction location: France, Transaction time: 2026-01-30 10:00 UTC.

[0249] In the past 30 days, all of this user's transactions have taken place in Japan, with an average transaction amount of $80 and a maximum transaction amount of $150.

[0250] Based on the information above, estimate the probability that this transaction is fraudulent, ranging from 0 to 1.

[0251] Please output only a single number; do not provide any other explanation. This step involves formatting numeric fields as strings and concatenating them into continuous natural language text in a specified order.

[0252] Step 9: The server calls a generative artificial intelligence model to make a judgment. The server inputs the prompt obtained in step 8 into the inference interface of the generative AI model. The input is a natural language prompt text, and the output is a numerical score (e.g., a decimal between 0 and 1) representing the probability of incorrectness. The server first segments or decomposes the prompt into sub-word units, and feeds them into the self-attention layer of the transformer structure through embedding mapping and positional encoding. After multiple layers of matrix multiplication, non-linear activation, and normalization operations, a context vector representing the complete semantics is obtained. At the output end, the server converts the context vector into a numerical prediction through a specific decoding head, and parses the string form of the model output into a floating-point number.

[0253] Step 10: The server records the judgment result and determines the risk level. The server writes the numerical score returned by the generative AI model as a judgment result to the database and performs risk classification based on preset thresholds. The input is the initial judgment score and preset threshold parameters; the output is the updated transaction record, including the initial judgment score field and the risk level field. The server performs a calculation by comparing the score with low-risk and high-risk thresholds. If the score is lower than the low threshold, it is marked as low-risk; if it is higher than the high threshold, it is marked as high-risk; and if it falls between the two, it is marked as medium-risk. The server writes this information back to the database for subsequent processing branch control.

[0254] Step 11: The server generates detailed alerts for medium- to high-risk transactions. When the risk level is medium or high, the server extracts more historical transaction details from the database, such as the most recent transaction records, and generates detailed prompts based on these records for final judgment. The input includes the current transaction record, statistics, and multiple historical transaction records; the output is detailed natural language text containing the current transaction, a list of historical transactions, and statistics. The server formats the location, amount, and time of each historical transaction into text entries, embedding them as a list into the prompts, generating, for example: "You are an assistant for detecting irregular settlements in banks."

[0255] Current transaction User ID: 12345 - Amount: $1000 Location: France - Time: 2026-01-30 10:00 UTC The last 10 transactions 1. Location: Japan, Amount: US$60, Time: 2026-01-28 09:00 UTC 2. Location: Japan, Amount: US$90, Time: 2026-01-27 14:30 UTC ... Statistical information - Frequently traded with: Japan - Average amount over the past 30 days: $80 - Largest amount in the past 30 days: $150 - History of irregular transactions: 0 Please determine whether the current transaction is an invalid transaction and provide a JSON output containing three items: 1) "label": "fraud" or "legit" 2) "action": "notify_only" or "block_transaction" 3) "reason": Briefly explain the reason. This step combines multiple data records to form a highly informative context, allowing the model to make more detailed judgments about time series patterns and behavioral biases.

[0256] Step 12: The server calls a generative artificial intelligence model to make the final judgment. The server uses the detailed prompt text generated in step 11 as input to call the generative AI model again to obtain the final output regarding whether the current transaction is incorrect and the recommended action. The input is the detailed prompt text, and the output is a text result containing a label (incorrect or normal), a recommended action (notification only or blocking), and a reason. The server parses the text generated by the model, extracts field values ​​through string lookup and structural parsing to obtain the final judgment label and recommended action, and converts this content into structured data and writes it back to the database. Internally, the server uses the model's classification head and text generation head to perform linear transformation and probability normalization on the embedded context vector to determine the incorrect label and generate the suggested action.

[0257] Step 13: The server updates the transaction status based on the final judgment result. The server updates the transaction status field based on the final judgment tag and recommended action. The inputs are the final tag, recommended action, and the current transaction record; the output is the transaction record with the updated status. Depending on whether the tag is incorrect or normal, the server sets the status to "Confirmed Incorrect," "Confirmed Normal," or "Requires Notification for Review," etc. If the recommended action is to block the transaction, the server sets the status to "Pending Freeze" or "Freeze Requested," and records the reason text. This data update operation is completed through database update statements, technically forming the basis for subsequent business control and auditing.

[0258] Step 14: The server generates a notification message and sends it to the terminal. When a user needs to be notified, the server generates a notification text containing details of the suspicious transaction and a confirmation option for the user. The input is the current transaction record (including fields such as amount, location, time, risk level, and reason), and the output is a notification message that can be displayed on the terminal. The server converts the transaction amount and location into a user-readable format, condenses the reason generated by the model into a brief explanation, and sends the notification to the bound terminal via a push service interface. The notification includes a transaction identifier to facilitate subsequent requests for details from the terminal.

[0259] Step 15: The terminal receives the notification and displays the transaction details. The terminal receives notification data sent by the server from the push service. The input is the notification message forwarded by the push service, and the output is a suspicious transaction alert displayed on the terminal interface. After the user clicks the notification, the terminal retrieves detailed transaction information by calling the server's query interface, displaying the transaction amount, location, time, and risk warning in text format, and providing interactive buttons for "This was your operation" and "This was not your operation." The terminal determines the record to query by parsing the transaction identifier in the notification.

[0260] Step 16: Users confirm suspicious transactions on the terminal. After observing the transaction details provided by the terminal, the user selects an interaction option based on their memory and actual actions. The input consists of the transaction details displayed on the terminal and system prompts; the output is the user's click action on the interface (confirmation successful or incorrect). The user makes a selection via the touchscreen or other input device, and the terminal prepares to send this selection as a confirmation result to the server.

[0261] Step 17: The terminal sends the user's confirmation result to the server. After receiving the user's selection, the terminal constructs a request message containing the user identifier, transaction identifier, and confirmation result, and sends it to the server's confirmation processing interface via HTTPS. The input is the user's selection result and the transaction identifier, and the output is an HTTP request containing the confirmation result. After sending, the terminal can wait for a server response to update the interface display status (e.g., displaying "Submitted for Review" or "Transaction Frozen").

[0262] Step 18: The server receives the user's confirmation and updates the transaction status. The server receives confirmation requests from the terminal and parses the transaction identifier and confirmation type. The input is the transaction identifier and confirmation result from the HTTP request, and the output is the updated transaction record and learning data record in the database. The server writes the user's confirmation result to a designated field in the transaction record table and changes the transaction status to "confirmed normally" or "confirmed incorrectly" based on the confirmation result. The server thus completes the final update of the transaction status.

[0263] Step 19: The server sends a control request to the external settlement processing device. When a transaction is deemed illegitimate by the user or needs to be blocked based on a recommended action, the server constructs a control request message and sends it to the external settlement processing device. The inputs are the amount from the transaction record, merchant information, and control instructions (continue, freeze, cancel, etc.). The output is the control request sent to the external device. The server interacts with the external system through a network interface, requesting to freeze the transaction or execute a refund, and upon receiving a response from the external system, records the processing result in the transaction record to form a complete technical execution chain.

[0264] Step 20: The server associates transaction data with prompts and confirmation results as learning data. The server will display a message indicating that the transaction has been completed. The initial judgment result, the final judgment result, and the user confirmation result are treated as a complete sample and written into the learning data table. The input consists of the prompt text, the model output, the user confirmation label, and the corresponding transaction identifier; the output is a new record in the learning data table. Through this data processing, the server transforms the original business transactions into training samples that can be used for supervised learning. In subsequent training, these samples are used to calculate the loss function and update the model parameters, thereby improving the accuracy and stability of the generative AI model in the error detection task.

[0265] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0266] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0267] As applications using generative artificial intelligence models to automatically judge, review, and inspect image and text information become more widespread, the following problems commonly exist in existing technologies.

[0268] First, existing systems mostly adopt fixed rules or fixed model structures to make single judgments on input data. They do not dynamically adjust the judgment process and model resource usage according to the complexity or probability of anomalies in the input data itself. This results in the use of high-load and high-computation models for processing even in scenarios with a large amount of normal data, which leads to waste of computing resources, increased response latency, and difficulty in meeting the real-time requirements of large-scale online services.

[0269] Second, existing judgment systems that utilize generative artificial intelligence models typically only input raw data directly into the model and let the model produce results. They lack a mechanism to automatically construct targeted prompts based on internal representations and cannot adaptively generate prompts for more refined analysis based on a single judgment result. As a result, the model input lacks the constraints of task context and judgment objectives, which can easily lead to unstable judgment results, poor interpretability, and difficulty in ensuring the accuracy and consistency of the review.

[0270] Third, existing technologies, after generating review results, mostly output them in the form of simple labels or numerical scores, without making full use of generative artificial intelligence models to generate natural language interpretations of structured review result data. This results in the final explanations presented to users being too brief, too technical, or difficult to understand. Users need to rely on human professionals for secondary interpretation, which increases communication and maintenance costs.

[0271] Fourth, under the existing architecture, the invocation of generative artificial intelligence models is usually single and static. That is, regardless of the result of a judgment, the same model of the same size and accuracy is invoked to perform subsequent analysis. This approach cannot dynamically balance and optimize between "processing time / resource consumption" and "judgment accuracy", making it difficult to improve system throughput and resource utilization efficiency while ensuring the quality of review.

[0272] Therefore, a new system architecture and processing flow are needed that can: (1) Convert the image and text information from the user terminal into an internal representation that can be processed by the generative artificial intelligence model; (2) Based on this internal representation, different levels of prompt statements for the first and final judgments are automatically constructed to make the model input more in line with the task requirements. (3) Based on the probability index of anomalies in a single judgment, adaptively select whether to perform a final judgment with a higher load, and call the appropriate generative artificial intelligence model during the final judgment. (4) Based on the generated structured review result data, further utilize generative artificial intelligence models to automatically generate user-oriented natural language descriptions, thereby improving the interpretability and user experience of the system without relying on a large amount of human intervention. To address the technical problems existing in current technologies, such as wasted computing resources, slow response, insufficient interpretability of judgment results, and difficulty in understanding by users, this technology aims to achieve refined control over the process of calling generative artificial intelligence models and improve the overall performance of computer technology.

[0273] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0274] In this invention, the server includes means for receiving information sent by a user from a user terminal, acquiring the information as image information or text information, and converting the image information and text information into an internal representation that can be processed by a generative artificial intelligence model; means for dynamically generating a prompt statement based on at least the internal representation, for outputting a first-order judgment content, whether an anomaly exists, and an anomaly probability index, and inputting the prompt statement and the internal representation into the generative artificial intelligence model to perform a first-order judgment process; means for adaptively adopting the result of the first-order judgment process as the final judgment content without detailed analysis, or generating a prompt statement containing detailed final judgment content including anomaly type, anomaly degree, and judgment reason based on the result of the first-order judgment process and the internal representation when detailed analysis is required, and inputting the prompt statement, the result of the first-order judgment process, and the internal representation into the generative artificial intelligence model to perform the final judgment process; and means for... The device generates structured review result data indicating the presence, type, and degree of anomalies based on the result of the final judgment process, converts the review result data into a user-viewable format, and sends it to the user terminal. The device further includes: a summary of the judgment content and anomaly type contained in the review result data based on the result of the final judgment process; a prompt statement using the summary result as explanatory input information; inputting the prompt statement and the explanatory input information into a generative artificial intelligence model to automatically generate user-oriented natural language explanatory text; associating the explanatory text with the review result data and sending it to the user terminal; and a device for switching the processing load or processing precision level of the generative artificial intelligence model used in the final judgment process based on an anomaly probability index in the result of the first judgment process, such that a low-load judgment process is selected when the anomaly probability index is less than a predetermined threshold, and a high-precision judgment process is selected when the anomaly probability index is greater than or equal to the predetermined threshold. This enables hierarchical control and dynamic scheduling of the generative AI model invocation process on the server side. On the one hand, by separating the initial judgment from the final judgment and adaptively switching the model load level, unnecessary high-load inference in large-scale normal data scenarios is reduced, overall processing time and computing resource consumption are lowered, and the throughput and response performance of information processing devices are improved. On the other hand, by automatically generating targeted prompts based on internal representations and automatically generating natural language explanatory text based on structured review results data, the model inference process is made more focused on the task objective and the interpretability and user comprehensibility of the judgment results are significantly improved. Thus, at the computer technology level, the performance and practicality of generative AI models in judgment and review business are improved.

[0275] A "system" refers to an entire system consisting of one or more information processing devices, storage devices, and communication devices with user terminals, used to perform a series of functions such as data reception, processing, judgment, and result output.

[0276] "Information processing device" refers to an electronic device including a processor, memory and communication interface, capable of parsing, converting, calculating and controlling the flow scheduling of received data, and is a hardware and software integrated entity used to realize the various functions of this invention.

[0277] "User terminal" refers to an electronic device operated by a user that can send and receive data with an information processing device, including but not limited to mobile terminals, fixed terminals or other computing terminals, used to send information to be processed and receive review results.

[0278] "Image information" refers to visual information represented in the form of bitmaps, vector graphics, or other graphic data, including but not limited to photographs, scans, screenshots, and other static or frame-stripped image data.

[0279] "Textual information" refers to text data consisting of character data that can be read and parsed by humans or machines, including but not limited to natural language sentences, symbol strings, tokenized text, and extracted document content.

[0280] "Internal representation" refers to the data expression form obtained by an information processing device based on image or text information through preprocessing, feature extraction, or encoding operations. This data form is suitable as input to generative artificial intelligence models, such as vectors, tensors, or embedded representations.

[0281] "Generative artificial intelligence models" refer to artificial intelligence models trained based on machine learning algorithms that can automatically generate output content based on input data and prompts, including but not limited to text generation models, image generation or analysis models, and multimodal generation models based on neural networks.

[0282] "Prompt statements" refer to input text or instructions designed to guide generative artificial intelligence models in performing specific tasks. They describe the judgment target, output format, or key points of interest, thereby guiding the model to generate expected output results.

[0283] "One-time judgment processing" refers to the process by which an information processing device inputs its internal representation and corresponding prompts into a generative artificial intelligence model to perform preliminary analysis on the input data and output results such as whether an anomaly exists and an index of the probability of an anomaly.

[0284] "Final judgment processing" refers to the process by which, based on the judgment result of a previous judgment, the information processing device decides whether to conduct a more detailed analysis and obtains detailed information including the type of anomaly, the degree of anomaly, and the reason for the judgment by inputting more detailed prompts and related data into the generative artificial intelligence model.

[0285] "Anomaly Probability Index" refers to the numerical value or level output by the generative artificial intelligence model during the judgment process, which is used to quantify the probability of anomalies in the input data, to indicate the strength of the anomaly tendency and to serve as the basis for subsequent process control.

[0286] "Abnormality type" refers to the category or type of abnormality identified during the judgment process. It is used to classify and describe the nature of the abnormality, such as appearance defects, non-compliance of content, structural abnormalities, etc.

[0287] "Abnormality level" refers to the result of quantifying or classifying the severity of detected abnormalities, used to indicate the degree of impact of the abnormality on the quality or compliance of the target object, such as minor, moderate, severe, or corresponding numerical values.

[0288] "Review result data" refers to data information generated based on the output of the final judgment process, which represents the judgment conclusion in a structured form. It includes at least whether an anomaly exists, the type of anomaly, the degree of anomaly, and related judgment information.

[0289] "Structured review results data" refers to review results data organized in a predefined data structure that facilitates storage, retrieval, and program processing, including but not limited to key-value pair records, tabular records, or hierarchical data structures.

[0290] "Natural Language Explanatory Text" refers to user-oriented explanatory text content written in natural language. It is generated by a generative artificial intelligence model based on review result data and related prompts, and is used to provide an easy-to-understand description of the judgment conclusions and their reasons.

[0291] "Processing load" refers to the amount of computing resources required by a generative artificial intelligence model when performing inference or generation tasks, including but not limited to processing time, processor utilization, storage usage, and bandwidth usage.

[0292] "Processing accuracy level" refers to the accuracy and detail of the output results of a generative artificial intelligence model in judgment or generation tasks. Different levels can be distinguished by differences in model size, number of parameters, inference strategy, or post-processing method.

[0293] "Preset threshold" refers to a benchmark value or level that is pre-set in the system for comparing anomaly probability indicators. When the anomaly probability indicator is compared with this benchmark, it is used to determine the subsequent processing path or model selection.

[0294] In one embodiment of the present invention, a server is configured as an information processing device in a computer system. This computer system may include a multi-core central processing unit, a graphics processing unit, a large-capacity main memory, and non-volatile storage media. The server can run a general-purpose operating system and deploy application frameworks, network service components, and generative artificial intelligence model inference components on it.

[0295] At the hardware level, servers can use multi-core processors with vectorized instruction sets to perform batch operations on image matrices and text vectors; at the acceleration unit level, servers can use graphics processing units to perform deep neural network inference operations to reduce the latency of matrix multiplication and convolution operations; at the storage level, servers can use high-speed cache storage structures to store frequently used model parameters and intermediate features in high-speed storage areas to reduce access latency.

[0296] At the software level, servers can use deep learning frameworks to implement generative artificial intelligence models. For example, they can use one type of tensor computation library to build neural network structures and another type of open-source deep learning library to implement the training and inference processes. At the communication level, servers can use web server software and application server software to receive requests from terminals and exchange data with terminals via a secure version of Hypertext Transfer Protocol (HTTP).

[0297] In one implementation, the server uses an image processing library to parse the image information uploaded by the terminal, decoding the image into a pixel matrix in the form of a multidimensional array, typically a floating-point tensor of height × width × color channels. The server performs data preprocessing operations such as normalization, scaling, and center cropping on this pixel matrix, and linearly maps the pixel values ​​from the original intervals to predetermined intervals according to the input requirements of the generative artificial intelligence model. For text information, the server uses a word segmentation tool or sub-word encoder to convert natural language text into sub-word units or word unit sequences, and uses an embedding matrix to map discrete indices into continuous vectors, thereby forming part of the internal representation.

[0298] In a preferred embodiment, the server employs a multimodal generative artificial intelligence model, which structurally includes an image encoding subnetwork, a text encoding subnetwork, and a cross-modal attention fusion subnetwork. In the image encoding subnetwork, the server may use a convolutional neural network or a visual transformation network to perform multi-layer convolution, pooling, or self-attention operations on the input image pixel matrix to obtain an image feature map, which is then further compressed into an image feature vector. In the text encoding subnetwork, the server may employ a transformation structure to encode the text sequence through multi-layer self-attention and feedforward networks, obtaining a context-dependent representation for each position and a sentence-level convergent representation. In the cross-modal attention subnetwork, the server inputs image features and text features into a multi-head attention module, calculates the correlation weights between different modalities, and obtains a fused feature representation, which constitutes another part of the internal representation.

[0299] When performing a decision, to ensure the generative AI model focuses on anomaly detection, the server constructs a prompt statement based on its internal representation. The server can combine the task description, output requirements, and possible anomaly types into natural language text and input this prompt statement along with the internal representation into the decoding part of the generative AI model. An example of a prompt statement that the server can use is as follows: "Based on the product image you input, please determine if there are any obvious appearance abnormalities (such as cracks, scratches, damage, contamination, etc.) and give a score for the probability of the abnormality (between 0 and 1)." Please review the following text for fraudulent or non-compliant language, and provide a preliminary assessment of 'risky' or 'generally normal', along with a risk score (between 0 and 1): "Please take into account the product images and their descriptions, and conduct a preliminary review of the sample. Determine if there are any quality issues or compliance risks, and provide an anomaly probability score." The server inputs the prompt statements, after word segmentation and embedding, into the text channel of the generative AI model. Combining image and text features from the internal representation, the model's output is constrained by the task's semantics. During each decision process, the server configures the output header to regress anomaly probability indicators and output a label indicating whether an anomaly exists. The server can use a logistic regression layer or a multilayer perceptron to receive fused features and output anomaly probability values; the server maps the anomaly probabilities to binary labels based on a predetermined threshold. Because the server explicitly injects the decision target into the model input through the prompt statements, the model's internal attention mechanism reinforces features related to anomalies, thereby improving the stability and interpretability of each decision.

[0300] After completing a judgment, the server controls the flow branches based on the probability of anomalies. If the indicator is below a predetermined threshold, the server directly uses the result as the final judgment, eliminating the need for complex subsequent reasoning and saving computational resources. If the indicator is above or close to the threshold, the server deems a more refined judgment necessary and triggers the final judgment process. In the final judgment, the server constructs more detailed prompts, such as: "Based on the previous preliminary assessment, this sample has a certain probability of being abnormal. Please conduct a more detailed analysis of the following input data, provide a final judgment (normal / abnormal), and explain the type of abnormality and the main basis for the judgment." "Please locate the specific abnormal area in the image, describe the type of defect in that area, and give a final judgment of pass / fail based on the severity." "Please analyze the following text paragraph by paragraph to determine whether there are any false promises, exaggerated benefits, or wording that violates regulations, and provide a final compliance assessment." When making the final decision, the server can choose generative AI models of different sizes: when the probability of an anomaly in a single decision is high, the server switches to a model variant with more parameters, more layers, or more attention heads in the model selection module to improve the fineness of the decision boundary; when the probability of anomalies is at a low or medium level, the server can choose a model with fewer parameters to reduce the computational load during inference. Through this indicator-based model routing mechanism, the server achieves non-habitual adaptive processing at the computation graph construction and inference scheduling levels, rather than simply calling a fixed model, thereby significantly reducing the average inference time and resource consumption while maintaining overall decision accuracy.

[0301] During the final judgment process, the server can apply additional feature extraction algorithms to the internal representation. For example, it can use attention weight visualization algorithms to identify regions with high contribution in the image, or use gradient class saliency map methods to analyze which pixel or text locations have the greatest impact on the anomaly label. Based on these features, the server generates anomaly region coordinates or anomaly text fragment indexes and records them in the structured review result data. Internally, the server can also compare the feature vector of the current sample with the feature vectors of historical normal and anomaly samples using clustering algorithms or similarity calculations, calculating metrics such as Mahalanobis distance and cosine similarity to provide a quantitative basis for anomaly severity classification.

[0302] When training generative AI models, servers can employ supervised learning, using a large amount of labeled image and text data as training sets. A loss function is used to quantify the error between the predicted output and the true label. In a single decision task, the server can use binary cross-entropy loss to classify the presence or absence of anomalies, while simultaneously using mean squared error loss for anomaly probability regression. In the final decision task, the server can use multi-class cross-entropy loss to train an anomaly classification head and sequence-to-sequence loss to train a decoder that generates anomaly description text. The server updates model parameters using gradient descent and its variants (e.g., optimization algorithms with adaptive learning rates) and employs data augmentation techniques during training, such as rotating, scaling, and color perturbations on images, and synonym replacement and random masking on text, to improve the model's robustness to input variations.

[0303] During runtime, the server does not retrain itself on user data; instead, it uses pre-trained model parameters to perform inference. When the server inputs the internal representation of user data into the model, the model performs linear transformations, non-linear activations, attention weighting, and normalization operations during forward propagation using the parameters learned during training. From a computer science perspective, the server adjusts the attention distribution through prompts, causing the model to allocate more weight to anomaly-related feature dimensions in the multi-head attention layer. This reduces the effective feature space and decreases sensitivity to irrelevant features, directly leading to a reduction in model output error and improved judgment accuracy.

[0304] After generating structured review result data, the server stores the presence or absence of anomalies, anomaly types, anomaly severity, and the location of anomaly regions or text as key-value pairs or hierarchical structures. The server assigns a unique identifier to each review record for easy retrieval and statistical analysis. The server further performs summarization processing on this structured data, extracting key fields as explanatory input information and constructing prompts for natural language generation, such as: "Based on the following review result fields, please generate a natural language description for end users, explaining whether the review passed or failed and the reasons for failure. The review result fields include: whether it is abnormal, the type of abnormality, the degree of abnormality, and the location of the abnormality." The server inputs the prompt and explanation information together into the text generation model. The text generation model can employ an encoder-decoder structure. The encoder embeds and contextualizes the explanation input information, while the decoder generates the explanation text word by word using an autoregressive mechanism. During decoding, the server can use a bundle search strategy to control the generation quality, avoiding inconsistencies or redundant content. The server appends the generated natural language explanation text to the review result and sends both together to the terminal.

[0305] After receiving the review results from the server, the terminal performs local parsing of the structured data, combining information such as review status and anomaly locations with the received natural language explanatory text for display. The terminal can highlight abnormal areas in the display module, such as drawing rectangles or semi-transparent masks on images, and using different colors to mark abnormal sentences in text. Through this visualization method, the terminal presents the complex feature analysis and judgment results from within the server to the user in an intuitive form, lowering the barrier to understanding for the user.

[0306] Users can observe the review results and explanatory text on the terminal to determine whether data needs to be re-collected, text content modified, or submitted for manual review. Users do not need to understand the deep learning details inside the server to operate based on the information presented on the terminal. Because the server internally implements adaptive judgment processes, model switching, and structured result generation, the terminal and user side only need to handle high-level result display, thus centralizing the complex data processing and computational load on the server, which is beneficial for unified system maintenance and expansion.

[0307] In another implementation, the server can employ different generative AI model architectures, such as using a decoder-only structure to process both the prompt statements and internal representations, injecting the internal representations as context-encoded vectors into the decoder's input or attention layer. In this architecture, the server can also achieve hierarchical processing of primary and final decisions by changing the prompt statements and model size. Furthermore, the server can adjust the anomaly probability threshold and model selection strategy under different business scenarios to achieve a finer-grained trade-off between performance and accuracy.

[0308] In another implementation, the server can be expanded into a multi-server cluster, with each server node loading generative AI models of different sizes or tasks. At the access layer, the server routes requests to appropriate nodes based on anomaly probability indicators and current system load, further improving system scalability and fault tolerance. Through this architecture, the server achieves fine-grained scheduling and dynamic resource allocation for generative AI model calls, thereby improving computational efficiency, judgment accuracy, and system stability at the technical level, going beyond simply automating manual review processes.

[0309] use Figure 13 The processing flow is explained.

[0310] Step 1: Users prepare and send the data to be reviewed on the terminal.

[0311] Inputs: raw image files (such as JPEG, PNG), raw text content (such as product descriptions, explanatory text), and data-related metadata (such as time, number).

[0312] Output: The request message transmitted to the server over the network.

[0313] Users select a local image or take a picture using the terminal's camera on the terminal interface, enter a text description in the text input box, and then click the "Submit for Review" button on the terminal. The terminal packages the selected image and input text along with user identifier and other metadata into a network request and sends it to the review interface provided by the server using the HTTPS protocol.

[0314] Step 2: The terminal packages and performs preliminary verification of user data before sending it to the server.

[0315] Input: Raw image and text data selected or entered by the user on the terminal.

[0316] Output: An HTTP / HTTPS request containing an image binary stream, a text string, and metadata.

[0317] The terminal reads the image file, checks its size and format (e.g., only JPEG and PNG are allowed), imposes simple restrictions on the text length, serializes this data into a request body (e.g., using multipart / form-data or application / json format), adds authentication information and content type tags to the request header, and sends the request to the specified URL of the server through the network communication module.

[0318] Step 3: The server receives the request and parses the raw data.

[0319] Input: The HTTP / HTTPS request message sent by the terminal.

[0320] Output: The original image byte array in memory, the original text string, and the extracted metadata object.

[0321] At the network layer, the server receives requests through the web server and application framework, parses the HTTP messages, reads the image and text fields in the request body, saves the image part as a byte array and the text part as a string object, and encapsulates metadata such as user ID and timestamp into structured data to provide basic input for subsequent processing.

[0322] Step 4: The server preprocesses images and text and generates internal representations.

[0323] Input: raw image byte array, raw text string.

[0324] Output: Internal representation data structures such as image feature tensors and text embedding tensors.

[0325] The server uses an image processing library to decode the image byte array into a pixel matrix. This matrix is ​​then scaled, centered, and normalized to map pixel values ​​to a specified range. The dimensions are rearranged to fit the input format of the generative AI model. The server uses a tokenizer or sub-word encoder to segment the text string into a sequence of tokens. Each token is then mapped to a continuous vector using an embedding matrix, forming an embedded representation of the text sequence. Subsequently, the server inputs the image pixel matrix into an image encoding network (such as a convolutional network or visual transformation network), extracting image feature vectors through operations such as convolution, pooling, or self-attention. The text embedding is input into a text encoding network (such as a transformation structure), where sentence-level feature vectors are calculated through multiple layers of self-attention. Finally, the image and text features are combined into a unified internal representation structure.

[0326] Step 5: The server constructs a prompt statement for a decision and prepares the input for the generative artificial intelligence model.

[0327] Input: Internal representation (image features, text features), task type information (e.g., image review only, text review only, or multimodal review).

[0328] Output: A single decision prompt text and a set of model input tensors (including prompt embeddings and internal representations).

[0329] The server automatically generates a suitable first-order judgment prompt based on the task type. For example, for image review, it generates "Please determine whether there are obvious appearance abnormalities (such as cracks, scratches, damage, contamination, etc.) based on the input product image, and give a score of the probability of abnormality (between 0 and 1)." For text review, it generates "Please review the following text for fraudulent or non-compliant language, and give a preliminary judgment of 'risk exists' or 'basically normal', and a risk score (between 0 and 1)." The server inputs this prompt into the text encoding module to obtain the prompt embedding vector, and encapsulates the prompt embedding and internal representation together into a model input tensor, which serves as the complete input for the first-order judgment inference stage.

[0330] Step 6: The server calls the generative artificial intelligence model to perform a single judgment.

[0331] Input: A single decision is made using embedded prompts, and tensors representing images and text.

[0332] Output: A single determination result, including a label indicating whether an anomaly exists, an anomaly probability index, and an intermediate feature vector.

[0333] The server feeds these input tensors into the encoding and fusion module of the generative AI model. It calculates the relevance weights between the prompt and its internal representation using a multi-head attention mechanism, assigning higher weights to feature dimensions relevant to the anomaly detection task. Then, it performs data computation using classification and regression heads at the output layer. In the classification head, the server obtains the probability distribution of anomalies / normalities through linear transformations and activation functions, selecting the category with the higher probability as the label for the presence or absence of an anomaly. In the regression head, it obtains the numerical value of the anomaly probability through linear output. The server also retains the last layer of hidden feature vectors for use in the final decision.

[0334] Step 7: The server determines whether to proceed to the final judgment based on the result of one judgment.

[0335] Input: Anomaly presence / absence label, anomaly probability index, and intermediate feature vector from a single judgment result.

[0336] Outputs: Process control decisions (direct output results or initiation of final decision), and context data for use in the final decision.

[0337] The server compares an anomaly probability index with a pre-set threshold. When the index is below the first threshold, the server marks the initial judgment as reliable and uses it directly as the final conclusion. When the index is above or close to the threshold, the server determines that further refined analysis is needed, packaging the internal representation and intermediate features along with the initial judgment label as input for the final judgment stage. This decision logic reduces unnecessary high-load reasoning, thereby achieving conditional branching and resource conservation at the data processing level.

[0338] Step 8: The server constructs a final decision using prompts and selects an appropriate generative artificial intelligence model.

[0339] Input: First-order label, anomaly probability index, internal representation, intermediate feature vector.

[0340] Output: Detailed prompt text for final decision, selected model configuration, and model input tensor for the final decision stage.

[0341] When the anomaly probability index is high, the server selects a generative AI model with larger parameter scale, more layers, or more attention heads to achieve higher accuracy in judgment. When the index is in a moderate range, a medium-sized model can be selected to balance performance and cost. The server generates more detailed prompts based on a single judgment result, such as, "Based on the previous preliminary judgment result, this sample has a certain probability of being abnormal. Please perform a more detailed analysis of the following input data, give a final judgment (normal / abnormal), and explain the anomaly type and main basis." or "Please locate the specific abnormal area in the image, explain the defect type of the area, and give a final judgment of qualified / unqualified based on the severity." The server encodes this prompt as an embedding and combines it with internal representations and intermediate features to form the complete input of the final judgment model.

[0342] Step 9: The server performs the final decision and generates detailed decision information.

[0343] Input: The final decision is based on the embedded prompt statement, internal representation, intermediate feature vector, and the selected high-precision or medium-precision model.

[0344] Output: Final judgment result, including final abnormal / normal label, abnormal type, abnormality degree, abnormal location or text fragment index, and related vector of judgment reason.

[0345] The server feeds the input tensor into the selected version of the generative AI model. It then re-analyzes the features using a deep multi-head attention and feedforward network, and simultaneously sets multiple output heads at the top layer of the model: a classification head outputs anomaly type labels (e.g., appearance cracks, contamination, structural damage), a regression head outputs anomaly severity scores (e.g., continuous values ​​from 0 to 1), and a localization head outputs the coordinates of the anomaly region or the index of the anomaly text fragment. The server uses these outputs to annotate the internal raw data, associating anomaly categories with specific locations, and internally records attention weights or contribution scores for subsequent interpretation, serving as the basis for the judgment.

[0346] Step 10: The server generates structured review result data.

[0347] Input: The final output values ​​(final label, anomaly type, anomaly degree, anomaly location, etc.), the first judgment result, and metadata.

[0348] Output: Structured audit result data records (e.g., key-value pairs or hierarchical data structures).

[0349] The server organizes the judgment result fields into a unified format, including fields such as "whether it is abnormal," "type of abnormality," "degree of abnormality," "coordinates or text index of the abnormality location," "first-time judgment index," and "final judgment model version," and stores them in a database or in-memory data structure. During this process, the server performs data shaping, field mapping, and type conversion, unifying the output from different model headers into standardized internal data types for subsequent generation of natural language descriptions and output to the terminal.

[0350] Step 11: The server generates natural language explanatory text based on the structured results.

[0351] Input: Structured review result data (abnormal status, type, degree, location, etc.).

[0352] Output: User-oriented natural language description text.

[0353] The server extracts a summary from the review results data, encapsulating key fields such as "anomaly type," "anomaly location," and "anomaly severity" into explanatory input information. It then constructs a prompt statement for generating the explanation, such as, "Based on the following review result fields, please generate a natural language explanation for the end-user, stating whether the review passed or failed and the reasons for failure. Review result fields include: whether it is anomaly, anomaly type, anomaly severity, and anomaly location." The server inputs this prompt statement along with the explanatory input information into a text-based generative AI model. The model's encoding part vectorizes the field content, and the decoding part generates explanatory text word by word under attention mechanism control, such as, "This product failed the appearance inspection. Reason: A significant crack was found in the lower right corner of the screen, indicating a high degree of severity; rework or replacement is recommended." The server binds the generated explanatory text with the structured results for unified return to the terminal.

[0354] Step 12: The server sends the review results and explanations to the terminal, which then displays them.

[0355] Input: Structured review result data generated by the server, and natural language description text.

[0356] Output: The review results page (image annotations, text descriptions, etc.) displayed on the terminal interface.

[0357] The server encapsulates the structured results and explanatory text into a response body and sends it back to the corresponding terminal via HTTPS. Upon receiving the response, the terminal parses the results in JSON or other formats, displays status information such as "pass / fail" on the interface, and draws highlight boxes or masks on images based on the location of the anomaly provided by the server, while color-coding the abnormal sentences in the text. Users can read the explanations on the terminal interface and intuitively understand the reasons for the system's judgment. The terminal uses rendering components to perform specific actions such as drawing and text layout, transforming the complex data processing results from the server's internal processes into an easily understandable visual presentation.

[0358] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0359] In existing computer-based communication support technologies, free text messages sent to service recipients (such as users, customers, guardians, etc.) are usually generated by simply matching templates or searching keywords to produce responses. These technologies suffer from the following problems: First, the processing flow is mostly based on a "fixed rules + manual writing" model, which cannot dynamically adjust the response structure according to different contexts and emotional states. This results in insufficient understanding of communication information by the system, making it difficult to reflect the true intentions and emotions of the service recipients in a timely and accurate manner. Second, existing systems lack structured modeling and data accumulation for the closed loop of "prompt statements - generative AI models - human review feedback." Prompt statements are mostly statically configured and cannot be automatically optimized over time. This leads to a disconnect between the calling method of the generation model and the upstream parsing results and downstream editing results, failing to fully leverage the expressive power of the generative AI model. Third, communication information, parsing results, prompt statements, model output, and the reviewer's editing results are not uniformly managed and linked for storage. The computer system cannot reuse historical information in subsequent sessions, and therefore cannot gradually improve the prompt statement generation logic and model control strategy at the algorithm level. This limits the scalability of the system in terms of response quality, consistency, and operational efficiency.

[0360] In other words, existing technologies lack a systematic solution to the computer technology problem of "how to enable computer systems to automatically and dynamically generate prompts that adapt to different contexts and emotions, and thereby efficiently drive generative artificial intelligence models to output high-quality response documents that can be quickly reviewed by humans." A technical solution is needed that can automatically parse communication information on the server side, construct high-quality prompts, call generative artificial intelligence models to generate candidate response documents, and provide structured feedback of the reviewer's editing results for continuous optimization of prompt generation and model control. This would improve the overall efficiency of human-computer collaborative processing and the quality of responses from the perspectives of computer architecture and data processing workflows.

[0361] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0362] In this invention, the server includes a device for acquiring and parsing communication information related to the service recipient; a device for dynamically generating prompt statements input to a generative artificial intelligence model based on the intent and emotion information obtained from the parsing; a device for generating response document candidates using the generative artificial intelligence model based on the prompt statements and the communication information; a device for presenting the response document candidates to an information processing device for reviewers and allowing reviewers to perform editing operations; and a device for recording the parsing results, the prompt statements, the output results of the generative artificial intelligence model, and the reviewers' editing results in accordance with the communication information and the response document candidates, and storing them in association as data that can be reused in subsequent prompt statement generation processing and generative artificial intelligence model control processing. This allows for a closed-loop computational process within the server, encompassing "communication information parsing—dynamic generation of prompt statements—invocation of generative AI models—human review and editing—historical data feedback and optimization." The computer system can automatically adjust the structure of prompt statements based on topic and sentiment information, precisely control the output of the generative AI model, and iteratively update the prompt statement generation rules and model control strategies using accumulated related data. This improves response generation quality at the algorithm and data processing architecture level, shortens the editing time for reviewers, enhances overall human-computer collaborative processing efficiency and system resource utilization efficiency, and ultimately improves the computer technology itself.

[0363] A "system" refers to an entire system consisting of one or more information processing devices, as well as the programs, storage devices, and communication interfaces running on them, used to perform a series of data processing operations such as parsing communication information, generating prompt statements, calling generative artificial intelligence models, and processing response documents.

[0364] A "server" refers to an information processing device that provides data processing and service functions in a network environment. It may include a processor, memory, and communication interface, and is used to centrally perform functions such as parsing communication information, generating prompt statements, calling generative artificial intelligence models, and recording results.

[0365] "Communication information" refers to information content in the form of text, voice, or images sent to the system by the service recipient through a terminal, used to express inquiries, feedback, requests, or other communication intentions.

[0366] "The service recipient" refers to any party that sends communication information to the system through a terminal and receives a response from the system, such as a user, customer, or guardian, without any specific attribute restrictions.

[0367] "Analysis" refers to the process of applying natural language processing or other data processing techniques to communication information to extract intent information, sentiment information, topic information, and other semantic features.

[0368] "Intent information" refers to abstract information obtained through semantic analysis of communication information, which represents the purpose or request type that the service recipient is trying to achieve, such as inquiries, complaints, confirmations, requests for explanations, etc.

[0369] "Emotional information" refers to characteristic information obtained through emotion analysis of communication information that indicates the psychological state or emotional tendency of the service recipient, such as peace of mind, anxiety, worry, anger, satisfaction, etc.

[0370] "Topic information" refers to structured information extracted from communication information that represents the main topic or content area involved in the communication, such as learning progress, service quality, and cost issues.

[0371] "Generative artificial intelligence models" refer to data processing models trained using machine learning algorithms and large-scale data that can automatically generate text or other forms of output based on input prompts, including but not limited to deep learning models used for natural language generation.

[0372] "Prompt statements" refer to the instructional text or structured inputs given to generative artificial intelligence models. These statements specify the goals, style, constraints, and contextual information of the content generated by the model, thereby guiding the generative artificial intelligence model to produce expected outputs.

[0373] "Dynamically generated prompt statements" refers to the generation process in which the server adjusts the text content, structure, and parameters of the prompt statements as needed based on the currently parsed intent information, sentiment information, and historical correlation data, so that the prompt statements are automatically updated as the context changes.

[0374] "Response document candidate" refers to a draft text content automatically generated by a generative artificial intelligence model after receiving prompts and relevant context information, which is used to reply to the service recipient and has not been sent as a final response before being edited by the reviewer.

[0375] "Auditor information processing device" refers to a terminal or computing device used by the auditor to receive and display candidate response documents and allow the auditor to edit them, such as a work computer or mobile terminal.

[0376] "Notified party information processing device" refers to a terminal or computing device used by the service recipient to receive and display the final response document sent by the system, such as a user terminal or client device.

[0377] "Editing operations" refers to the manual rewriting actions performed by the reviewer on candidate response documents, such as adding, deleting, modifying, reorganizing, or adjusting the format, in order to form a final response document that meets the actual communication needs.

[0378] "Record and retain" refers to the behavior of a server storing communication information, parsing results, prompts, outputs of generative artificial intelligence models, and editing results from reviewers in a searchable and reusable form in a storage medium, and then retrieving this data in subsequent processing flows.

[0379] "Associative storage" refers to a storage method in which a server links communication information, parsing results, prompt statements, candidate response documents, and editing results together through identifiers or indexes when storing various types of data, so that they can be accessed and utilized uniformly at the session or event level in subsequent processing.

[0380] "Prompt statement generation rules" refer to the set of logic or parameters used to generate prompt statements from parsing results and historical correlation data, including text templates, fill strategies, sentiment adjustment strategies, and control parameters used when interacting with generative artificial intelligence models.

[0381] "Control and processing of generative artificial intelligence models" refers to the technical processing by which the server sets or adjusts the model's input format, output length, generation style, temperature parameters, sampling strategy, etc., when calling a generative artificial intelligence model, so as to affect the model's output results.

[0382] "Progressive optimization" refers to the process by which the server continuously updates the rules for generating prompt statements and the model control strategy based on historical records and related data, so that the candidate response documents generated by the generative artificial intelligence model in subsequent sessions are continuously improved in terms of quality, adaptability and editability.

[0383] In the following embodiments, the server, terminal, and user are described as subjects. This invention is not limited to specific hardware or software platforms, but for clarity, an exemplary description is given below using a general-purpose server, a general-purpose terminal, and a common software framework.

[0384] In one embodiment, the server comprises a multi-core processor, main memory, non-volatile storage, and a network interface. The server runs an application program on an operating system. This application program uses a natural language processing library (e.g., a general-purpose natural language processing library), a data processing library (e.g., a data frame processing library), and a deep learning framework (e.g., a general-purpose tensor computation framework) to perform various data processing tasks of the present invention. The server deploys generative artificial intelligence models on local computing resources (such as a central processing unit and a graphics processing unit) or invokes remote inference services via the network interface.

[0385] In one embodiment, the terminal is a portable information terminal or a desktop information terminal, equipped with a display device, an input device, and a communication module. The terminal runs application software on an operating system. This application software is used to send communication information to the server, receive candidate response documents generated by the server, display and allow users to edit them, and return the final edited result to the server.

[0386] In one implementation, the user acts as either the reviewer or the recipient of the service. The user inputs free-text messages into the server via a terminal and manually reviews and modifies the candidate response documents generated by the server. These user actions provide training signals for the server to update subsequent prompt generation rules and optimize the control strategy of the generative artificial intelligence model.

[0387] In one implementation, the server receives communication information sent by the terminal via a network interface. The server uses a natural language processing library to perform text segmentation, part-of-speech tagging, dependency parsing, and named entity recognition, thereby extracting topic information, intent information, and sentiment information from the communication information. The server stores this information as structured data, for example, recording fields such as "topic tag," "intent tag," and "sentiment vector" in a relational data store with "session identifier" as the primary key. The server further utilizes a sentiment analysis model to classify the communication information according to emotion, outputting a sentiment vector containing multi-dimensional emotion intensity values, thus enabling subsequent prompts to be fine-tuned in tone and content for different emotions.

[0388] In one implementation, the server constructs prompt statements based on the parsed intent, topic, and sentiment information. The server maintains prompt statement generation rules in storage, which can take the form of a combination of templates and parameterized placeholders, such as a paragraph structure like "respect the other party's emotions—explain the current situation—offer suggestions." The server selects the corresponding template based on the intent type of the communication information and adjusts the strength of words, the degree of reassurance, and the length of the response based on the sentiment vector. Furthermore, the server adds more precise constraints to the prompt statements based on statistical information about user (reviewer) editing of the model output in historical sessions, such as "avoid using vague wording" and "add specific examples."

[0389] A specific example of the server generating a prompt statement in one implementation is as follows: "You are a generative artificial intelligence model that helps staff respond to service recipients."

[0390] The original message from the client: 'I'm worried that my child has been losing focus in class lately.' Analysis results: Intention = concern for classroom attention, Emotion = worry.

[0391] Please generate a draft response based on the above information, with the following requirements: 1. First, understand and soothe the other person's emotions; 2. Provide a detailed description of the child's actual performance in class; 3. Propose improvements that allow staff and service recipients to work together effectively; 4. The tone should be friendly and positive, and the word count should not exceed 300 words.

[0392] Please directly output a Chinese reply text that can be sent to the service recipient. In one implementation, the server inputs the aforementioned prompts and necessary contextual information into the generative AI model. The server employs a deep neural network architecture within the generative AI model, such as a sequence-to-sequence model based on self-attention. During model training, the server uses a large historical dialogue dataset, including user comments, human responses, and related intent and sentiment labels, to perform supervised learning on the model. During training, the server sets an objective function, such as a cross-entropy loss function, to measure the difference between the model-generated text and human responses, and updates the network weights using gradient descent. During training, the server can use data augmentation techniques, such as synonym replacement, sentence transformation, and noise injection, to improve the model's robustness to diverse expressions.

[0393] In one implementation, the server receives prompts during the inference phase, converts them into a labeled sequence, and inputs it into the encoder of the generative AI model. Within the model, the server performs multi-layered self-attention computations. Each layer extracts semantic features through matrix multiplication and nonlinear transformations, and then the decoder uses beam search or sampling strategies to progressively generate response text. The server controls the diversity and stability of the output text by adjusting control parameters of the generative AI model, such as temperature parameters, penalty repetition parameters, and maximum generation length. These control parameters are automatically adjusted by the server based on historical editing data, making the output text more aligned with the reviewer's preferences, thereby reducing subsequent editing workload.

[0394] In one implementation, the server associates the generated response document candidates with their corresponding session identifiers and stores them in a data storage device. The server then sends the response document candidates to the terminal via a communication interface for the user (reviewer) to view. Upon receiving the response document candidates, the terminal displays them as text on a display device and provides a text editing interface. The user can add, delete, modify, rearrange paragraphs, and change the tone of the text on the terminal. After the user completes editing and confirms, the terminal sends the final response document and the user's editing history (such as modified locations and added content types) back to the server.

[0395] In one implementation, the server receives the edited text from the terminal and compares the original candidate response document with the edited text. The server uses text comparison algorithms (such as longest common subsequence or vector-based difference detection) to identify frequently modified sentence structures, lexical features, or structural features, and records these statistical results in a prompt generation rule dataset. The server can thus identify "high-modification-rate segments" and adjust the constraints on the generative AI model in subsequent prompts. For example, the server can add requirements such as "Please use more specific examples" and "Avoid using vague words like 'maybe' or 'perhaps'" to reduce the frequency of these segments in the generated results.

[0396] In one implementation, the server continuously updates the prompt generation rules based on accumulated historical data. Periodically or when triggering conditions are met, the server performs statistical analysis on the stored five-tuple of "communication information—parsing result—prompt statement—response document candidate—editing result," using clustering or feature-weighted algorithms to summarize the optimal prompt statement patterns for different types of conversations. Through this data-driven rule update mechanism, the server ensures that subsequent prompt statements are more structurally and constrained to match the corresponding intent and sentiment type, thereby directly improving the matching accuracy of the response text during the generation stage, reducing manual editing, and enhancing the efficiency of the overall computation process.

[0397] In one implementation, the server integrates the parsing module, prompt generation module, generative artificial intelligence model invocation module, and editing feedback analysis module in a pipelined manner to form a highly efficient data flow structure. When receiving new communication information, the server can not only generate high-quality response document candidates for the current session, but also analyze historical records in parallel in the background and update the prompt generation rules. By leveraging this pipelined and parallel processing structure, the server effectively improves overall processing throughput and achieves the technical effect of maintaining low response latency even in large-scale session scenarios.

[0398] In one implementation, the server introduces a prompt message adjustment mechanism based on emotional information, changing the traditional approach of generating responses solely based on keywords or fixed templates. The emotional vector extracted by the server during the parsing phase not only influences the word choice in the prompt message but also affects the control parameter settings of the generative AI model. For example, for communication information detected as having "high negative emotion," the server can explicitly instruct the generative AI model in the prompt message to "prioritize expressing understanding and support, avoiding direct rebuttal," while simultaneously lowering the upper limit of the generated length to avoid information overload. Through this dynamic generation method of prompt messages bound to emotional information, the server enables the generative AI model to execute a specific strategy at the computational level that differs from traditional rule-based systems, thereby significantly improving the matching accuracy between responses and emotional states while maintaining efficiency.

[0399] In one implementation, the server associates and stores communication information, parsing results, prompts, model output, and editing results, allowing each response generation process to be viewed as a traceable training sample. The server uses these structured samples to incrementally train or fine-tune the prompt generation module and the generative AI model. Through this continuous learning mechanism, the server improves the model's adaptability to domain-specific data at the algorithmic level, reducing the generation error rate and the probability of inappropriate expressions, thereby achieving both improved accuracy and reduced error.

[0400] In one implementation, the terminal not only serves as the user interface but also handles some preprocessing tasks, reducing the communication load and computational burden on the server. The terminal performs simple format checks, length limits, and sensitive word filtering on user input locally before sending it to the server, thereby reducing invalid requests. When receiving candidate response documents, the terminal can request only the changed parts using differential updates, reducing network transmission volume. This client-server collaborative design achieves the technical effects of communication load reduction and resource utilization optimization at the system architecture level.

[0401] In one implementation, users edit candidate response documents generated by the server, enabling the system to obtain high-quality human-annotated data. The server feeds this data back to the prompt generation and model control modules, allowing the system to gradually develop a generation strategy that is not simply a mimicking of human writing, but rather automatically optimized using statistical patterns and deep features. The key feature of this strategy is that it doesn't merely automate human editing behavior, but rather creates a closed-loop learning process within the computer: "parsing—generation—feedback—regeneration," thereby fundamentally improving the computer's processing methods and performance in natural language generation tasks.

[0402] In another embodiment, the server can employ other types of generative artificial intelligence models, such as text generation models based on recurrent neural networks or convolutional sequence networks, or a hybrid structure combining an attention-based encoder with a recursive decoder. This invention does not limit the specific model structure; as long as the server can generate candidate response documents based on prompts and update the generation control strategy using editing feedback, it falls within the scope of this invention. In another embodiment, the server can also input modal information other than text (such as simple category labels or ratings) as auxiliary features into the generative artificial intelligence model to further improve the matching degree between the generated results and the context.

[0403] In summary, by tightly integrating communication information parsing, dynamic generation of prompts, generative artificial intelligence model inference, and editorial feedback learning, the server constructs a specific data structure and algorithmic process, resulting in technological improvements in response quality, processing speed, resource utilization, and scalability. This improvement is not merely an automated replacement of human writing work, but rather a structural control and adaptive optimization of the natural language generation process within the computer, representing an improvement and advancement in computer technology itself.

[0404] use Figure 14 The processing flow is explained.

[0405] Step 1: The terminal sends communication information to the server.

[0406] The terminal's input consists of text messages entered by the user on the interface, such as "I'm worried my child has been distracted in class lately," along with metadata such as session identifier and sending time. After performing basic checks on the text length and character encoding, the terminal packages the text and session identifier together into a request message. The terminal then sends this request to the server using a secure communication protocol. The terminal's output is a network request message containing the communication text and metadata.

[0407] Step 2: The server receives and stores communication information.

[0408] The server's input is a network request message from the terminal, containing the raw communication text and session identifier. The server first parses the message, extracting the communication text and session identifier fields, and then creates or updates a record for that session in the database. The server performs a uniform encoding conversion on the text (e.g., to UTF-8) and stores the raw text in a communication information table. The server's output is a communication information record with a unique session identifier written to the storage device.

[0409] Step 3: The server performs language preprocessing on the communication information.

[0410] The server's input consists of the raw communication text stored in step 2 and the session identifier. The server invokes a natural language processing library to perform word segmentation, stop word removal, part-of-speech tagging, and sentence segmentation on the text. The server uses algorithms to convert the text into word sequences and sentence structure information, constructing intermediate data structures (e.g., objects containing word lists, sentence boundaries, and part-of-speech tags). The server's output is a preprocessed text representation associated with the session identifier, used for subsequent semantic analysis.

[0411] Step 4: The server extracts topic information and intent information from the communication information.

[0412] The server's input is the preprocessed text representation generated in step 3. The server uses a trained classification model or rule engine to classify the text into topic and intent categories, mapping the text to several topic tags (e.g., "classroom performance," "learning status") and intent tags (e.g., "consultation," "concern," "request for explanation"). The server obtains the probability distribution for each topic and intent by weighting keywords, sentence structure, and contextual features, and selects the tag with the highest probability as the output. The server's output consists of structured topic and intent information, which is stored in association with the session identifier.

[0413] Step 5: The server performs sentiment analysis on the communication information to generate sentiment information.

[0414] The server's input is the raw communication text or its preprocessed representation. The server invokes a sentiment analysis model to classify the text by emotion and regress its intensity, calculating numerical values ​​for several emotional dimensions such as "worry," "anger," and "satisfaction." The server normalizes the emotion scores to form an emotion vector and determines the overall emotional tendency (e.g., "high worry") based on a preset threshold. The server's output is sentiment information representing the emotional state of the communication, including the emotion category and corresponding intensity, which is written to a sentiment information table along with the session identifier.

[0415] Step 6: The server integrates topic information, intent information, and sentiment information.

[0416] The server's input consists of the topic and intent information obtained in step 4, and the sentiment information obtained in step 5. Based on the session identifier, the server merges these three types of information into a unified feature structure, which includes the main topic tag, intent tag, and sentiment vector. The server can encode these features, for example, by converting them into key-value pairs or vector forms, to facilitate subsequent calls by the prompt generation module. The server's output is a parsing result object corresponding to the session identifier, representing the semantic and sentiment understanding of the communication information.

[0417] Step 7: The server generates a prompt statement based on the parsing results.

[0418] The server's input is the parsed result object obtained in step 6. The server queries the prompt statement generation rule base, selects the corresponding template based on the intent type (e.g., a "soothing response template"), determines the content focus based on topic tags (e.g., "classroom attention status"), and adjusts the tone intensity and length based on the sentiment vector. The server fills these information into placeholders in the template, combining them into a complete prompt statement text. The server's output is a prompt statement that clearly describes the generation task, such as the aforementioned instruction text "You are a generative AI model that helps staff respond to service recipients...", and records it along with the session identifier in the prompt statement table.

[0419] Step 8: The server constructs the input context for generative artificial intelligence models.

[0420] The server's input consists of the prompt statement generated in step 7 and the original communication text. The server concatenates the prompt statement with the original text according to a predetermined format, such as appending contextual information like "The original message from the service recipient: '..." to the end of the prompt statement. The server then tokenizes the concatenated text, converting it into a tokenized sequence or vector representation to meet the input requirements of the generative artificial intelligence model. The server's output is the model input sequence used for inference, which is stored in memory as a tensor or sparse vector.

[0421] Step 9: The server invokes a generative artificial intelligence model to generate candidate response documents.

[0422] The server's input is the model input sequence obtained in step 8. The server feeds this input sequence into a generative AI model deployed locally or remotely, performing forward inference computation. Within the model, the server performs matrix multiplication, normalization, and nonlinear transformations through multiple network structures (e.g., multi-head self-attention layers and feedforward networks) to progressively generate output tags. The server uses a beam search or sampling strategy to control the generation process, selecting the next tag based on the generation probability distribution until a termination tag or length limit is reached. The server decodes the generated tag sequence into natural language text, i.e., response document candidates. The server's output is a complete draft response text, which is stored in association with a session identifier.

[0423] Step 10: The server sends the response document candidate to the terminal.

[0424] The server's input consists of the response document candidate and session identifier generated in step 9. The server constructs a response message, placing the response document candidate in the response body, and attaching the session identifier and necessary status information. The server sends this response to the corresponding auditing terminal via the network interface. The server's output is a network response message containing the response document candidate.

[0425] Step 11: The terminal displays candidate response documents and allows users to edit them.

[0426] The terminal's input consists of candidate response documents and a session identifier returned by the server. The terminal displays the original communication information and candidate response documents on the interface, allowing users to visually compare and contrast them. The terminal provides text editing controls, allowing users to add, delete, modify, and rearrange the response text. Users perform these modifications via keyboard or touch. After the user completes editing and confirms, the terminal generates the final response document and can record editing differences (e.g., marking modified sentences). The terminal's output is a data structure containing the final response text and optional editing traces.

[0427] Step 12: The terminal sends the final response and edited information to the server.

[0428] The terminal's input consists of the final response document obtained in step 11, the edit difference information, and the session identifier. The terminal encapsulates this data into a request message and sends it to the server via a secure communication protocol. The terminal can indicate the completion time of this edit and the editor's identity in the message. The terminal's output is the final response and edit feedback request submitted to the server.

[0429] Step 13: The server stores the final response document and performs discrepancy analysis.

[0430] The server takes as input the final response document from the terminal, edit discrepancy information, and candidate original response documents. First, the server stores the final response document along with the session identifier in the response document table. Then, the server invokes a text comparison algorithm to align the original candidate text with the final text sentence by sentence or word by word, identifying replaced, deleted, and added segments. The server statistically analyzes these discrepancies, generating edit features that include the location, type, and frequency of modifications. The server outputs a set of edit feature records associated with each session.

[0431] Step 14: The server updates the rules for generating prompt statements based on historical editing characteristics.

[0432] The server's input consists of accumulated editing features from multiple sessions, along with corresponding prompts and parsing results. The server aggregates and analyzes this data, calculating which expressions and structures are frequently modified across multiple sessions. Based on the statistical results, the server updates the prompt generation rules, for example, by reducing the probability of using certain frequently modified phrases, adding explicit explanations or concrete examples, or incorporating constraints to "avoid ambiguous words" under specific intent types. The server writes the updated rules back to the rule base. The server's output is a set of updated prompt generation rules, used for generating prompts in subsequent sessions.

[0433] Step 15: The server adjusts the control parameters of the generative artificial intelligence model according to the update rules.

[0434] The server's input consists of the updated prompt generation rules from step 14 and the current control parameter configuration of the generative AI model. Based on these rules, the server adjusts the model's inference parameters. For example, for scenarios requiring more specific content, the server lowers the temperature parameter during generation, reduces the random sampling amplitude, and increases the selection weight of high-confidence words. For scenarios requiring a more reassuring tone, the server emphasizes a "friendly and positive tone" in the prompts and may restrict the use of words with strong negative connotations. The server persistently stores these parameter settings. The server's output is a new model control configuration, making the generation behavior during subsequent inference more consistent with the preferences reflected in historical editing feedback.

[0435] Step 16: The server stores the overall data structure in a coherent manner to support subsequent learning.

[0436] The server's input includes communication information for each session, parsing results (topic information, intent information, sentiment information), prompts, candidate response documents, the final response document, and editing features. The server establishes an associated index for this data using session identifiers and timestamps, storing it in structured data storage. The server provides a complete data chain for subsequent offline learning or online incremental training, enabling the system to retrieve full-process data from any session at different time periods. The server's output is a complete session dataset that can be used for retraining and analysis, providing a foundation for continuous optimization of generative AI models and prompt generation modules.

[0437] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0438] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0439] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0440] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0441] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0442] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0443] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0444] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0445] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0446] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0447] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0448] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0449] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0450] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0451] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0452] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0453] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0454] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0455] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0456] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0457] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0458] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0459] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0460] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0461] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0462] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0463] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0464] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0465] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0466] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0467] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0468] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0469] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0470] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0471] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0472] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0473] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0474] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0475] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0476] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0477] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0478] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0479] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0480] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0481] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0482] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0483] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0484] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0485] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0486] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0487] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0488] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0489] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0490] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0491] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0492] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0493] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0494] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0495] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0496] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0497] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0498] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0499] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0500] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0501] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0502] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0503] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0504] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0505] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0506] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0507] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0508] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0509] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0510] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0511] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0512] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0513] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0514] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0515] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0516] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0517] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0518] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0519] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0520] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0521] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0522] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0523] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0524] In addition, the following notes are provided in response to the above explanation.

[0525] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for obtaining payment information from an information processing device; A device for sending the payment information to a server using encrypted communication; A device for extracting payment characteristic information based on the payment information and historical transaction information in the server and generating a prompt statement containing the payment characteristic information; A device for inputting the prompt statement into a generative artificial intelligence model to obtain a single determination result of the risk of improper transactions; An apparatus for updating the prompt statement based on the initial judgment result and the payment characteristic information, and then re-inputting it into the generative artificial intelligence model to obtain the final judgment result of the risk of improper transactions; An apparatus for generating a payment review result based on the final judgment result and outputting the payment review result to the information processing device.

[0526] (Note 2) The information processing system according to Appendix 1 is characterized in that, The server adds judgment condition information to the prompt statement to indicate the role of the generative artificial intelligence model, the output format, and the evaluation benchmark, so that the first judgment result and the final judgment result are output with a predetermined risk distinction.

[0527] (Note 3) The information processing system according to Appendix 1 is characterized in that, The server stores the final judgment result and the payment characteristic information as record information, and uses the record information to improve the efficiency of detecting and processing improper transactions and shorten the processing time of payment review.

[0528] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A unit for receiving transaction information from a terminal and storing the transaction information and the user information corresponding to the transaction information; A unit for calculating statistical information from historical transaction records based on the transaction information and the user information, and for standardizing the transaction information; A unit for generating prompt statements that are input to a generative artificial intelligence model based on the standardized transaction information and the statistical information; A unit for inputting the prompt statement into the generative artificial intelligence model and obtaining a determination result of the probability of the transaction information being incorrect, output by the generative artificial intelligence model; A unit used to determine the risk level of the transaction information based on a comparison between the initial judgment result and a preset threshold; A unit for generating detailed prompt statements containing multiple transaction records related to the transaction information and the user information when the risk level is within a predetermined range, and then inputting them again into the generative artificial intelligence model to obtain a final judgment result on whether the transaction information is incorrect and recommended processing content; A unit for recording the first determination result and the final determination result in association with the transaction information; A unit for identifying the transaction information as a suspicious transaction based on the risk level and the final judgment result, generating a notification message containing details of the suspicious transaction and a confirmation option, and sending it to the terminal; A unit for receiving user confirmation results from the terminal regarding the suspicious transaction, updating the transaction information status based on the confirmation results, and sending a control request to an external settlement processing device regarding the continuation or termination of the transaction; A unit for associating the confirmation result with the transaction information and the prompt statement and storing it as learning data for use in the learning or updating of the generative artificial intelligence model.

[0529] (Note 2) According to the information processing system described in Appendix 1, the generative artificial intelligence model is configured to perform natural language parsing on at least a portion of the transaction amount, transaction location, transaction time, user's historical transaction tendencies, and transaction terminal information contained in the prompt statement, and output a numerical score of the probability of irregularity and a recommended processing content for the suspicious transaction, either transaction termination or simply sending a notification.

[0530] (Note 3) According to the information processing system described in Appendix 1, the system determines the risk level by automatically generating the prompt statement and based on the output of the generative artificial intelligence model, thereby shortening the time required for detecting and processing unfair transactions and sending notifications to users, while improving the effectiveness of preventing unfair transactions in advance and the efficiency of communication with users.

[0531] Example 2 (Note 1) An information processing system, characterized in that it comprises: An apparatus for receiving information sent by a user from a user terminal via an information processing device, acquiring the information as image information or text information, and converting the image information and the text information into an internal representation that can be processed by a generative artificial intelligence model; An apparatus for dynamically generating, based on at least the internal representation, a prompt statement for outputting a single determination, whether an anomaly exists, and an anomaly probability index, and inputting the prompt statement and the internal representation into a generative artificial intelligence model to perform a single determination process; The device is used to, based on the anomaly probability index contained in the result of the first determination process, adopt the result of the first determination process as the final determination content without detailed analysis, and, when detailed analysis is required, generate a prompt statement containing detailed final determination content including anomaly type, anomaly degree, and determination reason based on the result of the first determination process and the internal representation, and input the prompt statement, the result of the first determination process, and the internal representation into the generative artificial intelligence model to execute the final determination process. The information processing device generates structured review result data indicating whether an anomaly exists, its type, and its severity based on the final judgment processing result, and converts the review result data into a user-viewable format and sends it to the output device of the user terminal.

[0532] (Note 2) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to, based on the result of the final judgment processing, summarize the judgment content and anomaly types contained in the review result data, generate a prompt statement using the summary result as explanatory input information, input the prompt statement and the explanatory input information into a generative artificial intelligence model, thereby automatically generating user-oriented natural language explanatory text, and sending the explanatory text to the user terminal after associating it with the review result data.

[0533] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to switch the processing load or processing accuracy level of the generative artificial intelligence model used in the final judgment process based on the anomaly probability index in the result of the first judgment process, so that when the anomaly probability index is less than a predetermined threshold, a low-load judgment process is selected, and when the anomaly probability index is greater than or equal to the predetermined threshold, a high-precision judgment process is selected, thereby reducing the overall judgment processing time and resource consumption while maintaining the review accuracy.

[0534] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring and parsing communication information related to the object being served; A device for dynamically generating prompt statements for input to a generative artificial intelligence model based on the intent and emotion information obtained from the analysis; Apparatus for generating candidate response documents using the generative artificial intelligence model based on the prompt statement and the communication information; A device for presenting the candidate response document to the reviewer's information processing device and allowing the reviewer to perform editing operations; A means for sending the edited response document to the information processing device of the notified party after the editing operation; An apparatus for recording the parsing results and the output results of the generative artificial intelligence model in correspondence with the communication information and the response document candidates, and storing them as data that can be reused in subsequent prompt statement generation processing and generative artificial intelligence model control processing.

[0535] (Note 2) The information processing system according to Appendix 1 is characterized in that, The device for acquiring and parsing is configured to extract topic information and sentiment information from the communication information using natural language processing technology. The device for dynamic generation is configured to adjust the text structure of the prompt statement based on the extracted topic information and sentiment information, so that the candidate response document generated by the generative artificial intelligence model includes content that enhances reassurance and trust, thereby controlling the expression content of the candidate response document.

[0536] (Note 3) The information processing system according to Appendix 1 is characterized in that, The recording and storage device is configured to store the communication information, the parsing results, the prompt statements, the candidate response documents, and the reviewer's editing results in association with each other. The dynamic generation device is configured to update the generation rules of the prompt statements based on the associated stored information, thereby gradually optimizing the processing of the generative artificial intelligence model in generating the candidate response documents, and improving communication effectiveness while shortening the reviewer's editing time.

Claims

1. An information processing system, characterized in that, include: Processor, the processor being configured to: Receive and parse contact book contents sent by the guardian; Based on the analysis results, input the prompt text into the generative artificial intelligence model to generate a response text scheme; The generated response text scheme is presented to the teacher's terminal so that the teacher can modify the response text scheme.

2. The information processing system according to claim 1, characterized in that, The processor is configured to control the response text scheme generated by the generative artificial intelligence model to include content designed to enhance the guardian's sense of security and trust.

3. The information processing system according to claim 1, characterized in that, The processor is configured to reduce business processing time and maximize communication effectiveness by utilizing the generated response text scheme.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A