Information processing system
Patent Information
- Application Number
- CN202610332920.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-18
- Publication Date
- 2026-09-22
AI Technical Summary
通过上述技术方案,本发明能够在检测到文本规模或结构复杂度超过预定范围时自动触发生成式人工智能模型给出推敲方案,实现从文本解析、阈值判定到智能推敲建议生成和呈现的自动化处理流程,从而有效解决现有技术中长文本与复杂结构文本难以高效撰写和理解的问题
[0004]为解决上述技术课题,本发明提供一种信息处理系统,该系统包括处理器,所述处理器被配置为执行以下处理步骤:第一,处理器接收文本数据,所述文本数据可以来自用户终端的输入界面、编辑器或其他应用程序的接口,从而在无需用户额外操作的情况下自动获取待分析文本;第二,处理器对所接收的文本数据进行解析,至少检测所述文本数据的字符数以及箇条列表项的数量,具体地,处理器可以通过统计文本长度得到字符数,并通过逐行分析、识别行首标记等方式确定箇条列表项的数量,从而获得反映文本规模与结构复杂度的解析结果;第三,处理器基于所述解析结果设定阈值或者从预先存储的阈值配置中选取对应的阈值条件,并判断所述解析结果是否超过所设定的阈值,其中,该阈值可以包括字符数上限、箇条列表项数量上限或二者的组合;第四,当处理器判定解析结果超过所设定的阈值时,处理器生成用于指示生成式人工智能模型生成推敲方案的提示信息,所述提示信息至少包括原始文本数据以及与文本长度、结构复杂度相关的指示内容,从而使生成式人工智能模型能够在了解文本特点和优化目标的前提下生成合适的推敲方案;第五,处理器将所述提示信息输入至生成式人工智能模型,以使所述生成式人工智能模型基于提示信息生成针对所述文本数据的推敲方案,包括改写建议、结构调整建议或摘要性建议等。进一步地,处理器可以通过用户界面呈现所生成的推敲方案,使用户在终端上直接查看和参考该推敲方案进行文本修改;处理器还可以通过通信模块将所生成的推敲方案通知给投稿者,使投稿者在消息通知、邮件或其他通信渠道中获取推敲建议。通过上述技术方案,本发明能够在检测到文本规模或结构复杂度超过预定范围时自动触发生成式人工智能模型给出推敲方案,实现从文本解析、阈值判定到智能推敲建议生成和呈现的自动化处理流程,从而有效解决现有技术中长文本与复杂结构文本难以高效撰写和理解的问题。
Smart Images

Figure CN122797489A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.
[0003] In various online communication platforms, collaborative office systems, and social media, users are increasingly communicating via text. As communication content becomes more complex and specialized, individual texts are often lengthy and contain numerous lists or itemized explanations. When texts are too long or overly complex, the following problems can easily arise: First, contributors may struggle to grasp the overall structure and key information while writing, resulting in lengthy, repetitive, or illogical statements that reduce information delivery efficiency. Second, recipients need to spend considerable time understanding the core content and key points of long texts, leading to a heavy reading burden. Third, existing text assistance tools typically only offer simple spell checks or basic formatting corrections, lacking intelligent revision and rewriting suggestions that are automatically triggered based on text size and structure, and failing to fully utilize generative AI models for targeted text optimization. Furthermore, in practical applications, if generative AI models are needed to assist in text revision, users often need to manually determine whether the text needs revision, manually write prompts, and actively invoke the generative model, making the process cumbersome and lacking automation. Therefore, how to provide a system that can automatically receive text data, analyze and judge thresholds based on indicators such as the number of words and the number of entries in the text, and automatically trigger a generative artificial intelligence model to generate a refinement scheme when the preset threshold is exceeded, thereby reducing the user's writing and reading burden and improving text quality, has become the technical problem that this invention urgently needs to solve. Summary of the Invention
[0004] To address the aforementioned technical challenges, this invention provides an information processing system comprising a processor configured to perform the following processing steps: First, the processor receives text data, which may originate from the input interface of a user terminal, an editor, or the interface of another application, thereby automatically acquiring the text to be analyzed without additional user intervention; Second, the processor parses the received text data, at least detecting the number of characters and the number of list items. Specifically, the processor can obtain the number of characters by counting the text length and determine the number of list items by line-by-line analysis and identification of line start markers, thereby obtaining a parsing result reflecting the text size and structural complexity; Third, the processor sets a threshold based on the parsing result or selects from a pre-stored threshold configuration. The processor takes a corresponding threshold condition and determines whether the parsing result exceeds the set threshold. This threshold may include an upper limit on the number of characters, an upper limit on the number of list items, or a combination of both. Fourth, when the processor determines that the parsing result exceeds the set threshold, it generates a prompt message to instruct the generative AI model to generate a refinement plan. This prompt message includes at least the original text data and instructions related to text length and structural complexity, enabling the generative AI model to generate a suitable refinement plan based on an understanding of the text characteristics and optimization objectives. Fifth, the processor inputs the prompt message into the generative AI model, allowing it to generate a refinement plan for the text data based on the prompt message, including rewriting suggestions, structural adjustment suggestions, or summary suggestions. Furthermore, the processor can present the generated refinement plan through a user interface, allowing users to directly view and refer to the refinement plan on their terminals for text modification. The processor can also notify the contributor of the generated refinement plan through a communication module, enabling the contributor to receive refinement suggestions via message notifications, emails, or other communication channels. Through the above technical solution, the present invention can automatically trigger a generative artificial intelligence model to provide a revised solution when the size or structural complexity of the text exceeds a predetermined range. This realizes an automated processing flow from text parsing and threshold determination to the generation and presentation of intelligent revision suggestions, thereby effectively solving the problem of efficient writing and understanding of long and complex texts in the prior art.
[0005] A "system" refers to an overall device or platform consisting of multiple hardware components and software modules, used to receive and parse text data and automatically generate proposed solutions. It can be implemented in the form of a server, terminal device, or cloud service.
[0006] A "processor" is a computing unit that can execute program instructions to complete processing tasks such as data reception, parsing, threshold judgment, prompt information generation, and interaction with generative artificial intelligence models. It can be a single physical processor, a combination of multiple processors, or any processing chip or its virtualized instance, including CPU, GPU, NPU, etc.
[0007] "Text data" refers to content composed of characters that is stored or transmitted in digital form, including but not limited to natural language sentences, paragraphs, manuscripts, messages, posts, comments, and other textual information that can be received and processed by the system.
[0008] "Character count" refers to the total number of characters counted in the text data, which may include the length of one or more character sets, including letters, Chinese characters, numbers, punctuation marks, and spaces.
[0009] "Number of list items" refers to the number of rows or units that are identified as list items when parsing text data by line or by structure. These list items usually begin with a specific marker (such as bullet points, numbers, etc.) and are used to list information in bullet points.
[0010] "Parsing results" refers to the collection of various indicators and structural information obtained by the processor after analyzing text data. It includes at least the number of characters and the number of list items, and may also include other statistics or features related to the text structure or content.
[0011] "Threshold" refers to a preset or dynamically determined limit value used to compare with the parsing results to determine whether the text data needs to be further processed. It can be an upper limit on the number of characters, an upper limit on the number of list items, or a combination of both and other indicators.
[0012] "Generative AI models" refer to AI models that can automatically generate text content based on input prompts, including but not limited to deep learning-based language models, large-scale pre-trained models, or fine-tuned versions thereof.
[0013] "Prompt information" refers to the instructional or explanatory text constructed by the processor and input into the generative artificial intelligence model. It describes the original text data, constraints, optimization objectives, or generation requirements to guide the generative artificial intelligence model to output the corresponding proposed solution.
[0014] "Refinement suggestions" refer to optimization recommendations for original text data generated by generative artificial intelligence models based on prompts. These suggestions typically include rewriting suggestions, word choice adjustments, structural reorganization, summarization, or other textual information that helps improve the readability and expressiveness of the text.
[0015] A "user interface" refers to an interactive interface used to present information to users and receive user input on a terminal device, including graphical user interfaces, web interfaces, and mobile application interfaces.
[0016] A "communication module" refers to a collection of hardware and software components used to transmit data between the system and external devices or users. It may include network interfaces, message push services, email sending modules, and related communication protocol stacks.
[0017] "Contributor" refers to a user subject who submits text data to the system through a terminal. This can be an individual user, an organizational account, or other entity with text publishing permissions. Attached Figure Description
[0018] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0019] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0020] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0021] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0022] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0023] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0024] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0025] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0026] Figure 9 This represents an emotion map that maps multiple emotions.
[0027] Figure 10 This represents an emotion map that maps multiple emotions.
[0028] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0029] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0030] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0031] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0032] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.
[0033] First, let me explain the terminology used in the following instructions.
[0034] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0035] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0036] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0037] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0038] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0039] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0040] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0041] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0042] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0043] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0044] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0045] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0046] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0047] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0048] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0049] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0050] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0051] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0052] In existing text processing technologies, servers typically perform only simple length statistics or keyword searches on text data, leaving users to read and manually modify lengthy or structurally complex texts. This approach has the following problems: First, the server lacks fine-grained analysis of text structural features (such as enumeration format, number of items, etc.), making it unable to make targeted automatic processing decisions based on the specific structural characteristics of the text, resulting in low efficiency in processing complex text. Second, when calling generative AI models, servers often simply input the raw text directly into the model without dynamically generating prompts based on the actual statistical results and structural features of the text. This leads to the generative AI model's unclear understanding of the task objective, resulting in poor stability and controllability of the output results. Third, existing systems often completely delegate the burden of whether to call the generative AI model and how to specify the generation task objective to the user, lacking a mechanism on the server side to automatically determine "when to refine" and "how to instruct the model to generate a suitable refinement solution," thereby increasing interaction costs and reducing overall human-computer collaboration efficiency. Fourth, although some systems introduce artificial intelligence to generate text, they do not manage and display the parsing results, threshold judgment results and generation results in a unified manner. Users find it difficult to understand why the system provides a certain proposed solution, and the interpretability and reliability of the system are insufficient.
[0053] Therefore, it is necessary to provide a new system that performs structured parsing and quantitative statistics on text data on the server side, automatically determines whether text refinement is needed based on preset or variable judgment thresholds, and automatically generates prompt statements corresponding to the text structure and statistical results. These prompt statements, along with the original text, are then input into a generative artificial intelligence model. This allows the computer to automatically refine the text in a more precise and controllable manner using the generative artificial intelligence model, and integrates the parsing results and correction schemes to feed back to external devices. This improves the computer's processing flow, resource utilization efficiency, and human-computer interaction experience in text processing tasks.
[0054] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0055] In this invention, the server includes a device for acquiring textual information from an external device, a device for parsing the textual information to detect the total number of character symbols and the number of enumerated structural elements, a device for determining a judgment threshold based on the detection results and stored benchmark information, a device for determining whether text correction processing is required based on the judgment threshold and the detection results, a device for generating instruction text containing textual structural information and the judgment threshold when text correction processing is determined to be required, and constructing the instruction text into a prompt statement for inputting into a generative artificial intelligence model, a device for inputting the prompt statement and the textual information into the generative artificial intelligence model to generate a correction scheme, and a device for generating response information containing the correction scheme and the detection results and sending it to an external device. This allows for an integrated processing flow automatically completed on the server side, from text structured parsing and threshold determination to prompt statement construction and generative artificial intelligence model invocation. This closely links the input conditions of the generative artificial intelligence model with the objective statistical characteristics of the text, thereby improving the automation and controllability of text refinement processing, reducing the manual setting burden on users, and enhancing the processing performance and human-computer interaction efficiency of the computer system in text communication and document processing scenarios.
[0056] A "system" refers to a collection of computer-based devices that consist of one or more information processing devices and external devices connected to them, used for acquiring, parsing, judging, generating, and outputting textual information.
[0057] "Information processing device" refers to a computer device with a processor and memory, which can execute program instructions, perform calculations and logical control on input data, and realize functions such as text parsing, threshold determination, prompt statement generation, and interaction with generative artificial intelligence models.
[0058] "External device" refers to a terminal device or other computing device that sends and receives data with the information processing device through a communication network or communication interface, including but not limited to network nodes other than user terminals, client devices or servers.
[0059] "Textual information" refers to data content that is stored and processed by a computer, with characters as the basic unit. This includes natural language text, documents containing sentences and paragraphs, and message content recorded using symbols and punctuation.
[0060] "Total number of characters" refers to the statistical result of the number of all characters appearing in textual information, including letters, numbers, punctuation marks, whitespace marks, and other text units that can be displayed or processed.
[0061] "Number of enumerated structural elements" refers to the statistical result of the number of list items presented in textual information in the form of columns, bullets, or numbers. These list items are usually identified by symbols, numbers, or parentheses and separated by lines or paragraphs.
[0062] "Baseline information" refers to reference data for judgment that is pre-stored in a storage device, including configuration data or rule data for setting the threshold for the total number of character symbols, the threshold for the number of enumerated structural elements, and other judgment parameters.
[0063] "Judgment threshold" refers to the boundary value or condition determined based on the detection results and benchmark information to determine whether textual information needs to be corrected. This includes thresholds for the total number of character symbols and thresholds for the number of enumerated structural elements.
[0064] "Text correction processing" refers to the modification, simplification, rewriting, or structural adjustment of original text information to reduce redundancy, compress length, merge entries, or improve readability, thereby generating rewritten text content.
[0065] "Structural information" refers to data extracted from textual formal information that reflects the internal composition and organization of the text, including paragraph division, sentence boundaries, list structure, number of items, and related tagging information.
[0066] "Instructional text" refers to descriptive text content generated by an information processing device to instruct a generative artificial intelligence model to perform a specific generative task. This text contains constraints, target requirements, or processing instructions related to textual information and decision thresholds.
[0067] "Prompt statements" refer to instructional text information that is constructed from text to be suitable as input to a generative artificial intelligence model. Their function is to guide the generative artificial intelligence model to understand the task objective and generate corresponding output based on the task objective.
[0068] "Generative AI models" refer to AI models trained through machine learning or deep learning that can generate new text content or amendments based on input text, including neural network-based text generation models or dialogue models.
[0069] "Correction scheme" refers to text content generated by a generative artificial intelligence model based on prompts and textual information to replace or supplement the original text. This text content reflects the result of modification, simplification, or rewriting of the original text.
[0070] "Detection results" refers to the statistical or analytical data obtained by the information processing device after parsing textual information, including the total number of character symbols, the number of enumerated structural elements, and other analytical information related to the text structure.
[0071] "Response information" refers to the output data generated by the information processing device for external devices, which includes at least the correction scheme and detection results, and may also include additional information related to the judgment threshold, processing procedure or prompt message.
[0072] "Display interface" refers to the graphical user interface presented by the display control device on an external device or terminal device. It is used to display text information, correction schemes and related explanatory messages in a visual manner, and supports users to view and operate them.
[0073] "Explanatory messages" refer to text information generated based on judgment thresholds and detection results to explain to users the reasons, triggering conditions, or processing results of the system's processing, including explanations for exceeding text length limits, excessive number of columns, and automatic deduction execution.
[0074] The embodiments of the present invention will be described in conjunction with the appendix. In the following description, "server" refers to the information processing device that implements the main functions of the present invention, "terminal" refers to an external device operated by a user, and "user" refers to the operator who uses the terminal to input and view text.
[0075] In one embodiment, the server includes: a multi-core central processing unit, semiconductor storage devices (including main memory and non-volatile memory), a network interface card, and an application execution environment running on a general-purpose operating system (such as a UNIX-like operating system). Within this environment, the server runs a text parsing module, a threshold determination module, a prompt generation module, a generative artificial intelligence model interface module, and a response generation module. The server can further maintain concurrent connections with multiple terminals via reverse proxy software and application server software.
[0076] In one embodiment, the terminal includes an electronic device with a display panel and input devices, such as a smartphone, tablet computer, or personal computer. The terminal runs browser software or native applications, providing the user with a text input area, button controls, and a results display area through a graphical user interface. The terminal sends and receives data with a server via a network.
[0077] In one implementation, users input text information requiring verification via the terminal's input components, or copy text from an existing document and paste it into the terminal's text input area. Users can also select a default threshold mode or a custom threshold mode on the terminal interface, and choose whether to provide custom prompts.
[0078] In one implementation, the server uses a general-purpose programming language to implement various functional modules. In the text parsing module, it combines a general-purpose natural language processing library (such as a word segmentation and syntactic analysis library that can be deployed on the server) to perform structured processing on the textual information. The server stores the textual information as string data in memory, performs character-level traversal and line-level segmentation on the strings, and obtains a data structure for subsequent statistical and structural analysis. The server counts each character to obtain the total number of character symbols; by performing pattern matching on the starting symbol sequence of each line, the server identifies lines starting with list markers as enumerated structural elements, thereby counting the number of enumerated structural elements.
[0079] In one implementation, the server stores baseline information in a storage device. This baseline information can be in the form of a configuration file or a database record, including character count thresholds and enumeration item count thresholds describing different application scenarios, as well as corresponding processing strategy configurations. At runtime, the server reads the baseline information from the storage device and maps it to internally used threshold parameters. When needed, the server overrides or adjusts these parameters based on user-provided custom thresholds, thereby supporting multiple operating modes.
[0080] In one implementation, the server stores the detection results and threshold parameters in a structured data object and performs comparison calculations in the threshold determination module. The server obtains several Boolean determination results based on whether the total number of character symbols exceeds the corresponding threshold and whether the number of enumerated structural elements exceeds the corresponding threshold. These results are then combined to determine whether text correction processing is necessary. Because this determination is implemented on the server side through arithmetic comparison and logical combination, it can operate stably and at high speed in scenarios with large-scale concurrent text processing, thereby improving overall processing throughput.
[0081] In one implementation, when the server determines that text correction is needed, it uses the detection results along with text structure information to generate a prompt statement. The server, through a prompt statement generation module, embeds the following information into the instruction text: total number of characters, number of enumerated structural elements, corresponding judgment threshold, desired compression ratio, desired structural elements to retain, and whether to limit the number of items. The server uses predefined templates and conditional branching logic when generating the prompt statement to ensure that it accurately reflects the structural features and rewriting goals of the current text, thereby providing richer and more structured constraints on the input dimension for downstream generative AI models.
[0082] For example, in one implementation, the server can generate the following prompt statement: Please examine the text according to the following requirements: 1. Compress the overall length as much as possible, reducing the word count to approximately 50% of the original text; 2. Merge or delete overly fragmented entries, keeping the number of entries to no more than 5; 3. Retain the core purpose, key steps, and important precautions.
[0083] The following text requires further consideration: [Insert original text here] For example, when a server receives a user-defined prompt, it can directly use the text entered by the user as part of the instruction text, for example: "Please rewrite the following project description into a version suitable for sending to non-technical management, while retaining all technical terminology, and keep it within 300 words:" [Insert original text here] In another example, the server can use the following prompt statement: "The following team communication email is too lengthy. Please rewrite it into a more concise version suitable for sending via chat tools, limited to 300 characters:" [Insert original text here] In this way, the server concatenates the prompt statement with the original text to form a complete input sequence for processing by the generative artificial intelligence model.
[0084] In one implementation, the server communicates with externally or locally deployed generative AI models via a generative AI model interface module. The server can run a multi-layer neural network-based text generation model locally or in the cloud. In a typical implementation, this model employs a sequence-to-sequence architecture based on a self-attention mechanism. This architecture includes multi-layer encoders and decoders, each layer comprising a multi-head self-attention unit, a feedforward network, and a normalization unit. During training, the model uses a large-scale text corpus, updating weight parameters by minimizing the cross-entropy loss function between the predicted and target texts. Weight updates employ a gradient descent-based optimization algorithm, combined with an adaptive learning rate strategy and regularization techniques to improve the stability and generalization ability of the generated data.
[0085] In one implementation, the server performs word segmentation and sub-word encoding preprocessing on the input text data during the training phase, mapping the character sequence into a vector representation. During the inference phase, the server encodes the prompt and the original text into discrete token sequences, maps them into high-dimensional vector sequences through the model's embedding layer, and then performs nonlinear transformations through multiple attention and feedforward networks. When generating correction schemes, the model simultaneously focuses on the feature representations of both the prompt and the original text through the attention mechanism at the decoder end, thus balancing task constraints and the semantics of the original text in the output. The server can set temperature parameters, maximum generation length, and repetition penalty parameters to control the diversity and conciseness of the generated text.
[0086] In one implementation, the server doesn't simply "let the model generate freely," but rather injects constraint information derived from the detection results and thresholds into the prompt statements. For example, when the server detects that the number of entries significantly exceeds the threshold, it adds explicit requirements such as "merge entries" or "reduce the number of entries" to the prompt statements, and can internally assign weight labels to these requirements. The server can further perform rule validation on the output after generation, checking whether the number of entries in the generated text falls back to the target range. If not, it can trigger rewriting or simple rule compression. This closed-loop structure of "statistics – prompt statements – generation – validation" makes the entire processing flow controllable and technically sophisticated, rather than simply a one-time call to generate the model.
[0087] In one implementation, the server packages the correction scheme and the detection results together into a response message within the response generation module. This response message explicitly lists: the total number of original character symbols, the number of enumerated structural elements, the corresponding judgment threshold, whether automatic deduction was triggered, and the generated correction scheme text. Upon receiving this data, the terminal can simultaneously display the original text and the correction scheme on the display interface, and show the reason for exceeding the limit in graphical or textual marker form, thereby helping the user understand the system's processing logic.
[0088] In one implementation, the server achieves several technical benefits through the aforementioned structured internal data flow and modular division. First, by performing character-level and structure-level joint parsing of textual information on the server side and executing threshold determination before invoking the generative AI model, the server can only invoke the computationally expensive generative model when necessary, significantly reducing unnecessary model inference iterations and lowering server computational load and network traffic. Second, by encoding detection results and threshold information into prompts, the server ensures that the input to the generative model includes precise task conditions, thereby reducing irrelevant content and verbose output in multiple rounds of generation and improving the consistency and stability of the generated results with the target requirements. Third, by performing posterior verification on structural features such as the number of columns in the generated results, the server adds a verifiable step to the traditional "black box" generation process, thus improving the overall reliability of the system.
[0089] In one implementation, the terminal uses a display interface to present the original text and the revised proposal side-by-side, and displays a brief explanatory message based on the judgment information provided by the server, such as "The original text exceeds the preset threshold; a simplified version has been automatically generated for you." The terminal can also provide user interaction controls for choosing whether to adopt the revised proposal, whether to request a new revision, or whether to modify the prompt statement. Therefore, the terminal is not merely a simple presentation device, but rather works with the server to help users understand the server's internal joint processing results based on statistical and generative models at a visual level.
[0090] In one implementation, users can input custom prompts via a terminal to refine the rewriting goals of the generative AI model. In this case, the server retrieves the instruction text from an external device, using it as the main content or part of the prompt, and can still automatically add necessary constraint information based on the detection results, thereby achieving coordinated control between "user intent" and "statistical constraints." Thus, the server can not only automatically determine whether refinement is needed, but also flexibly adjust the generation strategy at the model input level according to different user needs.
[0091] Furthermore, in another embodiment of the invention, different types of generative artificial intelligence models can be employed, such as sequence generation models based on recurrent neural network structures, or hybrid structures combining convolutional neural networks and attention mechanisms. The server can also employ different feature extraction methods, such as introducing features like sentence length distribution, word frequency statistics, and repeated sentence detection during the text parsing stage to further refine the determination of redundancy and structural complexity. When training these models, the server can use supervised learning, semi-supervised learning, or self-supervised learning methods. The loss function can incorporate a structure penalty term in addition to cross-entropy to encourage the model to reduce redundant structures and repetitive patterns during the generation process.
[0092] Through the aforementioned implementation methods, the server internally performs a series of technical processes, including multi-dimensional quantitative analysis of textual information, automatic judgment based on configurable thresholds, conditional generation based on prompt statements, and result verification based on structured indicators. Compared to traditional methods that rely solely on manual reading and editing, this system automates and optimizes the complex text analysis process within the computer. Compared to traditional computer systems that only perform simple length statistics or directly call generative models, this invention, by introducing structured judgment and constraint-based prompt statements, enables generative artificial intelligence models to be scheduled and utilized in a more efficient and refined manner, thereby improving computer technology itself in multiple aspects such as processing speed, generation accuracy, communication load, and resource utilization efficiency.
[0093] use Figure 11 The processing procedure is explained.
[0094] Step 1: The terminal receives user input and sends requests.
[0095] The terminal displays a text input area and a send button on the screen.
[0096] Input: The raw text entered by the user on the terminal via keyboard or touch, as well as the threshold setting mode selected by the user on the interface and the option to use custom prompts.
[0097] Processing: The terminal stores the raw text input by the user as a string data in local memory, converts the settings selected by the user in the interface into parameter values (such as character threshold, column threshold, custom prompt text, etc.), and performs format checks on these data (such as checking whether the text is empty or whether the number of characters exceeds the terminal's local limit).
[0098] Output: The terminal constructs a request data object containing the original text and parameter settings, and sends the request data to the server through the network interface. The specific actions performed by the terminal at this time include: serializing the string and numerical parameters into a transmission format, encapsulating them into a network message via the communication protocol stack, and sending it to the server address through the network adapter.
[0099] Step 2: The server receives the request and parses the text.
[0100] After receiving request data from the terminal via the network interface, the server passes it to the application.
[0101] Input: The request data object sent by the terminal, which includes the raw text string, the user threshold parameter (if any), and the custom prompt statement (if any).
[0102] Processing: The server deserializes the received data, restoring it to its internal data structure. The server checks the completeness of fields in the request (including text fields and parameter fields). If any fields are missing or incorrectly formatted, the server generates an error message and prepares to return it to the terminal. For correctly formatted requests, the server saves the original text to its working memory and stores the parameter values in a data structure for subsequent calculations.
[0103] Output: The server outputs a set of preprocessed internal data objects, which include at least the original text string, the parsed user threshold parameters, and the custom prompt statement flags, for use in subsequent parsing and statistics modules.
[0104] Step 3: The server counts the total number of character symbols.
[0105] The server performs character-level statistics on the original text.
[0106] Input: The original text string that the server stored in memory in step 2.
[0107] Processing: The server performs a linear traversal of the string, counting each character regardless of its type, and accumulating all visible characters and whitespace characters to determine the total number of characters. The server maintains an integer counter during the traversal, incrementing the counter by one for each character processed. This process involves simple arithmetic operations with linear time complexity.
[0108] Output: The server outputs an integer value representing the total number of character symbols and writes this integer into the analysis result data structure for use by the subsequent threshold determination module.
[0109] Step 4: The server parses the text structure and counts the number of structural elements in the enumeration form.
[0110] The server performs line-level and structure-level parsing of the text to detect the column structure.
[0111] Input: The original text string and the analysis result data structure generated in step 3.
[0112] Processing: The server first splits the original text according to newline characters, resulting in a multi-line text list. The server then iterates through each line, determining if it contains a list-like structural element based on whether a specific pattern appears at the beginning of the line (e.g., hyphen, asterisk, number followed by a period, parentheses followed by a number, etc.). The server performs a pattern matching operation on each line, comparing the first few characters of the string with a preset pattern set to determine if the conditions are met. If they are met, the list counter is incremented. This processing is a pattern recognition operation based on string prefixes.
[0113] Output: The server outputs an integer value representing the number of enumerated structural elements, which is written into the analysis results data structure. Simultaneously, the server can output index information describing the starting position of each list item, which can be referenced in subsequent prompts using the text structure.
[0114] Step 5: The server determines the judgment threshold and generates judgment conditions.
[0115] The server determines the final threshold to be used based on the stored baseline information and user parameters.
[0116] Input: Baseline information stored in the storage device (default character threshold and default enumeration threshold), analysis results from steps 3 and 4, and user-defined threshold parameters (if any).
[0117] Processing: The server reads the default threshold from the storage device and determines whether the user has provided a custom threshold. If the user provides a value within the allowed range, the server replaces the default value with the custom value. The server stores the character threshold and the enumeration threshold in a threshold data structure and compares them with the total number of character symbols and the number of enumerations in the analysis results, generating multiple Boolean judgment results, such as "Does the character exceed the threshold?" or "Does the column exceed the threshold?" The server combines these Boolean values through a logical OR operation to obtain a final judgment result indicating whether text correction processing is needed.
[0118] Output: The server outputs a judgment data structure containing character threshold, enumeration threshold, character over-limit flag, column over-limit flag, and overall judgment result, providing a basis for generating subsequent prompt statements.
[0119] Step 6: The server generates a prompt message.
[0120] The server constructs prompts for generative artificial intelligence models based on the judgment data structure and text structure information.
[0121] Input: the original text string, the structure parsing result of step 4, the judgment data structure of step 5, and (if necessary) user-defined prompt text.
[0122] Processing: The server first determines whether the overall judgment result is "needs correction". If not, the server can choose to generate a simple prompt or skip the generative AI model call. If yes, the server selects the appropriate prompt template based on different combinations of whether characters or columns exceed the threshold, and fills the total number of characters, the number of enumerated items, and the corresponding threshold into the placeholder positions in the template. For example, for cases where two thresholds are exceeded simultaneously, the server generates a prompt similar to the following: Please examine the text according to the following requirements: 1. Compress the overall length as much as possible, reducing the word count to approximately 50% of the original text; 2. Merge or delete overly fragmented entries, keeping the number of entries to no more than 5; 3. Retain the core purpose, key steps, and important precautions.
[0123] The following text requires further consideration: [Insert original text here] When a user provides a custom prompt, the server uses that custom text as the base prompt and appends necessary supplementary explanations before and after it based on the judgment result, thereby overlaying system constraint information into the prompt. This process is essentially a string concatenation and conditional insertion operation.
[0124] Output: The server outputs a complete prompt string, which will be used as one of the control inputs for the generative artificial intelligence model.
[0125] Step 7: The server invokes a generative artificial intelligence model to generate a correction scheme.
[0126] The server uses a generative artificial intelligence model interface module to input prompts along with the original text into the model and obtain the output.
[0127] Input: The prompt string generated in step 6, the original text string, and the model call parameters (such as maximum output length, temperature parameter, penalty for repetition, etc.).
[0128] Processing: The server combines the prompt and the original text into a model input sequence. It then converts the text into a discrete labeled sequence through encoding and passes it to a generative AI model deployed locally or remotely. Internally, the model maps the labels to a vector space through embedding layers, extracts features and calculates patterns from the input sequence using multi-layer self-attention and feedforward networks, and generates output labels sequentially based on the task constraints contained in the prompt. After receiving the model's output labeled sequence, the server decodes it into readable text to obtain the correction scheme. This process involves numerical operations such as vector matrix multiplication, non-linear activation, attention weight calculation, and probability distribution sampling.
[0129] Output: The server outputs a text of the correction scheme, which is stored in memory as a character sequence and is ready to be passed to the result integration module.
[0130] Step 8: The server integrates the correction plan and analysis results and generates response information.
[0131] The server combines the revised solution with the statistical results into unified response data.
[0132] Input: Analysis results from steps 3 and 4, judgment data structure from step 5, and text of the correction scheme obtained from step 7.
[0133] Processing: The server constructs a response information data structure, filling predefined fields with the total number of original characters and symbols, the number of enumerated structural elements, the corresponding judgment threshold, the flag indicating whether automatic refinement is triggered, and the generated correction scheme. The server can also generate a brief explanatory message based on the judgment result, such as "The original character count exceeds the set threshold; the system has automatically generated a simplified version." This processing is a structured data assembly process, aggregating data from multiple sources into a single output object.
[0134] Output: The server outputs a complete response information object, which contains the original statistical indicators and the text of the correction plan, and will be sent as the response content to the terminal.
[0135] Step 9: The terminal receives and displays the correction result.
[0136] The terminal receives the response information from the server and presents it to the user.
[0137] Input: The response information object sent by the server, which includes the revised solution text, raw statistics, and explanatory messages.
[0138] Processing: The terminal parses the response information, extracting the original text, the revised text, and statistical indicators. The terminal creates two text display areas on the interface, displaying the original text and the revised text side-by-side; it also displays statistical information and explanatory messages at appropriate locations on the interface, such as "Original text character count: 1200, exceeding threshold: 500; Column count: 9, exceeding threshold: 5." The terminal also places interactive controls on the interface (such as "Accept Revision," "Continue Editing Original Text," and "Regenerate") and associates these controls with local event handling logic.
[0139] Output: The terminal outputs the visual interface status displayed on the screen, as well as new request data generated based on subsequent user operations, providing a foundation for further human-computer interaction and iterative refinement.
[0140] Step 10: Users make selections and take subsequent actions based on the displayed results.
[0141] Users decide on their next course of action after reading the original text and the revised proposal.
[0142] Input: The raw text displayed on the terminal, the revised text, statistical data, and explanatory messages.
[0143] Processing: Users visually compare the two texts and make a judgment based on their personal needs, such as choosing to directly adopt the revised solution, edit again based on the revised solution, or abandon the current revision. Users input new operation instructions into the terminal by clicking control buttons on the terminal interface or editing again in the text area. The terminal can then resend the request to the server, which may include new custom prompts or adjusted thresholds, thereby initiating a new round of processing.
[0144] Output: User actions generate new input conditions, causing the terminal to generate new request data objects, providing new input for subsequent server processing, thus forming an iterative text refinement loop.
[0145] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0146] In network services, especially in information processing systems containing user-generated content, servers typically only receive and store the text data entered by users at the terminal as is, or perform simple processing such as keyword filtering and sensitive word checks. This traditional technology suffers from the following problems: First, servers lack the ability to finely analyze document characteristics such as text length and enumeration structure, and cannot automatically determine at the system level whether the text is lengthy or structurally complex. This makes it difficult to suppress redundant information in a timely manner, leading to inefficient use of storage and bandwidth resources. Second, when generating auxiliary text (such as simplified or recommended versions), servers often rely on fixed templates or simple rules, lacking a dynamic prompt generation mechanism that collaborates with generative artificial intelligence models. They cannot adaptively adjust the instructions based on the current document characteristics, resulting in unstable quality of generated results and requiring multiple manual modifications, thus reducing overall processing efficiency. Third, servers lack the function of automatically adjusting prompts and threshold parameters based on user feedback and statistical results. This makes it impossible to adaptively optimize for different application scenarios (such as product reviews and Q&A replies) or to perform personalized system tuning for the reading habits and interaction patterns of different user groups, thus limiting the improvement of text processing systems in terms of scalability and intelligence. Fourth, when text characteristics do not meet the simplification requirements, traditional systems may still execute unnecessary model calls, wasting computational resources, or completely omit intelligent judgment logic, resulting in a lack of fine-grained control over computational resources. Therefore, how to achieve technical optimization of text processing workflows and computational resource utilization on the server side through improved text analysis, threshold control, prompt generation, and collaborative mechanisms with generative artificial intelligence models, without increasing the user's operational burden, has become a computer technology challenge that needs to be addressed.
[0147] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0148] In this invention, the server includes a processing unit for acquiring document information containing text information from a user terminal; a processing unit for analyzing the document information using an information processing program with string processing capabilities to calculate the number of characters and the number of listed structural elements, thereby extracting document characteristic information; a processing unit for comparing the document characteristic information with baseline condition information stored in a storage device and determining whether at least one of the number of characters or the number of listed structural elements exceeds a threshold; and a processing unit for generating a prompt statement containing document simplification and key point extraction instructions based on the document information and the document characteristic information when the threshold is exceeded, and generating a generative artificial intelligence model. The system includes a processing unit for generating request information, a processing unit for sending the generation request information to the computing resources running the generative artificial intelligence model via communication functions and obtaining the rewritten document information generated by the model, and a processing unit for converting the rewritten document information into display data that can be displayed on the user terminal and sending it to the user terminal. It may also include a processing unit for associating and storing the rewritten document information with the original document information and evaluation information, and performing statistical processing on them to generate control information for updating the prompt statements and / or the baseline condition information; and a control unit for skipping the generative artificial intelligence model call and directly outputting the original document information when the document characteristic information does not exceed the threshold information. This allows for automatic characteristic analysis and threshold judgment of input text on the server side, on-demand triggering and fine-tuning the call of the generative artificial intelligence model, dynamically generating prompt statements matching the current document characteristics, and adaptively optimizing the prompt strategy and threshold parameters based on user feedback. This reduces storage and communication overhead while improving the automation level and result quality of text generation and refinement processing, achieving efficient management of computing resources and information flow, and ultimately improving the overall technical performance of the computer system in processing user-generated content.
[0149] A "system" refers to a collection of devices and methods consisting of one or more information processing devices and programs running on them, used to process input information and interact with terminal devices via a network.
[0150] "User terminal" refers to an electronic device operated by a user and having display and input functions, including but not limited to computer equipment, mobile communication equipment, and other terminal equipment that can send and receive data with a server via a network.
[0151] "Document information" refers to data that is primarily text-based, sent by user terminals and received by the system. This includes natural language text, text with line breaks and enumeration structures, and metadata related to the text content.
[0152] "Textual information" refers to the language content represented in the form of characters in document information, including letters, numbers, punctuation marks, and other symbols used to form natural language sentences.
[0153] "String processing function" refers to the program functions used to parse, split, count, and replace text information, including character counting, line splitting, pattern matching, and format normalization capabilities.
[0154] "Information processing program" refers to a software program that runs on an information processing device and is used to perform operations such as string processing, comparison operations, data conversion, and network communication control. It includes application programs, script programs, and service programs on the operating system.
[0155] "Character count" refers to the numerical value obtained by counting the characters in the document information through string processing functions, including the number of letters, numbers, Chinese characters, and symbols within a predetermined range.
[0156] "List structure" refers to the structural form used to list content by item in document information, including item rows that begin with a newline and have numbers, symbols or marks, used to indicate an itemized or list-style organization of content.
[0157] "Number of elements in an enumeration structure" refers to the number of item rows or entries in a document that are identified as belonging to an enumeration structure. It is a value obtained by identifying and counting specific pattern rows.
[0158] "Document characteristic information" refers to a data set that reflects attributes such as document structure and length, obtained by the system based on document information analysis. It includes characteristic parameters such as the number of characters and the number of elements in the enumerated structure.
[0159] "Storage device" refers to a data storage medium used to store programs, parameters, document information and analysis results in a readable and writable manner, including semiconductor memory, magnetic storage medium, optical storage medium and combinations thereof.
[0160] "Baseline condition information" refers to a set of conditional data stored in a storage device for evaluating whether document characteristic information is suitable for a predetermined purpose, including one or more threshold information and corresponding applicable rules.
[0161] "Threshold information" refers to the benchmark value or condition used in the benchmark condition information to determine the document characteristic information, and is used to determine whether the number of text or the number of elements in the enumerated structure exceeds the predetermined range.
[0162] "Comparison operation" refers to the process of logically or arithmetically comparing document characteristic information with benchmark condition information, including operations such as greater than, less than, equal to, and logical combination judgments.
[0163] "Prompt statements" refer to text content constructed to provide processing instructions for generative artificial intelligence models. They are used to describe the requirements and constraints for operations such as simplifying, rewriting, or extracting key points from document information.
[0164] "Generation request information" refers to the data set used to issue generation instructions to generative artificial intelligence models, including at least prompt statements as well as parameter information and context information related to model inference.
[0165] "Generative artificial intelligence models" refer to generative models trained based on machine learning algorithms that can generate natural language text or other data based on input prompts, including large-scale language models and their variants.
[0166] "Computing resources" refers to the hardware and system resources used to run generative artificial intelligence models and information processing programs, including processors, graphics processing units, memory, network interfaces, and combinations thereof.
[0167] "Rewriting document information" refers to document text that has been modified in terms of content structure, length, or expression by a generative artificial intelligence model based on prompts and original document information.
[0168] "Communication function" refers to the network communication capability used for data transmission and reception between the system and external devices, including data packetization, transmission, reception and protocol control functions via wired or wireless networks.
[0169] "Display data" refers to formatted data that has been processed by the system and is suitable for presentation on the display device of the user terminal, including text, structured markup, and control information related to display layout.
[0170] A "recording device" refers to a storage device used to maintain, rewrite, and evaluate document information over a relatively long period of time. It may be physically the same as or different from a storage device and is used to achieve persistent data storage.
[0171] "Evaluation information" refers to feedback data related to the rewritten document information, reflecting the user's or system's feedback on the quality of the rewritten results, including information such as user selection, rating results, adoption status, and behavior logs.
[0172] "Statistical processing" refers to the aggregation analysis process based on document information, rewritten document information, and evaluation information accumulated in the recording device, used to calculate statistical quantities such as frequency, proportion, distribution characteristics, and correlation.
[0173] "Control information" refers to data generated based on statistical processing results and used to adjust system operating parameters, including settings used to update prompt statements, modify threshold information, or change trigger conditions.
[0174] "Output information for submission processing" refers to document data that the system outputs to external services for publication, archiving, or display without invoking generative artificial intelligence models or after careful consideration.
[0175] In one embodiment of the present invention, the server comprises: at least one processor, main memory, long-term storage device, network interface, and application server program serving as the access point for user terminals. The server can employ a general-purpose computer hardware architecture, such as a rack-mount server based on a multi-core central processing unit and random access memory. The server runs an application framework on an operating system, such as a scripting language-based web application framework, to implement functions such as text reception, analysis, interaction with generative artificial intelligence models, and result feedback.
[0176] The server stores information processing programs, string processing libraries, configuration data, threshold parameters, and communication modules for interacting with generative artificial intelligence model services in long-term storage. The server loads the information processing programs into main memory to perform document information analysis and control logic. The server communicates with multiple user terminals and with computing resources running the generative artificial intelligence model via a network interface.
[0177] In this invention, the terminal can be an electronic device with display components, input components, and network communication functions, such as a mobile terminal device or a desktop terminal device. The terminal runs a client application or browser application locally to provide a text input interface, display modified document information generated by the server, and receive user operation instructions. The terminal sends the user-inputted document information to the server via the network and receives display data returned from the server.
[0178] In this invention, users operate a text input interface via a terminal, inputting document information including comments, descriptions, questions, or other natural language content. Users can view the rewritten document information returned by the server on the terminal and, as needed, adopt, partially edit, or ignore the rewritten document information. The user's selection results can be sent to the server as evaluation information for subsequent statistical processing and system parameter adjustments.
[0179] The server uses string processing functions to process document information within its information processing program. It uses a character counting function to iterate through the character sequences in the document information and count the number of characters. The server then splits the document information into multiple lines of text using a line-segmentation algorithm. For each line, the server uses a pattern matching algorithm to determine if it begins with a predetermined enumeration marker, such as a hyphen, asterisk, dotted number, or parentheses. The server records the number of lines identified as enumeration structures as the number of elements in the enumeration structure. The server then generates document characteristic information containing both the number of characters and the number of enumeration structure elements.
[0180] The server pre-stores a set of baseline condition information in the storage device. In a simple embodiment, the server sets the maximum number of characters to 500 and the maximum number of enumerated structural elements to 5. In another embodiment, the server can configure different threshold sets for different business scenarios, such as a lower threshold for product review scenarios and a higher threshold for technical document summary scenarios. During runtime, the server compares the calculated document characteristic information with the baseline condition information, determines whether the number of characters exceeds the corresponding threshold, whether the number of enumerated structural elements exceeds the corresponding threshold, and generates a logical judgment result indicating whether the threshold is exceeded.
[0181] When the server determines that at least one feature exceeds a threshold, it constructs a prompt statement to invoke the generative artificial intelligence model. The server maintains multiple prompt statement patterns in the form of templates in a storage device. The server selects an appropriate template based on the document type, the reason for exceeding the threshold, and the desired simplification level, and embeds the user's original document information into the template to generate the complete prompt statement. In one embodiment, the server generates the following prompt statement: Please shorten the following product review to no more than 200 characters, retaining only the most important reasons for purchase and 1-2 key drawbacks, and output it in Simplified Chinese: "The user's original comment text" In another embodiment, the server generates the following prompt statement: Below is a product review containing numerous entries and repetitive information. Please condense it into 3-5 concise points, covering both the main advantages and the most significant disadvantages. Please use Simplified Chinese and ensure it's suitable for posting in an e-commerce review section. "The user's original comment text" The server can also generate the following prompt statements as needed: Please rewrite the following product review in a more conversational and easier-to-understand style, suitable for the average consumer, without changing the facts. Keep the total word count between 150 and 250 words. Only output the revised text: "The user's original comment text" After generating the prompt statement, the server constructs a generation request message, which includes at least the prompt statement, a target generation length parameter, a temperature parameter, a sampling strategy parameter, and a context identifier. The server sends the generation request message to the computing resources running the generative artificial intelligence model via a network interface.
[0182] In one optional implementation, the server employs a generative artificial intelligence model based on a transformer architecture. This model consists of a multi-layered self-attention encoder-decoder network, each layer including a multi-head self-attention sublayer and a feedforward neural network sublayer. During model training, the server uses a large-scale corpus of natural language text as training data, updating the model weights by minimizing the cross-entropy loss function. During error backpropagation, the server employs gradient descent or its variants to update the model's parameter matrix and can use techniques such as weight decay, dropped units, and learning rate scheduling to prevent overfitting. During training, the server can use data augmentation techniques, such as random truncation, synonym replacement, or sentence order perturbation of the training text, to improve the model's robustness to different text structures.
[0183] During the inference phase, the server encodes the prompts as a sequence of tags. Each tag is mapped to a vector representation through an embedding layer, and attention weights and intermediate representations are calculated via forward propagation in a multi-layer self-attention network. The server samples or selects the next output tag based on a probability distribution, iteratively generating rewritten document information. The server can control the maximum generation length, stop tags, and repetition penalty parameters to avoid outputting excessively long or highly repetitive content.
[0184] After receiving the output of the generative artificial intelligence model, the server performs post-processing on the rewritten document information. The server can confirm whether the target length requirement is met by recounting the number of characters, and truncate or append explanations if necessary. The server can also clean up extra leading whitespace, consecutive blank lines, and abnormal symbols. In some implementations, the server can apply simple rule filtering to the generated results, such as deleting sentences containing prohibited words or violating platform policies, thereby ensuring that the system output meets predetermined content security requirements.
[0185] After rewriting and post-processing the document information, the server converts both the rewritten and original document information into display data. In a simple implementation, the server can package both into a markup-formatted text file and add structured markup to distinguish between the "original content" and "recommended rewritten content" areas. The server then sends this display data to the user terminal via a network interface.
[0186] After receiving the display data, the terminal renders the interface locally, displaying both the original and rewritten document information in the display area. The terminal can provide users with buttons such as "Use recommended rewrite," "Editor's recommended rewrite," and "Continue using the original text." After the user selects one of these actions, the terminal updates the local text content according to the user's intention and subsequently sends the final submission text to the server or other backend services.
[0187] Users can compare the original and rewritten document information in the terminal interface. They can make minor manual modifications to the rewritten document to quickly create a well-structured, concise, and appropriately sized final text. Users can also send feedback to the server via the interface, such as whether the document was adopted, a brief evaluation, or a rating.
[0188] The server stores the original document information, rewritten document information, and evaluation information together in a recording device. The server periodically performs statistical processing on the data in the recording device, calculating metrics such as the adoption rate of different prompt templates across different document types, average number of edits, and user satisfaction scores. Based on these statistical results, the server can automatically adjust the prompt content, such as adding instructions to "highlight weaknesses" or shortening the target word count range. The server can also automatically adjust thresholds in the baseline conditions based on statistical results; for example, when many comments are around 400 words and users still expect further simplification, the maximum word count threshold can be lowered to trigger the rewriting process earlier.
[0189] Through this closed-loop statistical and adjustment mechanism, the server moves beyond simply executing fixed rules. Instead, it leverages the computer's statistical computing and model reasoning capabilities to continuously optimize internal parameter configurations. This optimization directly impacts the server's internal data processing paths and resource scheduling logic, thereby improving processing efficiency and output quality at the system level.
[0190] In another implementation, when document characteristic information does not exceed a threshold, the server does not invoke the generative AI model. Instead, it directly sends the original document information as output for submission processing to the backend business system. This on-demand invocation strategy avoids unnecessary occupancy of the generative AI model, significantly reducing computing resource consumption and network communication load. In multi-user concurrent scenarios, the server can handle more requests with the same hardware configuration, improving system throughput.
[0191] The server achieves the following technical effects through the aforementioned structure and algorithm: First, by using pre-defined feature analysis and threshold judgment mechanisms, the server reduces the number of calls to the generative AI model, thereby shortening the overall response time and improving resource utilization. Second, based on a combination of templated prompts and document feature information, the server enables more precise guidance of model output, making it easier to obtain results that meet length and structure requirements, thus reducing the burden of subsequent manual editing and secondary processing. Third, by continuously updating prompts and thresholds through recording devices and statistical processing modules, the server forms a dynamic and adaptive parameter adjustment mechanism, significantly improving the long-term processing performance and stability of the system in different application scenarios.
[0192] In various embodiments of this invention, the server employs a processing path different from traditional manual editing. Instead of simply simulating manual sentence-by-sentence deletion, the server quantitatively analyzes features such as text length and enumeration structure, encoding these features as input conditions into prompt statements. This drives a generative artificial intelligence model to perform global semantic reconstruction in the vector space, thereby achieving efficient compression while preserving the core semantics. This comprehensive processing approach based on feature-driven and probabilistic language models represents a novel text processing strategy within computers, rather than simply automating human rules.
[0193] Through the above implementation methods, the server, terminal, and user work together in the system, enabling the processes of receiving, analyzing, rewriting, and responding to text to be not only executed automatically, but also optimized at the algorithm level within the computer, thereby achieving technical improvements in processing speed, the structuring of generated results, communication load, and storage consumption.
[0194] use Figure 12 The processing procedure is explained.
[0195] Step 1: The user enters document information on the terminal.
[0196] Users type or paste natural language text such as comments and descriptions into the text input interface provided on the terminal, and select the "Submit" or "Next" button by touch or click.
[0197] Input: The raw text content entered by the user in the terminal.
[0198] Output: The document information data structure temporarily stored inside the terminal.
[0199] Based on user actions, the terminal organizes the text content into a data object containing text fields and basic metadata (such as timestamps and session identifiers), ready to send it to the server.
[0200] Step 2: The terminal sends document information to the server.
[0201] The terminal encapsulates a data object containing document information into a network request message through the network communication module, and sends the request to the specified interface address of the server using a predetermined application layer protocol.
[0202] Input: The document information data structure temporarily stored in the terminal.
[0203] Output: The request message transmitted to the server over the network.
[0204] Before sending, the terminal can attach information fields such as user identifier and terminal type so that the server can associate and record them during subsequent processing.
[0205] Step 3: The server receives and preprocesses the document information.
[0206] After receiving the request message from the terminal via the network interface module, the server parses the request body in the application and extracts the document information fields. The server performs preprocessing operations on the text content, including removing leading and trailing whitespace, standardizing line breaks, and merging consecutive blank lines.
[0207] Input: The original request message sent by the terminal (containing document information).
[0208] Output: The preprocessed docstring in the server's memory.
[0209] The server uses string processing functions to iterate through and replace the input text, storing the results in a new memory buffer as the basis for subsequent analysis steps.
[0210] Step 4: The server calculates the number of characters.
[0211] The server performs a character counting operation on the preprocessed document string. The server uses a character traversal algorithm to read the encoded value of each character and accumulates the number of valid characters using a counter.
[0212] Input: The preprocessed document string.
[0213] Output: An integer representing the number of characters.
[0214] During the counting process, the server can filter out certain control characters that are not included in the length count, and determine which characters to include in the statistics according to preset rules, thereby obtaining accurate text quantity information.
[0215] Step 5: The server identifies and counts the number of listed structural elements.
[0216] The server splits the preprocessed document string line by line, generating a list of lines. For each line, the server applies pattern matching logic, such as checking if the line begins with a hyphen, asterisk, period, or parentheses. The server counts the lines that meet the enumeration criteria to obtain the number of elements in the enumeration structure.
[0217] Input: The preprocessed document string.
[0218] Output: An integer representing the number of listed structural elements.
[0219] During the recognition process, the server performs string prefix comparison or regular expression matching operations and saves the index of the matching lines and the total number as part of the document characteristic information.
[0220] Step 6: The server generates document characteristic information.
[0221] The server combines the number of text characters and the number of listed structural elements into structured data, forming document characteristic information. The server then adds metadata such as document identifier, user identifier, and session identifier to this information for subsequent association and recording.
[0222] Input: Integer value for the number of characters, integer value for the number of listed structural elements.
[0223] Output: Document characteristic information data structure.
[0224] The server constructs records containing multiple fields in memory, and stores the analysis results and document source information in a unified manner for threshold judgment and statistical purposes.
[0225] Step 7: The server obtains baseline condition information.
[0226] The server reads the currently applicable baseline conditions from the storage device, including parameters such as the text quantity threshold and the enumerated structural element quantity threshold. The server selects the corresponding configuration record based on the business scenario identifier.
[0227] Input: Scene identifier or default configuration identifier.
[0228] Output: Baseline condition information data structure (containing multiple threshold fields).
[0229] After reading the data, the server loads the threshold parameters into the runtime environment and caches them in memory to reduce the overhead of frequent access to storage devices.
[0230] Step 8: The server performs a threshold comparison operation.
[0231] The server compares the values in the document characteristic information with the corresponding thresholds in the baseline condition information. The server performs greater than, less than, or equal to checks on the number of characters and the number of listed structural elements and the number of items, respectively, and generates a comprehensive judgment result based on logical rules to determine whether the limits are exceeded.
[0232] Input: Document characteristic information data structure, baseline condition information data structure.
[0233] Output: Exceedance flag (e.g., a boolean value) and information about the reason for the excess.
[0234] During the comparison process, the server can generate a description field indicating whether the number of characters exceeds the threshold, the number of listings exceeds the threshold, or both, so that an appropriate prompt statement template can be selected later.
[0235] Step 9: The server determines whether it needs to invoke a generative artificial intelligence model.
[0236] The server determines whether to generate rewrite document information based on the out-of-bounds flag and the reason for the out-of-bounds error. The server terminates the rewrite process if the conditions are not met, and enters the prompt statement construction phase if the conditions are met.
[0237] Input: Exceedance flag and reason for exceeding the limit.
[0238] Output: The decision results of the generative artificial intelligence model.
[0239] The server performs simple logical operations in this step, avoiding unnecessary use of external computing resources and thus reducing the overall system load.
[0240] Step 10: The server constructs a prompt statement template and populates the document information.
[0241] When the server determines that a generative AI model needs to be invoked, it selects an appropriate prompt template based on the reason for exceeding the limit, such as a template biased towards length compression, item merging, or conversational optimization. The server embeds the original document information into the placeholder portion of the template to generate the complete prompt string.
[0242] Input: Original document information, reason for exceeding the limit information, and a set of pre-stored prompt statement templates.
[0243] Output: The complete message text.
[0244] The server uses string concatenation and replacement operations to replace placeholders in the template with actual document content, and sets constraints such as target word count and number of key points according to different scenarios.
[0245] Step 11: The server generates the generation request information to be sent to the generative artificial intelligence model.
[0246] The server combines the prompt, target length parameter, temperature parameter, sampling strategy, etc., into a generated request message. The server attaches a session identifier, request number, and security authentication information to this request for interfacing with external model services.
[0247] Input: Prompt text, model parameter settings, and session-related metadata.
[0248] Output: Generate request information data structure.
[0249] In this step, the server encapsulates all necessary fields to form a request payload that conforms to the external model service interface specification.
[0250] Step 12: The server sends generation requests to the generative artificial intelligence model.
[0251] The server transmits the request information to the computing resources running the generative artificial intelligence model via a network interface. The server sends the request body to the model server endpoint using a defined communication protocol and interface path.
[0252] Input: Generate request information data structure.
[0253] Output: Network request messages sent to the generative artificial intelligence model service.
[0254] During the sending process, the server handles specific network operations such as connection establishment, data serialization, and error retries to ensure that requests arrive reliably.
[0255] Step 13: The server receives the rewritten results returned by the generative artificial intelligence model.
[0256] The server waits for a response message from the model service on the network interface. After the response arrives, the server parses the message content and extracts the rewritten document text and other information that may be included (such as generated confidence scores, length statistics, etc.).
[0257] Input: The response message returned by the generative artificial intelligence model service.
[0258] Output: Original generated text and related metadata.
[0259] In this step, the server performs parsing and data type conversion, transforming the data returned by the external service into text strings and numeric fields that can be processed internally.
[0260] Step 14: The server performs post-processing on the rewritten document information.
[0261] The server performs formatting standardization on the generated text, such as removing extra line breaks at the beginning, merging multiple spaces, and standardizing punctuation. The server can then recount the text to confirm if it falls within the target range; if it does, it can truncate or add supplementary prompts.
[0262] Input: The original rewritten text output by the generative artificial intelligence model.
[0263] Output: Rewritten document information text after editing.
[0264] During post-processing, the server applies operations such as string transformation, regular expression replacement, and length validation to ensure that the output text conforms to preset standards in form.
[0265] Step 15: The server generates display data for presentation purposes.
[0266] The server combines the original document information and the rewritten document information to construct a display data structure that includes the content and identification information of both. The server can add fields to represent labels such as "original text" and "recommended rewrite" so that the terminal can distinguish and display them on the interface.
[0267] Input: Original document information text, and rewritten document information text.
[0268] Output: Display data structure.
[0269] In this step, the server performs data packaging and structured markup processing, enabling the terminal to directly render the view based on this structure.
[0270] Step 16: The server sends display data to the terminal.
[0271] The server sends display data to the requesting terminal via its network interface. The server includes the necessary status codes and data payload in its response message so that the terminal can correctly parse it.
[0272] Input: Display data structure.
[0273] Output: The response message sent to the terminal.
[0274] In this step, the server handles data encoding and transmission control to ensure that the data for display arrives at the terminal intact and without errors.
[0275] Step 17: The terminal receives and displays the original and rewritten document information.
[0276] After receiving the server's response message, the terminal parses and displays the data. The terminal simultaneously displays the original document information and the rewritten document information in the user interface, along with operation controls such as buttons for "Use recommended rewrite," "Edit recommended rewrite," and "Continue using the original text."
[0277] Input: Display data sent by the server.
[0278] Output: The state of the interactive interface displayed on the terminal screen.
[0279] During the rendering process, the terminal determines the layout of the display area based on the tag field, enabling users to intuitively compare the original text with the rewritten content.
[0280] Step 18: Users select the processing method for rewriting document information on the terminal.
[0281] Users can view the original and rewritten text on the terminal. They can choose to use the rewritten text directly, edit it further, or leave the original text unchanged. Users can then perform corresponding clicks or editing operations on the terminal.
[0282] Input: The interface displaying the original and modified document information on the terminal.
[0283] Output: User selection result (adopt, adopt after editing, or reject) and possible edited text.
[0284] The terminal temporarily stores the user's final confirmation text and selection information locally, preparing to send it to the business backend or resubmit it to the server as evaluation information.
[0285] Step 19: The terminal sends evaluation information and final text (optional) to the server.
[0286] After user confirmation, the terminal can send the user's adoption status and the final text as evaluation information to the server. The terminal includes information in the data indicating whether the rewritten result was adopted, whether it was modified, and the user's rating.
[0287] Input: User's final confirmation text and rating selection.
[0288] Output: A request message containing evaluation information and the final text.
[0289] The terminal sends this request through the network interface to provide basic data for the server's statistical processing and parameter optimization.
[0290] Step 20: The server records information about rewritten documents and evaluations, and updates the statistical results.
[0291] After receiving the evaluation message from the terminal, the server associates and stores the original document information, rewritten document information, and evaluation information in the recording device. The server periodically or as needed performs statistical processing on the accumulated data, calculates the adoption rate and satisfaction rate under different combinations of prompt statements and thresholds, and updates the internal benchmark information and prompt statement template content accordingly.
[0292] Input: Request message containing evaluation information and final text, and previously stored records of original and rewritten document information.
[0293] Output: Updated recording device content, latest statistical results, and adjusted prompt statement templates or threshold parameters.
[0294] In this step, the server performs aggregation statistics, ratio calculations, and parameter update operations, thereby gradually optimizing system behavior and continuously improving the generative artificial intelligence model invocation strategy and prompt statement design.
[0295] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0296] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0297] In scenarios such as text communication, online editing, and electronic submission, using generative artificial intelligence models to automatically refine and suggest rewrites has gradually become a common practice. However, current technologies typically only input the user's original text directly into the generative AI model, which then outputs a rewrite result or general modification suggestions in a single step. This approach has the following shortcomings at the computer technology level: On the one hand, existing systems lack structured parsing and quantitative analysis of the text itself before sending it to generative artificial intelligence models. They cannot accurately assess the complexity of the text and the necessity of rewriting from the dimensions of symbolic information, structural information and other features. This leads to the use of the same calling strategy for both simple and complex texts, resulting in a waste of computing resources and an increase in response latency.
[0298] On the other hand, existing systems typically use fixed templates when constructing prompts, attaching the entire text to the prompts as is. They do not adaptively set evaluation conditions based on the parsed text features, nor do they dynamically control the content, granularity, and emphasis of the prompts. This makes it difficult for generative AI models to focus their computation on the parts that truly need optimization, affecting the readability and operability of the generated results.
[0299] Furthermore, existing systems, after receiving the results from generative artificial intelligence models, often directly display the raw output, lacking a process to structure and format the results according to their constituent elements. This results in the user interface being unable to efficiently present different types of modification suggestions, making it difficult for users to quickly understand and adopt them, thus reducing the efficiency of human-computer interaction.
[0300] Furthermore, existing technologies often neglect the security and protocol compatibility of text and prompts at the network transmission layer, and lack a mechanism for uniformly adopting encrypted communication and standardized communication protocol management during the transmission phase. This results in limitations in security, reliability, and scalability when the system is deployed on a large scale and integrated across platforms.
[0301] Therefore, the technical problem to be solved by this invention is how to introduce a complete and programmable text parsing and feature extraction process into a computer system, automatically set evaluation conditions based on the parsing results and generate prompts for generative artificial intelligence models, and perform structured organization, format conversion and security notification after obtaining the results, so as to improve the accuracy, efficiency, security and user usability of generative artificial intelligence model invocation at the system level.
[0302] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0303] In this invention, the server includes a device for receiving text information, a device for parsing the received text information to extract feature information including symbolic information and structural information, a device for setting evaluation conditions based on the parsing results, a device for generating prompt statements for a generative artificial intelligence model based on the evaluation conditions, a device for inputting the prompt statements and the text information into the generative artificial intelligence model to generate an article improvement plan, a device for organizing the article improvement plan according to its constituent elements and converting it into display format information according to predetermined rules, and a device for generating notification information containing the article improvement plan based on the display format information and sending it to the submitting entity through a communication path. Furthermore, encrypted communication is used when sending the text information and the prompt statements, and the article improvement plan is notified to the submitting entity as an electronic document or electronic message according to a communication protocol when sending the article improvement plan. This allows for feature-driven preprocessing and condition setting of text information on the server side in a programmatic manner. This enables generative AI models to efficiently generate structured article improvement solutions under controlled prompts. The processed results are then transmitted to the submitting entity in a format adapted for terminal display through a secure and standardized communication process. This improves the efficiency of computing resource utilization, generation quality, and human-computer interaction experience during the text refinement process, achieving overall technical improvements to the text processing system based on generative AI models in terms of processing flow, security, and usability.
[0304] “Text information” refers to sequences of characters, symbols and their combinations thereof that are input by users or other information sources and represented electronically, including natural language text, punctuation marks and control symbols used to represent paragraphs, lists and other structures.
[0305] "Analysis" refers to the process of programmatically processing text information to identify its internal structure and features, including analyzing characters, words, sentences, paragraphs, and list structures to extract relevant feature information.
[0306] "Symbol information" refers to attribute information extracted from text information related to characters, punctuation marks, control symbols, and their arrangement, such as the number of characters, the types and frequencies of punctuation marks, and the presence and number of list markers.
[0307] "Structural information" refers to attribute information extracted from text information that reflects the text's hierarchy and logical organization, such as sentence division, paragraph division, distinction between headings and body text, and the hierarchical relationship between lists and sublists.
[0308] "Feature information" refers to the set of statistical characteristics and pattern features of text in the symbolic and structural dimensions obtained through the parsing of text information, including various descriptive data such as symbolic information and structural information.
[0309] "Evaluation criteria" refers to a set of parameters, rules, or thresholds set based on the feature information obtained from parsing text information, used to guide generative artificial intelligence models in performing text quality evaluation and generating improvement suggestions. These criteria include settings for length, structural complexity, number of paragraphs, number of lists, etc.
[0310] "Generative artificial intelligence models" refer to artificial intelligence systems that use machine learning algorithms, especially models based on deep learning and large-scale parameter training, to automatically generate language-related output content based on input prompts and text information.
[0311] "Prompt statements" refer to input text constructed to guide generative artificial intelligence models to perform specific tasks. This text contains descriptions of task objectives, evaluation dimensions, output formats, and contextual information related to the text to be processed.
[0312] "Article improvement plan" refers to the suggestions for improving the quality of the original text in the output of the generative artificial intelligence model based on prompts and text information. These suggestions include explanations of problems in the original text, recommendations for modification directions, and rewriting examples.
[0313] "Components" refers to the different content units that can be distinguished and processed in the article improvement plan, including overall evaluation, problem list, specific modification suggestions, example rewritten paragraphs, and other categories of content.
[0314] "Display format information" refers to the descriptive information obtained by structuring and formatting an article in order to display the article improvement plan in a readable and distinguishable form on a terminal device or user interface. This includes paragraphs, sections, list markers, heading markers, and style instructions.
[0315] "Notification information" refers to electronic information generated based on display format information and sent to the submitting entity through a communication path. It contains all or part of the article improvement plan and is used to prompt the submitting entity to view and adopt the improvement suggestions.
[0316] "Communication path" refers to the network channel or communication link used to transmit notification information between the server and the submitting entity's terminal, including wired networks, wireless networks, and the various communication protocol channels they carry.
[0317] "Submitting entity" refers to the entity that submits text information to obtain article improvement suggestions. It can be a natural person user, a legal person entity, or other entity with an account identifier and the ability to receive notification information.
[0318] "Encrypted communication method" refers to a communication technology solution that encrypts transmitted data to prevent unauthorized access when sending text information, prompts, or article improvement plans through a communication path. This includes implementation methods based on symmetric encryption, asymmetric encryption, and secure transmission protocols.
[0319] A "communication protocol" refers to a set of predetermined communication rules used to ensure data format, transmission order, error detection, and retransmission mechanisms when transmitting electronic documents or messages through a communication path. These rules include application layer, transport layer, and network layer protocols.
[0320] "Electronic document" refers to a document entity that is stored and transmitted in the form of electronic data. Its content may include text, markup, and formatting information, and it can be viewed, edited, or archived in the form of a document on a terminal device.
[0321] "Electronic messaging" refers to short text or structured message units transmitted between different terminals or accounts through network services or messaging service systems, used to deliver information to the recipient in conversation, notification, or push scenarios.
[0322] This invention, in its embodiment, integrates a server, a terminal, and a user to form a text deduction system based on a generative artificial intelligence model. The specific implementation of this invention is described below in conjunction with its hardware structure, software modules, data structure, and algorithm flow, so that those skilled in the art can implement this invention accordingly.
[0323] The server may, in terms of hardware, include one or more computing devices equipped with a central processing unit (e.g., a multi-core general-purpose processor), main memory, persistent storage, and a network interface controller. In one embodiment, the server runs a general-purpose operating system, such as a Unix-like operating system, web server software (e.g., a reverse proxy server), an application server runtime environment (e.g., a runtime environment supporting scripting languages), and encrypted communication libraries (e.g., cryptographic libraries for implementing TLS / HTTPS). In another embodiment, the server may further include computing nodes equipped with graphics processing units or dedicated acceleration chips for performing large-scale matrix operations.
[0324] The terminal may include, in terms of hardware, a smartphone, tablet, or personal computer, and has a display device, input device (touchscreen, keyboard, etc.), local processing unit, storage unit, and wireless or wired network interface. In terms of software, the terminal runs an operating system (such as a mobile operating system or desktop operating system) and a browser or dedicated application. The terminal uses this browser or application to send text messages to the server and receive notification messages returned by the server.
[0325] The user inputs the text information to be considered through the terminal's graphical user interface. In one embodiment, the user inputs the following example text into the terminal's text input box: "The event was a great success, and we are all very happy." After confirming the input, the user triggers encrypted communication between the terminal and the server by operating the send control on the terminal. In one embodiment, the terminal uses the HTTPS protocol supporting TLS 1.2 or TLS 1.3 to encapsulate the text information and user account-related identification information into a request message, and sends it to the server through the network interface.
[0326] At the receiving end, the server uses the operating system's network stack and encryption library to decrypt and verify the integrity of received data packets. After decryption, the server forwards the HTTP request data to the application server module. At the application layer, the server uses a JSON parsing library or equivalent parsing mechanism to extract text information fields from the request payload and stores this text information in a string data structure in main memory. In one implementation, the server generates a unique request identifier for each text information reception operation and stores this identifier along with metadata such as text length and reception time in a log database for subsequent monitoring and optimization.
[0327] When parsing text information, the server uses a text parsing module running on the central processing unit to break down the text string into character sequences, punctuation sequences, and potential structural markers. In one embodiment, the server uses a regular expression engine to count the total number of characters, sentence boundaries, paragraph separators, and bullet points (such as "•", "-", "1.", etc.) in the text, thereby generating symbol information. Simultaneously, the server identifies the text's structural information, such as the number of sentences, paragraphs, list levels, and sentence distribution within each paragraph, through rule matching and segmentation algorithms. The server combines the symbol information and structural information into feature information records. This feature information can be stored using key-value mapping or a structure, for example, containing fields such as "total number of characters," "number of list items," "number of paragraphs," and "average sentence length."
[0328] After obtaining the feature information, the server generates evaluation criteria based on a preset algorithm or configurable rules. In one implementation, the server calculates a text complexity index based on the total number of characters and the number of list items, and compares this index with multi-level thresholds to determine the subsequent call parameters for the generative artificial intelligence model. For example, when the total number of characters is short and the number of list items is zero, the server sets the evaluation criteria to "focus on overall clarity and information content"; while when the text is long and the number of list items is large, the evaluation criteria are set to "focus on checking structural coherence and list logic consistency". The server encodes these criteria into an internal parameter set for subsequent construction of prompt statements.
[0329] When constructing the prompt statement, the server reads template text that matches the current evaluation criteria from the prompt template storage. In one embodiment, the server uses the following prompt statement template: "You are a professional Chinese writing editor. Please provide a detailed evaluation of the following text in terms of grammar, clarity of expression, logical coherence, and information content, and offer specific suggestions for revision. Also, please provide at least one complete rewritten example. The text is as follows: '{User Text}'. Please first give an overall evaluation, then list the revision suggestions one by one, and finally provide a rewritten example." The server replaces the placeholder "{user text}" with the text information submitted by the user, for example: "You are a professional Chinese writing editor. Please provide a detailed evaluation of the following text in terms of grammar, clarity of expression, logical coherence, and information content, and offer specific suggestions for revision. Also, please provide at least one complete rewritten example. The text is as follows: 'This event was a great success, and we were all very happy.' Please first give an overall evaluation, then list the revision suggestions one by one, and finally provide a rewritten example." In another embodiment, the server refines the prompts based on evaluation criteria. For example, when the text structure is complex, it adds instructions such as "Please focus on checking the connections between paragraphs and list items, and point out structural problems." In this case, the server uses evaluation criteria to drive the dynamic generation of prompt content, enabling the generative AI model to focus on different analysis dimensions under different text features, thereby improving the efficiency of computing resource utilization and the relevance of the generated results.
[0330] When the server invokes the generative AI model, it sends the constructed prompt statement along with necessary context parameters to the computing node hosting the generative AI model. In one embodiment, the generative AI model employs a deep neural network based on the Transformer architecture. The model includes multi-layer self-attention encoders and decoders, with each layer containing a multi-head attention sublayer, a feedforward sublayer, and a normalization sublayer. The server sends a request containing the prompt statement to the model service via an application programming interface (API). Internally, the model service first segments the input text into tokens, mapping the text to a discrete sequence of tokens, and then uses an embedding matrix to convert the token sequence into a continuous vector representation. During inference, the model performs matrix multiplication and addition operations, attention weight calculations, and nonlinear transformations on these vectors, thereby generating a high-dimensional representation reflecting the semantics and structure of the text.
[0331] During the training phase, the server can use a pre-built text dataset to perform supervised fine-tuning of the generative AI model. The server defines a loss function on this dataset, such as a language modeling loss based on cross-entropy, and calculates the gradients of the weights at each layer of the model using backpropagation. During training, the server updates the model parameters using optimization algorithms (such as adaptive learning rate optimization methods), and can employ gradient pruning, regularization, and data augmentation strategies (such as randomly masking some words and rearranging the order of some sentences) to improve the model's generalization ability and robustness. Through these training processes, the server enables the model to possess higher grammatical sensitivity and style discrimination capabilities when facing text reasoning tasks, thereby generating more accurate text improvement solutions during actual reasoning.
[0332] After receiving the output text from the generative artificial intelligence model, the server treats the output as an article improvement proposal. In one embodiment, the server performs secondary structuring processing on the article improvement proposal. The server segments the text into multiple constituent elements by matching specific tags in the model output (e.g., "Overall Evaluation:", "Issue 1:", "Suggestion 1:", "Rewrite Example:"). During the structuring process, the server assigns type labels to each constituent element, such as "overall_review", "issue_list", "suggestion_list", "rewrite_example", etc., and combines these labels with the corresponding text paragraphs to form records, which can be stored using a tree or list structure.
[0333] When generating display format information, the server maps the article improvement proposals to an appropriate display structure based on the terminal type and display capabilities. In one implementation, the server converts the overall evaluation into a title style, the problem list and suggestion list into ordered or unordered lists, and presents the rewrite examples as code blocks or highlighted paragraphs. During this process, the server performs text-to-markup language conversion operations, such as converting structured records into markup documents or mapping them to data structures that can be directly rendered by the terminal application. This structuring and formatting process reduces the terminal's burden of parsing complex text, thereby reducing terminal computational consumption and improving front-end rendering speed.
[0334] When sending notification information, the server selects a communication path based on the user's notification preferences and account information. In one embodiment, the server sends the notification information containing the article improvement proposal as an electronic document to the user's email address via an email protocol. During this process, the server constructs the email header and body, embedding content corresponding to the display format information within the email body. In another embodiment, the server sends the article improvement proposal as an electronic message to the user's instant messaging account via a message service interface. In all sending scenarios, the server uses encrypted communication methods, such as TLS at the transport layer and adheres to a unified communication protocol at the application layer, to ensure the confidentiality and integrity of the text information, prompts, and article improvement proposal during network transmission.
[0335] Upon receiving a notification, the terminal parses the email or message content using a local application. Based on display format information, the terminal presents the overall evaluation, a list of issues, a list of suggestions, and rewrite examples on the screen in a segmented manner. In one implementation, the terminal provides interactive controls, allowing users to switch between different components or collapse / expand the list, enabling them to quickly focus on the most relevant parts. In another implementation, the terminal allows users to insert rewrite examples into a local editor with a single click for further customization.
[0336] After reading the article's improvement plan, users can modify the original text based on the overall evaluation and specific suggestions. In one example, users can refer to the following model output example for modification: Overall assessment: The original text is concise, but lacks sufficient information. It lacks descriptions of the event content, participants, and specific results, making it difficult for readers to form a clear impression.
[0337] Suggested revisions: 1. Add the time, location, and participants of the activity.
[0338] 2. Describe the specific aspects or highlights of the activity.
[0339] 3. Increase participants' reactions or feedback to make emotional expression more specific.
[0340] Rewrite example: The event was a great success, with over 300 industry professionals in attendance. During the roundtable discussions and Q&A sessions, participants actively shared their experiences, creating a lively atmosphere. Many attendees expressed that they had gained a great deal and hoped to participate in similar exchange activities in the future. Users can use this to expand the original text content in the terminal's editing interface, thereby significantly improving text quality.
[0341] By introducing feature information extraction and evaluation condition setting during the text parsing stage, the server enables prompts to adaptively adjust to texts of varying complexity and structural features. The server reduces the processing burden on the terminal and minimizes human-computer interface ambiguity through structured organization and display format conversion. Furthermore, the server enhances data transmission security and reliability through encrypted communication and protocol standardization. These server-executed procedural processing techniques not only achieve automated refinement in functionality but also improve traditional systems in terms of computational resource utilization, response latency, output availability, and communication security. This constitutes an optimization of computer technology itself, rather than simply automating manual editing.
[0342] Servers can employ different generative AI model structures in different implementations. For example, a server can use a unified large language model, or it can use multiple sub-models optimized for grammar checking and style rewriting respectively. In this multi-model architecture, the server can select appropriate models or combinations based on evaluation criteria, thereby further reducing computational costs while ensuring quality. The server can also dynamically adjust model inference parameters (such as output length limits and sampling temperature) based on system load and network conditions to maintain overall response performance in multi-user concurrent scenarios.
[0343] Through the above embodiments and alternative configurations, the present invention provides a text deliberation system that can be implemented in existing general-purpose hardware and software environments. This system utilizes a generative artificial intelligence model and a prompt-driven control mechanism to achieve refined control of the text processing flow and overall improvement at the computer technology level.
[0344] use Figure 13 The processing procedure is explained.
[0345] Step 1: The user enters text information on the terminal and initiates a send request.
[0346] Input: User's natural language text (e.g., "This event was a great success, and we are all very happy.").
[0347] Output: A send command containing text information.
[0348] The user types the text to be reviewed into the text input box on the terminal. After completing the input, the user clicks "Send" or performs an equivalent operation on the terminal interface. This operation instructs the terminal to read the text data from the interface components and prepare for transmission over the network.
[0349] Step 2: The terminal constructs request data and sends it to the server via encrypted communication.
[0350] Input: User-input text information, user account identifier, and terminal environment information.
[0351] Output: Network request packets sent to the server via HTTPS.
[0352] The terminal reads a text string from the interface controls, assembles the text with the user identifier into a request message, and generates structured data containing text fields in its local memory. The terminal then calls the network interface provided by the operating system, serializes this structured data into a byte stream, establishes a secure connection via a TLS handshake, and sends it as an HTTPS request to the network address specified by the server.
[0353] Step 3: The server receives and parses the request messages from the terminal.
[0354] Input: Network data packets transmitted via TLS encryption.
[0355] Output: The parsed text information and related metadata.
[0356] The server receives data packets at the network interface, decrypts them using an encryption module, and recovers the original HTTP request. The server then calls the application-layer parsing component to parse the request message, extracting text fields and user identification fields. The server stores the text information as a string in main memory and records metadata such as text length and reception time.
[0357] Step 4: The server preprocesses and extracts features from the text information.
[0358] Input: The original text string and its corresponding metadata.
[0359] Output: Cleaned text information, symbol information, and structural information.
[0360] The server calls string processing functions on the text string, removing leading and trailing whitespace, standardizing newline character format, and eliminating invisible control characters. The server uses regular expressions and delimiters to count the total number of characters, punctuation types and quantities, sentence boundary positions, paragraph separators, list markers, etc., forming symbol information. Based on periods, newlines, and list markers, the server parses out the number of sentences, paragraphs, and list levels, forming structural information. The server stores the cleaned text, symbol information, and structural information into a feature data structure.
[0361] Step 5: The server generates evaluation criteria based on feature information.
[0362] Input: Symbol information and structural information.
[0363] Output: Set of evaluation criteria parameters.
[0364] The server reads fields from the feature data structure and calculates a text complexity metric, such as a weighted function based on the total number of characters, the number of sentences, and the number of list items. The server compares this metric with a preset threshold to categorize the text as "short text," "medium-length text," or "long text." Based on the category and structural features, the server sets evaluation priorities, such as labeling evaluation dimensions as "grammar-oriented," "structure-oriented," or "information-oriented," and generates corresponding parameter sets (including priority dimensions, output granularity, and whether multiple rewrite examples are needed), storing these parameters in an evaluation condition object.
[0365] Step 6: The server generates a prompt statement based on the evaluation criteria.
[0366] Input: Cleaned text information and a set of evaluation criteria parameters.
[0367] Output: The complete prompt string.
[0368] The server selects template text matching the evaluation criteria from the prompt template store. The server uses string formatting functions to embed the user's text into placeholder positions within the template, and simultaneously inserts or deletes certain prompt paragraphs based on the evaluation criteria (e.g., adding a note like "Please carefully check the connections between paragraphs and list items"). The server-generated prompt statement is as follows: "You are a professional Chinese writing editor. Please provide a detailed evaluation of the following text in terms of grammar, clarity of expression, logical coherence, and information content, and offer specific suggestions for revision. Also, please provide at least one complete rewritten example. The text is as follows: 'This event was a great success, and we were all very happy.' Please first give an overall evaluation, then list the revision suggestions one by one, and finally provide a rewritten example." The server stores the prompt statement as a single string in memory for subsequent calls to the generative artificial intelligence model.
[0369] Step 7: The server sends prompts to the generative artificial intelligence model and requests suggestions for improving the article.
[0370] Input: Prompt string and model call parameters (such as model name, maximum output length, etc.).
[0371] Output: Inference request sent to the model service.
[0372] The server constructs an inference request through the model interface client, using the prompt statement as model input, along with parameters such as model identifier, temperature coefficient, and maximum generation length. The server serializes the request into text or binary format and sends it to the server endpoint where the generative AI model is deployed via an encrypted communication channel. After sending, the server enters a waiting state, ready to receive the output generated by the model.
[0373] Step 8: The server receives the output of the generative artificial intelligence model and extracts suggestions for improving the article.
[0374] Input: Response data returned from the model service.
[0375] Output: Text of the original article's proposed improvement plan.
[0376] After receiving the response data from the model service via the network interface, the server decrypts and verifies the integrity of the response, and uses a parsing module to parse the response structure, extracting fields representing the model's output content. The server then saves the content of these output fields as the original text of the article improvement proposal in memory, typically a long text containing an overall evaluation, problem descriptions, modification suggestions, and rewriting examples.
[0377] Step 9: The server organizes the article improvement plan into a structure and generates display format information.
[0378] Input: Original article improvement proposal text.
[0379] Output: Display format data structure containing the types and content of constituent elements.
[0380] The server searches the text of the article improvement proposal for tagged phrases (such as "overall evaluation:", "revision suggestions:", "rewrite example:") or numbering patterns (such as "1.", "2.", etc.) and uses these to segment the original text into multiple fragments. The server assigns a type label to each fragment (such as "overall_review", "issue", "suggestion", "example") and records the fragment order and hierarchical relationship in structured data. The server then maps each type label to display style information (such as headings, list items, and body paragraphs) according to predefined style rules, ultimately forming a display format information structure for terminal rendering.
[0381] Step 10: The server generates a notification message and sends it to the user via the communication path.
[0382] Input: Display formatting information and user notification preference information.
[0383] Output: Notification information (electronic document or electronic message) sent over the network.
[0384] The server reads the user account's notification configuration to determine whether to send the result via email or messaging service. Based on the display format information, the server assembles the notification body, inserting the overall evaluation, a list of issues, suggested modifications, and rewrite examples into a predefined message template to generate highly readable notification content. In email scenarios, the server constructs an email message; in messaging service scenarios, it constructs a message body and sends it to the corresponding mail server or messaging service interface using an encrypted communication protocol. After sending, the server records the sending result and a timestamp.
[0385] Step 11: The terminal receives and parses the notification information sent by the server.
[0386] Input: Notification data received from the mail server or messaging service.
[0387] Output: Display data for article improvement solutions suitable for local interface rendering.
[0388] The terminal synchronizes messages with external services via a built-in email or messaging client. Upon detecting a new notification, the terminal downloads the message content and parses the message header and body according to the protocol. The terminal maps the text content in the message body to the internally supported display formats, mapping the overall evaluation, list items, and sample text to UI components (such as title areas, list controls, and text areas), and generates the display data structure for rendering in local memory.
[0389] Step 12: The terminal displays an article improvement plan on the display device, and the user modifies the text according to the plan.
[0390] Input: Displays data structures and user interaction operations.
[0391] Output: The user's revised version of the text or confirmation of the result.
[0392] The terminal presents segmented content on the screen based on the displayed data structure, including an overall evaluation area at the top, a list of issues and suggestions in the middle, and rewritten examples at the bottom. The terminal allows users to scroll through, expand, or collapse sections, and provides buttons for actions such as "copy example" and "open in editor." After reading the content, users can modify the original text in the terminal's editing interface, referring to the rewritten examples to expand sentences and enrich details. After completing the modifications, users can choose to send a new version of the text again through the terminal, thus re-triggering the above processing flow and creating multiple rounds of refinement.
[0393] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0394] In text-based electronic communication environments, users typically input comments, messages, or long documents via terminals. Traditional systems often rely solely on fixed rules for simple character limits or sensitive word filtering, lacking a comprehensive understanding of the complexity of text structure, emotional state, and the multi-party communication context, leading to the following technical problems: (1) Insufficient processing efficiency and resource utilization: Existing text processing workflows are mostly sequential and rule-based judgments, which cannot dynamically trigger deep processing based on the amount of text information and the scale of the enumeration structure. This makes it easy to waste computing resources on simple texts or to underprocess complex texts, making it difficult to achieve adaptive allocation and efficient utilization of computing resources.
[0395] (2) Coarse granularity of intelligent generation module calls: In some systems, even if a generative artificial intelligence model is introduced, the user's original text is often simply input directly into the model. There is a lack of task-level prompt statements based on the parsing results and evaluation thresholds, which leads to unstable and uncontrollable model output. It is difficult to ensure the relevance, consistency and interpretability of the deduction results at the computer level.
[0396] (3) Insufficient utilization of emotional state: Existing text rewriting or polishing systems generally ignore the emotional state of users and the emotional feedback of other readers, or only provide superficial prompts on the front-end interface. They do not use the emotional analysis results as core parameters in the server processing pipeline for threshold setting, prompt statement generation and candidate result adjustment, thus failing to achieve automatic mitigation of negative emotions or balance of multiple emotions at the system level.
[0397] (4) Insufficient sharing and management capabilities in multi-user scenarios: Traditional systems usually only provide rewriting suggestions to the original submitting entity, rarely centrally manage and control the sharing of revision candidate information, lack a unified management mechanism for multiple users, and cannot efficiently reuse high-quality revision results as reusable resources across multiple terminals and sessions, thus limiting the overall scalability and knowledge accumulation capabilities of the system.
[0398] (5) Limited user interaction experience and controllability: The user interface side often only sees a single rewriting result, and cannot compare different candidate versions side by side on the terminal. It also lacks a fine-grained interaction mechanism for local editing and selection of candidate versions, which is not conducive to users generating appropriate text quickly while ensuring the main intention, thus affecting the overall efficiency of modern human-computer interaction.
[0399] (6) Insufficient optimization of long text structure: For long texts containing a large number of paragraphs and entries, traditional systems only make superficial modifications to the content and fail to automatically generate summary and structured versions on the server side based on the deliberation of candidate information and sentiment analysis results. This makes it impossible to improve the structured processing capability of long texts and the performance of downstream applications (such as retrieval, recommendation and display) from the system architecture level.
[0400] Therefore, it is necessary to propose an improved computer implementation scheme, which introduces on the server side the parsing of text information and structural features, multi-dimensional estimation of emotional state, prompt statements constructed based on the parsing results, and fine-grained call control of generative artificial intelligence models. It also provides multi-user sharing management and multi-version output capabilities for refining candidate information, so as to improve resource scheduling, model call controllability, and human-computer interaction experience in the text processing process, thereby improving the overall system performance and availability from the perspective of computer technology.
[0401] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0402] In this invention, the server includes a device for receiving text information, a device for parsing the constituent elements of the received text information and detecting the amount of information and enumerating the structural scale, a device for setting an evaluation threshold based on the parsing result and emotional state information and determining whether text information exceeding the evaluation threshold needs to generate revision candidate information, a device for constructing a prompt statement based on the determination result and the parsing result to instruct a generative artificial intelligence model to generate the revision candidate information, a device for inputting the prompt statement into the generative artificial intelligence model and obtaining revision candidate information for the text information from the generative artificial intelligence model, a device for estimating the emotional state of the submitting subject and the reading subject of the text information through emotional parsing processing and adjusting the expression content and expression intensity of the revision candidate information based on the estimation result, and a device for prompting the revision candidate information to a single user and managing it as information that can be shared by multiple users. This allows for the construction of a hierarchical processing pipeline for text information on the server side: First, the text is parsed and filtered based on its information content and structural features. Only when the evaluation threshold is met is a generative artificial intelligence model invoked using constructed prompts, thereby reducing unnecessary deep reasoning overhead. Subsequently, the sentiment analysis results are integrated into the generation and adjustment process of candidate information, automatically mitigating negative emotions or balancing multiple emotions while ensuring semantic consistency. Through centralized management and multi-terminal sharing of candidate information, the reuse and unified control of high-quality candidate texts are achieved, thereby improving the resource utilization efficiency, model invocation controllability, long text structure optimization capabilities, and user interaction experience of the text processing system at the computer technology level.
[0403] “Text information” refers to digital data composed of characters, punctuation marks and their combinations, used to express semantic content, including short messages, comments, long documents, and various written contents containing paragraphs and paragraphs.
[0404] "Constituent elements" refer to the basic units and their structural relationships identified when parsing textual information, including sentences, words, punctuation marks, paragraphs, entries, and the hierarchy and dependencies between these units.
[0405] "Information content" refers to a quantitative indicator used to characterize the scale and richness of textual information content, including but not limited to the number of characters, words, sentences, and the number of semantic units contained therein.
[0406] "List structure scale" refers to the quantity and complexity of structures presented in the form of lists or entries in text information, including the number of entries, the number of nesting levels, and the logical relationships between entries.
[0407] "Emotional state information" refers to data obtained through sentiment analysis of textual information or related interactive data, used to characterize the type and intensity of a subject's emotions, including emotion labels such as anger, sadness, joy, and tension, and their corresponding scores.
[0408] "Evaluation threshold" refers to a set of benchmark values or conditions set based on the parsing results and emotional state information, used to determine whether to perform scrutiny candidate generation processing on text information. It may include one or more numerical boundaries related to information content, enumeration structure size, and emotional intensity.
[0409] "Refine candidate information" refers to one or more alternative or rewritten texts generated from the original text information. These texts are semantically related to the original text, but have been optimized or adjusted in terms of content, structure, tone, or length, and are used as candidates for user selection or further editing.
[0410] "Generative artificial intelligence models" refer to machine learning models that take natural language prompts as input and automatically generate text output. By training on large-scale corpora, they can generate thought-provoking candidate information, summaries, or structured text content according to given instructions.
[0411] "Prompt statements" refer to natural language instruction texts constructed by the server based on the parsing results and evaluation thresholds, used to explicitly describe the target task, constraints, and output format to the generative artificial intelligence model.
[0412] "Submitting Entity" refers to the entity that inputs or uploads text information to the system, which can be an individual user, a group of users, or an application instance.
[0413] "Viewing subject" refers to the party that reads or accesses the text information or candidate information that has been submitted and processed by the system, and may be the same as or different from the submitting subject.
[0414] "Sentiment analysis processing" refers to the process of analyzing text information and related interaction data to extract and quantify the subject's emotional type and intensity. It can be achieved using sentiment dictionaries, machine learning models, or deep learning models.
[0415] "Content" refers to the combination of words and sentence structure in text information that carries semantics and emotions, including specific word choices, sentence organization, and semantic focus.
[0416] "Expression intensity" refers to the degree of emotional expression, attitude, or tone of textual information, and is used to distinguish different levels of expression intensity such as neutral, slightly, and strong.
[0417] "Users" refers to the general term for entities that interact with the system, including entities that submit text information and entities that view or manipulate candidate information.
[0418] "Shared information" refers to candidate information or its derivative text that can be accessed, browsed or reused by multiple users after being managed by the system. This information is uniformly managed in terms of storage and access control.
[0419] "Display control unit" refers to the functional unit used to control the presentation of content in the display area of a terminal device, including software or hardware components that control the text position, style, arrangement order, and highlighting mode.
[0420] "Terminal device" refers to an electronic device used to communicate with a server and provide a human-computer interaction interface to the user, including but not limited to mobile communication terminals, portable computing terminals, desktop computing terminals and wearable display terminals.
[0421] "User interface processing" refers to the process of generating, updating, and managing graphical interface elements on a terminal device to receive user input and present system output to the user, including candidate list display, selection operation processing, and editing operation processing.
[0422] The "communication processing unit" refers to the functional unit used for sending and receiving data between the server and the terminal device, including the network protocol stack, session management module, and control logic related to external communication interfaces.
[0423] "Summary version information" refers to a short text generated based on the original text information, which compresses and simplifies the content while retaining key information, and is used to summarize the main points of the original text.
[0424] "Structured version information" refers to a text version generated based on the original text information and refined candidate information, which has been reorganized and optimized in terms of paragraph division, item organization, and logical order, in order to improve the overall structural clarity and readability.
[0425] The embodiments of this invention will be described in conjunction with the coordinated actions of the server, terminal, and user. The following embodiments are merely examples to illustrate how to implement this invention in a specific hardware and software environment, and are not intended to limit the scope of the claims.
[0426] I. Overall System Composition In this invention, the server operates as the core computing node. The server can be a rack-mounted computing device, a virtual machine instance, or a cloud computing node; typical hardware includes a multi-core central processing unit, a graphics processing unit, high-speed memory, and non-volatile storage devices. The server communicates with one or more terminals via wired or wireless networks.
[0427] In this invention, the terminal can be a mobile communication device, a portable computing device, a desktop computing device, or a wearable display device. The terminal includes a display unit, an input unit, and a communication interface. The terminal runs a user interface application to provide the user with a text input interface and a display interface for considering candidate information.
[0428] In this invention, users input text information, view and consider candidate information, and make selections or edits via a terminal, thereby triggering the program processing flow on the server side.
[0429] The server's software architecture includes: a communication processing unit, a text parsing module, a sentiment analysis module, a threshold setting module, a prompt generation module, a generative artificial intelligence model invocation module, a candidate selection and adjustment module, a shared management module, and a data storage module. These modules can be implemented by programs executed by one or more processors.
[0430] II. Server-side text parsing and data structure After receiving text information from the terminal, the server uses a text parsing module to perform structured processing on the text information. The text parsing module can be implemented based on a natural language processing library. For example, the server can use a language processing library (such as the general-purpose natural language processing library spaCy or NLTK) to build a sentence segmenter, tokenizer, and dependency parser.
[0431] Internally, the server maps each piece of text information to a set of data structures. For example, the server can construct a record structure for each piece of text information, which includes: - Text identifier; - User identifier; - Original character sequence; - A list of sentences (each sentence includes a start and end index and a syntax tree structure); - A list of words (each word includes its form, part-of-speech tag, and position in the sentence); - List structural information (e.g., starting position of bullet points, depth of item hierarchy, total number of items); - Basic statistical characteristics (e.g., number of characters, number of words, number of sentences, number of paragraphs).
[0432] The server uses the aforementioned data structure to calculate the information content and the size of the enumeration structure. Information content can include character counts, sentence counts, and the number of different words; the size of the enumeration structure can include the number of list entries, nesting levels, and the range of characters covered by the list. The server uses these values as one of the input features for the subsequent threshold setting module.
[0433] III. Server-side sentiment analysis and sentiment feature construction To analyze the sentiment state of both submitters and readers, the server incorporates a sentiment analysis module. This module can be implemented using traditional sentiment classification algorithms or deep learning models. For example, the server can use a sentiment classifier based on convolutional neural networks or bidirectional recurrent neural networks with attention mechanisms, taking a sequence of word vectors from the text as input and outputting a probability distribution of multiple sentiment categories.
[0434] The server can process text in the following ways in terms of feature construction: - Convert each word into a vector representation, which can be generated from pre-trained word embeddings or sub-word embeddings; - Concatenate sentence-level features with text-level features, including average word vectors, max-pooling vectors, sentence length information, etc. - Use the probability values corresponding to emotion labels (such as anger, sadness, joy, fear, etc.) as emotion intensity features.
[0435] The server outputs a structured result containing sentiment type and intensity through the sentiment analysis module, which serves as "sentiment state information." This information, along with the information content and the size of the enumeration structure, is then processed by the threshold setting module.
[0436] IV. Threshold Setting and Triggering Logic In the threshold setting module, the server sets the evaluation threshold based on the text parsing results and sentiment state information. The server can store a set of configurable parameters in the storage module, for example: - Information content threshold: When the number of characters exceeds a preset value, the text length is considered too long; - List size threshold: When the number of entries exceeds a preset value, the structure is considered complex; - Emotional intensity threshold: When the probability of a certain negative emotion category is greater than a preset threshold, the emotion is considered strong.
[0437] The server integrates the above parameters to form an evaluation logic. For example, the server can be set to trigger structural simplification analogy when "the text length exceeds a first threshold and the enumeration structure size exceeds a second threshold"; and to trigger emotion mitigation analogy when "the emotion intensity exceeds a third threshold". This design allows the server to internally implement selective control over the invocation of generative artificial intelligence models, thereby avoiding wasting computational resources on text that does not require complex processing.
[0438] Compared to traditional fixed rules or manual triggering methods, this triggering logic based on multi-feature evaluation thresholds can significantly reduce invalid inference calls under limited hardware resources, thereby improving processing speed and computational efficiency.
[0439] V. Prompt Statement Generation and Generative Artificial Intelligence Model Structure When the server determines that it needs to generate further deliberation candidate information, it uses a prompt generation module to construct prompts in natural language form. Internally, the server encodes the parsing results and sentiment state information into conditional descriptions within the prompts, for example: - Includes text-based uses (chat, comments, reports, etc.); - Includes sentiment analysis results and objectives (e.g., mitigating negative emotions, preserving factual content, compressing length, etc.); - Includes output requirements (number of candidates, language type, length control, style preference, etc.).
[0440] The server can generate the following types of prompt statements: - Example of a prompt statement 1 (in a scenario where emotions are calming down): The user is entering the following sentence: "Today's meeting made absolutely no progress, I'm about to break down." Please rewrite the sentence in a calmer and more constructive tone without changing the facts. Provide three different suggestions for rewriting it in Chinese.
[0441] - Example 2 of prompt statements (for simplifying long text): Below is a project progress report. Please compress the overall text to approximately 60% of its original length, while retaining key progress, risks, and next steps. Also, reorganize the bullet points structure for greater clarity. Please output the new report in Chinese. The original text is as follows: "(Original content)".
[0442] - Example of a prompt statement 3 (Scenario involving balancing the emotions of multiple parties): The submitted text is as follows: "Today's meeting made absolutely no progress, I'm about to break down." The reply from others was: "What you're saying makes me very sad. Actually, everyone is working very hard." Please generate two rewrite options while preserving the submitter's genuine feelings, making the expression more considerate of others' feelings, and using a polite and empathetic tone.
[0443] After generating the prompt, the server inputs the prompt into the generative AI model via the generative AI model invocation module. The generative AI model can employ a neural network model based on a transformer structure. This neural network typically includes multi-layer self-attention modules, feedforward networks, residual connections, and layer normalization structures.
[0444] When training this generative AI model, the server can use a large-scale text corpus, employ an autoregressive language modeling objective, use the next word prediction error as a loss function, and update the model parameters through backpropagation and optimization algorithms (such as adaptive learning rate methods). During the inference phase, the server uses strategies such as beam search or temperature sampling to generate candidate texts, thereby achieving a balance between diversity and stability.
[0445] By explicitly encoding task objectives and constraints in the prompt statements, the server shrinks and constrains the output space of generative artificial intelligence models, making the model output more controllable and more in line with system expectations, thereby achieving fine-grained control over model calling behavior within the computer.
[0446] VI. Refining Candidate Information Adjustment and Balancing Multi-Subject Emotions After obtaining raw candidate information from the generative artificial intelligence model, the server uses a candidate adjustment module to perform secondary processing on the candidate content. The server can recalculate the sentiment features for each candidate text and compare them with the sentiment features of the original text and other sources.
[0447] The following technical methods can be used when adjusting the server: - If the negative sentiment intensity of the candidate text is still higher than the preset range, the server can further soften the wording by calling the generative artificial intelligence model again or using a rule-based replacement dictionary. - If the candidate text does not significantly improve in length or enumeration structure complexity, the server can add constraints to reconstruct the prompt statement, prompting the generative artificial intelligence model to more strictly control the output length or structure complexity. If a significant difference in sentiment is detected between the submitter and the reader, the server can add the goal of "taking both parties' emotions into account" to the prompt, enabling the generative AI model to generate candidate text that is more emotionally balanced.
[0448] Through the above adjustments, the server no longer simply returns the model output directly, but introduces a multi-round evaluation and regeneration mechanism to make the candidate texts more in line with the technical goals in terms of sentiment, length and structure.
[0449] VII. Refine the sharing management and terminal display of candidate information In the shared management module, the server stores the candidate information obtained through the above process in a centralized data store, and attaches metadata to each candidate information, including: - Source text identifier; - Applicable scenario tags (emotional relief, structural simplification, multi-faceted balance, etc.); - Evaluation scores (e.g., emotional balance, structural clarity).
[0450] The server can mark some high-scoring candidate information as shareable resources, allowing multiple different users to reuse these candidate texts in similar scenarios, thereby reducing the computational cost of repeated generation. This sharing mechanism enables the server to build a candidate expression library over long-term operation, thus technically enabling the caching and reuse of model output results, further reducing the frequency of calls to generative artificial intelligence models and lowering the overall communication and computational load.
[0451] After receiving the candidate text information returned by the server, the terminal displays multiple candidate texts side-by-side through its display unit. For example, the terminal can display three candidate texts below the original text input box, with a selection button and an editing entry next to each candidate. The terminal can maintain a lightweight state structure locally to record the user's selection habits in different sessions, and the server can refer to this information to adjust preferences in subsequent calls.
[0452] After reviewing the candidate texts, users can choose one to directly replace the original text, or they can make partial modifications based on the candidate texts. When the user confirms the send, the terminal only submits the final confirmation text as the official output to the backend system, thus maintaining the user's control over the final expression.
[0453] VIII. Technical Effects and Causal Relationships The server achieves adaptive control over generative AI model calls by introducing evaluation thresholds based on parsing results and sentiment state information. Since the server only triggers deep generation processing when the text meets specific complexity or sentiment intensity conditions, and can use lightweight rules or directly access the original text in other cases, it reduces a large number of invalid deep inference calls in the overall system, thereby improving processing speed and saving computing resources from a computer technology perspective.
[0454] By using information content, structural features, and emotional features as inputs to construct prompts, the server explicitly constrains the generative AI model, making the output more stable and predictable in terms of emotional mitigation, structural simplification, and multi-faceted balance. Compared to the traditional method of "directly inputting the original text," this prompt-driven generation method reduces the randomness and uncertainty of the model's output, improves the efficiency of refining candidate information, and thus reduces the number of times users need to make repeated adjustments. At the human-computer interaction level, this translates to fewer operation steps and a lower error rate.
[0455] The server leverages a shared management module to centrally store and reuse high-quality candidate information across multiple terminals. This reduces redundant computational load in similar future scenarios, alleviating continuous pressure on network bandwidth and computing resources. Through this caching and reuse mechanism, the system can develop a self-optimizing representation resource library over long-term operation, thereby achieving a reduction in communication load and an overall shortening of response time on a macro level.
[0456] The server utilizes a multi-level modular processing pipeline structure (including parsing, sentiment assessment, threshold determination, prompt generation, model invocation, secondary adjustment, and shared management) to internally organize data flow and establish clear information interfaces between modules. This modular pipeline not only facilitates parallel processing and resource scheduling but also allows for the replacement or expansion of specific modules for different scenarios, such as replacing them with higher-precision sentiment models or more efficient language model architectures, thereby enhancing the system's scalability and maintainability.
[0457] In summary, through the specific modules, data structures, and algorithm configurations described above, the server not only automates text analysis but also establishes a technical mechanism independent of human operation in terms of generative artificial intelligence model invocation control, resource scheduling, sentiment balance, and result reuse. This results in improved processing efficiency, stable output quality, and optimized resource utilization at the computer technology level.
[0458] use Figure 14 The processing procedure is explained.
[0459] Step 1: Users input text information and initiate requests using the terminal.
[0460] Users type text (such as chat messages, comments, or report snippets) into the terminal's input interface and select "Send" or "Get Suggestions" via touch or key gestures. The terminal encapsulates the user-input string, user identifier, session identifier, and text type into request data. The input is the raw text string edited by the user in the input box along with related identification information; the output is a request data packet directed to the server. Based on this input, the terminal constructs a network message, encodes the text string into a unified character set format, appends metadata fields, and then sends it to the server through the communication interface.
[0461] Step 2: The server receives the requested data and performs initial storage.
[0462] The server receives request data packets from terminals via its communication processing unit, parses the network messages, and extracts user identifiers, session identifiers, text types, and raw text information. The server writes these fields into a data storage structure, such as creating records in a relational database table or key-value store. The input is the request data packet sent by the terminal, and the output is a text record registered in the server's internal storage structure. Based on this input, the server performs message parsing and field mapping operations, converting the bitstream into structured data rows.
[0463] Step 3: The server performs language parsing and structural analysis on the text information.
[0464] The server invokes the text parsing module to perform sentence segmentation, word segmentation, part-of-speech tagging, and dependency parsing on the stored raw text. The server uses a natural language processing library to iteratively scan the text, identify sentence boundaries, assign part-of-speech tags to each word based on a language model, and construct a dependency tree. The input is the raw text string, and the output is a structured representation containing a list of sentences, a list of words, dependency relations, paragraph boundaries, and entry information. Based on this input, the server performs symbolic and syntactic data processing, transforming the linear string into a hierarchical data structure.
[0465] Step 4: The server calculates the amount of information and the size of the enumerated structure.
[0466] The server reads the structured text representation generated in step 3, calculates the number of characters, words, sentences, and paragraphs, and counts the occurrence position and quantity of each entry marker, identifying the nesting level of the list. The input consists of a list of sentences, a list of words, and a set of entry items. The output is a numerical vector of information content indicators (such as total number of characters and total number of sentences) and enumeration structure scale indicators (such as number of entries and maximum nesting depth). Based on this input, the server performs counting and hierarchical traversal operations, compressing the structured representation into a set of numerical features that facilitate threshold comparison.
[0467] Step 5: The server performs sentiment analysis on the submitted text.
[0468] The server invokes the sentiment analysis module, feeding the original text and its word vector representations into the sentiment classification network to calculate the probability distribution of multiple predefined sentiment categories (such as anger, sadness, and joy). The input is the original text string and its vector sequence generated by the embedding layer; the output is sentiment state information composed of sentiment type and sentiment intensity. Based on this input, the server performs a forward propagation operation in the neural network, calculating the weighted sum of each layer and the activation function results, generating a probability vector in the output layer, and then selecting the sentiment label and its intensity value corresponding to the highest probability.
[0469] Step 6: The server analyzes the emotions of the readers in multi-subject scenarios.
[0470] When the server detects that the text belongs to a public communication context, it reads related replies, comments, or reactions from the database and sends these texts to the sentiment analysis module to calculate the sentiment distribution of other participants. The input is a set of response text strings associated with the original text, and the output is a synthesized sentiment state describing the overall sentiment trend of each reader. Based on this input, the server performs batch sentiment inference and statistical aggregation operations to calculate the average sentiment intensity or sentiment category distribution, thereby obtaining a multi-subject sentiment overview.
[0471] Step 7: The server sets an evaluation threshold and makes a judgment based on the analysis results and emotional state.
[0472] The server reads information content metrics, list size metrics, and sentiment state information from the threshold setting module, compares them with pre-stored threshold parameters, and generates a decision on whether to perform refinement processing. The input is a feature vector containing text length, list size, and sentiment intensity; the output is the refinement processing type and trigger flag (e.g., "emotional mitigation refinement," "structural simplification refinement," "multi-faceted emotional balance refinement," etc.). Based on this input, the server performs numerical comparisons and conditional logic operations to determine whether each feature exceeds the corresponding threshold and combines them to formulate a specific processing strategy.
[0473] Step 8: The server constructs prompts for generative artificial intelligence models.
[0474] In the prompt statement generation module, the server synthesizes a natural language instruction text based on the decision result of step 7 and the parsed data from steps 3-6, explicitly informing the generative AI model of the task to be performed, the constraints, and the output format. The input includes the parsed results (text purpose, structural features), sentiment state information, and a processing type label; the output is one or more complete prompt statement strings. Based on this input, the server performs template filling and conditional concatenation operations, embedding feature values into predefined sentence patterns to generate a semantically complete natural language description.
[0475] Step 9: The server invokes a generative artificial intelligence model to generate initial candidate information for consideration.
[0476] The server invokes the generative AI model module, inputting the prompt statement generated in step 8 into a pre-deployed language model based on a transformer structure, thus initiating the text generation process. The input is the prompt statement string, and the output is a sequence of candidate texts (each a piece of deliberation candidate information). Based on this input, the server performs self-attention calculations, multi-layer feedforward network operations, and probability sampling steps within the model to generate new text in lexical order until the termination condition is met.
[0477] Step 10: The server performs a second sentiment and structure assessment on the candidate information.
[0478] The server recalculates the information content, list structure size, and sentiment state of each generated candidate text using the text parsing and sentiment analysis modules, and compares these recalculations with the features of the original text. The input is a set of candidate texts for analysis, and the output is the structural feature vector and sentiment feature vector for each candidate text. Based on this input, the server performs parsing and classification operations similar to steps 3-6, with the addition of difference calculations to determine the changes in length, list complexity, and sentiment intensity between the candidate texts and the original text.
[0479] Step 11: The server adjusts candidate texts based on multi-subject sentiment and structural objectives.
[0480] In the candidate refinement module, the server utilizes the difference information obtained in step 10 and the multi-subject sentiment profile from step 6 to filter and regenerate the candidate set as necessary. The input consists of candidate texts and their difference features, and the output is the final set of candidate information after filtering, sorting, or regeneration. Based on this input, the server performs sorting operations (e.g., sorting by emotional intensity or structural simplicity), filters candidates that do not meet the constraints, and, when necessary, adjusts the prompt parameters to re-invoke the generative AI model to generate new candidates to meet predetermined technical objectives.
[0481] Step 12: The server will review candidate information, register it as a shareable resource, and manage access.
[0482] In the shared management module, the server generates a resource identifier for each final candidate text, records its source text identifier, applicable scenario tag, and evaluation score, and sets access permissions so that multiple users can reuse it in subsequent similar scenarios. The input is the final set of candidate information and its metadata, and the output is the resource index and management entries registered in the shared resource repository. Based on this input, the server performs index construction and metadata storage operations, mapping the candidate texts to an efficient and searchable data structure.
[0483] Step 13: The server will then finalize the candidate information and send it to the terminal.
[0484] The server selects candidate text relevant to this processing from the shared management module, packages the text content and its type tags into response data, and sends it to the requesting terminal and other terminals that need to be displayed synchronously through the communication processing unit. The input is the final candidate text and its display parameters, and the output is the response data packet for each terminal. Based on this input, the server performs data serialization and network packet encapsulation operations, compressing the transmission load as much as possible while ensuring data integrity.
[0485] Step 14: The terminal receives and displays the candidate information for consideration.
[0486] The terminal receives data packets returned by the server through a communication interface, parses the candidate text and tags within them, and presents them in the display section as a list or card. The input is the server's response data packet, and the output is multiple candidate text items and interactive controls displayed on the terminal screen. Based on this input, the terminal performs interface element generation and layout calculations, maps the candidate text to clickable interface components, and binds selection and edit event handling logic.
[0487] Step 15: Users select or edit candidate text on the terminal.
[0488] Users select one or more texts from a candidate list on the terminal interface via touch or pointer operation, and can make partial modifications to the candidate text, such as adding details or deleting parts of the text. The input consists of the candidate text displayed on the terminal and the user's operation event, and the output is the final text content confirmed by the user. Based on this input, the terminal performs string replacement, merging, or partial editing operations, writes the candidate text selected by the user into the input box, and submits the final text as the official content to subsequent business systems or resends it to the server for archiving after user confirmation.
[0489] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0490] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0491] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0492] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0493] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0494] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0495] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0496] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0497] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0498] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0499] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0500] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0501] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0502] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0503] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0504] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0505] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0506] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0507] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0508] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0509] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0510] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0511] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0512] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0513] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0514] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0515] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0516] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0517] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0518] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0519] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0520] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0521] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0522] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0523] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0524] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0525] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0526] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0527] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0528] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0529] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0530] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0531] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0532] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0533] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0534] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0535] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0536] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0537] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0538] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0539] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0540] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0541] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0542] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0543] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0544] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0545] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0546] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0547] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0548] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0549] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0550] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0551] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0552] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0553] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0554] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0555] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0556] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0557] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0558] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0559] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0560] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0561] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0562] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0563] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0564] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0565] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0566] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0567] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0568] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0569] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0570] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0571] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0572] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0573] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0574] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0575] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0576] In addition, the following notes are provided in response to the above explanation.
[0577] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for obtaining textual information from an external device by an information processing device; An apparatus for parsing the textual information by an information processing device, thereby detecting the total number of character symbols and the number of enumerated structural elements contained in the textual information; An apparatus for determining, by an information processing device, a threshold for determining the total number of character symbols and the number of enumerated structural elements based on the detection results and reference information stored in a storage device; An apparatus for determining whether text correction processing is required based on the determination threshold and the detection result by an information processing device; An apparatus for generating, when an information processing device determines that text correction processing is required, structural information containing the text form information and the determination threshold, and constructing the instruction text into a prompt statement input to a generative artificial intelligence model; An apparatus for inputting the prompt statement and the textual information into the generative artificial intelligence model by an information processing device, so that the generative artificial intelligence model generates a correction scheme for the textual information; A means for generating response information containing the correction scheme and the detection result by an information processing device, and sending the response information to the external device.
[0578] (Note 2) According to the information processing system described in Appendix 1, the information processing device displays the textual information and the correction scheme side by side through a display interface composed of a display control device, and simultaneously displays explanatory messages based on the judgment threshold and the detection results.
[0579] (Note 3) According to the information processing system described in Appendix 1, the information processing device obtains instruction text as input as a prompt statement from the external device, and by using the instruction text as the prompt statement, enables the content input to the generative artificial intelligence model to be changed according to the user.
[0580] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A means of obtaining document information containing text information from a user terminal; This is a means of processing the document information using an information processing program with string processing capabilities to calculate the number of characters and the number of elements in the enumerated structure, thereby extracting document characteristic information. Means for comparing the document characteristic information with reference condition information stored in the storage device and determining whether at least one of the number of characters or the number of elements of the enumerated structure exceeds the threshold information contained in the reference condition information. The means for generating a prompt statement containing instructions on document simplification and key point extraction based on the document information and the document characteristic information when the threshold information is determined to be exceeded, and for generating a generation request information for inputting the prompt statement into a generative artificial intelligence model. Means for sending the generation request information to the computing resources running the generative artificial intelligence model via communication functions, and for obtaining rewritten document information generated by the generative artificial intelligence model; Means for converting the rewritten document information into display data that can be displayed on the user terminal, and sending the display data to the user terminal.
[0581] (Note 2) The information processing system according to Appendix 1 is characterized in that, It also includes: means for storing the rewritten document information in a recording device in association with the document information, and for generating control information for updating at least one of the prompt statement or the baseline condition information by statistically processing the rewritten document information stored in the recording device in association with the evaluation information.
[0582] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system is configured to: if it is determined that the document characteristic information does not exceed the threshold information, not to perform the generation and sending processing of the generation request information for the generative artificial intelligence model, but to send the document information as output information for submission processing through the communication function.
[0583] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for receiving text information; A device for parsing received text information to extract feature information, including symbolic and structural information. A device for setting evaluation conditions based on the results of the analysis; A device for generating prompt statements for a generative artificial intelligence model based on the evaluation conditions; A device for inputting the prompt statement and the text information into the generative artificial intelligence model to generate an article improvement scheme; A device for organizing the article improvement scheme according to its constituent elements and converting it into display format information according to predetermined rules; A device for generating and sending notification information containing the article improvement plan to the submitting entity based on the display format information and via a communication path.
[0584] (Note 2) The information processing system according to Appendix 1 is characterized in that, This is used to display the article improvement scheme contained in the notification information in sections on the screen of the user's display device based on differentiation information.
[0585] (Note 3) The information processing system according to Appendix 1 is characterized in that, This is used to send the text information and the prompt statement to the generative artificial intelligence model using encrypted communication when sending the text information, and to notify the submitting entity of the article improvement plan as an electronic document or electronic message according to the communication protocol when sending the article improvement plan.
[0586] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for receiving text information; A device for parsing the constituent elements of received text information and detecting the amount of information and enumerating the structural scale. An apparatus for setting an evaluation threshold based on the parsing results and emotional state information, and for determining whether text information exceeding the evaluation threshold needs to generate deliberation candidate information. A device for constructing a prompt statement for instructing a generative artificial intelligence model to generate the deliberation candidate information based on the determination result and the analysis result; A device for inputting the prompt statement into the generative artificial intelligence model and obtaining refinement candidate information for the text information from the generative artificial intelligence model; An apparatus for estimating the emotional state of the submitting and reading subjects of the text information through sentiment analysis processing, and adjusting the expression content and intensity of the candidate information for consideration based on the estimation results; A device for prompting the proposed information to a single user and managing it as information that can be shared by multiple users.
[0587] (Note 2) The information processing system according to Appendix 1 is characterized in that, The system also includes a user interface processing device for displaying the proposed candidate information side-by-side on the display area of the terminal device via a display control unit, and for receiving selection and editing operations from the user.
[0588] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system also includes a device for generating summary version information and structured version information of text information based on the deliberation candidate information and the result of the sentiment analysis processing, and for sending the summary version information and the structured version information to the terminal device through a communication processing unit.
Claims
1. An information processing system, characterized in that, include: processor, The processor is configured to receive text data, parse the received text data to detect the number of characters and the number of list items, and set a threshold based on the parsing result. When the parsing result exceeds the set threshold, a prompt message is generated to instruct the generative artificial intelligence model to generate a proposed solution. The prompt message is then input into the generative artificial intelligence model to generate the proposed solution.
2. The information processing system according to claim 1, characterized in that, The processor is configured to present the generated deliberation scheme through a user interface.
3. The information processing system according to claim 1, characterized in that, The processor is configured to notify the contributor of the generated revised proposal via a communication module.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A