Information processing system
Patent Information
- Application Number
- CN202610254481.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-03
- Publication Date
- 2026-09-22
AI Technical Summary
传统的信息摘要方式多为人工编辑或固定规则自动摘要,存在如下问题:首先,系统难以及时、高效地从多种信息源自动获取信息并进行统一预处理,导致摘要生成流程效率低、适应性差;其次,现有自动摘要技术多采用简单截取、关键词提取等方式,无法充分利用生成式人工智能模型对文章内容进行深度理解与结构化表达,摘要质量和可读性有限,难以以用户易于理解的方式呈现要点;再次,现有系统往往仅提供单一风格或固定格式的摘要内容,缺乏与用户之间的交互机制,无法根据不同用户的理解水平、知识背景和偏好,对摘要表达进行动态调整,从而降低了信息获取的个性化程度;此外,面对多语言内容时,现有系统通常侧重于全文翻译,未能针对已摘要的信息进行高质量、多语言翻译,用户在跨语言环境中快速掌握核心信息的需求难以满足
[0005]进一步地,所述处理器在接收到生成式人工智能模型输出的摘要结果后,将所述已摘要信息整理和映射为预定的展示格式,例如按照标题、项目符号列表、时间信息、来源信息等字段进行结构化封装,并将整理后的摘要数据发送至用户的终端,以便终端在用户界面中进行呈现。为提升摘要结果对不同用户的适配性,所述处理器还被配置为与用户进行对话式交互:所述处理器基于用户通过终端输入的自然语言指令或选项设置,生成用于指示将摘要的表达调整为适配不同理解水平的提示词,并依据用户选择应用相应的摘要模板,使生成式人工智能模型能够输出符合“简明易懂”“专业详细”等不同层级和风格要求的摘要文本,从而实现摘要表达的个性化调整。此外,所述处理器还被配置为对已摘要信息执行多语言翻译处理,即所述处理器将生成好的摘要作为翻译对象,调用翻译模块或具备多语言能力的生成式人工智能模型,将所述已摘要信息翻译为不同语言,形成多语言版本的摘要结果,再发送至用户终端,从而使用户能够在跨语言环境下快速获取信息要点。通过上述结构与步骤,本发明的系统能够在统一架构下完成信息获取、预处理、生成式摘要、对话式调整及多语言翻译,有效解决现有技术中摘要质量有限、个性化程度不足以及跨语言使用不便等问题。
Smart Images

Figure CN122797464A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot speech in response to the user's speech.
[0003] In existing technologies, when users acquire information from multiple sources such as online media, news websites, blogs, and social media platforms, they typically need to read large amounts of lengthy and structurally complex text. Traditional information summarization methods are mostly manual editing or automatic summarization based on fixed rules, which have the following problems: First, the system struggles to automatically acquire and preprocess information from multiple sources in a timely and efficient manner, resulting in low efficiency and poor adaptability in the summary generation process. Second, existing automatic summarization technologies often employ simple truncation and keyword extraction methods, failing to fully utilize generative artificial intelligence models for in-depth understanding and structured expression of article content, resulting in limited summary quality and readability, and difficulty in presenting key points in a way that is easy for users to understand. Third, existing systems often only provide summaries in a single style or fixed format, lacking an interactive mechanism with users and failing to dynamically adjust the summary expression according to different users' understanding levels, knowledge backgrounds, and preferences, thus reducing the personalization of information acquisition. Furthermore, when faced with multilingual content, existing systems usually focus on full-text translation, failing to provide high-quality, multilingual translation of the summarized information, making it difficult to meet users' needs for quickly grasping core information in cross-language environments. In summary, the technical challenge this invention aims to address is to provide a system that can automatically acquire information source content, generate high-quality summaries using generative artificial intelligence models, and support conversational adjustment of summary content and multilingual translation of summaries, thereby improving information acquisition efficiency and personalized experience. Summary of the Invention
[0004] To address the aforementioned technical challenges, this invention proposes an information processing system. This system includes a processor that performs automatic summarization, interactive adjustment, and multilingual translation of information in the following manner: The processor first acquires information to be processed from at least one information source and converts the acquired information into text data. The text data is then preprocessed, including operations such as removing formatting tags, cleaning up noise, segmentation, and normalization encoding, to provide structured input for subsequent generative artificial intelligence (AI) model processing. Then, based on the preprocessed text data, the processor generates prompts indicating that the information should be summarized. These prompts, along with the text data, are input into the generative AI model, enabling the model to perform semantic analysis and content generation on the original information. In this invention, the processor preferably uses prompts instructing the generative AI model to analyze the information acquired from the information source and summarize key points in bullet point format, thereby generating a structured, hierarchical, and easily readable bullet point summary.
[0005] Furthermore, after receiving the summary results output by the generative artificial intelligence model, the processor organizes and maps the summarized information into a predetermined display format, such as structurally encapsulating it according to fields like title, bullet points, time information, and source information, and sends the organized summary data to the user's terminal for presentation in the user interface. To improve the adaptability of the summary results to different users, the processor is also configured to engage in conversational interaction with the user: based on the natural language commands or option settings input by the user through the terminal, the processor generates prompts to adjust the expression of the summary to suit different levels of understanding, and applies the corresponding summary template according to the user's selection, enabling the generative artificial intelligence model to output summary text that meets different levels and style requirements such as "concise and easy to understand" and "professional and detailed," thereby achieving personalized adjustment of the summary expression. Furthermore, the processor is configured to perform multilingual translation processing on the summarized information. Specifically, the processor uses the generated summary as the translation object, calls a translation module or a generative artificial intelligence model with multilingual capabilities to translate the summarized information into different languages, forming a multilingual summary result, which is then sent to the user terminal. This allows users to quickly obtain key information in a cross-language environment. Through the above structure and steps, the system of this invention can complete information acquisition, preprocessing, generative summarization, conversational adjustment, and multilingual translation under a unified architecture, effectively solving problems such as limited summary quality, insufficient personalization, and inconvenience in cross-language use in existing technologies.
[0006] A "system" refers to a collection of devices consisting of one or more hardware and / or software components, used to perform functions such as information acquisition, preprocessing, summarization, conversational adjustment, and multilingual translation.
[0007] A "processor" refers to a hardware and / or virtual computing unit that can execute program instructions, process data, and control the operation of various functional modules of the system, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and cloud computing-based virtual processing instances.
[0008] "Information source" refers to an external or internal data provider that provides the system with original information content, including but not limited to news websites, blog platforms, social media, databases, file storage systems, and other online or local content services.
[0009] “Text data” refers to textual data consisting of character sequences that can be processed and stored by a computer. It is usually obtained from the original content obtained from the information source after format conversion and cleaning.
[0010] "Preprocessing" refers to a series of data processing operations performed before text data is input into a generative artificial intelligence model in order to improve the processing effect and efficiency. These operations include, but are not limited to, removing format tags, cleaning up noisy content, sentence segmentation, paragraph segmentation, normalization encoding, and language detection.
[0011] "Cue words" refer to input instructions or text descriptions used to control and guide generative artificial intelligence models to perform specific tasks, which include requirements for summarization objectives, output format, language style, length limits, etc.
[0012] "Generative AI models" refer to AI models that can generate new text content based on input text data and prompts, using learned parameters. These include, but are not limited to, large language models, multimodal generative models, and deep learning networks that support text generation tasks.
[0013] "Abstract" refers to the result of compressing and reorganizing information content while keeping the main meaning of the original information unchanged, extracting key points and presenting them in a shorter form.
[0014] “Summarized information” refers to text that has been condensed in length and content relative to the original information after being processed by generative artificial intelligence models or other summarization processes, including bulleted summaries or other structured summaries.
[0015] "Predefined format" refers to the structure and style that the system pre-sets for outputting and displaying summary results, including but not limited to title fields, bulleted list fields, time fields, source fields, and layout rules suitable for display on user terminals.
[0016] "User terminal" refers to a device that is directly operated by the user to receive and display system output results and send user input, including but not limited to smartphones, tablets, personal computers, wearable devices and other electronic devices that support network connectivity.
[0017] "Dialogue mode" refers to an interactive mode in which information is exchanged between the user and the system through multiple rounds of interaction. Users can continuously input their needs in natural language text or interface options, and the system dynamically adjusts its processing strategy and output content based on these inputs.
[0018] "Conversational adjustment" refers to the process by which the system dynamically modifies the content, expression, length, style, or difficulty of understanding of a summary after receiving feedback or instructions from the user through dialogue.
[0019] "Different levels of comprehension" refers to various levels of comprehension based on the user's knowledge background and reading ability, including but not limited to different levels of complexity and professionalism of expression for primary and secondary school students, the general public, and professionals.
[0020] A "template" is a set of predefined rules or format models used to constrain and standardize the structure and language style of abstract output. The template can be applied according to the user's selection to generate abstract text that meets specific style or hierarchical requirements.
[0021] "Translation" refers to the process of converting text content expressed in one natural language into another natural language expression. In this invention, it is mainly used to convert summarized information from the source language into the target language.
[0022] "Different languages" refers to multiple independent natural language types, including but not limited to Chinese, English, Japanese, Korean, and other natural languages that can be supported by the system. Attached Figure Description
[0023] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0024] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0025] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0026] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0027] Figure 5This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0028] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0029] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0030] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0031] Figure 9 This represents an emotion map that maps multiple emotions.
[0032] Figure 10 This represents an emotion map that maps multiple emotions.
[0033] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0034] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0035] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0036] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0037] Hereinafter, an example of an implementation of the system to which the technology of this disclosure relates will be described with reference to the accompanying drawings.
[0038] First, let me explain the terminology used in the following instructions.
[0039] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0040] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0041] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0042] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0043] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0044] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0045] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0046] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0047] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0048] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0049] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0050] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0051] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0052] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0053] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0054] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0055] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0056] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0057] In existing computer information processing technologies, providing users with summaries of long electronic information usually employs simple rule extraction, fixed template compression, or one-time invocation of intelligent models. These solutions suffer from the following technical problems: First, the server-side has limited preprocessing capabilities for electronic information from network-based media, making it unable to uniformly parse and standardize various structured documents or hierarchical data. This leads to unstable input quality of summaries, affecting subsequent model inference performance. Second, existing systems often use static, single-instruction prompts for generative AI models, lacking the ability to dynamically adapt to multi-dimensional parameters such as user attributes, reading difficulty level, output language, and output format. This makes it difficult to achieve fine-grained control on the server side for different users and application scenarios. Third, existing conversational summarization systems typically place the dialogue logic in the terminal application, with the server passively forwarding requests. This lacks the ability to regenerate prompts and iteratively adjust summary results based on interactive operation information and natural language instructions within the server, resulting in low server resource utilization efficiency and complex and unstable state management. Fourth, in cross-language scenarios, summarization and translation are often implemented through independent modules. The lack of a unified mechanism for language detection, translation method selection, and integrated processing of summary results on the server side increases system integration complexity and latency, limiting the performance of generative AI models in cross-language summarization tasks.
[0058] Therefore, it is necessary to propose a system architecture and processing method that integrates high-quality preprocessing, dynamic prompt generation, iterative summary adjustment based on interactive information, and adaptive translation control on the server side. This would improve the overall processing flow and resource management of computers in long text summarization, conversational summarization, and cross-language summarization, thereby enhancing the controllability, consistency, and system operating efficiency of the summarization results. This would enable a more efficient and finer-grained computer implementation of generative artificial intelligence models.
[0059] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0060] In this invention, the server includes means for acquiring electronic information from an information providing medium via network communication; means for parsing the acquired structured documents or hierarchical data, removing display control information and decorative information, normalizing strings, and deleting useless components, thereby converting the electronic information into natural language text and performing preprocessing; means for dividing the natural language text into multiple parts based on the length and structure of the preprocessed natural language text and attaching metadata about the processing order and overall composition; and means for generating natural language prompts based on user attribute information and user needs information, including summary type, summary length, output language, output format, and reading difficulty level indications, embedding these parts into the data, and controlling the processing of the preprocessed natural language text. The system includes: a means for inputting prompt statements into a generative artificial intelligence (AI) model; a means for performing item enumeration, numbering, and hierarchical formatting on the summary result text obtained from the generative AI model and converting it into at least one of machine-readable and human-readable formats for transmission to a user device; a means for parsing interactive operation information and natural language instructions from the user device and regenerating updated prompt statements accordingly, and iteratively adjusting the summary result text by calling the generative AI model again; and a means for detecting inconsistencies between the language of the summary result text and the language of the user interface and selectively performing machine translation or using the generative AI model for translation to convert the summary result text into different languages. This allows for an integrated computational process within the server, encompassing electronic information acquisition, unified and standardized preprocessing, dynamic prompt statement generation, interactive summary iteration control, and cross-language translation. It optimizes the calling method and resource scheduling strategy of the generative AI model in long text summarization and conversational summarization scenarios, improves the adaptability of the summary results to user attributes and reading needs, reduces terminal-side logical complexity, and thus improves the processing efficiency and controllability of the entire computer system in information summarization and translation tasks.
[0061] "Information providing medium" refers to a data source platform or storage device that can provide electronic information to the outside world through a network, including but not limited to websites, application servers, database systems and file storage services.
[0062] "Electronic information" refers to content data stored or transmitted in digital form, including structured documents, hierarchical data, hypertext markup documents, markup language documents, and the parts of binary data that can be converted into text.
[0063] "Structured documents" refer to document data organized according to predefined tags, labels, or fields, including Hypertext Markup Language documents, Extensible Markup Language documents, and documents containing field name-value pairs.
[0064] "Hierarchical data" refers to data organized in a tree-like or nested structure, including key-value nested data exchange formats, markup language data, and other data structures represented in the form of hierarchical nodes.
[0065] "Display control information" refers to non-core content information used to indicate how content is displayed on the terminal, including tags, attributes, and codes used for layout, style, script execution, and interactive control.
[0066] "Decorative information" refers to content information that does not directly constitute the semantic subject of the text but is only used for visual presentation or beautification, including style tags, advertising blocks, icon tags, and decorative elements unrelated to the main text.
[0067] "String normalization" refers to the standardization of characters in text for the purpose of unifying encoding and format. This includes character encoding unification, whitespace character normalization, restoration of special escape characters, and regularization of full-width, half-width, uppercase, and lowercase characters.
[0068] "Removal of useless elements" refers to the process of removing content from text that does not affect the main semantic understanding or summary generation, including the removal of redundant whitespace, extra punctuation, advertising prompts, navigation text, and repetitive prompts.
[0069] "Natural language text" refers to a continuous sequence of characters written in human natural language, including sentences, paragraphs and their combinations, but excluding control tags and formatting instructions that are only for machine parsing.
[0070] "Preprocessing" refers to a series of standardization and cleaning operations performed before text input generative artificial intelligence models to improve the quality of subsequent analysis and generation, including parsing, extraction, normalization, segmentation, and noise removal.
[0071] "Partial data" refers to several sub-text units obtained by dividing the original natural language text according to length, structure, or semantic boundaries. Each unit can be used independently as part of the input of a generative artificial intelligence model.
[0072] Metadata refers to structured information attached to a portion of data to describe the attributes of that portion of data, including processing order, overall composition, source identifier, language, and content type.
[0073] "User attribute information" refers to information used to characterize a user's basic features and usage preferences, including age group, professional background, language preference, reading ability level, and terminal type.
[0074] "User requirements information" refers to the task objectives and output requirements expressed by users in specific sessions or operations, including the expected purpose of the summary, the required level of detail, the type of target audience, and time constraints.
[0075] "Abstract type" refers to the category of abstract task selected when generating abstracts from natural language text, including summary abstracts, general abstracts, role-specific abstracts, and abstracts with explanations, etc.
[0076] "Abstract length" refers to the quantitative constraints on the number of words, sentences, or key points in the generated abstract text, used to control the length and compactness of the output.
[0077] "Output language" refers to the type of natural language used in the generated summary text, including different natural languages or different writing conventions under the same language.
[0078] "Output format" refers to the way the summary results are organized in terms of structure and presentation style, including item list format, numbered list format, paragraph format, and machine-parseable structured format, etc.
[0079] "Reading difficulty level" refers to the grading index of the summary text in terms of vocabulary difficulty, syntactic complexity, and level of abstraction, which is used to match the reading needs of users with different comprehension abilities.
[0080] "Prompt statements" refer to instructional text issued to generative artificial intelligence models in the form of natural language. They are used to describe the task content, constraints, output requirements, and embed the text to be processed or part of its data.
[0081] "Generative artificial intelligence models" refer to artificial intelligence models built on machine learning and deep learning technologies that can automatically generate natural language text or other target content based on input prompts.
[0082] "Summary results text" refers to natural language text generated by a generative artificial intelligence model based on prompts and input text, used to summarize the key points and main content of the original electronic information.
[0083] "Item list structure" refers to a structure that organizes the summary text into multiple relatively independent list items, with each item representing a key point or sub-key point.
[0084] "Numbered structure" refers to a structure that assigns sequential identifiers to each item based on an item enumeration structure, including the use of Arabic numerals, letters, or other ordered symbols.
[0085] "Hierarchical structure" refers to the hierarchical organization of items or paragraphs in the summary text, so that they logically form a hierarchical structure with a main-subordinate, general-to-specific, or layered relationship.
[0086] "Machine-readable format" refers to a data representation format that is suitable for automatic parsing and processing by computer programs, including structured data formats and markup language formats.
[0087] "Human-readable format" refers to a text display format that is suitable for direct human reading and understanding, including natural language paragraphs, lists, headings, and other formats that are presented visually.
[0088] "User device" refers to a terminal device that allows users to access the system and present summary results, including computing terminals, mobile terminals, display terminals and other devices with network communication capabilities.
[0089] "Interactive operation information" refers to data with semantic control intent generated by users operating on the interface through user devices, including operation records such as button clicks, option switching, slider adjustments, and menu selections.
[0090] "Natural language instructions" refer to the user's control needs or preferences expressed through natural language input, including textual instructions on summary content, style, length, language, and other attributes.
[0091] "Update prompts" refer to prompts that are adjusted or reconstructed based on the original prompts after obtaining new interactive operation information or natural language instructions from the user, and are used to iteratively update the summary results.
[0092] "Language inconsistency" refers to a situation where the natural language used in the summary text is different from the natural language set in the user interface or user preferences.
[0093] "Machine translation processing" refers to the process of converting text between different natural languages using specialized automatic translation algorithms or systems, without relying on the dialogue context of the generative artificial intelligence models currently in use.
[0094] "Translation processing of generative artificial intelligence models" refers to the process of inputting prompts containing translation instructions into a generative artificial intelligence model, which then performs the conversion from the source language to the target language based on its generative capabilities.
[0095] In this embodiment of the invention, the server, terminal, and user each assume different functional roles, collaboratively achieving summarization, dialogic adjustment, and cross-language conversion for long texts and multi-source electronic information. The system of this invention can be deployed in a data center or cloud computing environment. The server can be a general-purpose computer equipped with a multi-core central processing unit and a graphics processing unit, and the operating system can be a kernel-based server operating system, such as a Unix-like server operating system. The server can run backend programs using an interpreted programming language, for example, using an interpreter to execute data processing code and call third-party software libraries, such as network communication libraries, hypertext markup parsing libraries, and artificial intelligence interface libraries. The terminal can be a computer terminal, mobile terminal, or embedded terminal, running a graphical user interface program or browser for exchanging data with the server and presenting summary results.
[0096] In this invention, the server communicates with external information providers via a network interface card and a transmission control protocol / Internet protocol stack. The server uses a network communication library to obtain electronic information from websites, application servers, or database interfaces, such as hypertext markup documents, hierarchical data, or structured records. After receiving the electronic information, the server temporarily stores it in main memory or a non-volatile storage device for subsequent text extraction and preprocessing.
[0097] During text preprocessing, the server uses a hypertext markup parsing library (such as a tree-structured parsing library) to perform syntactic analysis on the acquired structured document. The server removes display control and decorative information from the markup, including scripts, styles, advertising areas, and navigation bars, retaining only nodes relevant to the main text content. The server performs normalization using string processing algorithms, such as unifying character encoding to a unified encoding format, merging multiple consecutive whitespace into a single whitespace, removing invisible control characters, and converting special escape entities to their corresponding characters. The server further removes content deemed useless, such as general copyright notices, recommended link lists, and duplicate titles, to reduce noise interference with subsequent model inference. At this stage, the server uniformly converts multi-source electronic information into natural language text, thus forming structured, less noisy summary input data.
[0098] To improve the computational efficiency and summarization accuracy of generative AI models, the server automatically segments the preprocessed natural language text. Based on text length, paragraph structure, and logical boundaries, the server divides the text into multiple parts. The server attaches metadata to each part, including its sequence number in the original text, its section identifier, estimated semantic importance score, and language tag. The server uses a specific data structure (e.g., a combination of lists and key-value maps) to store the parts and their metadata. This allows the server to flexibly combine different blocks when calling the generative AI model, improving the utilization of computational resources.
[0099] When constructing the prompt, the server first receives user attribute information and user request information from the terminal. The terminal, through a graphical user interface, packages parameters such as the user's age, professional background, intended use, expected summary length, target language, and reading difficulty preference into a structured request and sends it to the server. Based on these parameters, the server selects a corresponding template from a set of templates stored in a database or configuration file. For example, the server determines the appropriate text style and output constraints from multiple templates such as "children's," "general," and "professional," based on the user's selection and attribute information. The server then replaces the placeholders in the template with specific indications such as summary type, summary length, output language, and reading difficulty level, and inserts some data sequentially into the designated positions in the prompt, generating a complete natural language prompt.
[0100] In one implementation, the server can generate the following example prompt statement: Please summarize the following article into 5 key points in simplified Chinese, each no more than 30 characters.
[0101] Output format requirements: 1. Use a numbered list: 1. 2. 3. … 2. Highlight the time, place, main participants, and key results.
[0102] The article content is as follows: [Insert preprocessed text here] In another implementation, the server can generate prompts for middle school students, for example: "If you are a teacher, please rewrite the following article into a point summary suitable for middle school students."
[0103] Require: 1. Use Simplified Chinese.
[0104] 2. Sentences should be short and vocabulary should be simple.
[0105] 3. The total length should not exceed 300 characters.
[0106] Original content: [Insert preprocessed text here] In cross-language summarization scenarios, the server can generate prompts that include both translation and summarization instructions, for example: Please read the following English article first, then write a point-by-point summary in Simplified Chinese for readers with basic programming knowledge. Requirements: 1. First, summarize the topic in one sentence; 2. Then, use 3-5 key points to explain the main conclusions and key figures; 3. Use the original English terminology as much as possible and provide a brief explanation in parentheses.
[0107] The following is the article content: [Insert original English text here] When the server needs to adjust the difficulty of an existing digest, it can generate the following update prompt statement: "Below is the previously generated article summary. Please rewrite the summary to a version suitable for middle school students, using shorter sentences and more common vocabulary, while retaining the bullet-point structure, without omitting important information."
[0108] Original abstract: [Insert previous summary text here] The server takes the aforementioned prompt as text input, along with corresponding data, and sends it to the inference service of the generative AI model via a network request. The generative AI model can be deployed on a dedicated inference server, employing a neural network architecture with multi-layered encoders and decoders, such as a transformer architecture based on a combination of self-attention mechanisms and feedforward networks. During offline training, the model utilizes a large-scale text document dataset for supervised and self-supervised learning, employing a cross-entropy loss function to measure the difference between the generated sequence and the labeled sequence, and updating network weights through gradient descent and backpropagation algorithms. The model can introduce a multi-head attention mechanism to capture long-distance dependencies between different positions in the text and use positional encoding to represent the relative positional information of words in the sequence, thereby achieving high modeling capabilities in long text processing. During training, the model can also use data augmentation techniques, such as randomly deleting, replacing, or rearranging parts of sentences, to enhance robustness to noisy text.
[0109] During the inference phase, the server inputs prompts and partial data into the generative AI model. Internally, the model first maps the input tokens into high-dimensional vector representations through an embedding layer. Then, matrix multiplication, linear transformation, normalization, and nonlinear activation operations are performed within a multi-layered encoder-decoder structure. Based on explicit constraints in the prompts regarding summary type, summary length, and output format, the model compresses information and extracts key points from the input text, generating a summary text sequence that conforms to the instructions. During generation, the model can use temperature parameters, sampling strategies, or bundle search strategies to control the diversity and determinism of the output. The server receives the output token sequence from the model's inference interface and converts it into a natural language string.
[0110] After retrieving the summary text, the server performs formatting and structuring processing. Based on the output structure agreed upon in the prompt, the server splits the text line by line, removes redundant prefixes automatically added by the model (e.g., "The following is the summary:"), and renumbers and groups it as needed. The server can represent an itemized structure as an ordered list, or a hierarchical structure as a nested list or multi-level headings, thus facilitating both human reading and further machine parsing. The server determines whether to return a summary in machine-readable, human-readable, or a combination of both formats based on the terminal request.
[0111] In a multilingual environment, the server checks whether the language of the summary text matches the language of the user interface. The server can use language recognition algorithms to perform statistical feature analysis or model-based classification to obtain language labels. When an inconsistency is found, the server chooses a translation path: on the one hand, it can call a dedicated machine translation engine to perform text translation; on the other hand, it can input prompts containing translation instructions into a generative artificial intelligence model, requesting the model to complete the natural language conversion from the source language to the target language. For example, the server can construct the following translation prompt: Please translate the following English abstract into simplified Chinese, ensuring natural and accurate word choice and maintaining the original bullet point structure: [Insert English summary here] By implementing language detection and translation selection uniformly on the server side, the server can achieve integrated processing of summarization and translation without increasing the burden on the terminal.
[0112] In this invention, the terminal performs display and interaction functions. The terminal communicates with the server via an interface through a browser or local application, receiving formatted summary text. The terminal's display module presents the summary to the user in the form of a numbered list, an indented list, or cards, based on the structure returned by the server. The terminal also provides multiple interactive controls in the interface, such as a slider to change the summary length, a drop-down menu to select the reading difficulty level, a button to switch the output language, and a text box for inputting natural language instructions. After user interaction, the terminal encapsulates the interactive operation information and natural language instructions into a request and sends it to the server.
[0113] In this invention, users can read summary results, adjust summary styles, and submit additional requests via a terminal. When reading lengthy reports, technical documents, or cross-language materials, users can obtain summary versions of varying complexity through this system, thereby reducing the time spent reading word by word. Users can also quickly switch between summary results in multiple languages to improve information comprehension efficiency in multilingual work environments.
[0114] The system of this invention does not merely automate the manual summarization process superficially; rather, it optimizes the computer processing itself within the server through specific data structures, prompt generation strategies, and model invocation methods. During the preprocessing stage, the server significantly reduces noise and format differences in the input of generative AI models through unified parsing and normalization, enabling the models to reason on a more stable input distribution, thereby improving the stability and accuracy of the summarization results. By dividing long texts into portions with sequential and structural metadata, the server avoids the context truncation problem caused by feeding excessively long sequences to the model at once, reducing unnecessary redundant calculations and lowering memory usage and computational load during inference.
[0115] During the prompt generation process, the server combines user attribute information, user requirement information, and template information to automatically construct refined prompts. Compared to the traditional static parameter calling method, this mechanism can precisely adjust the output style, length, and difficulty through input control without changing the model parameters, thereby improving the controllability and reusability of model calls. By parsing interactive operation information and natural language instructions, the server regenerates and updates prompts on the server side and iteratively calls the model, centralizing the dialogue logic within the server. This simplifies the terminal logic and reduces network communication overhead, while enabling the server to execute more intelligent reuse and caching strategies based on context history and metadata, further improving the overall system response speed.
[0116] This invention employs a multi-layered self-attention structure and a large-scale parameter matrix in its generative artificial intelligence model. By introducing multi-task learning and instruction fine-tuning during the training phase, it enhances the model's ability to understand various task instructions. Internally, the model performs semantic decomposition and importance modeling of the input text through multi-head attention weights and intermediate representations output by a feedforward network. Compared to traditional rule-based methods based on keyword matching or sentence scoring, it possesses higher feature representation dimensions and non-linear combination capabilities. This invention explicitly encodes the summary type and output format requirements in the prompt statements, enabling the model to generate summaries along specific "task subspaces" during inference, thereby improving the structure and information focus of the summaries. The server combines this internal model structure with the external prompt statement control mechanism to form a technical solution that improves both the accuracy and efficiency of the computer system in long text summarization tasks.
[0117] Through the aforementioned structure and processing flow, the server can reduce the number of repeated network calls under the same hardware resources, and reduce the latency of a single inference through text segmentation and result caching mechanisms. When the server concurrently processes summary requests in a multi-user environment, it can deduplicate and share some data based on metadata, thereby reusing intermediate summary results among multiple similar documents, further reducing the overall computational load. Since the server implements unified formatting and translation management internally, the terminal only needs to handle simple interface logic, reducing the overall system maintenance cost and error rate.
[0118] This invention can also be implemented in various variations. For example, the server can use different types of generative artificial intelligence models, including models with smaller parameter sizes suitable for deployment in edge computing environments; the server can also assign summarization and translation tasks to different models and select models based on resource availability through a scheduling module. The server can employ a combination of fixed-length chunking, semantic boundary chunking, or title-paragraph-based chunking strategies to adapt to different document types. In translation processing, the server can choose between using a general generative model or a dedicated translation engine based on the statistical characteristics of different language pairs, thus striking a trade-off between translation quality and response time.
[0119] Through the aforementioned implementation, the collaborative interaction of the server, terminal, and user within the system of this invention enables the generative artificial intelligence model's capabilities to be strategically scheduled through specific prompting strategies and data processing flows, achieving the technical integration of long text summarization, conversational summarization adjustment, and cross-language summarization. This system not only improves information processing efficiency and summarization quality but also enhances the overall performance and scalability of the computer system in natural language processing tasks through improvements to the computational process, data structure, and model invocation methods.
[0120] use Figure 11The processing flow is explained.
[0121] Step 1: The server acquires electronic information from the information providing medium. The input consists of request parameters sent by the terminal (including the target URL, interface identifier, authentication information, etc.) and pre-configured data source information. Based on the input, the server sends a request message to the target data source via the network interface card and network communication protocol, and receives response data from the website, application server, or database interface. The server parses the response message, extracts the main content as raw electronic information, and outputs a raw data set containing structured documents or hierarchical data. The data processing performed by the server in this process includes: decoding network messages, decompressing compressed content, identifying and filtering error status codes, and storing successfully acquired content in a temporary cache.
[0122] Step 2: The server performs structural parsing and text extraction preprocessing on the raw electronic information. The input is the raw data set output from step 1. The server first determines the data format based on the content type tag: is it a hypertext markup document, hierarchical data, or another format? When dealing with hypertext markup documents, the server uses a parsing library to construct a document tree structure, traversing nodes layer by layer, and removing display control and decorative information such as scripts, styles, and ad blocks. When dealing with hierarchical data, the server extracts important attributes such as title, body text, and time fields based on the key-value structure. Afterward, the server performs normalization processing on the extracted text, including unifying character encoding, merging consecutive whitespace, removing control characters and common noise segments. Through the above data processing, the server converts the messy raw data into continuous natural language text and forms document metadata (such as title, time, and source). The output is a set of preprocessed natural language text and corresponding metadata records.
[0123] Step 3: The server segments the preprocessed natural language text and appends metadata. The input is a set of natural language text and its metadata output from step 2. The server calculates the number of segments and the target length of each segment based on the length and paragraph structure of each text; the server aligns the segmentation positions as closely as possible to paragraph boundaries to maintain semantic continuity. For each text, the server divides it into multiple parts, appending metadata fields to each part, including the original document identifier, segment number, start and end positions in the document, inferred importance score, and language tag. Internally, the server uses lists and mapping structures to store the association between the parts and the metadata. Data processing in this step includes string slicing, boundary correction of the segmentation results, and calculation of the number of characters and sentences for each segment. The output is a list of parts carrying detailed metadata, used for subsequent generation of prompts and invocation of generative artificial intelligence models.
[0124] Step 4: The terminal collects user attribute information and user requirement information and sends them to the server. Input consists of user actions and text input on the terminal interface, including the user's selected reading difficulty, target summary length, target language, summary purpose (e.g., quick browsing, report preparation), and the user's natural language instructions. The terminal organizes this information into a structured request, adds device and session identifiers, and sends it to the server over the network. Output is a request message containing user attribute information and requirement parameters, which influences how the server subsequently generates prompts. Data processing performed by the terminal during this process includes: reading the state of user interface components, encoding text input, and performing simple validity checks.
[0125] Step 5: The server generates prompt statement templates and populates partial data based on user attribute and requirement information. The input consists of the partial data list from step 3 and the user request message from step 4. First, the server matches appropriate text templates based on user attributes (e.g., age group, professional level) and template fragments related to reading difficulty and output format. Then, based on parameters such as summary type (point list, overview summary, etc.), summary length, and output language, the server combines these template fragments into a complete prompt statement framework. Next, the server selects an appropriate number of partial data from the partial data list and embeds their content sequentially into the reserved content paragraph positions in the prompt statement. The server may indicate "Part X / of Part Y" in the prompt statement when necessary, to help the generative AI model understand the overall structure. The output is a complete set of natural language prompt statements, each associated with its corresponding partial data, ready to be sent to the generative AI model. Data processing in this step includes: string concatenation, template placeholder replacement, and sorting of partial data based on metadata.
[0126] Step 6: The server invokes a generative AI model to generate the summary text. The input consists of the prompts output in step 5 and their corresponding partial data. The server uses a network interface and an AI service interface to send each prompt as input to the model, along with parameters such as the model name, maximum output length, and sampling temperature to the inference service. In the inference server, the generative AI model encodes the input labeled sequence and performs matrix multiplication, attention weight calculation, and nonlinear transformations within a multi-layered self-attention structure to progressively generate the output labeled sequence. Upon receiving the output, the server decodes the labeled sequence into a natural language string as the summary text. The data processing performed by the server in this step includes: encoding the prompts, parsing the data structure returned by the model, and performing preliminary cleaning of the generated text (e.g., removing extra blank lines). The output is a list of summary texts that correspond one-to-one with the input partial data.
[0127] Step 7: The server merges and formats multiple summary results. The input is the list of summary result texts output from step 6, along with metadata for the corresponding data. The server sorts and groups the summary results based on document identifiers and block numbers in the metadata, merging multiple summary results belonging to the same original document in the correct order. During the merging process, the server eliminates duplicate content across blocks and adjusts the connection relationships of linking sentences to ensure overall summary coherence. Subsequently, the server reconstructs the summary results into a numbered list, hierarchical list, or paragraph format according to the user-specified output format, adding entry prefixes and indentation information. The server can also convert the merged summary into a structured format for terminal interface rendering or subsequent processing by other systems. The output is one or more structured and formatted summary result texts, each corresponding to one original document. This data processing step includes: text similarity comparison for deduplication, block number-based sorting and merging, and insertion of structure tags.
[0128] Step 8: The server performs language consistency checks and translates as needed. The input is the summary text output from step 7 and the user interface language setting. The server calls the language recognition module to calculate a language type label for each summary text. The server compares this label with the user interface language or the user's preferred language to determine if there is a language inconsistency. If an inconsistency is found, the server decides whether to use a machine translation engine or re-invoke a generative AI model to perform the translation, based on a preset strategy. If a generative AI model is selected, the server constructs a prompt containing translation instructions and embeds the summary text within it. For example, the server can generate: "Please translate the following summary into Simplified Chinese, maintaining the original point structure: [Summary Text]". The server sends this prompt to the model service and receives the translated summary text. The output is the target summary text with a language consistent with the user interface. This step's data processing includes: language label calculation, conditional branch selection of the translation path, and format alignment of the text before and after translation.
[0129] Step 9: The server sends the final summary result to the terminal. The input is the summary text and its structure information from step 7 or 8, which has already been translated and formatted as needed. The server encapsulates these results into a response data packet, including links to the original documents, generation timestamps, and language tags, and returns it to the terminal via the network interface. The server may compress the response before sending, depending on network conditions, to reduce bandwidth usage. The output is a response message arriving at the terminal, containing summary data that can be displayed and used for subsequent interaction. Data processing includes serialization, optional compression, and field trimming if necessary.
[0130] Step 10: The terminal displays the summary results and collects subsequent interaction information. The input is the summary result message returned by the server in step 9. The terminal parses the message structure and draws the summary content on the user interface according to numbering, hierarchy, or paragraph format; the terminal also displays summary-related control components, such as a slider to adjust the summary length, a drop-down box to switch reading difficulty, a language switch button, and a text box for inputting natural language instructions. The user views the summary content on the terminal and can further request the summary by clicking components, dragging controls, or entering text. The terminal records these operations, organizes them into new interactive operation information and natural language instructions, and prepares to send them to the server. The output is new request data for the server, which includes the user's adjustment requests for the current summary. The data processing performed by the terminal in this process includes: interface rendering, event listening, operation log collection, and simple data packaging.
[0131] Step 11: The server regenerates update prompts based on user interaction information and iteratively adjusts the summary. Inputs include the user interaction information and natural language instructions output in step 10, as well as the currently generated summary text and related metadata. The server parses the user interaction information, identifying user-requested changes, such as changing the summary difficulty from "professional" to "general," shortening the summary length from 5 points to 3 points, or requesting the addition of risk analysis. Based on these changes, the server selects new text styles and output constraints from the template information set, inserting the original summary text as input into the new prompts. For example, the server could generate: "Below is the previously generated article summary. Please rewrite the summary to a version suitable for middle school students, using shorter sentences and more common vocabulary, while retaining the bullet-point structure, without omitting important information. Original summary: [Original summary text]." The server sends the update prompts to the generative AI model to obtain the new summary text. Output is the updated summary text that matches the user's latest needs. The data processing in this step includes: mapping interaction parameters to templates, reconstructing prompt statements, embedding old summary content, and managing old and new summary versions.
[0132] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0133] In existing technologies, for text information from large-scale information sources such as online media and databases, summaries are typically generated using only fixed rules or simple models, which presents the following technical problems: First, the server-side processing flow for raw text is rather crude, lacking a preprocessing mechanism that dynamically adjusts based on text attributes and user context. This results in excessive text redundancy in the input generative AI model, consuming computational resources and affecting the quality of the summary, thus reducing the overall processing efficiency of the computer system in the summary task.
[0134] Second, existing systems often use static templates when constructing prompts to drive generative artificial intelligence models. They cannot dynamically generate differentiated prompts based on users' real-time operations, comprehension levels, language preferences, and areas of interest. As a result, the model's output summary results are difficult to adapt to different users and different terminal scenarios in terms of structure, style, and readability, thus limiting the effective use of artificial intelligence reasoning capabilities in computer systems.
[0135] Third, many systems only return a continuous text summary without performing structured parsing and formatting of the generated results. They cannot automatically decompose the summary into a column structure suitable for continuous vertical scrolling display on the terminal, and cannot uniformly reshape it with multimedia information such as images and videos. This results in the client needing to perform additional complex processing at the rendering and interaction layers, leading to low overall human-computer interaction performance and user-side rendering efficiency.
[0136] Fourth, in existing solutions, the generation of summaries usually lacks a closed-loop feedback mechanism for user interaction. The server has difficulty updating the information acquisition strategy and prompt generation conditions in a timely manner based on data such as the user's browsing history, dwell time, and click behavior. This results in the summary generation process being disconnected from the recommendation process, making it impossible to continuously optimize information retrieval priorities and model calling strategies at the system level. Consequently, it is difficult to fully leverage the computer system's capabilities in personalized information processing.
[0137] Fifth, in terms of cross-language processing, traditional technologies often separate translation processing from abstract generation, lacking an integrated processing flow that uniformly manages "original text – abstract – multilingual versions" on the server side. This results in multiple redundant text transmissions and conversions, increasing latency and resource consumption, and reducing the overall system performance in a multilingual environment.
[0138] Therefore, it is necessary to provide a new system that improves the intelligent summarization and display capabilities of large-scale text information at the computer technology level by introducing a dynamic prompt generation mechanism based on a generative artificial intelligence model, a structured summary parsing and multimedia shaping mechanism, an interactive feedback-driven acquisition strategy adjustment mechanism, and a unified cross-language processing mechanism on the server side, thereby enhancing the system's processing efficiency, scalability, and user experience.
[0139] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0140] In this invention, the server includes: a device for acquiring information from an information source and filtering the information based on the user's areas of interest and usage history; a device for preprocessing the acquired information as text data to remove useless elements, perform sentence segmentation, and segment or truncate long texts; a device for generating prompts for inputting into a generative artificial intelligence model based on text attributes and specified summary expression styles, comprehension levels, language, and categories received from the user, and controlling the input of the prompts and the text data into the generative artificial intelligence model; and a device for parsing the summary results output by the generative artificial intelligence model and splitting the summary results into a column structure, integrating them with titles, image information, and video information to form a pre-reading format suitable for continuous vertical scrolling. The system includes a device for defining a display format, a device for sending summaries in a predetermined display format to a user terminal and providing a user interface for displaying the summaries page by page, a device for acquiring input operations or natural language input from the user terminal and interactively modifying the prompt template, summary length, style, comprehension level, and language accordingly, and then calling the generative artificial intelligence model again to regenerate the summary results after modification, a device for performing translation processing between different languages on the summary results and the corresponding original text and sending summaries or original texts in multiple languages to the user terminal according to user selection, and a device for learning user interest levels based on user browsing history, selection operations, and dwell time, and updating the priority of information retrieval from information sources and prompt generation conditions accordingly. This allows for the formation of a computational process on the server side centered around a generative artificial intelligence model, encompassing dynamic prompt generation, structured summary parsing, multimedia integrated shaping, interactive feedback-driven retrieval and generation strategy adjustment, and unified cross-language processing. This fundamentally optimizes the data flow and control flow of text summarization tasks in the computer system, reduces invalid computation and redundant data transmission, significantly improves the efficiency and response speed of large-scale information processing, and enhances the overall performance of terminal display and human-computer interaction.
[0141] "Information source" refers to a data provider or storage medium that can be accessed by the system through a communication network to provide information to be processed, including but not limited to network media services, database systems, local storage devices, and external application programming interfaces.
[0142] "User terminal" refers to an electronic device operated by a user, which interacts with a server through a communication network to display summary results and receive user input. It includes, but is not limited to, mobile terminals, wearable terminals, and fixed terminals.
[0143] "Text data" refers to textual information that can be processed by a computer and is represented as a sequence of characters after being obtained from an information source and converted into a format. This includes, but is not limited to, natural language sentences, paragraphs, and document content.
[0144] "Preprocessing" refers to the procedural processing operations performed on text data before it is input into a generative artificial intelligence model, in order to improve the efficiency and quality of subsequent processing. These operations include removing useless elements, sentence segmentation, paragraphing, and long text truncation.
[0145] "Useless elements" refer to text parts that do not contribute substantially to the summary generation task and may affect the generation quality or increase the computational burden, including but not limited to advertising statements, copyright notices, redundant marks, format control symbols, and duplicate content.
[0146] "Generative artificial intelligence models" refer to models built based on machine learning methods that can automatically generate or rewrite natural language content based on input prompts and text data, including but not limited to language models based on deep neural networks and sequence models for text generation.
[0147] "Prompt statements" refer to a type of instructional text constructed by the server and used as input to a generative artificial intelligence model. It is used to indicate the generation goal, generation method, or constraints of the generative artificial intelligence model, including information such as the number of summary entries, text type, target comprehension level, and language.
[0148] "Summary results" refers to the text content output by the generative artificial intelligence model after receiving prompts and text data, which is a compression and summary of the original information, including but not limited to key points, summary paragraphs, and supplementary explanations in headings.
[0149] "Block structure" refers to a structured representation of the summary results by breaking them down into multiple independent points and arranging them one by one, so that each point is presented as an independent item.
[0150] "Preset display format" refers to the data structure and layout style of the summary results that are pre-set on the server side for display on the user's terminal, including titles, bullet points, image information, video information, and layout parameters for vertical continuous scrolling display.
[0151] "User interface" refers to the interactive interface presented by the user terminal, used to display summary results and receive user operation instructions, including but not limited to list interface, detail interface, settings interface, and search interface.
[0152] "Interactive editing" refers to a processing method in which the server adjusts parameters such as prompt message templates, summary length, style, comprehension level, or language based on the user's operation or natural language input submitted through the user terminal, either in real time or in near real time.
[0153] "Translation processing" refers to the procedural operation of automatically converting the summary results or the corresponding original text between different languages, so as to generate another target language text while keeping the original semantics basically unchanged.
[0154] "User interest level" refers to the result of quantifying or qualitatively representing user preferences based on user behavior data in the system, including but not limited to the intensity of preference for specific topics, categories, or styles of content.
[0155] "Browsing history" refers to the record of a user's access to summary results or original text on the user terminal within a predetermined time range, including access time, access frequency, and access order.
[0156] "Selection action" refers to the interactive action performed by the user on the user interface through the user terminal to express preferences or issue control commands, including clicking, swiping, selecting categories, selecting text styles, specifying keywords, and switching languages.
[0157] "Dwell time" refers to the duration for which a user remains browsing a particular summary result, original text page, or interface, reflecting the user's level of attention to the content.
[0158] "Image information" refers to static visual data or its descriptive data associated with the summary results or original information, including image files, image addresses, and metadata used to present the image.
[0159] "Video information" refers to dynamic visual data or its descriptive data associated with the summary results or original information, including video files, video addresses, and metadata used to present the video.
[0160] "Category" refers to the attribute tag used to classify information in an information source. It indicates the subject or field to which the information belongs, including but not limited to types such as science and technology, finance, sports and culture.
[0161] "Style" refers to the stylistic features of the abstract in terms of language expression, including but not limited to formal style, casual style, and explanatory style aimed at a specific audience.
[0162] "Comprehension level" refers to the cognitive level of the target audience used to set the complexity and professionalism of the summary results, including but not limited to beginner, intermediate and advanced levels.
[0163] "Language" refers to the types of natural languages used to present textual information, including but not limited to Chinese, English, and other human natural languages.
[0164] "Vertical continuous scrolling display" refers to a display method in which multiple summary results are arranged sequentially from top to bottom in the user interface of a user terminal, allowing users to continuously browse multiple summary items by scrolling up or down.
[0165] In one embodiment of the invention, the server is deployed on a computer cluster or cloud computing platform. The server includes one or more computing devices equipped with a central processing unit (CPU), a graphics processing unit (GPU), main memory, persistent storage, and a network interface. The server runs a general-purpose operating system, such as a Unix-like operating system, and installs on it a scripting runtime environment (e.g., an interpreted language-based runtime environment), natural language processing libraries (e.g., libraries for word segmentation, sentence segmentation, and entity recognition), a deep learning inference framework (e.g., an inference framework based on the Transformer architecture), and client programs for calling generative artificial intelligence models. The server communicates with one or more database servers and multiple user terminals via a communication network.
[0166] In one embodiment of the present invention, the server stores program modules in a storage device, including: an information acquisition module, a text preprocessing module, a prompt statement generation module, a generative artificial intelligence model invocation module, a summary parsing and structuring module, a multimedia association and shaping module, an interaction parameter management module, a translation processing module, and a user interest learning and priority update module, etc. The server loads the above program modules into the main memory and executes these modules sequentially or in parallel on the processor to complete data processing such as information filtering, summary generation, display format shaping, and interactive feedback.
[0167] In a more specific implementation, the server uses a natural language processing library (such as spaCy or a similar library) to perform sentence segmentation, part-of-speech tagging, and removal of useless elements from the raw text obtained from the information source. The server uses a regular expression engine and text filtering rules to remove useless elements such as advertising logos, copyright notices, duplicate headers and footers, and footnote links. During preprocessing, the server stores the main text of an article as an ordered list of sentences and uses this list as the input data structure for the subsequent summary generation process. When processing very long articles, the server sets thresholds based on the number of sentences and the total number of characters. For example, when the number of sentences exceeds a predetermined limit, the server only retains the first few paragraphs, or extracts a portion of paragraphs based on paragraph importance scores (calculated by a simple term frequency-inverse document frequency algorithm or a lightweight classification model) to control the maximum length of input to the generative artificial intelligence model, thereby reducing the computational burden.
[0168] In one embodiment of the invention, the server uses a generative artificial intelligence model based on the Transformer architecture as the core for summarization. This model consists of a multi-layer self-attention encoder and a multi-layer self-attention decoder, or an autoregressive language model consisting only of stacked decoders. During model training, the server uses a large-scale corpus, where each training sample includes the source document and its manually written summary. The server uses a cross-entropy loss function as the error function to measure the difference between the predicted output tag sequence and the target summary sequence, updating the model parameters through stochastic gradient descent or its variants (e.g., adaptive learning rate optimization algorithms). During training, the server uses data augmentation techniques, such as sentence rearrangement, synonym substitution, and noise insertion, to increase the model's robustness to different expressions. During the inference phase, the server fixes the model parameters and only performs forward propagation to avoid the uncertainties introduced by online learning.
[0169] In one embodiment of the invention, the server does not directly input the original article into the generative artificial intelligence model. Instead, it generates prompts to control the model's behavior. The server searches for a corresponding prompt template in a storage device based on the category, style, comprehension level, and language preference selected by the user on the terminal. The server inserts constraints into the prompt template, such as the maximum number of summary entries, the maximum character length of each entry, and the target audience's comprehension level. The server concatenates the prompt template with the preprocessed main text to form a continuous input sequence, which is then used as input to the generative artificial intelligence model. Thus, without modifying the model's structure and parameters, the server controls the model's generation strategy through textual instructions, thereby avoiding reliance on a single static summarization algorithm.
[0170] In a specific example, when a user selects the "Technology" category and prefers a formal writing style, the server generates the following Chinese prompt: Please summarize the following technology news article into 3-5 bullet points in simplified Chinese, each no more than 30 characters, maintaining a formal and objective tone: 『Article text…』 In another example, the user specifies "targeting middle school students, with a relaxed tone" in the terminal, and the server generates the following prompt: Please summarize the following news article into 4-6 bullet points in simplified Chinese, in a way that is easy to understand, suitable for middle school students, and fun to read. Provide brief explanations of technical terms if necessary: 『Article text…』 In an English-speaking environment, the server can generate the following prompt: Please summarize the following news article in 3-5 key points. Use concise, neutral English expressions suitable for busy professionals: "Article body..." In this way, the server can generate differentiated prompts for different user needs under different implementation forms. The prompts contain structured constraints such as the number of summary entries, word limit, text type, and comprehension level. The generative artificial intelligence model adjusts the generation strategy according to these instructions during the decoding stage.
[0171] In one embodiment of the present invention, after receiving the text output by the generative artificial intelligence model, the server performs summary parsing and structuring processing. The server identifies bullet points (e.g., "•", "-", "1.", etc.) using string parsing rules and splits the generated text into multiple summary points. When necessary, the server uses sentence boundary detection algorithms and delimiter matching algorithms to correct for inconsistencies in the model's output format, ensuring that each summary point is an independent sentence or phrase. The server stores these points in a data structure, such as a list or array, where each element contains the point text, length information, importance score, and other fields. According to pre-defined display specifications, the server packages the title, bullet points, and multimedia link addresses into structured records for use in terminal interface rendering.
[0172] In one embodiment of the invention, the server reshapes the summary results along with image and video information. The server retrieves multimedia resources matching the article title and text keywords from an information source or media database. The server can employ a simple keyword matching algorithm or a lightweight vector retrieval algorithm to represent the text as vectors and search for the most similar image or video entry in the multimedia index. The server writes the address and thumbnail information of the found multimedia resources into a data record associated with the summary results. During the reshaping process, the server generates a data layout suitable for vertical continuous scrolling display. For example, it generates an information block for each article containing a title, several columns of highlighted keywords, and one or more preview images for direct rendering on the terminal side, eliminating the need for complex layout calculations on the terminal.
[0173] In one embodiment of the invention, the terminal is a mobile terminal or a wearable terminal. The terminal runs a local application and establishes an encrypted connection with the server through a network module provided by the system. After receiving the summary data returned by the server, the terminal uses a local user interface framework (e.g., a list component provided by the mobile operating system) to render the data into a vertically scrollable list interface. The terminal displays a title, a set of bulleted summaries, and a corresponding image or video preview in each list item. The terminal implements view reuse and lazy loading strategies as the user scrolls, reducing the repeated loading of image resources, thereby reducing communication load and memory consumption, and improving interface responsiveness.
[0174] In one embodiment of the invention, the user provides constraints to the server through a terminal interface. The user selects different information categories in the category selection area, formal or casual style in the style selection area, sets the comprehension level to general reader or beginner in the comprehension level area, and sets the display language in the language selection area. The user can provide free-form requests via text or voice input, such as "only read AI-related technology news, with brief summaries and a casual tone." The terminal encodes these parameters into structured request data and sends it to the server via the network. After parsing the request, the server updates the prompt generation conditions and information retrieval conditions, thus forming a closed-loop control based on user interaction.
[0175] In one embodiment of the invention, the server not only adjusts the prompts based on the user's immediate instructions but also learns the user's interests based on accumulated behavioral data. The server records the user's browsing history, click selections, and dwell time on specific summary pages in a local or independent storage system. Through statistical analysis or simple machine learning models, the server calculates the attractiveness of different categories and topics to the user, forming a user interest vector. When subsequently retrieving information from information sources, the server adjusts the search ranking priority based on the interest vector and adds keywords related to the user's highly interested topics to the prompts, making the generative artificial intelligence model more inclined to highlight these contents. Because the server centrally performs these analysis and update operations on the computing nodes, the terminal does not need to bear the complex computational burden, thereby reducing terminal resource consumption.
[0176] In another embodiment of the invention, the server performs translation processing, enabling the system to support multilingual summaries and original text display. After generating the summary, if the user's target language differs from the text language, the server invokes the translation module. The translation module can be implemented based on a statistical machine translation model or a neural network-based translation model. The server inputs the summary key points or the original text into the translation model sentence by sentence and obtains the corresponding target language result. The server stores the translated text and the original text summary together in a data structure, allowing the terminal to quickly switch the display when the user switches languages without needing to request summary generation or translation again, thereby reducing network communication load and improving response speed.
[0177] In terms of technical effectiveness, this invention introduces a prompt-driven generative artificial intelligence model invocation method, structured summary parsing, multimedia shaping, and user interest feedback mechanism on the server side. This transforms text summarization processing from a simple automated replication of human reading behavior into a deep optimization of the computer's internal data and control flows. The server preprocesses ultra-long texts by segmenting and truncating them, reducing redundant content input to the model, thereby lowering inference time and memory usage, and increasing processing throughput. By embedding parameters such as character limits, number of entries, and target comprehension levels into the prompts, the server ensures the generated results better match the terminal's display needs, reducing the complexity of secondary processing on the client side. The server dynamically updates the information acquisition strategy and prompt generation conditions through a user interest learning module, enabling adaptive optimization of the subsequent summary generation process at both the source data selection and generation control levels, thus achieving a higher effective information output ratio under the same hardware conditions.
[0178] Unlike traditional rule-based summarization, the generative AI model in this invention is not based on manually preset fixed extraction rules. Instead, it automatically learns the long-distance dependencies and importance distribution within the text through a multi-layer attention mechanism. During the training phase, the model updates the weights of each layer using an error backpropagation algorithm to minimize the difference between the predicted and target summaries in the training samples. During the inference phase, the server utilizes this internal structure to generate summarizing sentences that meet the given prompt constraints step-by-step through autoregressive decoding. The output quality and flexibility are superior to simple keyword extraction or sentence scoring and ranking methods. Because the server incorporates unconventional constraints into the prompts (e.g., simultaneously limiting the number of entries, characters, style, and audience comprehension level), the generative AI model follows a set of machine optimization rules that differ from human intuitive writing habits during inference, thereby increasing information density while reducing redundant output.
[0179] In other optional embodiments of the present invention, the server can employ different types of generative artificial intelligence models, such as sequence-to-sequence models with encoder-decoder structures, autoregressive models with decoder-only structures, or extended models including external memory modules. The server can also select different preprocessing algorithms and abstraction strategies based on the application scenario; for example, adding chapter recognition and clause numbering in legal document scenarios, and adding terminology recognition and formula placeholders in scientific paper scenarios. The terminal can also be a head-mounted display terminal or an in-vehicle terminal, displaying the abstract results in an interface adapted to the screen size and interaction method in these scenarios. The present invention can be implemented with different hardware and software combinations. Its core lies in the server's collaborative control between prompts and generative artificial intelligence models, combined with feedback-based priority updates based on user behavior, thereby achieving structured acceleration and accuracy improvement of text summarization tasks within the computer.
[0180] use Figure 12 The processing flow is explained.
[0181] Step 1: The server receives the request and parses the user parameters.
[0182] Input: A request message sent by the terminal over the network, which includes user identifier, target category, expected text style, comprehension level, target language, and optional keywords.
[0183] The server receives the request using a network interface. The server calls the communication protocol stack and application layer framework to parse the HTTP / HTTPS message and deserialize JSON or other structured data into an internal object.
[0184] Based on the parsing results, the server extracts fields such as user identifier, category, style, comprehension level, language, and keywords from the request object and generates a "request context record" in memory.
[0185] Output: A request context record containing user parameters and request timestamps, for use by subsequent modules.
[0186] Step 2: The server loads user interests and historical information based on the user identifier.
[0187] Input: The request context record generated in step 1, which contains the user identifier.
[0188] The server accesses the user profile database or key-value storage system, using the user identifier as an index to query data such as the user's browsing history, common selection categories, preferred text styles, and average dwell time distribution.
[0189] The server performs statistical calculations on the retrieved behavioral data, such as calculating click-through rates, average dwell time, and recent activity levels for each category, to form user interest vectors or interest distributions.
[0190] Output: A user interest data structure bound to the current request, used to guide information source retrieval and prompt generation.
[0191] Step 3: The server retrieves candidate articles from the information source.
[0192] Input: The category, keywords, and time range settings in the request context record, as well as the user interest data obtained in step 2.
[0193] The server constructs database query conditions or search engine query statements, including category filtering, time range restrictions, keyword matching, and priority sorting parameters.
[0194] The server sends query requests to information sources (such as news databases or indexing systems) through database clients or search engine clients, and receives several article records that meet the criteria.
[0195] The server reorders the returned results based on the user's interest vector, for example, increasing the weight of articles in categories of high user interest and increasing the weight of recent content.
[0196] Output: A list of candidate articles sorted by priority. Each record contains fields such as article ID, title, body text, category, language, and publication time.
[0197] Step 4: The server performs text preprocessing on the candidate articles.
[0198] Input: The list of candidate articles output in step 3.
[0199] The server iterates through the candidate articles one by one, and uses the text processing module to perform the following data processing on the main text: (1) Use regular expressions and HTML parsers to remove useless elements such as tags, advertising slogans, copyright notices, and duplicate headers and footers; (2) Use a natural language processing library to segment the text into sentences, and convert an article into a sentence sequence; (3) Truncate or sample the excessively long text based on the number of characters and sentences, for example, retaining the first few paragraphs or extracting some sentences according to sentence importance; (4) Check if the language field of the article is consistent with the user's target language. If not, call the translation module to translate the text into the target language.
[0200] The above process takes the original text as input and transforms the noisy text into a structured list of sentences with controlled length through string matching, sentence boundary recognition, and optional translation inference calculation.
[0201] Output: A list of preprocessed article data, with each article accompanied by cleaned main text, sentence sequences, and language tags for subsequent summary generation.
[0202] Step 5: The server generates prompts for controlling generative artificial intelligence models.
[0203] Input: A list of preprocessed article data, request context records (category, style, comprehension level, target language), and user interest data.
[0204] The server selects appropriate prompt templates from a stored set of templates based on style, comprehension level, and language, such as formal templates, casual templates, or beginner-friendly templates.
[0205] The server fills the template placeholders with parameters such as the maximum number of summary entries, the character limit per entry, the target audience description, and the category description, and appends the pre-processed body text to the end of the template to form a complete prompt statement.
[0206] The server performs the above string concatenation operation on each article, merging the template text and the article body into an input sequence.
[0207] Output: A list of pairs of prompt statements and corresponding body text, each containing the model input string of "prompt statement + body text".
[0208] Step 6: The server invokes a generative artificial intelligence model to generate a list summary.
[0209] Input: The list of model input strings output in step 5.
[0210] The server feeds each input string into the generative artificial intelligence model via an inference framework or a remote interface.
[0211] Generative AI models internally use embedding layers to map text tags into vectors, use multi-layer self-attention mechanisms to compute context-related representations, and then use decoders to generate output tag sequences step by step.
[0212] When the server is invoked, parameters such as temperature, maximum generation length, and sampling strategy are set for the model to control the diversity and length of the generated results.
[0213] The server receives the complete text output by the model, which contains several summary sentences separated by bullet points or line breaks.
[0214] Output: A list of original abstract texts, with a corresponding generated text containing key points for each article.
[0215] Step 7: The server parses and generates text and constructs a list structure.
[0216] Input: The list of raw summary texts output from step 6.
[0217] The server uses string parsing rules to detect bullet points or line breaks, segments the generated text, and breaks it down into multiple summary points.
[0218] The server performs length statistics and simple cleaning (removing redundant prefix symbols, whitespace characters, etc.) on each key point, and stores the key point text in a list structure.
[0219] The server may optionally perform importance estimation on key points, such as calculating scores based on keyword weights or positional features, for subsequent sorting or filtering of redundant key points.
[0220] Output: Structured summary data, with each article containing an ordered array of summary points, along with the length of each point and an optional score.
[0221] Step 8: The server associates image and video information into a summary and reshapes the data for display.
[0222] Input: Structured summary data and candidate article metadata from step 3.
[0223] The server performs queries from the media resource library based on article titles and keywords in the text. It can use keyword matching or vector similarity retrieval to compare text features with multimedia feature indexes.
[0224] For each article, the server selects one or more associated image or video resources and records their addresses, thumbnail information, resolution, and other parameters in the corresponding summary structure.
[0225] The server combines the title, summary key points array, image information, video information, and timestamp into a complete display record according to the predetermined display format.
[0226] Output: A collection of display data that can be directly used for front-end rendering. Each record contains all the necessary fields for vertical scrolling display on the terminal.
[0227] Step 9: The server sends display data to the terminal, and the terminal renders a vertically scrolling interface.
[0228] Input: The set of display data generated in step 8.
[0229] The server encodes the displayed data into a response message via a network interface and transmits it to the terminal using HTTP / HTTPS.
[0230] After receiving the response, the terminal uses the local parsing library to parse the display data and maps it into the interface data model.
[0231] The terminal calls the user interface component to create a vertical list, which arranges the title, column summary and associated image or video preview of each article in sequence, and sets the scrolling behavior and click event handling logic.
[0232] Output: A vertically scrollable summary list interface rendered on the terminal screen for the user to browse.
[0233] Step 10: Users can perform interactive operations and adjust summary conditions through the terminal.
[0234] Input: The summary list displayed in the terminal interface and the interactive controls provided by the terminal.
[0235] Users can perform actions such as clicking, swiping, selecting categories, adjusting text styles, selecting comprehension levels, switching languages, or entering keywords on the terminal, or input natural language commands via voice input.
[0236] The terminal converts these operations into a new set of request parameters and packages them together with the user identifier into a request message.
[0237] Output: The updated request message, containing new summary control conditions and filtering conditions, for the next round of server processing.
[0238] Step 11: The server dynamically updates the conditions for generating prompts and the information retrieval strategy based on user interaction.
[0239] Input: The update request message generated in step 10 and the historical user interest data recorded on the server.
[0240] The server parses the new request parameters, determines whether the user has changed the category, style, comprehension level, or language, and updates the request context record.
[0241] The server adjusts the prompt template selection logic based on new preferences, such as switching from a formal template to a relaxed template, or from a general audience template to a beginner template.
[0242] The server also incrementally updates the user's interest vector, for example, by increasing the weight of a category based on the fact that it is frequently selected.
[0243] Based on the updated interest weights and conditions, the server applies new parameters in subsequent information source retrieval and prompt statement construction processes.
[0244] Output: The updated prompt generation strategy and information retrieval priority settings provide a new calculation basis for the next execution of steps 3 to 7.
[0245] Step 12: The server performs translation processing and supports multilingual display.
[0246] Input: Structured summary data, original text, and the user's target language settings.
[0247] The server checks whether the summary language matches the target language. If they do not match, the server inputs the summary key points or original sentences into the translation model in sequence.
[0248] The translation model is based on an encoder-decoder architecture. It encodes the source language sentence into vectors and then decodes it in the target language space to generate the translation. The server stores the generated target language text and the corresponding source text.
[0249] The server encapsulates the multilingual summaries along with the original text into a display data structure, adding language tags so that the terminal can directly select the corresponding text to display when the user switches languages.
[0250] Output: Includes extended display data of multilingual summaries and the original text, allowing the terminal to switch between multiple languages without requesting the server again.
[0251] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0252] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0253] In existing information retrieval and summarization systems, servers typically extract text from information sources based on simple user-inputted search queries and then directly perform summarization according to fixed rules. This presents the following technical problems: First, the server's preprocessing capabilities for raw text are limited, lacking a unified mechanism for structuring and noise removal of multi-source heterogeneous data. This results in the input text to generative AI models containing a large amount of redundant information, increasing the computational burden on the models and reducing summarization quality. Second, when calling generative AI models, servers mostly only input raw content and simple prompts, lacking fine-grained control prompts that can dynamically control the summary format, length, and target domain according to user needs, leading to inconsistent generated results. First, it is difficult to adapt to the different levels of understanding and focus of different users. Second, existing technologies generally treat the summary generation process as a one-time process and lack a mechanism for iterative updates based on multiple rounds of user dialogue commands. The server cannot use dialogue history and summary adjustment history to adaptively optimize the subsequent generation process. Third, in cross-language scenarios, summary results are often processed through independent translation steps. The server side does not tightly integrate machine translation with summary generation, resulting in semantic distortion and redundant consumption of computing resources. Fourth, the display form of summary results on the terminal side is mostly a simple list or pagination, lacking a structured layout design for vertical scrolling reading interfaces, which is not conducive to efficiently presenting information in limited display environments such as mobile terminals.
[0254] Therefore, it is necessary to propose a technical solution that performs structured preprocessing of multi-source text on the server side, combines generative artificial intelligence models to programmably control prompts, supports conversational summary adjustment and integrates language conversion functions, and optimizes the output structure for vertical scrolling display on the terminal. This solution aims to improve the efficiency of computing resource utilization and user information understanding from the perspectives of system architecture and algorithm flow, thereby achieving a substantial improvement in computer technology.
[0255] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0256] In this invention, the server includes: a processing unit for acquiring information from an information source and converting the information into text information, and performing preprocessing including structuring and noise removal; a generation unit for generating control prompts based on the preprocessed text information and prompts received from the user, wherein the control prompts instruct a generative artificial intelligence model on the form, length, and target domain of the summary content, and further instruct the selection of relevant information units from multiple information sources and the generation of hierarchical list summaries; a reasoning unit for inputting the control prompts and the preprocessed text information into the generative artificial intelligence model and performing reasoning to generate summary information including key point extraction results; and a server for attaching identification information and reference location information corresponding to the information source to the summary information, and shaping the summary information to be scrolled vertically on the user terminal. The system includes: a shaping unit for displaying hierarchical and paragraph structures dynamically; a sending control unit for sending the shaped summary information to the user terminal via a communication interface and making the user terminal present the summary information in visual form; a dialogue control unit for acquiring additional prompts input from the user terminal in a dialogue format, generating correction prompts based on the additional prompts to change the summary's level of detail, target domain, and expression, and then calling the generative artificial intelligence model again to update the summary information; a translation unit for performing machine translation on the text information used to generate or update the summary information and the summary information itself to output translations in different languages specified by the user; and a learning control unit for recording dialogue history and summary adjustment history divided by user, and adaptively optimizing them when generating the control prompts and correction prompts later. This allows for efficient summarization and cross-language presentation of multi-source heterogeneous texts by finely controlling the input context and generation constraints of generative artificial intelligence models on the server side. Combined with an adaptive adjustment mechanism based on dialogue history and a structured output design oriented towards vertical scrolling display, it significantly reduces invalid computation, improves the readability and relevance of the summary results, and thus improves the overall technical performance of the computer system in the information acquisition, processing and display process.
[0257] "Information source" refers to various data sources used to provide information to be obtained, including but not limited to websites, databases, document storage, file systems, and other storage and service systems that can provide text, images, or multimedia content.
[0258] “Textual information” refers to the content extracted or converted by the server from the original information and represented in the form of a character sequence, including natural language text, tokenized text, and encoded symbol sequences.
[0259] "Preprocessing" refers to the collective process performed on text information before it is input into a generative artificial intelligence model, including format unification, structuring, noise removal, speech recognition, segmentation, and length truncation.
[0260] "Structured" refers to analyzing raw text information and dividing it into multiple information units with hierarchical relationships, such as titles, paragraphs, sections, and entries, according to predetermined rules, for subsequent processing and display.
[0261] "Noise removal" refers to removing content from text information that is irrelevant to user needs or summary generation, including advertisements, script code, navigation text, repeated paragraphs, and formatting control characters.
[0262] "Prompt statements" refer to natural language or semi-structured instruction text input by users or servers into generative artificial intelligence models to indicate task objectives, output requirements, or key points of focus.
[0263] "Control prompts" refer to instruction text automatically generated by the server based on preprocessed text information and prompts entered by the user. These texts are used to explicitly specify control parameters such as the form, length, target domain, structure, and display method of the summary to the generative artificial intelligence model.
[0264] "Correction prompts" refer to instruction text generated by the server based on the user's needs to add or modify during the conversation. These instructions are used to adjust the level of detail, target area, expression style, or language form of the summary based on existing summary information.
[0265] "Generative artificial intelligence models" refer to models built on machine learning and deep learning algorithms that can automatically generate natural language text based on input prompts and text information, including but not limited to generative models that use multi-layer neural networks and attention mechanisms.
[0266] "Summary information" refers to the core content representation generated by a generative artificial intelligence model based on control prompts and text information, which is a compressed and reorganized representation of the original information. It may include a list of key points, segmented descriptions, or hierarchical text.
[0267] "Key point extraction" refers to the process of identifying and extracting the facts, conclusions, events, or viewpoints that are most important to the user's goals from the original information during the summary generation process, and using them as the main content units of the summary information.
[0268] An "information unit" refers to the smallest content fragment with relatively independent semantics obtained after structured processing, including sentences, paragraphs, sections, or entries, which are used as the basic elements for summary generation and display.
[0269] "Hierarchical list format" refers to organizing multiple information units into a multi-level list structure with parent-child relationships according to the importance or logical relationship of the content, and presenting them in the form of entries and sub-entries.
[0270] "Vertical scrolling display" refers to a display style on the user terminal's display interface that continuously displays content by scrolling up and down, allowing users to browse summary information vertically through swipe gestures.
[0271] "Identification information" refers to a mark used to uniquely identify an information source or information unit, including but not limited to identifiers, titles, source names, timestamps, or index numbers.
[0272] "Reference location information" refers to location information used to indicate the correspondence between summary information and original information, including but not limited to link addresses, document position offsets, paragraph numbers, or anchor marks.
[0273] "Shaping" refers to the process of reorganizing and formatting summary information according to predetermined display and interaction requirements, including its structure, paragraph division, list format, and tagging style.
[0274] "User terminal" refers to an electronic device used to receive and display summary information sent by a server and to interact with the server, including but not limited to mobile terminals, tablet devices, desktop devices, or other computing devices with network communication and graphical interfaces.
[0275] "Dialogue format" refers to the exchange of information between users and servers through multi-turn natural language interaction. Users can input new prompts or adjustment instructions based on the displayed results, and the server dynamically updates the output according to the input.
[0276] Machine translation refers to the technical process of converting one natural language text into another natural language text using automatic language conversion processing performed by computing devices.
[0277] A “translation unit” refers to a functional module in the system that implements machine translation, used to perform cross-language conversion processing of textual or summary information.
[0278] The “learning control unit” refers to a functional module or program component used to record and analyze user dialogue history and summary adjustment history, and to adaptively optimize the generation strategy of subsequent control prompts and correction prompts based on this historical information.
[0279] "User attribute information" refers to a set of parameters related to the user and used for personalized summary generation and display, including but not limited to the user's level of understanding, field of expertise, language used, display medium type, and interaction preferences.
[0280] "Explanation templates" refer to pre-designed text patterns with fixed structures and styles, used to organize summary content at a specified level of abstraction and style. These include hierarchical explanation templates, bullet point explanation templates, and question-and-answer explanation templates.
[0281] The embodiments of this invention will combine the collaborative work of the server, terminal and user, and provide a detailed description of the system structure, data structure, algorithm flow and internal composition of the generative artificial intelligence model, so that those skilled in the art can implement this invention and understand the improvements of this invention at the computer technology level.
[0282] I. Overall System Composition The server is the core processing device of this invention. The server includes: at least one computing device with a multi-core processor and large-capacity memory, at least one graphics processing device for neural network inference computation, and several storage devices. The server can use general-purpose rack-mount hardware, such as computing nodes equipped with multi-core processors (e.g., general-purpose multi-core processors), 64GB or more of memory, and solid-state storage. The server can run an operating system, such as a Linux kernel-based server operating system. The server can connect to a wide area network (WAN) or local area network (LAN) via a network interface.
[0283] The server's software architecture includes: a network communication module (e.g., using Nginx or other web server components), an application service module (e.g., a backend framework running in a Python, Java, or JavaScript runtime environment, such as Python-based FastAPI, Java-based Spring, or JavaScript-based Express framework), a text preprocessing module, a generative artificial intelligence model inference module, a machine translation module, a learning control module, and a data storage module. The server can also integrate logging, monitoring, and caching modules.
[0284] A terminal is a display and input device used for user interaction. A terminal can be a mobile terminal, tablet terminal, or desktop terminal. A terminal includes a processor, memory, display device, touch input device, and network communication module. The terminal can run a general-purpose mobile operating system or a desktop operating system. The terminal runs client applications on its software, which can be native applications or browser-based web applications. Terminal applications use UI frameworks (such as mobile UI component libraries) to build vertically scrolling display interfaces.
[0285] Users are the users of the system. They input prompts through the terminal application, view the summary information returned by the server, and adjust the summary results through a dialog box.
[0286] II. Server-side program modules and data structures The server uses a data storage module to manage various data structures. The server can store the following data in a relational data management system or document database: The server encapsulates each user request into a request record. A request record includes: a request identifier, a user identifier, the original prompt, the language, a timestamp, and a list of information source entry identifiers associated with the request. The server stores these information source entries as document records. A document record includes at least: a document identifier, the information source type, the original URL or path, the original content, the cleaned content, the language, a timestamp, and a topic tag vector.
[0287] The server records the dialogue history as a dialogue record. The dialogue record includes: dialogue identifier, user identifier, round number, original prompt statement, server-generated control prompt statements, correction prompt statements, and the corresponding generated summary information version identifier. The server stores the summary information in the form of summary records. Summary records include: summary identifier, corresponding request identifier, summary text content, structured representation (e.g., paragraph list, item tree structure), target language, summary granularity level, and the model configuration parameters used.
[0288] Internally, the server can represent the hierarchical structure of the summary as a tree data structure. The server can assign each summary entry an entry node identifier, parent node identifier, entry text, and a list of associated information unit identifiers. The server uses this tree structure to map the vertical scrolling display order and collapsible / expanded state on the terminal.
[0289] III. Structure and Training of Generative Artificial Intelligence Models The server deploys at least one transformer-based neural network model in the generative artificial intelligence model inference module. The server can implement this neural network using open-source deep learning frameworks, such as those based on PyTorch or TensorFlow. The server stores the model parameters in high-speed storage during deployment and loads them into the graphics processing unit's video memory at runtime.
[0290] The generative AI model used by the server can employ a multi-layered encoder-decoder structure, with each layer including a multi-head self-attention sub-layer and a feedforward neural network sub-layer. The server can be configured to have a certain number of layers, attention heads, and hidden layer dimensions. At the model input, the server uses a tokenizer to encode text information and prompts into discrete token sequences, such as using sub-word unit encoding. The server then maps these token sequences into vector sequences, which serve as input to the neural network.
[0291] During the model training phase, the server pre-trains the model using a large-scale corpus and fine-tunes it using a dataset relevant to the summarizing task. The server can employ a teacher-forced approach, using a reference summary as the target output during training. The server defines a loss function, such as cross-entropy loss, comparing the predicted label distribution with the target labels. The server computes gradients using backpropagation and updates the model weights using optimization algorithms (such as adaptive moment estimation). The server can employ data augmentation strategies, such as sentence rearrangement, synonym substitution, and noise injection, to improve the model's robustness to multi-source heterogeneous text.
[0292] During the inference phase, the server uses bundle search or sampling strategies to generate summary text. The server can limit parameters such as output length, control diversity, and deduplication penalties. These parameters are explicitly stated in control prompts, such as requiring the output to consist of a certain number of bullet points, each not exceeding a specified word count. In this way, the server translates user requirements into specific generation constraints, thus forming not simple rule matching, but programmable control over the model's internal generation process.
[0293] IV. Text Preprocessing and Information Unit Construction The server implements the text preprocessing process in the application service module. The server uses text parsing libraries (such as HTML parsing components) to extract tags and scripts from web pages or documents. The server executes language detection algorithms (such as language recognition models based on statistical features or lightweight neural networks) to determine the language tags for each segment of content.
[0294] The server segments the text into sentences and paragraphs. It can use sentence boundary recognition functionality from a natural language processing library to identify sentence terminators. The server divides the text into paragraphs based on headings, blank lines, and punctuation patterns. The server assigns an information unit identifier to each sentence and paragraph. The server calculates an importance score for each information unit based on keyword density, sentence position, and entity distribution, which is used for subsequent information unit selection.
[0295] During the structuring process, the server uses document titles as root nodes, first-level subheadings as child nodes, and paragraphs or sentences as leaf nodes to construct a tree structure. The server can calculate the similarity of information units between different documents, for example, by encoding sentences into vectors and then calculating cosine similarity. The server clusters information units based on similarity and timestamps, merging similar information units into the same topic cluster. The server incorporates information from these topic clusters into the generation context within the control prompts, thereby reducing redundant input and lowering the computational burden on the model.
[0296] V. Generation of Control and Correction Prompt Statements The server generates control prompts based on the user's original prompt and preprocessed text information. Users can enter the following example prompts in the terminal: "Please summarize the latest AI technology development trends in simple terms." "Help me extract the key tech news from the past week." "I am a beginner. Please explain the basic principles and applications of generative artificial intelligence in layman's terms." "Please summarize the key points of recent English academic articles on 'Generative AI in Healthcare' in Chinese." The server parses keywords, time range, domain description, and comprehension level description from user prompts. It can use word segmentation and named entity recognition technologies to transform concepts such as "latest AI technology," "this week," "tech news," "beginner," and "medical field" into internal parameters. Based on these parameters, the server explicitly specifies in the control prompts: target domain (e.g., artificial intelligence, medical technology), summary granularity (e.g., high-level overview or detailed explanation), target reader's comprehension level (e.g., beginner or professional reader), and output format (e.g., list of items or segmented description).
[0297] The server uses natural language in its control prompts to constrain the generative AI model. For example, it might say, "Based on the following information, please summarize 3 to 5 of the most important developments related to AI technology in the past week in simplified Chinese, in bullet points, each no more than two sentences, suitable for non-specialist readers." The server inputs this control prompt along with the selected information unit into the model. In this way, the server maps abstract requirements to concrete generation rules that can be parsed by the model, thereby improving the inefficient traditional method of "simple keywords + full text input."
[0298] During the dialogue, the server generates correction prompts based on new user needs. Users can input commands such as, "Please make it shorter, under 200 words," "Keep only the medical-related content," and "Rewrite it in more accessible language, suitable for those completely unfamiliar with AI." The server then generates correction prompts based on the existing summary and the new commands, such as, "While keeping the main conclusions unchanged, please compress the above summary to approximately 200 words, removing technical details and retaining the core conclusions and application examples." The server inputs these correction prompts along with the existing summary as context into the model, allowing the model to perform secondary editing based on the original generated result, rather than completely regenerating it, thereby reducing computational load and improving response speed.
[0299] VI. Machine Translation and Cross-Language Processing The server can deploy a neural machine translation model within its machine translation module. The server can use a sequence-to-sequence network based on a transformer architecture. The server trains this model on texts in different languages, enabling it to map source language sentences to target language sentences. During training, the server aligns source and target language sentences using parallel corpora and optimizes the results using a cross-entropy loss function.
[0300] The server determines the user's desired output language during runtime. If the original or summary information does not match the user's language, the server invokes the machine translation module. The server can choose one of two strategies: translate the original text before summarizing, or translate the summary results uniformly after summarizing. The server can choose based on model performance and target language quality. The server aligns the translation results with the corresponding positions in the original text, records the mapping relationship, and allows the user to jump to the original text location via the terminal when needed.
[0301] By integrating machine translation with summarization, the server reduces the number of times users need to copy the translation, avoiding network latency and redundant computation caused by multiple independent calls to external translation services. Simultaneously, by sharing text encoding vectors internally, the server reduces repetitive word segmentation and encoding calculations, improving overall processing efficiency.
[0302] VII. Learning Control and Adaptive Optimization The server collects and analyzes user interaction data within the learning control module. It maintains user attribute information for each user, including their self-reported level of expertise, commonly used languages, preferred terminal types, and past preferences for summary granularity. The server records user adjustments to summaries in the dialogue history, such as frequent requests for "more concise" or "more detailed." This historical data is then used to build a user preference model.
[0303] The server can use simple statistical analysis or lightweight machine learning models (such as gradient boosting trees or small neural networks) to predict which summary length and expression style users will prefer in the future. When generating control prompts or correction prompts, the server embeds the predicted preferences as default parameters, making the model closer to user expectations from the initial generation, thereby reducing the number of subsequent interaction rounds and lowering the computational and communication load.
[0304] Through this learning control mechanism, the server enables the system to gradually develop adaptive behavior for each user, thus improving the shortcomings of traditional fixed-template summary systems that cannot be optimized according to changes in user habits.
[0305] VIII. Terminal-side display and interactive interface The terminal runs the client application and renders the summary information sent by the server. The terminal parses the structured summary data provided by the server into an internal view model, including a list of paragraphs and a hierarchical item structure. The terminal uses a vertical scrolling view component to create a scrollable interface. The terminal sets indentation or font styles based on the hierarchy of nodes to help users intuitively understand the information hierarchy.
[0306] The terminal detects the user's vertical swipe gestures via the touchscreen and converts them into scroll events. When processing scroll events, the terminal only renders a few summary nodes in the currently visible area, avoiding drawing all nodes at once, thus reducing the graphics processing burden and saving power. Since the server has already optimized the paragraph length and structure when shaping the summary, the terminal can maintain a relatively stable line height and paragraph spacing during scrolling, thereby avoiding frequent layout recalculations and improving rendering efficiency.
[0307] The terminal retains an input box at the bottom of the interface for users to enter new prompts to initiate the dialogue. Users can request adjustments to the current summary at any time without having to re-enter the original query. The terminal includes a current summary identifier with each new prompt, enabling the server to update based on the correct summary version.
[0308] IX. Technical Effects and Causal Relationships By introducing structured preprocessing and information unit selection, the server compresses the original information source content into a smaller, highly relevant text set, significantly reducing the number of labels input to the generative AI model, thereby lowering the computational load and inference latency of the graphics processing device. Furthermore, by embedding explicit constraints on the output format, length, structure, and target domain in control and correction prompts, the server changes the traditional black-box generation method, making the generation process controllable and repeatable. This "programmable generation" represents an improvement to text generation algorithms, rather than simply automating the task of manually editing summaries.
[0309] The server incorporates user feedback into the generation process through multi-round conversational correction and models historical dialogue data using a learning control module. This ensures that subsequent generation processes are closer to the final answer from the initial stage, reducing redundant generation and repeated requests. Consequently, the amount of network communication data and the number of inference calls on the server side are reduced at the system level. The server shares encoding results between summary generation and machine translation through a unified internal representation, avoiding repeated word segmentation and encoding, and improving overall processing throughput.
[0310] Furthermore, the server employs a specially designed summary structure for vertical scrolling on the terminal side, enabling the terminal to efficiently present key content within limited screen space and reducing frequent pagination and search operations for users. Since the summary content is already organized in a hierarchical column format, the terminal only needs to perform simple list rendering and scrolling management, eliminating the need for complex layout calculations and thus reducing the computational burden on the terminal side.
[0311] Through the aforementioned hardware and software co-design, this invention does not simply automate the human reading and translation process, but rather achieves comprehensive improvements in algorithms and system architecture across multiple technical levels, including generative artificial intelligence model input control, text preprocessing, model inference optimization, cross-language processing, and terminal display structure. This results in an overall improvement in technical performance in terms of processing speed, summary quality, cross-language consistency, resource utilization efficiency, and terminal rendering performance.
[0312] use Figure 13 The processing flow is explained.
[0313] Step 1: Users input prompts via the terminal and send them to the server.
[0314] Users launch the client application on their terminal and type natural language prompts in the text input box, such as: "Please summarize the latest AI technology development trends in simple language." or "I am a beginner, please introduce the basic principles and applications of generative artificial intelligence in layman's terms."
[0315] The terminal takes the user's input prompt as input, reads the character sequence from the interface controls, and adds metadata such as language, user identifier, and timestamp to form a request object. The terminal calls a serialization library to convert the request object into a text format (such as a JSON string), and then sends it to the server's specified interface address as an HTTPS request via the network communication module.
[0316] In this step, the input consists of a prompt entered by the user and the terminal's current environment information, while the output is a structured request message transmitted to the server over the network. The terminal performs data processing operations such as encoding, packaging, and network transmission on the input character sequence to ensure that the server can correctly parse and process the request.
[0317] Step 2: The server receives and parses requests from the terminal.
[0318] After receiving an HTTPS request at the network layer, the server forwards the request to the application service module. The server reads text data from the request message body, calls a parsing library to deserialize the text data into an internal request object, and extracts prompts, language information, user identifiers, and other additional parameters. The server records metrics such as request arrival time and request length in the logging module for subsequent performance analysis.
[0319] In this step, the input is a structured request message sent by the terminal, and the output is a request object residing in the server's memory and the parsed prompt text. The server performs data operations such as decryption, decoding, and parsing on the input message, converting the network layer representation into an internal data structure that can be processed by subsequent modules.
[0320] Step 3: The server determines the retrieval strategy based on the prompts and retrieves the original information from the information source.
[0321] The server reads the parsed prompts, performs word segmentation and keyword extraction, and identifies the time range, domain terms, and user goals. For example, from "highlights of recent technology news," the time constraint is "one week," the domain is "technology," and the task type is "summary." The server then constructs search query expressions based on these analysis results.
[0322] The server calls a search interface or database query interface, takes the constructed query expression as input, sends a retrieval request to an external information source or internal storage system, and retrieves several raw document data related to the topic of the prompt statement. The server performs preliminary filtering on the search results, such as selecting the most relevant documents based on timestamps, source credibility, or keyword matching.
[0323] In this step, the input consists of prompts from the server and the keywords they parse, and the output is a set of raw documents containing titles, body text, source addresses, and time information. The server uses keyword matching, filtering, and sorting to perform data calculations, selecting the raw data most relevant to the user's intent from a massive amount of information, thus narrowing down the data range for subsequent summarization processing.
[0324] Step 4: The server preprocesses and structures the raw documents it retrieves.
[0325] The server takes the content of each original document as input, uses a text parsing library to remove noise such as HTML tags, scripts, and advertisements, and retains the main text. The server executes a language detection algorithm to label the language type of each piece of content. If the document language differs from the user-specified language, the server can call a translation module at this stage to convert the original text to a unified language, or simply mark it for later processing reference.
[0326] The server segments the cleaned text into sentences and paragraphs, breaking down long texts into lists of sentences and paragraphs. The server assigns an information unit identifier to each sentence and paragraph and calculates an importance score based on the sentence's position in the document, the number of keywords it contains, and entity density. The server organizes the text into a tree structure according to the hierarchical structure of headings and subheadings, attaching paragraphs and sentences to their corresponding nodes.
[0327] In this step, the input is a collection of unstructured raw documents, and the output is a collection of structured information units with language tags, hierarchical structure, and importance scores. The server processes the data through format cleaning, sentence segmentation, hierarchical partitioning, and importance calculation, transforming the loose text into a computable tree-like data structure, thereby improving the effectiveness and controllability of subsequent model inputs.
[0328] Step 5: The server selects the information unit and generates control prompts.
[0329] The server takes a set of structured information units and the user's original prompts as input. It first filters representative information units from multiple documents based on importance scores and similarity calculations. The server can perform similarity calculations on sentence vectors, merging or discarding sentences with highly repetitive content, thereby reducing the input size and redundancy.
[0330] The server parses the task type and output constraints in the prompt, such as "3 to 5 key points," "each no more than two sentences," and "suitable for non-specialist readers." Based on these constraints, the server constructs a control prompt, describing the generation requirements in natural language, for example: "Based on the following information, please summarize 3 to 5 of the most important developments related to artificial intelligence technology in the past week in simplified Chinese, in item form, with each item no more than two sentences, suitable for non-specialist readers." The server adds parameters such as domain, summary granularity, language, and display structure to the control prompt.
[0331] In this step, the input consists of a set of structured information units and user prompt text, and the output consists of one or more control prompts and a subset of selected information units. The server uses data operations such as similarity calculation, importance ranking, and rule matching to compress a large number of information units into small, highly relevant ones, and translates abstract user needs into explicit instructions that can be executed by the generative artificial intelligence model.
[0332] Step 6: The server will control the input of prompts and information units into the generative artificial intelligence model and generate summary information.
[0333] The server concatenates the control prompts with the selected information unit text to form the model input sequence, and calls a tokenizer to convert the text into a tokenized sequence. The server maps the tokenized sequence into a vector representation and sends it to the generative artificial intelligence model deployed on a graphics processing device.
[0334] The server performs forward propagation of a multi-layer transformer network within the model. The attention layer calculates relevance weights for the labels in the input sequence using a self-attention mechanism, while the feedforward layer performs non-linear transformations on the intermediate representations. During the decoding phase, the server selects the next label based on the probability distribution of the generated labels, progressively generating summary text using a beam search or sampling strategy. During generation, the server prunes candidate sequences that do not meet the constraints specified in the prompts, such as maximum length, number of entries, and tone style.
[0335] In this step, the input consists of control prompts and selected text information units, and the output is an initial summary text in natural language form. The server performs numerical operations on the input data, such as word vector mapping, matrix multiplication, attention weight calculation, and probability sampling, and uses the reasoning ability of a generative artificial intelligence model in a high-dimensional vector space to generate a summary result that meets the constraints.
[0336] Step 7: The server restructures the generated summary information and adds associated metadata.
[0337] The server takes the initial summary text as input and, combined with the previously defined information unit tree structure, performs secondary parsing and segmentation of the summary text. The server identifies entry and paragraph boundaries in the summary, dividing it into several entry nodes and paragraph nodes. For each entry node, the server labels its hierarchical depth, display order, and corresponding list of information unit identifiers.
[0338] The server adds source identification information and reference location information to the summary as a whole and each entry, such as indicating the main source document identifier, original paragraph number, and URL location for each entry. Based on the vertical scrolling limitations of the terminal, the server checks the length of each paragraph or entry; if it exceeds the limit, it performs sentence-level splitting or folding configuration. The server ultimately generates a structured summary object for terminal rendering.
[0339] In this step, the input is the summary text and information unit tree structure output by the generative artificial intelligence model, and the output is structured summary data containing an entry tree, paragraph list, and source mapping table. The server performs operations such as paragraph segmentation, node attribute assignment, and metadata appending to ensure that the summary corresponds semantically to the original text and is formatted to fit the front-end display components.
[0340] Step 8: The server performs machine translation based on the user's language requirements (if necessary).
[0341] The server reads the target language specified in the user request and the language identifier of the current summary, taking the structured summary data as input. If the two do not match, the server invokes the machine translation module on the summary text field. The server feeds the text of each paragraph or entry into the neural machine translation model line by line, and the model maps the source language representation to the target language representation through an encoder-decoder structure.
[0342] The server maintains the original summary structure after translation, replacing only the text content and updating the language identification information. The server can also detect translation quality based on model confidence levels, and can mark sentences with insufficient confidence for subsequent manual review or retranslation.
[0343] In this step, the input is structured summary data and target language settings, and the output is structured summary data converted to the target language. The server uses sequence-to-sequence mapping and probability maximization to convert the summary content into another natural language while preserving the structural and semantic correspondence.
[0344] Step 9: The server sends the final summary data to the terminal for display.
[0345] The server takes the structured digest object as input and calls the serialization module to convert it into a text representation suitable for network transmission. The server returns the result data to the terminal via the network communication module in the form of an HTTPS response, and also returns the necessary status code and timestamp so that the terminal can determine when to display it.
[0346] In this step, the input is a summary data object containing text content, hierarchical structure, and metadata, and the output is a response message sent to the terminal over the network. The server performs processing such as serialization and encrypted transmission to ensure that the summary result is delivered to the user in a reliable and low-error format.
[0347] Step 10: The terminal receives and displays summary information in a vertical scrolling manner, while also allowing the user to continue entering prompts to engage in dialogue.
[0348] After receiving the server's response at the network layer, the terminal passes the response text as input to the parsing module, which parses it to obtain structured summary data. Based on the hierarchical relationships and item order, the terminal constructs a tree of interface components, creating a vertically scrollable list view. The terminal sets a display area for each item, draws the text content on the screen, and sets indentation or style variations based on node depth.
[0349] The terminal presents the rendered interface to the user, who can browse various items and paragraphs by swiping up and down on the touchscreen. The terminal monitors the user's input during the reading process. When the user enters a new prompt in the bottom input box, such as "Please make it shorter, keep it under 200 words" or "Only keep content related to medical care," the terminal again executes the sending actions described in steps 1 and 2, sending the new prompt and the current summary identifier to the server.
[0350] In this step, the input is the structured summary data returned by the server, and the output is a visual summary interface that scrolls vertically on the screen, along with requests for new user prompts. The terminal processes the data through data parsing, interface construction, and event response, transforming the abstract structured data into concrete display frames and user interaction events, thus completing the user-side closed loop of the entire information processing cycle.
[0351] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0352] In existing computer information processing technologies, although servers can obtain text data from network information sources and use general natural language processing algorithms or generative artificial intelligence models to summarize or translate the text, the following technical problems still exist: First, existing systems typically generate summaries based on fixed rules or simple templates. The server side lacks the ability to dynamically adjust the summarization strategy according to different users' comprehension levels, reading preferences, and language abilities. This makes it difficult to achieve a balance between readability and information content in the generated results. Users on the terminal side need to repeatedly manually filter and understand a large amount of content, resulting in inefficient use of computing resources and network bandwidth.
[0353] Second, existing generative AI models are mostly static calls of "single request - single output". When generating prompts, the server fails to make full use of the user's constantly changing needs and contextual information during the interaction. It is difficult to achieve dialogic and iterative optimization for the summary style, level of detail and information classification. As a result, the same original information needs to be processed independently multiple times, which increases the number of model inferences and the overall processing latency.
[0354] Third, the existing system lacks a unified control mechanism at the prompt level for generative artificial intelligence models in the summarization and translation process. The server cannot simultaneously fine-tune key parameters such as "summary format (e.g., bullet-style output)," "number of items," "output language," and "tone and style" in the integrated process. As a result, the uncertainty of the model output is relatively large, the post-processing burden is increased, and the stability and scalability of the system are affected.
[0355] Fourth, although some systems have used sentiment analysis technology to identify user emotions, these sentiment signals are usually only used to recommend content types and are not deeply integrated into the prompt generation logic on the server side. Therefore, it is impossible to dynamically control tone, difficulty, and information classification selection conditions at the input stage of generative artificial intelligence models. As a result, sentiment adaptation only stays at the recommendation level and does not achieve end-to-end adaptation of user emotional state in the underlying text generation process.
[0356] Fifth, while existing information display terminals can provide user interfaces such as vertical scrolling, the server side still lacks a unified design for "how to organize column summaries, visual information, and translation results in a structured manner and support subsequent conversational reprocessing." This results in a loose data structure between the terminal and the server, making it difficult to support complex scenarios such as low latency, multi-turn interaction, and multi-language switching, thereby limiting the overall system performance and user experience.
[0357] In summary, how can we improve the server-side architecture and processing flow? (1) The way generative artificial intelligence models are invoked is achieved through programmable prompts to achieve unified control over the summary / translation / style; (2) Integrate various contextual information such as user comprehension level, interactive input, and emotional state into the process of prompt statement generation and data selection; (3) Using column-structured data as the core output format, improve the usability and scalability of summary results in multi-round interaction and multi-terminal display scenarios; The present invention aims to optimize the overall process of information acquisition, processing, and presentation at the computer technology level.
[0358] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0359] In this invention, the server includes: a data acquisition unit for acquiring electronic information from an information source; a text preprocessing unit for converting the electronic information into text information and performing preprocessing to generate target text information for summarization; a prompting unit for generating prompting statements based on the target text information, user attribute information, and specified information related to the user's comprehension level, for instructing a generative artificial intelligence model to perform summarization processing and / or translation processing, and inputting the prompting statements to the generative artificial intelligence model; a summary shaping unit for parsing the summary results output by the generative artificial intelligence model, organizing the summary points into structured data in the form of bullet points, and associating them with visual information and dynamic image-related identification information to generate display data for display; a sending unit for sending the display data to a user information processing terminal through a communication channel; and a unit for... The information processing terminal receives user conversational input, regenerates prompt statements for changing the summary style, summary detail, and information classification, and re-inputs the prompt statements to the generative artificial intelligence model to obtain updated summary results; an interactive adjustment unit generates prompt statements containing translation processing instructions based on user specified input regarding different language displays and the updated summary results, inputs the prompt statements to the generative artificial intelligence model and / or translation processing device to obtain translated summary results, and sends the translated summary results to the information processing terminal; and an emotion adaptation unit dynamically changes the tone, difficulty, and information classification selection conditions in the prompt statements based on user emotional state information obtained from the information processing terminal or an external computing device to generate summary results corresponding to the emotional state. This allows for fine-grained orchestration of the summarization and translation processes of generative AI models on the server side, using prompts as the control core. It integrates user comprehension levels, interaction commands, and emotional states into the model's input construction and information selection strategies, resulting in a unified output of structured data with multimodal identifiers. This reduces redundant reasoning and post-processing overhead, lowers network transmission and terminal rendering burdens, and improves the responsiveness, controllability, and user experience of the entire information processing chain, achieving a substantial improvement in computer information processing technology.
[0360] A "system" refers to an integrated information processing unit consisting of multiple electronic devices interconnected through a communication network, used to perform information acquisition, processing, generation, and provision.
[0361] "Information source" refers to any data provider or storage medium that provides electronic information, including websites, application programming interfaces, databases, and local or remote storage devices.
[0362] "Electronic information" refers to various types of information data stored or transmitted in digital form, including text data, image data, audio data, video data, and combinations thereof.
[0363] "Data acquisition unit" refers to a functional module or program component in the system that is used to receive electronic information from the information source through a communication interface and perform preliminary storage or forwarding.
[0364] “Textual information” refers to text data extracted from electronic information that can be processed by natural language processing algorithms, including character sequences, word sequences, sentences, and paragraphs.
[0365] "Preprocessing" refers to the cleaning and standardization of textual information, including removing irrelevant symbols, segmenting sentences, segmenting words, deleting noisy content, and standardizing formats, in order to generate data suitable as input for subsequent summaries.
[0366] "Target text information" refers to the set of text information that has been preprocessed and is used as input to generative artificial intelligence models.
[0367] "User attribute information" refers to static or semi-static information related to users, including age group, language ability, professional background, and areas of interest, which are used to determine the style and difficulty of the summary or translation output.
[0368] "Comprehension level" refers to the degree to which a user can understand information content in terms of knowledge depth and language complexity, corresponding to different reading levels such as beginner, intermediate, and advanced.
[0369] "Specified information" refers to parameters that are pre-set by the user or system and are related to the user's level of understanding, output style, or content preferences, used to guide the generation strategy of summaries and translations.
[0370] "Generative artificial intelligence models" refer to algorithmic models built on machine learning and deep neural networks that can automatically generate output text from input text, including language models, dialogue models, and text generation models that can perform summarization and translation tasks.
[0371] "Prompt statements" refer to the control text inputted to a generative artificial intelligence model to explicitly specify the processing goal, output format, language style, and constraints, in order to guide the model to generate results in the expected manner.
[0372] The "prompt statement generation unit" refers to a functional module or program component that automatically constructs and adjusts prompt statements based on target text information, user attribute information, and specified information in the system, and provides them to the generative artificial intelligence model.
[0373] "Summary results" refers to the simplified text content generated by the generative artificial intelligence model based on prompts and target text information, which is a compressed and reorganized version of the original electronic information.
[0374] "Block format" refers to an output format that lists items one by one, usually using bullet points or numbers to identify each point, and is used to clearly present multiple independent information points.
[0375] "Structured data" refers to a data representation with clearly defined fields and hierarchical relationships, allowing information such as abstracts, titles, and media identifiers to be accessed and manipulated through field names.
[0376] "Visual information" refers to non-textual information presented through images, graphics, icons, or other visual elements, used to assist or enhance the understanding of summary results.
[0377] "Motion images" refer to visual content that changes over time and consists of multiple consecutive frames, including video clips, animations, and GIFs.
[0378] "Identification information" refers to metadata or reference information used to uniquely or explicitly indicate visual information or dynamic images, including Uniform Resource Locators, index numbers, tags, etc.
[0379] "Display data" refers to a comprehensive set of data used for display on user terminals, which includes structured summary points, visual information identifiers, and other display parameters.
[0380] "Sending unit" refers to a functional module or program component responsible for transmitting display data, updated summary results, or translation results to the user information processing terminal through a communication channel.
[0381] "Information processing terminal" refers to an electronic device operated by a user to receive data sent by a server and to present and interact with it, including portable terminals, desktop terminals and other computing devices with network connectivity.
[0382] "Conversational input" refers to an input format in which users gradually raise their needs or make modifications through multiple rounds of interaction, using natural language text, voice commands, or interface controls.
[0383] The “interactive adjustment unit” refers to a functional module that regenerates or modifies prompts based on dialogic input sent from an information processing terminal, thereby changing the abstract style, level of detail, and information classification, and obtaining updated abstract results accordingly.
[0384] "Abstract style" refers to the form of abstract text in terms of language style, tone, formality, etc., such as different expressions for children, professionals, or the general public.
[0385] "Abstract detail level" refers to the level of detail in the abstract results in terms of content compression ratio and information granularity, reflecting whether the abstract tends to be a brief overview or a detailed explanation.
[0386] "Information classification categories" refers to the categories of electronic information based on themes, fields, or uses, such as different classification types like science and technology, sports, finance, and art.
[0387] The "translation processing unit" refers to a functional module that generates prompts containing translation instructions based on the user-specified language and updated summary results, calls a generative artificial intelligence model and / or translation processing device to perform translation, thereby obtaining the translated summary results and sending them to the terminal.
[0388] "Translation processing device" refers to a hardware or software entity specifically designed to convert text from one natural language to another, including translation systems implemented using rule-based or statistical and neural network methods.
[0389] "User emotional state information" refers to the category or intensity of emotion inferred from a user's facial expressions, voice, text, or interactive behavior through emotion recognition technology, such as information like joy, sadness, relaxation, and excitement.
[0390] The “emotional adaptation unit” refers to a functional module that dynamically changes the tone, difficulty, and information classification selection conditions of the prompts based on the user’s emotional state information, so that the generated summary results match the user’s current mood in terms of content and style.
[0391] The embodiments of the present invention concretize the structured processing flow of collaboration among the server, terminal and user. By finely controlling the invocation of generative artificial intelligence models on the server side, an integrated technical solution for text acquisition, preprocessing, summarization, translation and sentiment adaptation is achieved, thereby achieving technical improvements in data structure organization, model invocation methods and communication load control within the computer.
[0392] In the following embodiments, the server generally consists of a computer device with a multi-core central processing unit, storage device, and network interface, runs a general-purpose operating system (such as a UNIX-like operating system or a general-purpose server operating system), and has the following software components installed: Natural language processing libraries (such as spaCy, NLTK), deep learning inference frameworks (such as text generation frameworks based on transformer architectures), HTTP service frameworks (such as REST-based service frameworks), and communication modules for calling external translation service interfaces and sentiment analysis interfaces.
[0393] Terminals are typically composed of smart terminal devices that run mobile operating systems and provide user interfaces and sensor access capabilities through application frameworks (such as cross-platform mobile interface frameworks).
[0394] (1) Server-side module structure and data structure In this embodiment, the server comprises multiple functional modules, which exchange data via intermediate data structures in memory. The main internal data structures used by the server include: The table includes a title field (string), a body field (long string), a summary array (string array), a media identifier field (Uniform Resource Identifier string), a user context structure (a set of key-value pairs containing user attributes, comprehension level, sentiment state, historical preferences, etc.), a prompt string, and a structured summary object (a dictionary structure containing a title, summary array, multilingual version mapping, media identifier, etc.).
[0395] The server comprises a data acquisition module, a text preprocessing module, a prompt statement generation module, a model inference interface module, a summary shaping module, a translation processing module, a sentiment adaptation module, and a communication module. These modules sequentially or in parallel read and write the aforementioned data structures in memory. The server maintains a session context object for each user session, recording the user's comprehension level, current language preference, sentiment state, and recently generated summary key points. The session context object is stored in a cache or session database in the form of a key-value mapping.
[0396] (2) Structure and learning methods of generative artificial intelligence models The generative artificial intelligence model used by the server in this embodiment is an autoregressive language model based on the Transformer architecture. This model uses a large-scale text corpus during the pre-training phase and is trained through masked language modeling or next-tag prediction tasks. Its core structure includes multi-layer self-attention encoder-decoder units, a multi-head attention mechanism, feedforward network layers, and layer normalization units.
[0397] During the inference phase, the server only uses the forward propagation portion of the model and does not update parameters. During offline training, the server uses the cross-entropy loss function as the error function to measure the difference between the model's output word probability distribution and the true word sequences in the training set. During training, the server updates weights using gradient descent or variant optimization algorithms (such as adaptive moment estimation algorithms). The server can use data augmentation strategies during training, such as random masking, synonym substitution, and fragment rearrangement, to improve the model's robustness to semantics under different representations.
[0398] When using a generative AI model for summarizing or translating tasks, the server imposes constraints on the model's generation process by explicitly specifying the task type (summarizing or translating), output format (bulk format, limit on the number of items), target reader level, and target language in the prompt statements. During inference, the server can set temperature parameters, sampling strategies (such as top-k or top-p sampling), and maximum generation length to control the diversity and length of the output. This prompt-based control method differs from traditional rule concatenation or template replacement. By changing the control sequence of the model input without altering the model weights, it can obtain outputs with different styles and levels of detail without retraining, thereby improving processing efficiency at the system level.
[0399] (3) Server-side rule-based generation and adjustment of prompt statements In this embodiment, the prompt statement generation module of the server constructs prompt statements using a combination of rules and context. The server encodes user attribute information (such as age group, professional background), comprehension level parameters (such as "junior high school student" or "general adult"), user-specified output preferences (such as "3-5 items" or "each item no more than 30 characters"), and current emotional state (such as "relaxed," "excited," or "sad") into several control clauses. The server combines these control clauses into a complete prompt statement using a predefined template fragment library.
[0400] For example, when generating summaries of science and technology news that are easy for middle school students to understand, the server can generate the following prompt: Please summarize the key points of the following technology news in Simplified Chinese: 1. Designed for junior high school students, avoid using complex technical terms; 2. List the key information using 3-5 bullet points; 3. Each entry should not exceed 30 characters.
[0401] The full text of the news article is as follows: {Article text} When a user is relaxed and wants to read sports news easily, the server can generate the following prompt: "The user is currently in a relaxed state. Please use Simplified Chinese in a light and pleasant tone." Summarize the following sports news into 3-5 key points. Requirements: 1. The language is conversational and easy to understand; 2. Not exaggerated, not inaccurate; 3. Each entry should not exceed 25 characters.
[0402] The full text of the news article is as follows: {Article text} When processing the translation of English scientific articles and the summaries written by middle school students, the server can generate the following prompt: "Below is an English science article. Please process it in two steps:" 1. First, translate the entire text into simplified Chinese; 2. For junior high school students, summarize the key points using 3-5 bullet points.
[0403] The English article is as follows: {English article}” Through this rule-based prompt generation process, the server explicitly encodes the "format constraints," "reader level constraints," "tone and style constraints," and "emotional adaptation constraints," which were originally implicit in the caller's logic, into the prompt string. Since the prompts themselves are in text form, the server can cache, version-manage, and dynamically combine them, reducing the overhead of repeated generation. It can also optimize template design based on statistical results, achieving technical improvements at the model invocation level.
[0404] (4) Server-side text preprocessing and column-based structured summary generation In the text preprocessing module, the server uses a natural language processing library to clean and structure the raw electronic information obtained from the information source. The server uses specific algorithms such as word segmentation, sentence segmentation, and denoising to decompose and normalize long texts, generating target text information. The server can use sentence importance assessment algorithms (such as scoring based on term frequency-inverse document frequency, sentence position features, and title similarity) to rank sentences and guide the model to further compress the text by providing a prompt such as "The following is the main content of the original text." This preprocessing helps reduce model input redundancy, lowers inference computation, and thus shortens the overall response time.
[0405] After receiving the summary results from the generative artificial intelligence model, the server uses the summary shaping module to parse the text: the server splits the output content line by line, identifies the bullet points or serial numbers in the prefix of each line, and organizes them into an array-like data structure. The server then encapsulates this array together with the article metadata (title, original link) and relevant visual information identifiers (image URLs, video URLs) into a structured summary object.
[0406] The server uniformly adopts a structured summary representation in a columnar format, allowing the terminal to simply render according to the array order without requiring additional parsing logic. This unification and simplification of the summary result data structure directly reduces the terminal's parsing burden and rendering computation, while also facilitating the server to modify, rewrite, or extend to multiple languages for individual points in subsequent rounds of interaction, thus improving the system's maintainability and scalability.
[0407] (5) Multilingual translation and model collaboration mechanism of the server The server can employ two collaborative strategies in the translation processing module: One approach involves using an external translation processing device to perform machine translation on the original or abstract text, then inputting the translation result into a generative AI model for secondary summarization or style adjustment via prompts. The second approach is to directly use prompts to instruct the generative AI model to simultaneously complete translation and summarization in the same round of reasoning. Both methods can be controlled via prompts.
[0408] For example, when a server wants to generate a Chinese translation summary of an English article, it can use the following prompt: "Below is an English news article. Please translate it into simplified Chinese first, and then summarize it using 3-5 key points, targeting a general adult reader:" {English article}” In practice, the server can dynamically select between the two strategies mentioned above based on internal network bandwidth, external translation service latency, and the performance metrics of the generative AI model. When the external translation device has high latency, the server can prioritize the method of direct model translation combined with integrated summary prompts, reducing one network call, thereby reducing communication load and overall processing latency, and thus improving throughput at the system level.
[0409] (6) Server's emotional adaptation and content selection logic In the emotion adaptation module, the server obtains user emotional state information from the terminal or external computing device. The terminal captures user images and voice clips through a camera and microphone. The server sends this data to the emotion analysis device, which uses convolutional neural networks or temporal neural network structures to classify the image and voice features, outputting emotion labels and their probability distributions. The server selects the emotion category with the highest probability as the emotional state of the current session.
[0410] The server stores the user's emotional state along with their interests and browsing history in the user's context structure. When selecting information sources and generating prompts for the same user later, the server adjusts the information categorization criteria and tone based on this emotional state. For example, when the emotion is "excitement" and the user's interests include "sports," the server is more likely to select articles from sports-related sources and specify a lively and positive tone in the prompts; when the emotion is "sadness," the server, based on the same news content, constrains the tone to be calm and soothing through the prompts, thereby reducing the reinforcement of negative emotions without altering the facts.
[0411] Because the server embeds emotional information into the prompts, it constrains the input space of the generative AI model, and the model's output summary naturally reflects emotional fit in style. Unlike using emotions only at the content recommendation level, this implementation delves into the text generation layer, directly influencing the model's generation path. This substantially reduces users' subjective aversion to inappropriate content or tone, increasing user trust and continued usage. This end-to-end emotional adaptation approach, compared to traditional methods that only recommend categories based on emotions, introduces new control variables and decision paths into the computer's internal processing flow, representing a technical improvement to the text generation process.
[0412] (7) Terminal interface rendering and interactive information feedback In this embodiment, the terminal displays structured summaries through a graphical user interface framework. After receiving the structured summary object from the server, the terminal presents the title as a text control, arranges each point in the array in a list format, and loads the corresponding image or video based on the media identifier. The terminal uses a vertical scrolling container to organize multiple summary cards, and users can browse large amounts of information by swiping up and down using touch gestures.
[0413] When a user enters a new request, the terminal provides a text input box or a predefined option button. For example, a user can enter: "Keep it short, no more than three points." "Rewrite this summary into a version suitable for elementary school students." Please translate the key points into English. The terminal encodes the aforementioned text or options into conversational input and sends it to the server along with the identifier of the currently selected summary. In its interaction adjustment module, the server regenerates prompts based on this input and invokes a generative artificial intelligence model to produce a new summary result. Through multiple rounds of interaction, the user can gradually converge to a summary version that meets their individual needs.
[0414] This interactive mode shifts the work that originally required users to read and rewrite themselves to the programmable prompts on the server side. It replaces a large amount of traditional manual editing work with multi-round model reasoning, and ensures the efficiency and consistency of the whole process through structured data and a unified text generation interface.
[0415] (8) Technical effects and causal relationships The server achieves significant technical improvements in the following aspects by using a unified prompt generation mechanism, a list-based structured summary data structure, sentiment adaptation control logic, and a multilingual translation collaboration mechanism: 1. Improved processing speed: The server performs text trimming and sentence importance ranking during the text preprocessing stage, filtering out redundant information. This significantly shortens the input length of the generative AI model, thereby reducing the attention computation overhead in the model's transformer structure and lowering inference time.
[0416] 2. Improved accuracy and relevance: The server explicitly provides task requirements, number of items, and target reader level in the prompts, constraining the model to output only a limited number of key points. It also filters source articles through sentiment and interest context, reducing information irrelevant to the user's actual needs, thereby improving the matching degree between the summary content and the user's needs.
[0417] 3. Reduced communication load and terminal burden: The server structures the summary results into compact column data and pre-associates them with media identifiers, so that the terminal only needs to transmit the necessary content and identifiers to complete the rendering, without transmitting the full text and complex typesetting information, thereby reducing network transmission volume and reducing the computational burden of terminal parsing and rendering.
[0418] 4. Optimization of computing resource utilization: The server introduces a variety of prompt statement templates and strategies, and can choose whether to use an external translation device based on the current load and latency, thereby dynamically balancing translation and summarization and making the allocation of computing resources among different modules more reasonable.
[0419] 5. Reduced errors and uncertainty: By standardizing the prompt statements, the server constrains the model's output range within the expected format (e.g., fixed number of items, column format), significantly reducing post-processing errors caused by inconsistent model output formats, thereby improving the overall predictability and stability of the system.
[0420] Through the collaborative work of the above modules, the system of the present invention does not simply automate human summarization and translation work, but systematically transforms the calling method of generative artificial intelligence models and the organization of information flow within the computer by introducing a unified prompt statement control layer, structured data representation layer and emotion adaptation control layer inside the server. This achieves comprehensive technical improvements in processing speed, accuracy, resource utilization and user experience.
[0421] use Figure 14 The processing flow is explained.
[0422] Step 1: Users input their needs and preferences on the terminal.
[0423] Users can select the information category (e.g., science, sports, finance), target language (e.g., Simplified Chinese), and comprehension level (e.g., elementary school student, junior high school student, general adult) in the terminal interface, and can enter natural language commands in the text input box (e.g., "Please summarize this science article in a way that even a junior high school student can understand"). Users can also choose whether to enable the sentiment analysis function.
[0424] Input: User's category selection, language selection, comprehension level selection, and natural language command text.
[0425] Based on these inputs, the terminal constructs a request object in its local memory, containing the user ID, selection parameters, and instruction text, and prepares to send it to the server. The terminal's specific actions include: reading the current values of UI controls, encoding the text into a UTF-8 string, and packaging multiple fields into structured data.
[0426] Output: A request data structure containing user needs and preferences.
[0427] Step 2: The terminal collects user emotion-related data and sends it to the server.
[0428] With the user's consent, the terminal accesses the camera and microphone interfaces to acquire one or more frames of facial images and a short audio segment. The terminal compresses and encodes the image and audio into binary data suitable for network transmission and merges it with the request data structure generated in step 1.
[0429] Input: User's image data and / or audio data, and the request data structure from step 1.
[0430] The terminal performs simple preprocessing on the image and audio (such as resolution scaling and audio sampling rate unification), and then sends a request containing text parameters and multimedia data to the server's interface address via a network protocol.
[0431] Output: The complete request message sent to the server.
[0432] Step 3: The server obtains the original electronic information from the information source.
[0433] After receiving a request from the terminal, the server parses fields such as information category and user language preference from the request message. Based on the category field, the server selects a preset network interface or data source (such as a news API URL) and sends a query request to the information source using a network request library.
[0434] Input: Category parameters and other filtering conditions in the request message sent by the terminal.
[0435] The server performs JSON or HTML parsing on the response data, extracting the article title, body content, publication time, and media URL, and then packs multiple original electronic documents into an internal list structure. This data processing includes error checking of the network response and type conversion of data fields.
[0436] Output: A list structure containing the content of multiple original articles and metadata.
[0437] Step 4: The server preprocesses the raw text information.
[0438] The server iterates through the original article list obtained in step 3, extracts the main text of each article, and calls a natural language processing library to clean the text, including removing HTML tags, scripts and advertising paragraphs, standardizing punctuation, and merging redundant spaces. The server then performs sentence segmentation and word segmentation operations, dividing the long text into sentence arrays and word arrays.
[0439] Input: The text string of the original article and related metadata.
[0440] During the preprocessing process, the server also uses simple statistical methods (such as word frequency-inverse document frequency and sentence length filtering) to roughly assess the importance of sentences and remove obviously irrelevant sentences, thereby shortening the input length of subsequent generative artificial intelligence models.
[0441] Output: The target text information object for each article, including the cleaned text, sentence array, and basic importance score.
[0442] Step 5: The server analyzes the user's emotional state and updates the user's context.
[0443] The server reads image and audio data from the terminal request and passes it to the sentiment analysis device or sentiment recognition model interface. After receiving the sentiment analysis output, the server parses out the sentiment label (e.g., joy, sadness, relaxation, excitement) and its probability.
[0444] Input: User image / audio data from the terminal.
[0445] The server selects the sentiment tag with the highest probability as the current sentiment state, and writes it, along with the user ID, areas of interest, and browsing history, into or updates the user context object. Data processing includes converting the sentiment probability vector into discrete tags and storing them in key-value format.
[0446] Output: The updated user context data structure.
[0447] Step 6: The server selects articles that match the user's emotions and interests.
[0448] The server reads the target text information list generated in step 4 and the user context generated in step 5, and filters and sorts the candidate articles based on the user's sentiment tags and interest categories. For example, when the sentiment is "relaxed," the server lowers the priority of articles with negative themes, and when the interest is "sports," the server increases the weight of sports articles.
[0449] Input: List of target text information, user context (including emotions and interests).
[0450] The server calculates a matching score for each article. The scoring function combines category matching, sentiment fit, and content novelty. Then, it selects the articles with the highest scores as the target for the abstract.
[0451] Output: A selected subset of articles for subsequent abstracting.
[0452] Step 7: The server generates prompts for controlling generative artificial intelligence models.
[0453] The server selects a suitable template from a predefined template library based on a curated subset of articles, the user's comprehension level, language preferences, and emotional state, and inserts control clauses to form a complete prompt statement. For example, when the comprehension level is "junior high school student" and the category is "technology," the server generates the following prompt statement: Please summarize the key points of the following technology news in Simplified Chinese: 1. Designed for junior high school students, avoid using complex technical terms; 2. List the key information using 3-5 bullet points; 3. Each entry should not exceed 30 characters.
[0454] The full text of the news article is as follows: {Article text} Input: User context (comprehension level, sentiment, language preference), target text information of selected articles.
[0455] When generating the prompt, the server performs string concatenation and placeholder replacement on the text, embeds the target text information into a specified position in the prompt, and adds restrictions on the number of items and tone requirements.
[0456] Output: A pre-constructed prompt string for each article or collection of articles.
[0457] Step 8: The server invokes a generative artificial intelligence model to generate summary results.
[0458] The server takes the prompt generated in step 7 as input and passes it to the generative AI model through the model inference interface. The server sets model parameters, such as maximum output length, temperature value, and sampling strategy, to control the generation behavior.
[0459] Input: Prompt string and generated parameters.
[0460] The generative AI model internally encodes the prompts and the main text, performs multi-layered self-attention and feedforward operations, and progressively outputs summary text. The server receives the text sequence output by the model, removes redundant blank lines or echoed prompts, and retains only the summary paragraphs.
[0461] Output: A summary text string corresponding to each prompt statement.
[0462] Step 9: The server parses the summary text into structured data in column format.
[0463] The server splits the summary text from step 8 line by line, identifies the bullet points or numbers in each line, filters out blank lines, and stores the valid lines into a string array. The server assigns an array index to each point to form a programmable list of points.
[0464] Input: The summary text output by the generative artificial intelligence model.
[0465] The server simultaneously combines the corresponding article title, original link, and image or video URL with the key points array to generate a structured summary object. Data processing at this stage includes field mapping, array construction, and media identifier binding.
[0466] Output: A list of structured summary objects, each containing fields such as title, summary array, and media identifier.
[0467] Step 10: The server performs multilingual translation processing based on the user's language requirements.
[0468] The server reads the target language preference from the user's context. If the target language does not match the current language of the summary, it invokes a translation processing device or guides a generative AI model to translate using prompts. For example, the server can construct the following prompts: Please translate the following key points from Chinese into English, maintaining the bullet point format: {Key Points List Text} Input: An array of key points from the structured summary and target language information.
[0469] The server concatenates the key points array into text and sends it to the translation device or model interface. After receiving the translation results, it splits the text into new key points arrays line by line and stores them in the multilingual field of the structured summary object.
[0470] Output: An updated structured summary object containing an array of multilingual key points.
[0471] Step 11: The server sends the structured summary object to the terminal.
[0472] The server combines multiple structured summary objects into a response data structure, including fields such as article title, summary array, multilingual versions, media URL, and original link. The server serializes this data structure and sends it to the terminal's interface address via a network protocol.
[0473] Input: A list of structured summary objects.
[0474] The server can compress or prune fields before sending, retaining only the fields needed for terminal rendering, thereby reducing the size of the data packet and reducing network bandwidth usage.
[0475] Output: The response message sent to the terminal.
[0476] Step 12: The terminal receives and renders the summary content.
[0477] After receiving the server's response, the terminal parses the data structure and maps each structured summary object to a UI component: the title is displayed as text, the key points array is displayed as a list control, and the media URL corresponds to an image or video component. The terminal places multiple summary cards in a vertically scrollable container.
[0478] Input: The response message sent by the server.
[0479] During the rendering process, the terminal maps strings to the properties of UI controls, submits the image URL to the image loading library for caching and display, and sets scroll area parameters so that users can smoothly scroll up and down to browse.
[0480] Output: A view of multiple summaries displayed in a column format on the terminal screen.
[0481] Step 13: Users can perform secondary interactions on the terminal to adjust the summary.
[0482] After reading the summary, users can issue new commands through the terminal interface, such as: "Make it shorter, no more than three points," "Change it to a version suitable for elementary school students," and "Translate the key points into English."
[0483] Input: New natural language commands or button selections from the user.
[0484] The terminal sends these inputs, along with the identifier of the current digest object, back to the server. The terminal's specific actions include: reading the ID of the currently selected digest, encapsulating the user input into a new request, and transmitting it via the network interface.
[0485] Output: A secondary request message containing adjustment instructions.
[0486] Step 14: The server updates the prompt statement and regenerates the summary based on secondary interaction.
[0487] After receiving the request from step 13, the server extracts the latest level of understanding and emotional state from the user's context, and regenerates the prompt statement based on the new instruction content. For example, when the user requests a shorter prompt, the server adds the constraint "maximum 3 points, each no less than X words" to the prompt statement; when the user requests a primary school level prompt, the server adds the instruction "use very simple vocabulary and avoid technical terms".
[0488] Input: Secondary interaction request sent by the terminal (including new user instructions and current summary identifier), user context, and original target text information.
[0489] Based on the new prompt, the server calls the generative artificial intelligence model and / or translation device again, repeating the summary generation and translation processing steps 8 to 10 to obtain the updated summary text and structured data, and sends it to the terminal again.
[0490] Output: An updated summary structured object that meets the user's new requirements.
[0491] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0492] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0493] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0494] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0495] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0496] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0497] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0498] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0499] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0500] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0501] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0502] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0503] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0504] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0505] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0506] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0507] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0508] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0509] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0510] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0511] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0512] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0513] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0514] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0515] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0516] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0517] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0518] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0519] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0520] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0521] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0522] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0523] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0524] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0525] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0526] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0527] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0528] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0529] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0530] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0531] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0532] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0533] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0534] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0535] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0536] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0537] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0538] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0539] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0540] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0541] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0542] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0543] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0544] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0545] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0546] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0547] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0548] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0549] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0550] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0551] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0552] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0553] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0554] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0555] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0556] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0557] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0558] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0559] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0560] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0561] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0562] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0563] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0564] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0565] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0566] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0567] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0568] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0569] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0570] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0571] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0572] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that executes specific processes by executing software, i.e., a program. Furthermore, processors can include, for example, FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are dedicated circuits with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0573] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0574] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0575] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0576] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0577] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0578] In addition, the following notes are provided in response to the above explanation.
[0579] Example 1 (Note 1) An information processing system, characterized in that it comprises: Means for obtaining electronic information from an information-providing medium via network communication; Means for parsing structured documents or hierarchical data in acquired electronic information, removing display control information and decorative information, normalizing strings and deleting useless components, thereby converting the electronic information into natural language text and preprocessing the natural language text; A means for dividing a preprocessed natural language text into multiple parts based on its length and structure, and for attaching metadata about the processing order and overall composition to each part of the text. A means for determining the summary type, summary length, output language, output format and reading difficulty level of the natural language text based on user attribute information and user demand information, generating a natural language prompt statement that includes the instruction content and embeds the partial data, and controlling the input of the prompt statement into a generative artificial intelligence model. Means for performing formatting processing on the summary result text obtained from the generative artificial intelligence model, assigning an entry listing structure, a numbering structure and a hierarchical structure, converting the summary result text into at least one of a machine-readable format and a human-readable format, and sending the summary result text to a user device. A means for parsing interactive operation information and natural language instructions obtained from the user device, and accordingly changing at least one of the summary type, summary length, output language and reading difficulty level, regenerating a prompt statement for updating, and iteratively adjusting the summary result text by inputting the updated prompt statement back into the generative artificial intelligence model; A means of detecting inconsistencies between the language of the summary result text and the language of the user interface, and selectively performing machine translation processing or using the generative artificial intelligence model for translation processing based on the detection results, thereby converting the summary result text into different languages.
[0580] (Note 2) According to the information processing system described in Appendix 1, the generative artificial intelligence model is controlled to: parse the natural language text according to the specified summary type and output format contained in the prompt statement, and generate a summary result text in the form of an item list containing multiple key points including time information, location information, subject information, and result information.
[0581] (Note 3) According to the information processing system described in Appendix 1, the prompt statements input to the generative artificial intelligence model are automatically generated based on template information of at least one of user attribute information, reading difficulty level, and expression style. The template information includes multiple text styles and output constraints defined for different comprehension levels, such as children, ordinary users, and professional users. The template information is switched according to the user's selection operation, so that the expression of the summary result text for the same natural language text can be adapted to different comprehension levels.
[0582] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring information from information sources and filtering the information based on the user's areas of interest and usage history; A device for preprocessing acquired information as text data, performing tasks such as removing useless elements, sentence segmentation, and segmentation or truncation of long texts on the text data. A means for generating prompts for inputting into a generative artificial intelligence model based on the attributes of the text data and the specified summary expression style, comprehension level, language, and category received from the user, and for controlling the input of the prompts and the text data into the generative artificial intelligence model. An apparatus for parsing the summary results output by the generative artificial intelligence model, splitting the summary results into a column structure, and organizing the title, summary entries, image information, and video information into a predetermined display format; A device for sending the summary results, which are shaped into the predetermined display format, to the user terminal as a list that can be continuously scrolled vertically on the user terminal, and for providing a user interface that allows the user terminal to display the summary results page by page according to the article; An apparatus for acquiring input operations or natural language input from the user terminal, interactively modifying the template, summary length, style, comprehension level or language of the prompt statement based on the input, and calling the generative artificial intelligence model again to regenerate the summary result based on the modified conditions; An apparatus for performing translation processing on the summary results and the corresponding original text in different languages, and sending the summary results or original text in multiple languages to the user terminal according to the user's selection; An apparatus for learning a user's level of interest based on their browsing history, selection actions, and dwell time, and for updating the priority of obtaining information from the information source and the generation conditions of the prompt statement based on the learning results.
[0583] (Note 2) According to the information processing system described in Appendix 1, the device for generating prompt statements is configured to include in the prompt statements input to the generative artificial intelligence model instructions to summarize the acquired information into a predetermined number of key points, the maximum number of characters for each key point, the text type, and the target comprehension level, so as to control the structure and expression style of the summary results output by the generative artificial intelligence model.
[0584] (Note 3) According to the information processing system described in Appendix 1, the user terminal is configured to: display the title, key points, and related images or videos in a vertically scrolling manner when displaying summary results, and provide an operation interface for receiving users to select categories, specify keywords, select text styles, select comprehension levels, and select languages, and send the content input in the operation interface as the generation conditions for generating the prompt statement and the acquisition conditions for obtaining information from the information source to the processor.
[0585] Example 2 (Note 1) An information processing system, characterized in that it comprises: A means of obtaining information from information sources; A means for converting acquired information into text information and performing preprocessing, including structuring and noise removal; A means for generating control prompts that instruct generative artificial intelligence models on the form, length, and target domain of summary content, based on preprocessed text information and prompts received from the user; Means for using the control prompts and the preprocessed text information as input to execute the generative artificial intelligence model to generate summary information including key point extraction; Means for assigning identification information corresponding to the information source and information position for reference to the generated summary information, and for shaping the summary information into a hierarchical structure and paragraph structure suitable for display in a vertical scrolling manner on a user terminal; Means for sending shaped summary information to a user terminal that can display information in a vertically continuous manner, and for the user terminal to display the summary information as visual information; Means for acquiring additional prompts input from a user terminal in the form of a dialogue, generating corrective prompts based on the additional prompts to change the level of detail, target domain, and expression of the summary, and executing the generative artificial intelligence model again to update the summary information; Machine translation means used to translate the text information used to generate or update summary information, as well as the summary information itself, into different languages specified by the user. A learning control method used to record the dialogue history and summary adjustment history for each user, and to reflect this when generating the control prompts and correction prompts.
[0586] (Note 2) The information processing system according to Appendix 1 is characterized in that, When the control prompt statement is input into the generative artificial intelligence model, information units will be selected based on the degree of interrelation from information obtained from multiple information sources. The control prompt statement will instruct the generation of summary information in a hierarchical column format for the selected information units, and at the same time specify the display order and summary granularity corresponding to the vertical scrolling display of the user terminal.
[0587] (Note 3) The information processing system according to Appendix 1 is characterized in that, When generating the correction prompt, based on user attribute information representing the user's level of understanding, professional field, and type of display medium, as well as the dialogue history and summary adjustment history, a template is selected according to its fitness from the description templates with multiple levels of abstraction and multiple expression styles. While embedding the selected template as the input context into the generative artificial intelligence model, the summary information is regenerated.
[0588] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A data acquisition unit used to obtain electronic information from information sources; A text preprocessing unit for converting the electronic information into text information and performing preprocessing to generate target text information for summarization; The system is used to generate prompt statements for instructing the generative artificial intelligence model to perform summarization processing and / or translation processing based on the target text information, user attribute information, and specified information related to the user's comprehension level, and to input the prompt statements into the prompt statement generation unit of the generative artificial intelligence model. The summary results output by the generative artificial intelligence model are used to parse the summary points, organize the summary points into structured data in the form of columns, and associate them with identification information related to visual information and dynamic images to generate summary shaping units for display data. A sending unit for sending the display data to a user information processing terminal via a communication channel; An interactive adjustment unit is used to regenerate prompt statements for changing the summary style, summary detail, and information classification type based on user conversational input received from the information processing terminal, and to re-input the prompt statements into the generative artificial intelligence model to obtain updated summary results. A translation processing unit is used to generate a prompt statement containing translation processing instructions based on the user's specified input regarding different languages and the updated summary results, and input the prompt statement into the generative artificial intelligence model and / or translation processing device to obtain the translated summary results and send the translated summary results to the information processing terminal. An emotion adaptation unit is used to dynamically change the tone, difficulty, and information classification selection conditions of the prompt statement based on the user's emotional state information obtained from the information processing terminal or external computing device, so as to generate a summary result corresponding to the emotional state.
[0589] (Note 2) The information processing system according to Appendix 1 is characterized in that, The prompt statement generation unit is configured to explicitly define the output format and the number of entries in the prompt statement, so that the generative artificial intelligence model can extract important information from the target text information and generate summary results in a list format distinguished by entries.
[0590] (Note 3) The information processing system according to Appendix 1 is characterized in that, The data acquisition unit and the interaction adjustment unit are configured to select the category of the electronic information to be acquired based on the user's emotional state information, browsing history information and interest area information acquired from the information processing terminal, and apply the processing of each unit described in Note 1 to the electronic information belonging to the selected category, thereby preferentially generating summary results corresponding to the user's emotional state and interest area.
Claims
1. An information processing system, characterized in that, include: processor; The processor is configured as follows: Obtain information from information sources; The acquired information is converted into text data and the text data is preprocessed. Based on the preprocessed text data, prompt words are generated to indicate the summarization of the information, and the prompt words are input into a generative artificial intelligence model; The summarized information output by the generative artificial intelligence model is organized into a predetermined format and sent to the user's terminal; Adjust the summary of the information in a conversational manner; The summarized information will be translated into different languages.
2. The information processing system according to claim 1, characterized in that, The processor is configured to use prompt words to instruct a generative artificial intelligence model to parse information obtained from an information source and summarize key points in bulleted form.
3. The information processing system according to claim 1, characterized in that, The processor is configured to generate prompts based on user input to adjust the expression of the summary to suit different levels of understanding, and to apply appropriate templates according to the user's selection.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A