Information processing system and information processing method

The information processing system addresses the integration challenges of large-scale language models by generating translation guidelines using reader and stylistic features, ensuring high convenience, usefulness, and reliability in natural language processing.

WO2026033395A1PCT designated stage Publication Date: 2026-02-12SEMICON ENERGY LAB CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/057942
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-08
Filing Date
2025-08-05
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

The integration and operation of large-scale language models are hindered by equipment and cost constraints, necessitating the use of external services, which may compromise convenience and reliability in natural language processing applications.

Method used

An information processing system comprising components that utilize a large-scale language model to generate translation guidelines based on positive and negative examples, reader characteristics, and stylistic preferences, ensuring accurate and reliable translation generation.

Benefits of technology

The system provides highly convenient, useful, and reliable translation guidelines by incorporating reader characteristics and stylistic features, enhancing the quality and appropriateness of translated documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025057942_12022026_PF_FP_ABST
    Figure IB2025057942_12022026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a novel information processing system that has enhanced convenience, usefulness, and reliability. An information processing device according to the present invention is composed of three components. The first component receives a positive example list and transmits the positive example list to the third component. The positive example list includes features of intended readers and features of a preferred writing style. The second component performs processing by using a large language model, and extracts features according to a first instruction sentence. The third component receives the positive example list and translation guidelines and shares this information. Additionally, the device extracts the features of the intended readers and the features of the preferred writing style, generates translation guidelines, and provides the translation guidelines including recommendations.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system and information processing method

[0001] One embodiment of the present invention relates to an information processing system, an information processing method, or a semiconductor device.

[0002] Note that one embodiment of the present invention is not limited to the above technical field. The technical field of one embodiment of the invention disclosed in this specification relates to an object, a method, or a manufacturing method. Alternatively, one embodiment of the present invention relates to a process, a machine, a manufacture, or a composition of matter. Therefore, more specifically, examples of the technical field of one embodiment of the present invention disclosed in this specification include a data processing device, a semiconductor device, a memory device, a driving method thereof, or a manufacturing method thereof.

[0003] In recent years, the development of language models using neural networks has been actively carried out, and large-scale language models (LLMs) in particular have attracted attention. A large-scale language model is a natural language processing model trained using a large amount of data. A large-scale language model can realize a dialogue model that responds to user instructions, for example. Non-Patent Document 1 discloses GPT-4 (Generative Pre-trained Transformer 4) (registered trademark) as a large-scale language model, and also discloses ChatGPT as a dialogue model.

[0004] The use of large-scale language models has significantly increased the capabilities of natural language processing models. However, as language models become larger, it is difficult to incorporate and operate language models in-house due to the equipment and cost involved. Therefore, one way to use language models is to use external services that provide language models.

[0005] Summary of ChatGPT / GPT-4 Research and Perspective Towards the Future of Large Language Models, Yiheng Liu et al. (Submitted on 4 Apr 2023, [online], Internet <URL: https: / / arxiv.org / abs / 2304.01852>

[0006] An object of one embodiment of the present invention is to provide a novel information processing system with excellent convenience, usefulness, or reliability, or to provide a novel information processing method with excellent convenience, usefulness, or reliability, or to provide a novel information processing system, a novel information processing method, or a novel semiconductor device.

[0007] Note that the description of these problems does not preclude the existence of other problems. Note that one embodiment of the present invention does not necessarily solve all of these problems. Note that problems other than these will become apparent from the description of the specification, drawings, claims, etc., and it is possible to extract other problems from the description of the specification, drawings, claims, etc.

[0008] (1) One aspect of the present invention is an information processing device having a first component, a second component, and a third component.

[0009] The first component has a function of receiving a list of positive examples and transmitting it to the third component. The list of positive examples includes characteristics of the intended reader and characteristics of a preferred writing style.

[0010] The second component has a function of receiving the first directive and transmitting the characteristics of the intended reader and the preferred writing style to the third component, and a function of performing processing using a large-scale language model, where the large-scale language model has a function of extracting the characteristics of the intended reader and the preferred writing style in accordance with the first directive.

[0011] The third component has a function for receiving and sharing within the third component a list of positive examples, characteristics of the intended readers, characteristics of the preferred writing style, and translation guidelines. The third component also has subcomponents.

[0012] The subcomponent has the functionality to create and send a first directive and a second directive to the second component.

[0013] The first instruction includes a first instruction and a positive example list, and the first instruction includes a procedure for extracting characteristics of an intended reader and characteristics of a preferred writing style from the positive example list.

[0014] The second instruction includes second instructions, characteristics of the intended reader, and preferred stylistic features. The second instructions include a procedure for generating translation guidelines, and the translation guidelines include recommendations. The recommendations also include characteristics of the intended reader and preferred stylistic features.

[0015] (2) In accordance with another aspect of the present invention, the first component may receive a negative example list and transmit the negative example list to the third component. The negative example list may include stylistic features that are to be avoided.

[0016] The second component has a function of receiving the third directive and transmitting the avoided writing style features to the third component, and the large-scale language model has a function of extracting the avoided writing style features in accordance with the third directive.

[0017] The third component has a function of accepting the negative example list and the avoided stylistic features and sharing them within the third component.

[0018] The subcomponent has the function of creating a third instruction statement and sending it to the second component.

[0019] The third instruction sentence includes a third instruction and a negative example list, and the third instruction includes a procedure for extracting the characteristics of the objectionable writing style from the negative example list.

[0020] The second directive includes stylistic features to be avoided. The translation guidelines also include terms to be avoided, and the terms to be avoided include stylistic features to be avoided.

[0021] This makes it possible to generate translation guidelines from a list of positive examples that include characteristics of the expected reader and preferred stylistic features. Also, it is possible to generate translation guidelines from a list of negative examples that include avoided stylistic features. Also, it is possible to generate translation guidelines that take into account characteristics of the expected reader and preferred stylistic features. Also, it is possible to generate translation guidelines that take into account avoided stylistic features. As a result, it is possible to provide a novel information processing system that is highly convenient, useful, and reliable.

[0022] (3) Another aspect of the present invention is the information processing device described above, wherein the first component has a function of receiving an original document and transmitting it to the third component, and a function of receiving a translated document and providing it.

[0023] The second component has a function of receiving the fourth instruction sentence and transmitting the translated document to the third component. Note that the large-scale language model has a function of generating the translated document in accordance with the fourth instruction sentence.

[0024] The third component has a function of receiving an original document and sharing it within the third component, and a function of receiving a translated document and sending it to the first component.

[0025] The subcomponent has the function of creating a fourth directive.

[0026] The fourth instruction includes a fourth instruction, a translation guideline, and an original document. The fourth instruction includes a procedure for generating a translated document from the original document in accordance with the translation guideline.

[0027] (4) Another aspect of the present invention is the above-mentioned information processing device, wherein the second component has a function of receiving a fifth instruction sentence and sending an inappropriate list to the third component, and a function of receiving a sixth instruction sentence and sending a translated document to the third component.

[0028] The large-scale language model has a function of generating an inappropriate list according to a fifth instruction sentence, and a function of generating a translated document according to a sixth instruction sentence.

[0029] The subcomponent has the functionality to create the fifth and sixth directives.

[0030] The fifth instruction includes a fifth instruction, a translation guideline, an original, and a translated document. The fifth instruction includes a procedure for evaluating whether the translation from the original to the translated document complies with the translation guideline, outputting true if the translation is evaluated as conforming, and generating an inappropriate list if the translation is evaluated as not conforming. The inappropriate list also includes the original text and translation that are evaluated as not conforming.

[0031] The sixth instruction includes a sixth instruction, a translation guideline, an original, a translated document, and an inappropriate list. The sixth instruction includes a procedure for generating a translated document from the original in accordance with the translation guideline so as to avoid translations listed on the inappropriate list.

[0032] This makes it possible to generate a translated document from the original in accordance with the translation guidelines. It is also possible to generate a translated document that takes into account the characteristics of the expected reader and the preferred stylistic features. It is also possible to generate a translated document that does not have any stylistic features that are objectionable. It is also possible to proofread the translated document based on the translation guidelines. As a result, it is possible to provide a novel information processing system that is highly convenient, useful, and reliable.

[0033] (5) One aspect of the present invention is an information processing device having a first component, a second component, and a third component.

[0034] The first component has the functionality to accept and provide questions, and the functionality to accept and send answers to the third component.

[0035] The second component has a function of receiving the instruction sentence and transmitting the first writing style features and the second writing style features to the third component, and a function of performing processing using a large-scale language model.

[0036] The large-scale language model has a function of extracting first and second stylistic features according to the instruction sentence.

[0037] The third component has a function of receiving the answer and sharing it within the third component, and the third component also has a first subcomponent, a second subcomponent, and a third subcomponent.

[0038] The first subcomponent has the function of creating an instruction statement and sending it to the second component.

[0039] The instruction includes an instruction, a first document, and a second document, and the instruction includes a procedure for extracting and listing first style features from the first document and a procedure for extracting and listing second style features from the second document.

[0040] The second subcomponent has a function of selecting two of the n typical documents to create a combination, and a function of designating one of the combinations as the first document and the other as the second document.

[0041] The third subcomponent provides the functionality of generating a question, which asks whether the first stylistic feature or the second stylistic feature is more suitable for the profile of the intended reader.

[0042] When the answer is a first writing style characteristic, the third subcomponent has a function of adding the expected reader characteristic and the first writing style characteristic to a positive example list, and a function of adding the expected reader characteristic and the second writing style characteristic to a negative example list. Also, the first writing style characteristic is added as a preferred writing style characteristic, and the second writing style characteristic is added as an undesired writing style characteristic.

[0043] As a result, for example, a first stylistic feature can be added to a positive example list as a style that is appropriate for the characteristics of the expected reader. Furthermore, for example, a second stylistic feature can be added to a negative example list as a style that is not appropriate for the characteristics of the expected reader. Furthermore, using a large-scale language model, stylistic features of the first document can be extracted and listed as first stylistic features. Furthermore, stylistic features of the second document can be extracted and listed as second stylistic features. Furthermore, by comparing the first stylistic features with the second stylistic features, a user of the information processing system, for example, can evaluate whether the style is appropriate for the characteristics of the expected reader. As a result, a novel information processing system that is highly convenient, useful, and reliable can be provided.

[0044] (6) In accordance with another aspect of the present invention, the second subcomponent is configured to select two of the n typical documents and create all unique combinations. The n typical documents are classified into different clusters. The n typical documents each have an embedding closest to the center of gravity of the cluster.

[0045] The third subcomponent has a function for creating questions for each of all combinations.

[0046] As a result, for example, a feature of a first writing style can be added to the positive example list as it is deemed to be a writing style that is suitable for the characteristics of the expected reader. Also, for example, a feature of a second writing style can be added to the negative example list as it is deemed to be a writing style that is not suitable for the characteristics of the expected reader. Furthermore, the positive example list and the negative example list can be created using n typical documents classified into different clusters. As a result, a novel information processing system that is highly convenient, useful, and reliable can be provided.

[0047] (7) According to another aspect of the present invention, the information processing device further includes a fourth component that receives an original document and transmits an embedded representation to the third component, and performs processing using the embedded model.

[0048] The embedding model has a function of converting n+1 typical documents and the original document into embedded representations. Among the n+1 typical documents, n-1 are classified into different clusters and have embedded representations closest to the centers of gravity of their respective clusters. Furthermore, all of the n+1 typical documents are classified into clusters different from the original document. Furthermore, two of the n+1 typical documents are classified into the same cluster as the original document and have embedded representations selected in descending order of proximity to the center of gravity of the cluster.

[0049] The first component has a function of accepting the original document and transmitting it to the third component.

[0050] The third component has a function of receiving the original and sending it to the fourth component, and a function of receiving the embedded representation and sharing it within the third component.

[0051] The second subcomponent has a function of selecting two documents from n+1 typical documents and creating all unique combinations.

[0052] The third subcomponent provides the functionality to generate questions for each of all combinations.

[0053] As a result, for example, a first stylistic feature can be added to the positive example list as a writing style that is appropriate for the characteristics of the expected reader. Furthermore, for example, a second stylistic feature can be added to the negative example list as a writing style that is not appropriate for the characteristics of the expected reader. Furthermore, a positive example list and a negative example list can be created using n-1 typical documents classified into a different cluster from the original. Furthermore, by using two typical documents classified into the same cluster as the original, minor differences in stylistic features between similar documents can be detected by adding, for example, a first stylistic feature to the positive example list. Furthermore, for example, a second stylistic feature can be added to the negative example list. As a result, a novel information processing system that is highly convenient, useful, and reliable can be provided.

[0054] (8) One aspect of the present invention is an information processing method having a first phase.

[0055] The first phase comprises steps 1 to 11.

[0056] In the first step of the first phase, the first component accepts a list of positive examples and a list of negative examples and sends them to the second component.

[0057] In the second step of the first phase, the second component receives the positive example list and the negative example list and shares them within the second component, where the second component includes the first subcomponent.

[0058] In the third step of the first phase, the first subcomponent creates a first instruction statement and sends it to the third component, the first instruction statement including a first instruction and a list of positive examples, and the first instruction includes a procedure for extracting characteristics of the intended reader and preferred writing style from the list of positive examples.

[0059] In the fourth step of the first phase, the third component accepts the first directive and uses a large-scale language model to extract characteristics of the intended reader and preferred writing style.

[0060] In the fifth step of the first phase, the third component transmits the characteristics of the intended reader and the preferred writing style to the second component.

[0061] In a sixth step of the first phase, the first subcomponent creates and sends to the third component a second instruction statement, the second instruction statement including a second instruction and a list of negative examples, and the second instruction includes a procedure for extracting avoidable writing style features from the list of negative examples.

[0062] In the seventh step of the first phase, the third component accepts the second directive and extracts avoidable stylistic features using a large-scale language model.

[0063] In the eighth step of the first phase, the third component sends the avoided stylistic features to the second component.

[0064] In a ninth step of the first phase, the first subcomponent creates a third instruction statement and sends it to the third component. The third instruction statement includes a third instruction, expected reader characteristics, preferred stylistic characteristics, and avoided stylistic characteristics. The third instruction statement also includes a procedure for generating translation guidelines. The translation guidelines also include recommendations and avoidance items. The recommendations include expected reader characteristics and preferred stylistic characteristics, and the avoidance items include avoided stylistic characteristics.

[0065] In the tenth step of the first phase, the third component accepts the third directive sentence and generates translation guidelines using a large-scale language model.

[0066] In an eleventh step of the first phase, the third component sends the translation guidelines to the second component, thereby completing the first phase.

[0067] This makes it possible to generate translation guidelines from a list of positive examples that include characteristics of the expected reader and preferred stylistic features. Also, it is possible to generate translation guidelines from a list of negative examples that include avoided stylistic features. Also, it is possible to generate translation guidelines that take into account the characteristics of the expected reader and preferred stylistic features. Also, it is possible to generate translation guidelines that take into account avoided stylistic features. As a result, it is possible to provide a novel information processing method that is highly convenient, useful, and reliable.

[0068] (9) Another aspect of the present invention is an information processing method having the first and second phases described above, wherein the second phase follows the first phase and includes the first to twelfth steps.

[0069] In the first step of the second phase, the first component accepts the original and sends it to the second component.

[0070] In the second step of the second phase, the second component accepts the original and shares it within the second component.

[0071] In a third step of the second phase, the first subcomponent creates and sends to the third component a fourth instruction statement, the fourth instruction statement including a fourth instruction, a translation guideline, and the original document, and the fourth instruction includes a procedure for generating a translated document from the original document in accordance with the translation guideline.

[0072] In the fourth step of the second phase, the third component accepts the fourth directive and generates a translated document using a large-scale language model.

[0073] In the fifth step of the second phase, the third component sends the translated document to the second component.

[0074] In the sixth step of the second phase, the first subcomponent creates a fifth instruction statement and sends it to the third component. The fifth instruction statement includes a fifth instruction, a translation guideline, an original document, and a translated document. The fifth instruction also includes a procedure for evaluating whether the translation from the original document to the translated document conforms to the translation guideline, outputting true if the translation conforms, and generating an inappropriate list if the translation does not conform. The inappropriate list also includes the original text and the translated text that are evaluated as not conforming.

[0075] In the seventh step of the second phase, the third component accepts the fifth directive and uses a large-scale language model to output truth or generate an irrelevant list.

[0076] In the eighth step of the second phase, the third component sends the true or false list to the second component.

[0077] In the ninth step of the second phase, if the second component receives true, the process ends, and if the second component receives an inappropriate list, the process proceeds to the tenth step of the second phase.

[0078] In a tenth step of the second phase, the first subcomponent creates and sends to the third component a sixth instruction statement, the sixth instruction statement including a sixth instruction, a translation guideline, the original, a translated document, and an inappropriate list, and the sixth instruction includes a procedure for generating a translated document from the original in accordance with the translation guideline so as to avoid translations listed on the inappropriate list.

[0079] In the eleventh step of the second phase, the third component receives the fifth directive and generates a translated document using a large-scale language model.

[0080] In the twelfth step of the second phase, the third component sends the translated document to the second component, and the process proceeds to the sixth step of the second phase.

[0081] This makes it possible to generate a translated document from the original in accordance with the translation guidelines. It is also possible to generate a translated document that takes into account the characteristics of the expected reader and the preferred stylistic features. It is also possible to generate a translated document that does not have any stylistic features that are objectionable. It is also possible to proofread the translated document based on the translation guidelines. As a result, it is possible to provide a novel information processing method that is highly convenient, useful, and reliable.

[0082] (10) Another aspect of the present invention is an information processing method having the first and second phases described above, wherein the first phase follows the second phase, and the second phase includes the first to ninth steps.

[0083] In the first step of the second phase, the second subcomponent selects two of the n representative documents and creates all unique combinations. The second subcomponent then identifies one document in each combination as a first document and the other as a second document, where the second subcomponent is included in the second component. The n representative documents are classified into different clusters. The n representative documents have embeddings that are closest to the centroids of their respective clusters.

[0084] In a second step of the second phase, the first subcomponent selects one of the combinations not yet selected, creates a seventh instruction statement, and sends it to the third component, the seventh instruction statement including a seventh instruction, a first document, and a second document, and the seventh instruction includes a procedure for extracting and listing first stylistic features from the first document and a procedure for extracting and listing second stylistic features from the second document.

[0085] In a third step of the second phase, a third component receives the seventh directive and generates first and second stylistic features using a large-scale language model.

[0086] In a fourth step of the second phase, the third component transmits the first stylistic features and the second stylistic features to the second component.

[0087] In the fifth step of the second phase, the third subcomponent, which is included in the second component, creates a question and sends it to the first component, asking whether the first stylistic feature or the second stylistic feature is more suitable for the characteristics of the intended reader.

[0088] In the sixth step of the second phase, the first component presents a question and waits for an answer.

[0089] In the seventh step of the second phase, the first component accepts the response and sends it to the second component.

[0090] In the eighth step of the second phase, the third subcomponent accepts an answer, and if the answer is a first style characteristic, adds the expected reader characteristic and the first style characteristic to a positive example list, and adds the expected reader characteristic and the second style characteristic to a negative example list, where the first style characteristic is added as a preferred style characteristic and the second style characteristic is added as an undesired style characteristic.

[0091] In the ninth step of the second phase, if there is a combination that does not produce the seventh directive, the first subcomponent proceeds to the second step of the second phase; otherwise, the second phase ends.

[0092] As a result, for example, a feature of a first writing style can be added to the positive example list as it is deemed to be a writing style that is suitable for the characteristics of the expected reader. Also, for example, a feature of a second writing style can be added to the negative example list as it is deemed to be a writing style that is not suitable for the characteristics of the expected reader. Furthermore, a list of positive examples and a list of negative examples can be created using n typical documents classified into different clusters. As a result, a novel information processing method that is highly convenient, useful, and reliable can be provided.

[0093] (11) Another aspect of the present invention is an information processing method having the first and second phases described above, wherein the first phase follows the second phase, and the second phase includes the first to thirteenth steps.

[0094] In the first step of the second phase, the first component accepts the original and sends it to the second component.

[0095] In the second step of the second phase, the second component accepts the master copy and transmits it to the fourth component.

[0096] In the third step of the second phase, the fourth component accepts the original text and converts it into an embedded representation using the embedding model.

[0097] In the fourth step of the second phase, the fourth component sends the embedded representation to the second component.

[0098] In the fifth step of the second phase, the second subcomponent selects two of the n+1 typical documents, creates all unique combinations, and designates one of the combinations as the first document and the other as the second document. The second subcomponent is included in the second component. Of the n+1 typical documents, n-1 are classified into different clusters, have embedded representations closest to the center of gravity of their respective clusters, and are all classified into clusters different from the original. Of the n+1 typical documents, two are classified into the same cluster as the original. The n+1 typical documents have embedded representations selected in order of proximity to the center of gravity of the cluster.

[0099] In a sixth step of the second phase, the first subcomponent selects one of the combinations that have not yet been selected, creates a seventh instruction statement, and sends it to the third component, the seventh instruction statement including a seventh instruction, a first document, and a second document, and the seventh instruction includes a procedure for extracting and listing first stylistic features from the first document and a procedure for extracting and listing second stylistic features from the second document.

[0100] In a seventh step of the second phase, the third component accepts the seventh directive sentence and generates first and second stylistic features using the large-scale language model.

[0101] In an eighth step of the second phase, the third component transmits the first stylistic features and the second stylistic features to the second component.

[0102] In the ninth step of the second phase, the third subcomponent, which is included in the second component, creates a question and sends it to the first component, asking whether the first stylistic feature or the second stylistic feature is more appropriate for the characteristics of the intended reader.

[0103] In the tenth step of the second phase, the first component presents a question and waits for an answer.

[0104] In the eleventh step of the second phase, the first component accepts the response and sends it to the second component.

[0105] In a twelfth step of the second phase, the third subcomponent accepts an answer, and when the answer is a first writing style characteristic, adds the expected reader characteristic and the first writing style characteristic to a positive example list, and adds the expected reader characteristic and the second writing style characteristic to a negative example list, where the first writing style characteristic is added as a preferred writing style characteristic and the second writing style characteristic is added as an undesired writing style characteristic.

[0106] In the thirteenth step of the second phase, if there is a combination that does not produce the seventh directive, the first subcomponent proceeds to the sixth step of the second phase; otherwise, the second phase ends.

[0107] As a result, for example, a first stylistic feature can be added to the positive example list as a writing style that is appropriate for the characteristics of the expected reader. Furthermore, for example, a second stylistic feature can be added to the negative example list as a writing style that is not appropriate for the characteristics of the expected reader. Furthermore, a positive example list and a negative example list can be created using n-1 typical documents classified into a different cluster from the original. Furthermore, by using two typical documents classified into the same cluster as the original, minor differences in stylistic features between similar documents can be detected by adding, for example, a first stylistic feature to the positive example list. Furthermore, for example, a second stylistic feature can be added to the negative example list. As a result, a novel information processing method that is highly convenient, useful, and reliable can be provided.

[0108] One embodiment of the present invention can provide a novel information processing system with excellent convenience, usefulness, or reliability, or a novel information processing method with excellent convenience, usefulness, or reliability, or a novel information processing system, a novel information processing method, or a novel semiconductor device.

[0109] Note that the description of these effects does not preclude the existence of other effects. Note that one embodiment of the present invention does not necessarily have all of these effects. Note that effects other than these will become apparent from the description in the specification, drawings, claims, etc., and it is possible to extract other effects from the description in the specification, drawings, claims, etc.

[0110] FIG. 1 is a diagram illustrating the configuration of an information processing system according to an embodiment. FIG. 2 is a diagram illustrating the configuration of components used in the information processing system according to an embodiment. FIGS. 3A and 3B are diagrams illustrating the configuration of data used in the information processing system according to an embodiment. FIGS. 4A, 4B, and 4C are diagrams illustrating the configuration of directives used in the information processing system according to an embodiment. FIG. 5 is a diagram illustrating the configuration of data used in the information processing system according to an embodiment. FIGS. 6A, 6B, and 6C are diagrams illustrating the configuration of directives used in the information processing system according to an embodiment. FIG. 7 is a diagram illustrating the configuration of an information processing system according to an embodiment. FIG. 8 is a diagram illustrating the configuration of components used in the information processing system according to an embodiment. FIG. 9 is a diagram illustrating the configuration of directives used in the information processing system according to an embodiment. FIG. 10 is a diagram illustrating the configuration of an information processing device used in the information processing system according to an embodiment. FIG. 11 is a diagram illustrating an information processing method according to an embodiment. FIG. 12 is a diagram illustrating the information processing method according to an embodiment. FIG. 13 is a diagram illustrating the information processing method according to an embodiment. FIG. 14 is a diagram illustrating the information processing method according to an embodiment. FIG. 15 is a sequence diagram illustrating the information processing method according to an embodiment. Fig. 16 is a sequence diagram illustrating an information processing method according to an embodiment. Fig. 17 is a sequence diagram illustrating an information processing method according to an embodiment. Fig. 18 is a sequence diagram illustrating an information processing method according to an embodiment.

[0111] An information processing apparatus according to one aspect of the present invention includes a first component, a second component, and a third component.

[0112] The first component has a function of receiving a list of positive examples and sending it to the third component, the list of positive examples including characteristics of an intended reader and characteristics of a preferred writing style.

[0113] The second component has a function of receiving the first instruction sentence and transmitting the characteristics of the intended reader and the characteristics of the preferred writing style to the third component, and a function of performing processing using a large-scale language model, which has a function of extracting the characteristics of the intended reader and the characteristics of the preferred writing style in accordance with the first instruction sentence.

[0114] The third component has a function of receiving the positive example list, the characteristics of the expected reader, the characteristics of the preferred writing style, and the translation guidelines, and sharing them within the third component. The third component also has a subcomponent, and the subcomponent has a function of creating the first instruction sentence and the second instruction sentence and sending them to the second component.

[0115] The first instruction includes a first instruction and a list of positive examples. The first instruction includes a procedure for extracting characteristics of an intended reader and preferred stylistic features from the list of positive examples. The second instruction includes a procedure for generating second instructions, the characteristics of an intended reader, and preferred stylistic features. The second instruction includes a procedure for generating translation guidelines, the translation guidelines including recommendations, the recommendations including the characteristics of an intended reader and preferred stylistic features.

[0116] This makes it possible to generate translation guidelines from a list of positive examples that include characteristics of the expected reader and preferred stylistic features. Also, it is possible to generate translation guidelines from a list of negative examples that include avoided stylistic features. Also, it is possible to generate translation guidelines that take into account characteristics of the expected reader and preferred stylistic features. Also, it is possible to generate translation guidelines that take into account avoided stylistic features. As a result, it is possible to provide a novel information processing system that is highly convenient, useful, and reliable.

[0117] The embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description, and it will be readily understood by those skilled in the art that various changes in form and details can be made without departing from the spirit and scope of the present invention. Therefore, the present invention should not be interpreted as being limited to the description of the embodiments shown below. In the configuration of the invention described below, the same parts or parts having similar functions will be denoted by the same reference numerals in different drawings, and repeated explanations will be omitted.

[0118] In this specification, ordinal numbers such as "first" and "second" are used to avoid confusion between components and do not limit the number of components or the order of the components (for example, the order of processes or the order of stacking). Furthermore, even if a term does not have an ordinal number in this specification, an ordinal number may be added in the claims to avoid confusion between the components. Even if a term has an ordinal number in this specification, a different ordinal number may be added in the claims. Even if a term has an ordinal number in this specification, the ordinal number may be omitted in the claims.

[0119] In the drawings accompanying this specification, components are classified by function and shown as block diagrams that are independent of each other, but in reality, it is difficult to completely separate components by function, and one component may be involved in multiple functions.

[0120] Embodiment 1 In this embodiment, an information processing system according to one embodiment of the present invention will be described with reference to FIGS.

[0121] FIG. 1 is a diagram illustrating a configuration of an information processing system according to one embodiment of the present invention.

[0122] FIG. 2 is a diagram illustrating the configuration of components used in the information processing system according to the embodiment.

[0123] FIG. 3A is a diagram illustrating the configuration of a positive example list used in the information processing system according to the embodiment, and FIG. 3B is a diagram illustrating the configuration of a negative example list used in the information processing system according to the embodiment.

[0124] 4A to 4C are diagrams illustrating the configuration of instruction statements transmitted and received within an information processing system according to one aspect of the present invention.

[0125] FIG. 5 is a diagram illustrating the configuration of a translation guideline used in an information processing system according to an embodiment of the present invention.

[0126] 6A to 6C are diagrams illustrating the configuration of instruction statements transmitted and received within an information processing system according to one aspect of the present invention.

[0127] FIG. 7 is a diagram illustrating a configuration of an information processing system according to one embodiment of the present invention.

[0128] FIG. 8 is a diagram illustrating the configuration of components used in the information processing system according to the embodiment.

[0129] FIG. 9 is a diagram illustrating the configuration of instruction statements transmitted and received within an information processing system according to an aspect of the present invention.

[0130] FIG. 10 is a block diagram illustrating a configuration of an information processing device that can be used in an information processing system of one embodiment of the present invention.

[0131] <Configuration Example 1 of Information Processing System> The information processing system described in this embodiment includes a component 110, a component 130, and a component 120 (see FIG. 1 ). Note that the information processing device that performs the function of the component 110, the information processing device that performs the function of the component 130, and the information processing device that performs the function of the component 120 each include a calculation device and a communication device. Furthermore, the respective communication devices are connected using, for example, a network 51 to configure an information processing system according to one embodiment of the present invention.

[0132] <Configuration Example 1 of Component 110> The component 110 has a function of receiving a positive example list PEL and transmitting it to the component 120. The positive example list PEL includes expected reader characteristics CIR and preferred stylistic characteristics CPS (see FIG. 3A ). For example, reader attributes can be used as the expected reader characteristics CIR. Specifically, experts, students, children, etc. can be used as the expected reader characteristics CIR. Furthermore, characteristics that readers desire in documents can be used as the expected reader characteristics CIR. Specifically, readers who desire familiarity in documents, readers who desire formal expression in documents, etc. can be used as the expected reader characteristics CIR. Furthermore, colloquial language characteristics, literary language characteristics, familiar language characteristics, formal language characteristics, language characteristics used in academic papers, language characteristics used in patent applications, language characteristics used in legal documents, language characteristics used in novels, language characteristics used in personal letters, or language characteristics used by a particular author can be used as the preferred stylistic characteristics CPS. It should be noted that the stylistic features that can be used for the desirable stylistic features CPS can also be used for the inappropriate stylistic features CAS.

[0133] For example, a user 99 of the information processing system inputs the positive example list PEL to the component 110. Specifically, the user of the information processing system inputs the positive example list PEL to the component 110 using an input device such as a keyboard, a mouse, an eye-gaze input device, or a microphone.

[0134] Furthermore, for example, a user 99 of the information processing system prepares a positive example list when starting translation. Specifically, when starting translation, the user imagines the intended reader and then prepares the positive example list. This allows the characteristics CIR of the intended reader and stylistic characteristics CPS that are appropriate for the characteristics CIR of the intended reader to be collected in the positive example list PEL.

[0135] <Configuration Example 1 of Component 130> The component 130 has a function of receiving a directive Pt11 and transmitting the expected reader characteristics CIR and the preferred stylistic characteristics CPS to the component 120 (see FIG. 1 ). The component 130 also has a function of performing processing using a large-scale language model LLM.

[0136] <<Configuration Example 1 of Large Scale Language Model LLM>> The large scale language model LLM has a function of extracting the expected reader characteristics CIR and the preferred stylistic characteristics CPS from the positive example list PEL in accordance with the directive Pt11.

[0137] <Configuration Example 1 of Component 120> The component 120 has a function of receiving the positive example list PEL, the expected reader characteristics CIR, the preferred stylistic characteristics CPS, and the translation guidelines TGL, and sharing them within the component 120. The component 120 also has a subcomponent 120A.

[0138] <<Configuration Example 1 of Subcomponent 120A>> The subcomponent 120A has a function of creating directive statements Pt11 and Pt2 and transmitting them to the component 130.

[0139] [Configuration Example of Directive Pt11] The directive Pt11 includes a directive g11() and a positive example list PEL (see FIG. 4A).

[0140] Instruction g11( ) includes a procedure for extracting the characteristics of the assumed reader CIR and the characteristics of the preferred writing style CPS from the positive example list PEL.

[0141] For example, the following paragraph of text can be used as directive Pt11.

[0142] "#Writing style characteristics: {CPS_1} Intended reader characteristics: {CIR_1} #Writing style characteristics: {CPS_2} Intended reader characteristics: {CIR_2} Please list the common characteristics between these."

[0143] The step "Please list the common features of these" corresponds to the procedure of extracting the expected reader features CIR and the preferred stylistic features CPS from the positive example list PEL.

[0144] [Configuration Example 1 of Directive Pt2] Directive Pt2 includes directive g2(), characteristics of the intended reader CIR, and preferred stylistic characteristics CPS (see FIG. 4C). Directive g2() also includes a procedure for generating translation guidelines TGL.

[0145] The translation guidelines TGL include recommendations RM (see FIG. 5). The recommendations RM include characteristics CIR of the intended readers and characteristics CPS of the preferred writing style.

[0146] For example, the following paragraph of text can be used as directive Pt2.

[0147] "We will request a translation from a translator under the following conditions. To ensure a highly accurate translation, please create guidelines that clearly state the intended reader and preferred writing style. #intended reader {intended reader characteristics CIR} #preferred writing style {preferred writing style characteristics CPS} #guidelines"

[0148] Note that "We are requesting a translation staff member to translate under the following conditions. To ensure a highly accurate translation, please create guidelines that clearly state the intended reader and preferred writing style." corresponds to instruction g2(). Also, "#intended reader" and "#preferred writing style" are headings, "{intended reader characteristics CIR}" is the part where the intended reader characteristics CIR are inserted, and "{preferred writing style characteristics CPS}" is the part where the preferred writing style characteristics CPS are inserted. Also, "#guideline" is a heading that prompts the large-scale language model LLM to output.

[0149] <Configuration Example 2 of Component 110> The component 110 also has a function of receiving a negative example list NEL from, for example, a user 99 of the information processing system and transmitting it to the component 120. The negative example list NEL includes avoidable stylistic features CAS (see FIG. 3B ).

[0150] Furthermore, for example, a user 99 of the information processing system prepares a negative example list when starting translation. Specifically, when starting translation, the user imagines the intended reader and then prepares the negative example list. In this way, stylistic characteristics CAS that are not suitable for the assumed reader characteristics CIR can be collected in the negative example list NEL.

[0151] <Configuration Example 2 of Component 130> The component 130 has a function of receiving the directive sentence Pt12 and transmitting the avoidable writing style characteristics CAS to the component 120 (see FIG. 1).

[0152] <<Configuration Example 2 of Large Scale Language Model LLM>> The large scale language model LLM has a function of extracting avoidable stylistic features CAS in accordance with the directive sentence Pt12.

[0153] <Configuration Example 2 of Component 120> The component 120 has a function of receiving the negative example list NEL and the avoided writing style features CAS, and sharing them within the component 120.

[0154] <<Configuration Example 2 of Subcomponent 120A>> The subcomponent 120A has a function of creating a directive Pt12 and transmitting it to the component 130.

[0155] [Configuration Example of Directive Pt12] The directive Pt12 includes a directive g12( ) and a negative example list NEL (see FIG. 4B). The directive g12( ) includes a procedure for extracting avoidable stylistic features CAS from the negative example list NEL.

[0156] For example, the following paragraph of text can be used as directive Pt12.

[0157] "#Writing style characteristics: {CAS_1} Intended reader characteristics: {CIR_1} #Writing style characteristics: {CAS_2} Intended reader characteristics: {CIR_2} Please list the common characteristics between these."

[0158] The step "Please list the common characteristics of these" corresponds to the procedure of extracting the characteristics CIR of the assumed reader and the characteristics CPS of the avoidable writing style from the negative example list NEL.

[0159] [Configuration Example 2 of Directive Pt2] Directive Pt2 includes directive g2(), expected reader characteristics CIR, and avoided stylistic characteristics CAS (see FIG. 4C). Directive g2() also includes a procedure for generating translation guidelines TGL.

[0160] The translation guidelines TGL include avoidable items AM, which include avoidable stylistic features CAS (see FIG. 5).

[0161] For example, the following paragraph of text can be used as directive Pt2.

[0162] "We will ask our translation staff to avoid the following translations. To ensure that translations are of high accuracy, please create guidelines for avoidable writing styles. #Avoidable Writing Styles {Characteristics of Avoidable Writing Styles CAS} #Guidelines"

[0163] Note that "We will ask the translator not to translate the following. Please create guidelines for avoidable writing styles to ensure high-accuracy translations." corresponds to instruction g2(). Also, "#Avoided writing styles" is a heading, and "{Avoided writing style characteristics CAS}" is the part where the avoided writing style characteristics CAS are inserted. Also, "#Guideline" is a heading that prompts the large-scale language model LLM to output.

[0164] This makes it possible to generate translation guidelines TGL from a positive example list PEL that includes expected reader characteristics CIR and preferred stylistic features CPS. Furthermore, it is possible to generate translation guidelines TGL from a negative example list NEL that includes avoided stylistic features CAS. Furthermore, it is possible to generate translation guidelines TGL that take into account expected reader characteristics CIR and preferred stylistic features CPS. Furthermore, it is possible to generate translation guidelines TGL that take into account avoided stylistic features CAS. As a result, it is possible to provide a novel information processing system that is highly convenient, useful, and reliable.

[0165] <Configuration Example 3 of Component 110> The component 110 has a function of receiving an original ODoc from a user 99 of the information processing system and transmitting it to the component 120, for example.

[0166] The component 110 also has a function of receiving, for example, a translation document TDoc from the component 120 and providing it to, for example, a user 99 of the information processing system. Specifically, the component 110 provides it to the user 99 of the information processing system using an output device such as a display device, speaker, printer, or storage device.

[0167] <Configuration Example 3 of Component 130> The component 130 has a function of receiving the instruction statement Pt3 and transmitting the translation document TDoc to the component 120.

[0168] <<Configuration Example 3 of Large Scale Language Model LLM>> The large scale language model LLM has a function of generating a translation document TDoc in accordance with a directive Pt3.

[0169] <Configuration Example 3 of Component 120> Component 120 has a function to accept an original ODoc and share it within component 120, and a function to accept a translated document TDoc and send it to component 110, and subcomponent 120A has a function to create instruction text Pt3.

[0170] [Configuration Example of Directive Statement Pt3] Directive statement Pt3 includes instruction g3(), translation guidelines TGL, and original ODoc (see FIG. 6A).

[0171] Instruction g3() includes a procedure for generating a translation document TDoc from the original ODoc in accordance with the translation guidelines TGL. Specifically, by referring to the recommendations RM and the avoided items AM, the translation document TDoc is generated from the original ODoc in a style that includes desirable stylistic features CPS for readers with the expected reader characteristics CIR. The translation document TDoc is also generated from the original ODoc so as not to include avoided stylistic features CAS.

[0172] For example, the following paragraph of text can be used for directive Pt3.

[0173] Please adhere to the following guidelines and translate accurately from the original text. #Guidelines {Translation Guidelines TGL} #Original {Original ODoc}

[0174] Note that "Please observe the following guidelines and translate accurately from the original text" corresponds to instruction g3(). Also, "#Guidelines" and "#Original" are headings. Also, "{Translation Guidelines TGL}" is the part where the translation guidelines TGL are inserted, and "{Original ODoc}" is the part where the original ODoc is inserted.

[0175] <Configuration Example 4 of Component 130> The component 130 has a function of receiving the instruction Pt4 and transmitting the inappropriate list IL to the component 120. The component 130 also has a function of receiving the instruction Pt5 and transmitting the translated document TDoc to the component 120.

[0176] <<Configuration Example 4 of Large Scale Language Model LLM>> The large scale language model LLM has a function to generate an inappropriate list IL according to a directive Pt4. The large scale language model LLM also has a function to generate a translated document TDoc according to a directive Pt5.

[0177] <<Configuration Example 3 of Subcomponent 120A>> The subcomponent 120A has a function of creating directives Pt4 and Pt5.

[0178] [Configuration Example of Directive Statement Pt4] The directive statement Pt4 includes a directive g4(), a translation guideline TGL, an original ODoc, and a translation document TDoc (see FIG. 6B).

[0179] Instruction g4() includes a procedure for evaluating whether the translation from the original ODoc to the translated document TDoc conforms to the translation guidelines TGL, outputting true if the translation conforms, and generating an inappropriate list IL if the translation does not conform. The inappropriate list IL includes the original text and translation that are evaluated as not conforming.

[0180] For example, the following paragraph of text can be used for directive Pt4.

[0181] "If the translated document complies with the translation guidelines, please state "TRUE"; if not, please list the issues you would like to address. #Translation guidelines {Translation guidelines TGL} #Original {Original ODoc} #Translated document {Translated document TDoc}"

[0182] Note that "If the translated document complies with the translation guidelines, please state "true"; if not, please list the issues to be addressed." corresponds to instruction g4(). Also, "#guidelines," "#original," and "#translated document" are headings. Also, "{translation guidelines TGL}" is the part where the translation guidelines TGL are inserted, "{original ODoc}" is the part where the original ODoc is inserted, and "{translated document TDoc}" is the part where the translated document TDoc is inserted.

[0183] [Configuration Example of Directive Pt5] The directive Pt5 includes a directive g5(), a translation guideline TGL, an original ODoc, a translated document TDoc, and an inappropriate list IL (see FIG. 6C).

[0184] The instruction g5( ) includes a procedure for generating a translation document TDoc from the original ODoc in accordance with the translation guideline TGL so as not to translate the content listed in the inappropriate list IL.

[0185] For example, the following paragraph of text can be used as directive Pt5.

[0186] Please adhere to the following guidelines and translate accurately from the original text. Also, please translate in a way that avoids the following issues: #Guidelines {Translation Guidelines TGL} #Original {Original ODoc} #Improper Translation {Inappropriate List IL}

[0187] Note that "Please observe the following guidelines and translate accurately based on the original text. Also, please translate in a way that avoids the following points being raised" corresponds to instruction g5(). Also, "#guidelines," "#original," and "#points raised" are headings. Also, "{translation guidelines TGL}" is the part where the translation guidelines TGL are inserted, "{original ODoc}" is the part where the original ODoc is inserted, and "{inappropriate list IL}" is the part where the inappropriate list IL is inserted.

[0188] This makes it possible to generate a translated document TDoc from an original ODoc in accordance with the translation guidelines TGL. It is also possible to generate a translated document TDoc that takes into account the expected reader characteristics CIR and the preferred stylistic characteristics CPS. It is also possible to generate a translated document TDoc that does not include the avoidable stylistic characteristics CAS. It is also possible to proofread the translated document TDoc based on the translation guidelines TGL. As a result, it is possible to provide a novel information processing system that is highly convenient, useful, and reliable.

[0189] <Configuration Example 2 of Information Processing System> The information processing system described in this embodiment includes a component 110, a component 130, and a component 120 (see FIG. 7).

[0190] <Configuration Example 4 of Component 110> The component 110 has a function of receiving a question Q from the component 120 and providing the question Q to, for example, a user 99 of the information processing system. The component 110 also has a function of receiving an answer Ans from, for example, the user 99 of the information processing system and transmitting the answer Ans to the component 120.

[0191] <Configuration Example 5 of Component 130> The component 130 has a function of receiving the directive sentence Pt6 and transmitting the stylistic feature CS1 and the stylistic feature CS2 to the component 120 (see FIG. 1 ). The component 130 also has a function of performing processing using the large-scale language model LLM. Note that the stylistic feature CS1 is a stylistic feature of the sentences written in document Doc1, and the stylistic feature CS2 is a stylistic feature of the sentences written in document Doc2.

[0192] <<Configuration Example 5 of Large Scale Language Model LLM>> The large scale language model LLM has a function of extracting the stylistic feature CS1 and the stylistic feature CS2 in accordance with the directive sentence Pt6.

[0193] <Configuration Example 4 of Component 120> The component 120 has a function of receiving an answer Ans and sharing it within the component 120. The component 120 also has subcomponents 120A, 120B, and 120C.

[0194] <<Configuration Example 4 of Subcomponent 120A>> The subcomponent 120A has a function of creating a directive Pt6 and transmitting it to the component 130.

[0195] [Configuration Example of Directive Statement Pt6] Directive statement Pt6 includes instruction g6(), document Doc1, and document Doc2 (see FIG. 9).

[0196] Instruction g6( ) includes a procedure for extracting and listing stylistic features CS1 from document Doc1, and a procedure for extracting and listing stylistic features CS2 from document Doc2.

[0197] For example, the following paragraph of text can be used for directive Pt6.

[0198] "#Document A {Document Doc1} #Document B {Document Doc2} Please concisely list the characteristics of the text of Document A and Document B. Please do not mention the content. Also, please do not include phrases like Document A and Document B in the characteristics."

[0199] Note that "#DocumentA" and "#DocumentB" are headings. Also, {Document Doc1} is the part where Document Doc1 is inserted, and {Document Doc2} is the part where Document Doc2 is inserted. Also, "Please concisely list the characteristics of the text of Document A and Document B. Please do not mention the content. Also, please do not include wording like Document A and Document B in the characteristics" is the part corresponding to instruction Pt6.

[0200] <Configuration Example 1 of Subcomponent 120B> The subcomponent 120B has a function of selecting two documents from n typical documents Docs and creating a combination. The subcomponent 120B also has a function of designating one document of the combination as document Doc1 and the other as document Doc2.

[0201] For example, typical documents Docs may include academic papers, patent applications, legal documents, novels, or personal letters.

[0202] <<Configuration Example 1 of Subcomponent 120C>> The subcomponent 120C has a function of creating a question Q. Note that the question Q asks whether the stylistic feature CS1 or the stylistic feature CS2 is a more suitable style for the expected reader characteristic CIR.

[0203] For example, the following paragraph of text can be used for question Q:

[0204] "Which style, {stylistic feature CS1} or {stylistic feature CS2}, is more suitable for a reader with {intended reader characteristics CIR}?"

[0205] Note that {stylistic feature CS1} is the part where stylistic feature CS1 is inserted, and {stylistic feature CS2} is the part where stylistic feature CS2 is inserted. Also, {intended reader's feature CIR} is the part where the intended reader's feature CIR is inserted. Specifically, when assuming children as readers, a question Q can be used to ask whether a friendly style or a formal style is more appropriate.

[0206] <<Configuration Example 2 of Subcomponent 120C>> For example, when the answer Ans is a stylistic feature CS1, subcomponent 120C adds the expected reader feature CIR and the stylistic feature CS1 to the positive example list PEL (see FIG. 3A ). The stylistic feature CS1 is added as a preferred stylistic feature CPS. Specifically, when targeting children as the reader, if the stylistic feature CS1 is a friendly style and the stylistic feature CS2 is a formal style, subcomponent 120C can create a new record by storing “children” in the field storing the expected reader feature CIR and storing the stylistic feature CS1 in the field storing the preferred stylistic feature CPS. Subcomponent 120C can then add the record to the positive example list PEL.

[0207] Subcomponent 120C also adds the expected reader characteristic CIR and the stylistic characteristic CS2 to the negative example list NEL (see FIG. 3B ). Note that stylistic characteristic CS2 is added as an avoidable stylistic characteristic CAS. Specifically, when assuming children as the reader, if stylistic characteristic CS1 indicates a friendly style and stylistic characteristic CS2 indicates a formal style, subcomponent 120C can create a new record by storing "children" in the field storing the expected reader characteristic CIR and storing stylistic characteristic CS2 in the field storing the avoidable stylistic characteristic CAS. Subcomponent 120C can also add the record to the negative example list NEL.

[0208] As a result, for example, stylistic feature CS1 can be added to the positive example list PEL as a writing style that is appropriate for the expected reader characteristic CIR. Furthermore, for example, stylistic feature CS2 can be added to the negative example list NEL as a writing style that is not appropriate for the expected reader characteristic CIR. Furthermore, using the large-scale language model LLM, the stylistic features of document Doc1 can be extracted and listed as stylistic feature CS1. Furthermore, the stylistic features of document Doc2 can be extracted and listed as stylistic feature CS2. Furthermore, by comparing stylistic feature CS1 and stylistic feature CS2, a user of the information processing system can evaluate whether the writing style is appropriate for the expected reader characteristic CIR. As a result, a novel information processing system that is highly convenient, useful, and reliable can be provided.

[0209] <<Configuration Example 2 of Subcomponent 120B>> The subcomponent 120B has a function of selecting two documents from n typical documents Docs and creating all unique combinations. The number of combinations created by the subcomponent 120B is n × (n - 1) ÷ 2.

[0210] Furthermore, the n typical documents Docs are classified into different clusters, and the n typical documents Docs have embedding representations that are closest to the centroids of the respective clusters.

[0211] For example, a document set containing n or more documents is classified into n clusters. From each cluster, a document closest to the center of gravity of the cluster can be extracted and used as n typical documents Docs.

[0212] <<Configuration Example 3 of Subcomponent 120C>> The subcomponent 120C has a function of generating questions Q for all combinations.

[0213] As a result, for example, a stylistic feature CS1 can be added to the positive example list PEL as a style that is appropriate for the expected reader characteristics CIR. Furthermore, for example, a stylistic feature CS2 can be added to the negative example list NEL as a style that is not appropriate for the expected reader characteristics CIR. Furthermore, the positive example list PEL and the negative example list NEL can be created using n typical documents Docs classified into different clusters. As a result, a novel information processing system that is highly convenient, useful, and reliable can be provided.

[0214] <Configuration Example 3 of Information Processing System> The information processing system described in this embodiment includes a component 140 (see FIG. 7).

[0215] <Configuration Example of Component 140> The component 140 has a function of receiving the original ODoc and transmitting the embedded representation EE to the component 120, and a function of performing processing using the embedded model EM.

[0216] <<Configuration Example of Embedding Model EM>> The embedding model EM has a function of converting n+1 typical documents Docs and original documents ODoc into embedded representations. For example, natural language processing models such as BERT (Bidirectional Encoder Representations from Transformers) and Word2vec (Word to vector) can be used for the embedding model EM. Specifically, a multilingual text embedding model such as Multilingual-E5 can be used for the embedding model EM.

[0217] <Configuration Example 5 of Component 110> The component 110 has a function of receiving an original ODoc from a user 99 of the information processing system and transmitting it to the component 120, for example.

[0218] <Configuration Example 5 of Component 120> The component 120 has a function of receiving an original ODoc and transmitting it to the component 140. The component 120 also has a function of receiving an embedded expression EE and sharing it within the component 120.

[0219] <Configuration Example 3 of Subcomponent 120B> The subcomponent 120B has a function of selecting two documents from n+1 typical documents Docs and creating all unique combinations. The number of combinations created by the subcomponent 120B is (n+1) × n ÷ 2.

[0220] Furthermore, n-1 of the n+1 typical documents Docs are classified into different clusters, and have the embedded representations closest to the centers of gravity of their respective clusters. Each of these documents is also classified into a different cluster from the original ODoc.

[0221] Two of the n+1 typical documents Docs are classified into the same cluster as the original ODoc, and are provided with embedded representations selected in order of proximity to the center of gravity of the cluster.

[0222] For example, a document set containing n or more documents is classified into n clusters. The cluster to which the original ODoc belongs is removed from the n clusters, leaving (n-1) clusters, from which the documents closest to the center of gravity of the clusters are extracted. Furthermore, two documents from the same cluster as the original ODoc are selected in descending order of their proximity to the center of gravity of the cluster, and a total of n+1 documents can be used as typical document Docs.

[0223] <<Configuration Example 4 of Subcomponent 120C>> The subcomponent 120C has a function of generating questions Q for all combinations.

[0224] As a result, for example, a stylistic feature CS1 can be added to the positive example list PEL as a writing style that is appropriate for the expected reader characteristics CIR. Furthermore, for example, a stylistic feature CS2 can be added to the negative example list NEL as a writing style that is not appropriate for the expected reader characteristics CIR. Furthermore, the positive example list PEL and the negative example list NEL can be created using n-1 typical documents Docs that are classified into a different cluster from the original ODoc. Furthermore, by using two typical documents Docs that are classified into the same cluster as the original ODoc, for example, a stylistic feature CS1 can be added to the positive example list PEL in relation to minor differences in stylistic features between similar documents. Furthermore, for example, a stylistic feature CS2 can be added to the negative example list NEL. As a result, a novel information processing system that is excellent in convenience, usefulness, and reliability can be provided.

[0225] <Configuration Example 4 of Information Processing System> The information processing system described in this embodiment includes a component 110, a component 120, and a component 130 (see FIG. 7).

[0226] For example, an information processing system according to an embodiment of the present invention can be configured with an information processing device that performs the functions of component 110, an information processing device that performs the functions of component 120, and an information processing device that performs the functions of component 130. Note that the number of information processing devices that configure the information processing system according to an embodiment of the present invention is one or more. Furthermore, for example, the information processing system according to an embodiment of the present invention can be configured by connecting a plurality of information processing devices using a network 51.

[0227] When an information processing system according to one embodiment of the present invention is configured using a plurality of information processing devices, the load related to information processing can be distributed.

[0228] <Configuration Example 1 of Information Processing Device> Configuration Example 1 of the information processing device described in this embodiment can be used for the component 110. Configuration Example 1 of the information processing device can also be called a client computer, etc. For example, a desktop computer can be used for the component 110.

[0229] The information processing device according to the first exemplary configuration can accept data input by a user of the information processing system according to an embodiment of the present invention. The information processing device according to the first exemplary configuration can also provide the user with data output by the information processing system according to an embodiment of the present invention.

[0230] For example, dedicated application software, a web browser, etc., run on the component 110. A user of the information processing system according to an embodiment of the present invention can access the information processing system via either of these components, thereby enjoying services using the information processing system according to an embodiment of the present invention.

[0231] <Configuration Example 2 of Information Processing Apparatus> Configuration example 2 of the information processing apparatus described in this embodiment can be used for the component 120. For example, the component 120 can be a workstation, a server computer, a supercomputer, or the like.

[0232] Moreover, it is preferable that the information processing device in configuration example 2 has a function as a parallel computer. By using it as a parallel computer, it is possible to perform large-scale calculations necessary for learning and inference of artificial intelligence (AI), for example.

[0233] Furthermore, configuration example 2 of the information processing device can perform processing using a natural language model that uses AI.

[0234] For example, it is preferable to be able to perform processing using natural language models such as GPT-3 (registered trademark), GPT-3.5, GPT-4 (registered trademark), LaMDA, Llama2, and Llama3.

[0235] <Configuration Example 3 of Information Processing Device> For example, configuration example 3 of the information processing device described in this embodiment can be used for the component 130. Note that the component 130 is larger in scale and has higher computing power than the component 120. For example, a large computer such as a server computer or a supercomputer can be used for the component 130.

[0236] Moreover, it is preferable that the information processing device in configuration example 3 has a function as a parallel computer. By using it as a parallel computer, it is possible to perform large-scale calculations necessary for AI learning and inference, for example.

[0237] Furthermore, the information processing device according to the third exemplary configuration can perform processing using a natural language model that uses AI. In particular, the information processing device can perform processing using a general-purpose language model that can perform various natural language processing tasks.

[0238] For example, it is possible to perform processing using natural language models such as GPT-3 (registered trademark), GPT-3.5, GPT-4 (registered trademark), LaMDA, Llama2, and Llama3. In particular, it is preferable to be able to perform processing using GPT-4 (registered trademark). For example, if it is possible to perform processing using a large-scale language model that is larger than conventional natural language models, it is possible to realize more natural document generation or dialogue.

[0239] <Configuration Example 4 of Information Processing Apparatus> For example, configuration example 4 of the information processing apparatus described in this embodiment can be used for the component 140. For example, a large computer such as a server computer or a supercomputer can be used for the component 140.

[0240] Moreover, it is preferable that the information processing device in configuration example 2 has a function as a parallel computer. By using it as a parallel computer, it is possible to perform large-scale calculations necessary for learning and inference of artificial intelligence (AI), for example.

[0241] Furthermore, the configuration example 4 of the information processing device can perform processing using a natural language model that uses AI.

[0242] For example, it is possible to execute processing using natural language processing models such as BERT (Bidirectional Encoder Representations from Transformers) and Word2vec (Word to vector). Specifically, it is preferable to execute processing using a multilingual text embedding model such as Multilingual-E5.

[0243] Note that a person who provides a service using an information processing system according to one embodiment of the present invention does not necessarily have to own the information processing device of Configuration Example 3. For example, a service provider can use part of a service provided by another business or the like using the information processing device of Configuration Example 3.

[0244] <Configuration Example of Network 51> The network 51 that can be used in the information processing system of one embodiment of the present invention can connect multiple information processing devices. This allows the connected multiple information processing devices to transmit and receive data to and from each other. In addition, the load related to information processing can be distributed.

[0245] When wireless communication is performed, communication standards such as the fourth generation mobile communication system (4G), fifth generation mobile communication system (5G), and sixth generation mobile communication system (6G), or specifications standardized by the IEEE such as Wi-Fi (registered trademark) and Bluetooth (registered trademark), can be used as communication protocols or communication technologies.

[0246] For example, a local network can be used for the network 51. Also, an intranet or an extranet can be used for the network 51. Also, a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a global area network (GAN), etc. can be used for the network 51.

[0247] Furthermore, for example, a global network can be used for the network 51. Specifically, the Internet, which is the foundation of the World Wide Web (WWW), can be used.

[0248] Furthermore, a person who provides a service using the information processing system according to one embodiment of the present invention can provide the service using the information processing method according to one embodiment of the present invention via the network 51, for example.

[0249] When the information processing system according to an embodiment of the present invention is built within a local network, the possibility of confidential information leaking can be reduced, for example, compared to when the Internet is used.

[0250] <Configuration Example 5 of Information Processing Device> An information processing device 20 that can be used in an information processing system of one embodiment of the present invention has, for example, an input unit 21, a memory unit 22, a processing unit 23, an output unit 24, and a transmission path 25 (see Figure 10).

[0251] In the drawings attached to this specification, the components are classified by function and shown as independent blocks in the block diagrams. However, in reality, it is difficult to completely separate the components by function, and one component may be involved in multiple functions. For example, part of the processing unit 23 may function as the input unit 21. Also, one function may be involved in multiple components. For example, the processing performed by the processing unit 23 may be executed by different information processing devices depending on the processing.

[0252] <<Input Unit 21>> The input unit 21 can receive data from outside the information processing device. For example, the input unit 21 receives data via a network 51. Specifically, a device such as a personal computer equipped with a communication port or a communication function can be used.

[0253] The input unit 21 supplies the received data to one or both of the storage unit 22 and the processing unit 23 via a transmission path 25 .

[0254] <<Storage Unit 22>> The storage unit 22 has a function of storing a program executed by the processing unit 23. The storage unit 22 can also have a function of storing data generated by the processing unit 23 (e.g., calculation results, analysis results, inference results), data accepted by the input unit 21, etc.

[0255] The storage unit 22 may have a database. Furthermore, the information processing device may have a database separate from the storage unit 22. The information processing device may have a function to retrieve data from a database that exists outside the storage unit 22, outside the information processing device, or outside the information processing system. Furthermore, the information processing device may have a function to retrieve data from both its own database and an external database.

[0256] Either or both of a storage and a file server can be used as the memory unit 22. Also, a database that records paths of files stored in a file server can be used as the memory unit 22.

[0257] The storage unit 22 includes at least one of a volatile memory and a non-volatile memory. Examples of the volatile memory include a dynamic random access memory (DRAM) and a static random access memory (SRAM). Examples of the non-volatile memory include a resistive random access memory (ReRAM), a phase change random access memory (PRAM), a ferroelectric random access memory (FeRAM), a magnetoresistive random access memory (MRAM), and a flash memory. The storage unit 22 may include at least one of NOSRAM (registered trademark) and DOSRAM (registered trademark). The storage unit 22 may include a recording media drive. Examples of the recording media drive include a hard disk drive (HDD) and a solid state drive (SSD).

[0258] NOSRAM is an abbreviation for "Nonvolatile Oxide Semiconductor Random Access Memory (RAM)." NOSRAM refers to a memory in which memory cells are two-transistor (2T) or three-transistor (3T) gain cells and transistors (also called OS transistors) that use metal oxide in their channel formation regions. OS transistors have an extremely small leakage current, i.e., a current that flows between the source and drain in an off state. NOSRAM can be used as a nonvolatile memory by retaining a charge corresponding to data in the memory cell using its extremely small leakage current characteristic. In particular, NOSRAM can read stored data without destroying it (nondestructive readout), making it suitable for arithmetic processing in which only data read operations are repeated a large number of times. NOSRAM can increase its data capacity by stacking layers, and therefore can be used as a large-scale cache memory, main memory, or storage memory to improve the performance of semiconductor devices.

[0259] DOSRAM is an abbreviation for "Dynamic Oxide Semiconductor RAM" and refers to a RAM having 1T (transistor) 1C (capacitor) type memory cells. DOSRAM is a DRAM formed using OS transistors, and is a memory that temporarily stores information sent from an external device. DOSRAM is a memory that takes advantage of the low off-state current of OS transistors.

[0260] In this specification and the like, a metal oxide refers to an oxide of a metal in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), oxide semiconductors (also referred to as oxide semiconductors or simply as OSs), and the like. For example, when a metal oxide is used for a semiconductor layer of a transistor, the metal oxide may be referred to as an oxide semiconductor.

[0261] The metal oxide contained in the channel formation region preferably contains indium (In). When the metal oxide contained in the channel formation region contains indium, the carrier mobility (electron mobility) of the OS transistor is increased. For example, indium oxide (InOx) or indium gallium zinc oxide (In—Ga—Zn oxide, also referred to as “IGZO”) can be used for the channel formation region. The metal oxide contained in the channel formation region is preferably an oxide semiconductor containing an element M. The element M is preferably at least one of aluminum (Al), gallium (Ga), and tin (Sn). Other elements that can be used for the element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), and tungsten (W). However, the element M may be a combination of two or more of the above elements. The element M is, for example, an element having a high bond energy with oxygen. For example, the element M is an element having a higher bond energy with oxygen than indium. Furthermore, the metal oxide contained in the channel formation region is preferably a metal oxide containing zinc (Zn). Metal oxides containing zinc may be more likely to crystallize.

[0262] The metal oxide contained in the channel formation region is not limited to a metal oxide containing indium, but may be, for example, a metal oxide containing zinc but not indium, such as zinc tin oxide or gallium tin oxide, a metal oxide containing gallium, or a metal oxide containing tin.

[0263] <<Processing Unit 23>> The processing unit 23 has a function of performing processes such as calculation, analysis, and inference using data supplied from one or both of the input unit 21 and the storage unit 22. The processing unit 23 can supply generated data (e.g., calculation results, analysis results, and inference results) to one or both of the storage unit 22 and the output unit 24.

[0264] The processing unit 23 has a function of acquiring data from the storage unit 22. The processing unit 23 can also have a function of recording or registering data in the storage unit 22.

[0265] The processing unit 23 may include, for example, an arithmetic circuit, a central processing unit (CPU), a graphics processing unit (GPU), or a neural processing unit / neural network processing unit (NPU).

[0266] The processing unit 23 may include a microprocessor such as a DSP (Digital Signal Processor). The microprocessor may be implemented by a PLD (Programmable Logic Device) such as an FPGA (Field Programmable Gate Array) or an FPAA (Field Programmable Analog Array). The processing unit 23 may also include a quantum processor. The processing unit 23 can perform various data processing and program control by interpreting and executing instructions from various programs using the processor. Programs that can be executed by the processor are stored in at least one of the memory area of ​​the processor and the storage unit 22.

[0267] The processing unit 23 may include a main memory. The main memory may include at least one of a volatile memory such as a RAM and a non-volatile memory such as a ROM (Read Only Memory). The main memory may also include at least one of the above-mentioned NOSRAM and DOSRAM.

[0268] The RAM may be, for example, a DRAM or an SRAM, and a virtual memory space is allocated and used as a working space for the processing unit 23. The operating system, application programs, program modules, program data, lookup tables, etc. stored in the storage unit 22 are loaded into the RAM for execution. The data, programs, and program modules loaded into the RAM are each directly accessed and operated by the processing unit 23.

[0269] The ROM can store a BIOS (Basic Input / Output System), firmware, etc., which do not require rewriting. Examples of ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), and EPROM (Erasable Programmable Read Only Memory). Examples of EPROMs include UV-EPROMs (Ultra-Violet Erasable Programmable Read Only Memories), which allow stored data to be erased by exposure to ultraviolet light, EEPROMs (Electrically Erasable Programmable Read Only Memories), and flash memories.

[0270] The processing section 23 can include one or both of an OS transistor and a transistor having silicon in a channel formation region (a Si transistor).

[0271] The processing unit 23 preferably includes an OS transistor. Because an OS transistor has an extremely small off-state current, using the OS transistor as a switch for retaining charge (data) flowing into a capacitor functioning as a memory element can ensure a long data retention period. By using this characteristic in at least one of the register and cache memory of the processing unit, the processing unit can be operated only when necessary, and can be turned off in other cases by saving information from the previous processing in the memory element. In other words, normally-off computing is possible, and the power consumption of the information processing system can be reduced.

[0272] It is preferable that the information processing device uses AI for at least some of its processing.

[0273] It is particularly preferable that the information processing device uses an artificial neural network (ANN, hereinafter simply referred to as a neural network). A neural network is realized by a circuit (hardware) or a program (software).

[0274] In this specification, a neural network refers to a general model that mimics the neural circuit network of a living organism, determines the connection strength between neurons through learning, and has problem-solving capabilities. A neural network has an input layer, an intermediate layer (hidden layer), and an output layer.

[0275] In this specification and the like, when discussing neural networks, determining the connection strengths (also called weighting coefficients) between neurons from existing information may be referred to as "learning."

[0276] In this specification and the like, the act of constructing a neural network using connection strengths obtained by learning and deriving a new conclusion from it may be referred to as "inference."

[0277] <<Output Unit 24>> The output unit 24 can output at least one of the calculation results, analysis results, and inference results in the processing unit 23 to the outside of the information processing device. For example, the output unit 24 can transmit data via the network 51. Specifically, a device such as a personal computer equipped with a communication port or a communication function can be used. Furthermore, a device equipped with a communication function may be used for the input unit 21 and the output unit 24.

[0278] <<Transmission Path 25>> The transmission path 25 has a function of transmitting data. Data can be transmitted and received between the input unit 21, the storage unit 22, the processing unit 23, and the output unit 24 via the transmission path 25. Specifically, a LAN or the Internet can be used.

[0279] Note that this embodiment mode can be appropriately combined with other embodiment modes described in this specification.

[0280] Embodiment 2 In this embodiment, an information processing method according to one embodiment of the present invention will be described with reference to FIGS.

[0281] FIG. 11 is a flowchart illustrating an information processing method according to one embodiment of the present invention.

[0282] FIG. 12 is a flowchart illustrating an information processing method according to one embodiment of the present invention.

[0283] FIG. 13 is a flowchart illustrating an information processing method according to one embodiment of the present invention.

[0284] FIG. 14 is a flowchart illustrating an information processing method according to one embodiment of the present invention.

[0285] FIG. 15 is a sequence diagram illustrating an information processing method according to one embodiment of the present invention.

[0286] FIG. 16 is a sequence diagram illustrating an information processing method according to one embodiment of the present invention.

[0287] FIG. 17 is a sequence diagram illustrating an information processing method according to one embodiment of the present invention.

[0288] FIG. 18 is a sequence diagram illustrating an information processing method according to one embodiment of the present invention.

[0289] <Example 1 of Information Processing Method> An information processing method according to one embodiment of the present invention is an information processing method having a phase PH1 (see FIG. 11).

[0290] <Example of Phase PH1> Phase PH1 includes steps S1 to S11.

[0291] <<Step S1>> In step S1 of phase PH1, the component 110 accepts the positive example list PEL and the negative example list NEL and transmits them to the component 120. For example, a user of the information processing system inputs the positive example list PEL and the negative example list NEL. Note that step S1 corresponds to the arrows extending from (1) and (2) in FIG. 15 .

[0292] <Step S2> In step S2 of phase PH1, the component 120 receives the positive example list PEL and the negative example list NEL and shares them within the component 120. The component 120 includes a subcomponent 120A.

[0293] <Step S3> In step S3 of phase PH1, the subcomponent 120A creates an instruction statement Pt11 and sends it to the component 130. Note that step S3 corresponds to the arrows extending from (3) and (4) in FIG. 15 .

[0294] The directive Pt11 includes a directive g11() and a positive example list PEL. The directive g11() includes a procedure for extracting the expected reader characteristics CIR and the preferred stylistic characteristics CPS from the positive example list PEL.

[0295] <Step S4> In step S4 of phase PH1, the component 130 receives the directive Pt11 and extracts the expected reader characteristics CIR and the preferred writing style characteristics CPS using the large-scale language model LLM.

[0296] Step S5 In step S5 of phase PH1, the component 130 transmits the expected reader characteristics CIR and the preferred stylistic characteristics CPS to the component 120. Note that step S5 corresponds to the arrow extending from (5) in FIG. 15 .

[0297] <Step S6> In step S6 of phase PH1, the subcomponent 120A creates a directive Pt12 and sends it to the component 130. Note that step S6 corresponds to the arrows extending from (6) and (7) in FIG. 15 .

[0298] The directive Pt12 includes a directive g12( ) and a negative example list NEL. The directive g12( ) includes a procedure for extracting avoidable stylistic features CAS from the negative example list NEL.

[0299] <Step S7> In step S7 of phase PH1, the component 130 receives the directive sentence Pt12 and extracts avoidable stylistic features CAS using the large-scale language model LLM.

[0300] <Step S8> In step S8 of phase PH1, the component 130 transmits the avoidable stylistic features CAS to the component 120. Note that step S8 corresponds to the arrow extending from (8) in FIG.

[0301] <Step S9> In step S9 of phase PH1, the subcomponent 120A creates a directive Pt2 and sends it to the component 130. Note that step S9 corresponds to the arrows extending from (9) and (10) in FIG. 15 .

[0302] The instruction Pt2 includes an instruction g2(), expected reader characteristics CIR, preferred stylistic features CPS, and avoided stylistic features CAS. The instruction g2() includes a procedure for generating a translation guideline TGL, and the translation guideline TGL includes a recommendation RM and avoided items AM. The recommendation RM includes expected reader characteristics CIR and preferred stylistic features CPS, and the avoided items AM include avoided stylistic features CAS.

[0303] <Step S10> In step S10 of phase PH1, the component 130 receives the directive Pt2 and generates a translation guideline TGL using the large-scale language model LLM.

[0304] <<Step S11>> In step S11 of phase PH1, component 130 transmits the translation guidelines TGL to component 120, thereby terminating phase PH1. Note that step S11 corresponds to the arrow extending from (11) in Fig. 15. Also, as indicated by the arrow extending from (12), component 120 can transmit the translation guidelines TGL to component 110 and provide them to, for example, users of the information processing system.

[0305] This makes it possible to generate translation guidelines TGL from a positive example list PEL that includes expected reader characteristics CIR and preferred stylistic features CPS. Furthermore, it is possible to generate translation guidelines TGL from a negative example list NEL that includes avoided stylistic features CAS. Furthermore, it is possible to generate translation guidelines TGL that take into account expected reader characteristics CIR and preferred stylistic features CPS. Furthermore, it is possible to generate translation guidelines TGL that take into account avoided stylistic features CAS. As a result, it is possible to provide a novel information processing method that is highly convenient, useful, and reliable.

[0306] <Example 2 of Information Processing Method> An information processing method according to one aspect of the present invention is an information processing method having a phase PH1 and a phase PH2 (see FIG. 12).

[0307] <Example of Phase PH2> Phase PH2 follows phase PH1 and includes steps S1 to S12.

[0308] Step S1: In step S1 of phase PH2, the component 110 receives an original ODoc and transmits it to the component 120. For example, a user of the information processing system inputs the original ODoc. Step S1 corresponds to the arrows extending from (1) and (2) in FIG. 16 .

[0309] <Step S2> In step S2 of phase PH2, the component 120 receives the original ODoc and shares it within the component 120.

[0310] <Step S3> In step S3 of phase PH2, the subcomponent 120A creates a directive Pt3 and sends it to the component 130. Note that step S3 corresponds to the arrows extending from (3) and (4) in FIG. 16 .

[0311] The instruction Pt3 includes an instruction g3(), a translation guideline TGL, and an original ODoc. The instruction g3() includes a procedure for generating a translation document TDoc from the original ODoc in accordance with the translation guideline TGL, in a style that includes desirable stylistic features CPS and does not include undesirable stylistic features CAS, for readers with the expected reader characteristics CIR.

[0312] <Step S4> In step S4 of phase PH2, the component 130 receives the directive Pt3 and generates a translation document TDoc using the large-scale language model LLM.

[0313] <Step S5> In step S5 of phase PH2, the component 130 transmits the translation document TDoc to the component 120. Note that step S5 corresponds to the arrow extending from (5) in FIG.

[0314] <Step S6> In step S6 of phase PH2, subcomponent 120A creates directive Pt4 and sends it to component 130. Note that step S6 corresponds to the arrows extending from (6) and (7) in Fig. 16. Steps S6 to S12 correspond to the part enclosed by the rectangle "loop" in Fig. 16.

[0315] The instruction Pt4 includes instruction g4(), translation guidelines TGL, original ODoc, and translated document TDoc. Instruction g4() includes a procedure for evaluating whether the translation from the original ODoc to the translated document TDoc conforms to the translation guidelines TGL, outputting true if the translation conforms, and generating an inappropriate list IL if the translation does not conform. The inappropriate list IL also includes original text and translated text that are evaluated as not conforming.

[0316] <Step S7> In step S7 of phase PH2, the component 130 receives the directive Pt4 and outputs true or generates an improper list IL using the large-scale language model LLM.

[0317] <<Step S8>> In step S8 of phase PH2, the component 130 transmits true or the inappropriate list IL to the component 120. Note that outputting true corresponds to the arrow extending from (12) in Fig. 16, and generating the inappropriate list IL corresponds to the arrow extending from (8) in Fig. 16. Steps S8 to S12 correspond to the part enclosed by the rectangle alt in Fig. 16.

[0318] <Step S9> If the component 120 receives true in step S9 of phase PH2, the process ends. If the component 120 receives the inappropriate list IL, the process proceeds to step S10 of phase PH2.

[0319] <Step S10> In step S10 of phase PH2, the subcomponent 120A creates a directive Pt5 and sends it to the component 130. Note that step S10 corresponds to the arrows extending from (9) and (10) in FIG. 16 .

[0320] The instruction Pt5 includes an instruction g5(), a translation guideline TGL, an original ODoc, a translated document TDoc, and an inappropriate list IL. The instruction g5() includes a procedure for generating a translated document TDoc from the original ODoc in accordance with the translation guideline TGL so as not to translate the documents listed in the inappropriate list IL.

[0321] <Step S11> In step S11 of phase PH2, the component 130 receives the directive Pt4 and generates a translation document TDoc using the large-scale language model LLM.

[0322] <Step S12> In step S12 of phase PH2, the component 130 sends the translated document TDoc to the component 120, and the process proceeds to step S6 of phase PH2. Note that step S11 corresponds to the arrow extending from (12) in FIG.

[0323] This makes it possible to generate a translated document TDoc from an original ODoc in accordance with the translation guidelines TGL. It is also possible to generate a translated document TDoc that takes into account the expected reader characteristics CIR and the preferred stylistic characteristics CPS. It is also possible to generate a translated document TDoc that does not include the avoidable stylistic characteristics CAS. It is also possible to proofread the translated document TDoc based on the translation guidelines TGL. As a result, it is possible to provide a novel information processing method that is highly convenient, useful, and reliable.

[0324] <Example 3 of Information Processing Method> An information processing method according to one aspect of the present invention is an information processing method having a phase PH1 and a phase PH3 (see FIG. 13). Note that phase PH1 follows phase PH3.

[0325] <Example of Phase PH3> Phase PH3 includes steps S1 to S9.

[0326] Step S1: In step S1 of phase PH3, subcomponent 120B selects two documents from the n typical documents Docs and creates all unique combinations. Subcomponent 120B also designates one document in each combination as document Doc1 and the other as document Doc2. Step S1 corresponds to the arrow extending from (1) in Figure 17.

[0327] Note that subcomponent 120B is included in component 120. Furthermore, the n typical documents Docs are classified into different clusters, and the n typical documents Docs have embedded representations that are closest to the centers of gravity of the respective clusters.

[0328] <<Step S2>> In step S2 of phase PH3, subcomponent 120A selects one of the combinations that have not yet been selected, creates instruction statement Pt6, and sends it to component 130. Note that step S2 corresponds to the arrows extending from (2) and (3) in FIG. 17 .

[0329] The instruction Pt6 includes instruction g6(), document Doc1, and document Doc2. Instruction g6() includes a procedure for extracting and listing stylistic features CS1 from document Doc1, and a procedure for extracting and listing stylistic features CS2 from document Doc2.

[0330] <Step S3> In step S3 of phase PH3, the component 130 receives the directive sentence Pt6 and generates stylistic features CS1 and CS2 using the large-scale language model LLM.

[0331] <Step S4> In step S4 of phase PH3, the component 130 transmits the stylistic feature CS1 and the stylistic feature CS2 to the component 120. Note that step S4 corresponds to the arrow extending from (4) in FIG.

[0332] <Step S5> In step S5 of phase PH3, the subcomponent 120C creates a question Q and sends it to the component 110. Note that step S5 corresponds to the arrows extending from (5) and (6) in FIG. 17 .

[0333] Note that subcomponent 120C is included in component 120. Question Q asks whether stylistic feature CS1 or stylistic feature CS2 is a more suitable style for the assumed reader characteristic CIR.

[0334] <Step S6> In step S6 of phase PH3, the component 110 presents a question Q and waits for input of an answer Ans. For example, the component 110 presents the question Q to a user of the information processing system. Note that step S6 corresponds to the arrow extending from (7) in FIG. 17 .

[0335] <<Step S7>> In step S7 of phase PH3, the component 110 accepts the answer Ans and transmits it to the component 120. For example, a user of the information processing system inputs the answer Ans. Note that step S7 corresponds to the arrow extending from (8) in FIG. 17 .

[0336] <Step S8> In step S8 of phase PH3, subcomponent 120C receives answer Ans. If answer Ans is a stylistic feature CS1, subcomponent 120C adds the expected reader feature CIR and the stylistic feature CS1 to the positive example list PEL. Subcomponent 120C also adds the expected reader feature CIR and the stylistic feature CS2 to the negative example list NEL. Note that stylistic feature CS1 is added as a preferred stylistic feature CPS. Note that stylistic feature CS2 is added as an avoided stylistic feature CAS. Note that step S8 corresponds to the arrows extending from (9) and (10) in FIG. 17 .

[0337] <Step S9> In step S9 of phase PH3, if there is a combination for which directive statement Pt6 has not been created, the subcomponent 120A proceeds to step S2 of phase PH3; otherwise, it ends phase PH3. Steps S1 to S9 correspond to the part enclosed by the rectangle loop in FIG. 17 .

[0338] As a result, for example, a stylistic feature CS1 can be added to the positive example list PEL as a writing style that is appropriate for the expected reader characteristics CIR. Furthermore, for example, a stylistic feature CS2 can be added to the negative example list NEL as a writing style that is not appropriate for the expected reader characteristics CIR. Furthermore, the positive example list PEL and the negative example list NEL can be created using n typical documents Docs classified into different clusters. As a result, a novel information processing method that is highly convenient, useful, and reliable can be provided.

[0339] <Example 4 of Information Processing Method> An information processing method according to one aspect of the present invention is an information processing method having a phase PH1 and a phase PH3 (see FIG. 14). Note that phase PH1 follows phase PH3.

[0340] <Example of Phase PH3> Phase PH3 includes steps S1 to S13.

[0341] <<Step S1>> In step S1 of phase PH3, the component 110 accepts an original ODoc and transmits it to the component 120. For example, a user of the information processing system inputs the original ODoc. Note that step S1 corresponds to the arrows extending from (1) and (2) in FIG. 18 .

[0342] <Step S2> In step S2 of phase PH3, the component 120 accepts the original ODoc and transmits it to the component 140. Note that step S2 corresponds to the arrows extending from (2) and (3) in FIG.

[0343] Step S3: In step S3 of phase PH3, the component 140 receives the original ODoc and converts it into an embedded representation EE using the embedded model EM.

[0344] <Step S4> In step S4 of phase PH3, the component 140 transmits the embedded representation EE to the component 120. Note that step S4 corresponds to the arrow extending from (4) in FIG.

[0345] Step S5: In step S5 of phase PH3, subcomponent 120B selects two documents from the n+1 typical documents Docs and creates all unique combinations. Subcomponent 120B also designates one document in each combination as document Doc1 and the other as document Doc2. Step S5 corresponds to the arrow extending from (5) in Figure 18.

[0346] Note that subcomponent 120B is included in component 120. Furthermore, n-1 of the n+1 typical documents Docs are classified into different clusters, have embedded representations closest to the centers of gravity of their respective clusters, and are all classified into clusters different from the original ODoc. Furthermore, two of the n+1 typical documents Docs are classified into the same cluster as the original ODoc, and have embedded representations selected in order of proximity to the centers of gravity of the clusters.

[0347] <Step S6> In step S6 of phase PH3, subcomponent 120A selects one of the combinations that have not yet been selected, creates instruction statement Pt6, and sends it to component 130. Note that step S6 corresponds to the arrows extending from (6) and (7) in FIG.

[0348] The instruction Pt6 includes instruction g6(), document Doc1, and document Doc2. Instruction g6() includes a procedure for extracting and listing stylistic features CS1 from document Doc1, and a procedure for extracting and listing stylistic features CS2 from document Doc2.

[0349] <Step S7> In step S7 of phase PH3, the component 130 receives the directive sentence Pt6 and generates stylistic features CS1 and CS2 using the large-scale language model LLM.

[0350] <Step S8> In step S8 of phase PH3, the component 130 transmits the stylistic feature CS1 and the stylistic feature CS2 to the component 120. Note that step S8 corresponds to the arrow extending from (8) in FIG.

[0351] <Step S9> In step S9 of phase PH3, the subcomponent 120C creates a question Q and sends it to the component 110. Note that step S6 corresponds to the arrows extending from (9) and (10) in FIG. 18 .

[0352] Note that subcomponent 120C is included in component 120. Question Q asks whether stylistic feature CS1 or stylistic feature CS2 is a style that is more suitable for the assumed reader characteristic CIR.

[0353] <<Step S10>> In step S10 of phase PH3, the component 110 presents a question Q and waits for input of an answer Ans. For example, the component 110 presents the question Q to a user of the information processing system. Note that step S10 corresponds to the arrow extending from (11) in FIG. 18 .

[0354] <<Step S11>> In step S11 of phase PH3, the component 110 accepts the answer Ans and transmits it to the component 120. For example, a user of the information processing system inputs the answer Ans. Note that step S11 corresponds to the arrow extending from (12) in FIG. 18 .

[0355] <Step S12> In step S12 of phase PH3, subcomponent 120C receives answer Ans. If answer Ans is a stylistic feature CS1, subcomponent 120C adds expected reader feature CIR and stylistic feature CS1 to the positive example list PEL. Subcomponent 120C also adds expected reader feature CIR and stylistic feature CS2 to the negative example list NEL. Note that stylistic feature CS1 is added as a preferred stylistic feature CPS. Note that stylistic feature CS2 is added as an avoided stylistic feature CAS. Note that step S12 corresponds to the arrows extending from (13) and (14) in FIG. 18.

[0356] <Step S13> In step S13 of phase PH3, if there is a combination for which directive statement Pt6 has not been created, the subcomponent 120A proceeds to step S6 of phase PH3; otherwise, it ends phase PH3. Steps S5 to S12 correspond to the part enclosed by the rectangle loop in FIG. 18 .

[0357] As a result, for example, a stylistic feature CS1 can be added to the positive example list PEL as a writing style that is appropriate for the expected reader characteristics CIR. Furthermore, for example, a stylistic feature CS2 can be added to the negative example list NEL as a writing style that is not appropriate for the expected reader characteristics CIR. Furthermore, the positive example list PEL and the negative example list NEL can be created using n-1 typical documents Docs that are classified into a different cluster from the original ODoc. Furthermore, for example, a stylistic feature CS1 can be added to the positive example list PEL as a way to identify minor differences in stylistic features between similar documents using two typical documents that are classified into the same cluster as the original ODoc. Furthermore, for example, a stylistic feature CS2 can be added to the negative example list NEL. As a result, a novel information processing method that is highly convenient, useful, and reliable can be provided.

[0358] Note that this embodiment mode can be appropriately combined with other embodiment modes described in this specification.

[0359] AM: avoidance item, Ans: answer, CAS: feature, CIR: feature, CPS: feature, Docs: typical document, EE: embedded representation, EM: embedded model, g11: instruction, g12: instruction, IL: inappropriate list, LLM: large-scale language model, NEL: negative example list, ODoc: original document, PEL: positive example list, Pt11: instruction statement, Pt12: instruction statement, RM: recommendation, TDoc: translated document, TGL: translation guideline, 20: information processing device, 21: input unit, 22: memory unit, 23: processing unit, 24: output unit, 25: transmission path, 51: network, 99: user, 110: component, 120: component, 120A: subcomponent, 120B: subcomponent, 120C: subcomponent, 130: component, 140: component

Claims

1. A translation system comprising a first component, a second component, and a third component, wherein the first component has a function of receiving a list of positive examples and sending it to the third component, the positive example list including characteristics of an intended reader and preferred stylistic features, the second component has a function of receiving a first instruction statement and sending the characteristics of the intended reader and the preferred stylistic features to the third component, and a function of performing processing using a large-scale language model, the large-scale language model has a function of extracting the characteristics of the intended reader and the preferred stylistic features in accordance with the first instruction statement, the third component has a function of receiving the list of positive examples, the characteristics of the intended reader, the preferred stylistic features, and translation guidelines, and sharing them within the third component, the third component has a subcomponent, wherein the subcomponent has a function of creating the first instruction statement and a second instruction statement and sending them to the second component, the first instruction statement including a first instruction and the positive example list, An information processing system, wherein the first instructions include a procedure for extracting characteristics of the intended reader and characteristics of the preferred writing style from the list of positive examples; the second instruction sentence includes second instructions, characteristics of the intended reader, and characteristics of the preferred writing style; the second instructions include a procedure for generating the translation guidelines; the translation guidelines include recommendations; and the recommendations include characteristics of the intended reader and characteristics of the preferred writing style.

2. The information processing system of claim 1, wherein the first component has a function of receiving a negative example list and sending it to the third component, the negative example list including avoided stylistic features, the second component has a function of receiving a third instruction and sending the avoided stylistic features to the third component, the large-scale language model has a function of extracting the avoided stylistic features in accordance with the third instruction, the third component has a function of receiving the negative example list and the avoided stylistic features and sharing them within the third component, the subcomponent has a function of creating the third instruction and sending it to the second component, the third instruction including a third instruction and the negative example list, the third instruction including a procedure for extracting the avoided stylistic features from the negative example list, the second instruction including the avoided stylistic features, the translation guideline including avoided matters, and the avoided matters including the avoided stylistic features.

3. The information processing system of claim 2, wherein the first component has a function of accepting an original document and sending it to the third component, and a function of accepting and providing a translated document; the second component has a function of accepting a fourth instruction sentence and sending the translated document to the third component; the large-scale language model has a function of generating the translated document in accordance with the fourth instruction sentence; the third component has a function of accepting the original document and sharing it within the third component, and a function of accepting the translated document and sending it to the first component; the subcomponent has a function of creating the fourth instruction sentence; the fourth instruction sentence includes a fourth instruction, the translation guideline, and the original document; and the fourth instruction includes a procedure for generating the translated document from the original document in accordance with the translation guideline.

4. The second component has a function of receiving a fifth instruction sentence and sending an inappropriate list to the third component, and a function of receiving a sixth instruction sentence and sending the translated document to the third component; the large-scale language model has a function of generating the inappropriate list according to the fifth instruction sentence and a function of generating the translated document according to the sixth instruction sentence; the subcomponent has a function of creating the fifth instruction sentence and the sixth instruction sentence; the fifth instruction sentence includes a fifth instruction, the translation guideline, the original document, and the translated document; the fifth instruction includes a procedure of evaluating whether the translation from the original document to the translated document complies with the translation guideline, outputting true if the translation is evaluated as complied with, and generating the inappropriate list if the translation is evaluated as not complied with; the inappropriate list includes the original text and the translated text that are evaluated as not complied with; the sixth instruction sentence includes a sixth instruction, the translation guideline, the original document, the translated document, and the inappropriate list; The information processing system according to claim 3 , wherein the sixth instruction includes a procedure for generating the translated document from the original in accordance with the translation guidelines so as to avoid translations listed in the inappropriate list.

5. A system comprising a first component, a second component, and a third component, wherein the first component has a function of receiving and providing a question and a function of receiving and transmitting an answer to the third component, the second component has a function of receiving an instruction sentence and transmitting first and second stylistic features to the third component, and a function of processing using a large-scale language model, the large-scale language model has a function of extracting the first and second stylistic features according to the instruction sentence, the third component has a function of receiving the answer and sharing it within the third component, the third component has a first subcomponent, a second subcomponent, and a third subcomponent, the first subcomponent has a function of creating the instruction sentence and transmitting it to the second component, the instruction sentence including an instruction, a first document, and a second document, the instructions include a procedure for extracting and listing the first stylistic features from the first document and a procedure for extracting and listing the second stylistic features from the second document; the second subcomponent has a function for selecting two from n typical documents to create a combination and a function for designating one of the combinations as the first document and the other as the second document; the third subcomponent has a function for creating the question, the question asking which of the first stylistic features or the second stylistic features is a style that is more appropriate for the characteristics of an expected reader; when the answer is the first stylistic feature, the third subcomponent has a function for adding the characteristics of the expected reader and the first stylistic features to a positive example list and a function for adding the characteristics of the expected reader and the second stylistic features to a negative example list; the first stylistic feature is added as a preferred stylistic feature, and the second stylistic feature is added as an avoided stylistic feature.

6. The information processing system of claim 5, wherein the second subcomponent has a function of selecting two of the n representative documents and creating all unique combinations, the n representative documents are each classified into a different cluster, and the n representative documents have an embedding representation that is closest to the center of gravity of each cluster, and the third subcomponent has a function of creating the question for each of all combinations.

7. A system having a fourth component, wherein the fourth component has a function of receiving an original document and transmitting an embedded representation to the third component and a function of performing processing using an embedding model, wherein the embedded model has a function of converting the n+1 representative documents and the original document into embedded representations, respectively, wherein n-1 of the n+1 representative documents are classified into different clusters and have embedded representations closest to the centers of gravity of the respective clusters, and each is also classified into a different cluster from the original document, wherein two of the n+1 representative documents are classified into the same cluster as the original document and have embedded representations selected in order of proximity to the center of gravity of the cluster, wherein the first component has a function of receiving the original document and transmitting it to the third component, wherein the third component has a function of receiving the original document and transmitting it to the fourth component, and a function of receiving the embedded representation and sharing it within the third component, wherein the second subcomponent has a function of selecting two of the n+1 representative documents and creating all unique combinations, The information processing system according to claim 5 , wherein the third subcomponent has a function of generating the questions for all combinations.

8. An information processing method having a first phase, the first phase comprising first to eleventh steps, wherein in the first step of the first phase, a first component receives a positive example list and a negative example list and transmits them to a second component, and in the second step of the first phase, the second component receives the positive example list and the negative example list and shares them within the second component, the second component comprising a first subcomponent, and in the third step of the first phase, the first subcomponent creates a first instruction statement and transmits it to a third component, the first instruction statement comprising a first instruction and the positive example list, and the first instruction comprising a procedure for extracting characteristics of an intended reader and characteristics of a preferred writing style from the positive example list, In the fourth step of the first phase, the third component receives the first instruction and extracts the characteristics of the intended reader and the characteristics of the preferred writing style using a large-scale language model; in the fifth step of the first phase, the third component transmits the characteristics of the intended reader and the characteristics of the preferred writing style to the second component; in the sixth step of the first phase, the first subcomponent creates a second instruction and transmits it to the third component, the second instruction including second instructions and the negative example list, and the second instructions including a procedure for extracting avoided stylistic features from the negative example list; in the seventh step of the first phase, the third component receives the second instruction and extracts the avoided stylistic features using the large-scale language model; and in the eighth step of the first phase, the third component transmits the avoided stylistic features to the second component. In the ninth step of the first phase, the first subcomponent creates and sends a third instruction to the third component;An information processing method, wherein the third instruction sentence includes a third instruction, characteristics of the intended reader, characteristics of the preferred style, and characteristics of the avoided style; the third instruction sentence includes a procedure for generating translation guidelines; the translation guidelines include recommendations and avoidance items; the recommendations include characteristics of the intended reader and the preferred style features; and the avoidance items include characteristics of the avoided style; in the tenth step of the first phase, the third component accepts the third instruction sentence and generates the translation guidelines using the large-scale language model; and in the eleventh step of the first phase, the third component sends the translation guidelines to the second component, thereby completing the first phase.

9. An information processing method having a first phase and a second phase, wherein the second phase follows the first phase, and the second phase comprises first to twelfth steps, wherein in the first step of the second phase, the first component accepts an original document and sends it to the second component, and in the second step of the second phase, the second component accepts the original document and shares it within the second component, and in the third step of the second phase, the first subcomponent creates a fourth instruction statement and sends it to the third component, and the fourth instruction statement includes a fourth instruction, the translation guideline, and the original document, and the fourth instruction includes a procedure for generating a translated document from the original document in accordance with the translation guideline, and in the fourth step of the second phase, the third component accepts the fourth instruction statement and generates the translated document using the large-scale language model, In the fifth step of the second phase, the third component sends the translated document to the second component; in the sixth step of the second phase, the first subcomponent creates a fifth instruction sentence and sends it to the third component; the fifth instruction sentence includes a fifth instruction, the translation guideline, the original document, and the translated document; the fifth instruction includes a procedure for evaluating whether the translation from the original document to the translated document complies with the translation guideline, and outputting true if the translation is evaluated as conforming, and generating a reluctance list if the translation is evaluated as not conforming; the reluctance list includes the original document and the translated document that are evaluated as not conforming; in the seventh step of the second phase, the third component accepts the fifth instruction sentence and uses the large-scale language model to output true or generate the reluctance list; and in the eighth step of the second phase, the third component sends the true or the reluctance list to the second component.In the ninth step of the second phase, if the second component accepts true, the process ends, and if the second component accepts the reluctance list, the process proceeds to the tenth step of the second phase; in the tenth step of the second phase, the first subcomponent creates a sixth instruction statement and sends it to the third component, the sixth instruction statement including a sixth instruction, the translation guideline, the original, the translated document, and the reluctance list, the sixth instruction including a procedure for generating the translated document from the original in accordance with the translation guideline so as not to translate the documents listed in the reluctance list; in the eleventh step of the second phase, the third component accepts the fifth instruction statement and generates the translated document using the large-scale language model, 9. The information processing method of claim 8, wherein in the twelfth step of the second phase, the third component sends the translated document to the second component, and the process proceeds to the sixth step of the second phase.

10. An information processing method having a first phase and a second phase, wherein the first phase follows the second phase, and the second phase comprises steps 1 to 9, wherein in the first step of the second phase, a second subcomponent selects two from n representative documents, creates all unique combinations, and designates one of the combinations as a first document and the other as a second document, the second subcomponent is included in the second component, the n representative documents are classified into different clusters, and the n representative documents have embedded representations that are closest to the centers of gravity of their respective clusters, and in the second step of the second phase, the first subcomponent selects one of the combinations that has not yet been selected, creates a fourth instruction statement, and sends it to the third component, the fourth instruction statement including a fourth instruction, the first document, and the second document, the fourth instructions include a procedure for extracting and listing first stylistic features from the first document and a procedure for extracting and listing second stylistic features from the second document; in the third step of the second phase, the third component accepts the fourth instruction sentence and generates the first stylistic features and the second stylistic features using the large-scale language model; in the fourth step of the second phase, the third component sends the first stylistic features and the second stylistic features to the second component; in the fifth step of the second phase, the third subcomponent creates a question and sends it to the first component; the third subcomponent is included in the second component; the question asks whether the first stylistic feature or the second stylistic feature is a stylistic feature that is appropriate for the characteristics of the intended reader; in the sixth step of the second phase, the first component provides the question and waits for input of an answer; In the seventh step of the second phase, the first component accepts the response and sends it to the second component;9. The information processing method of claim 8, wherein in the eighth step of the second phase, the third subcomponent accepts the answer, and when the answer is characteristic of the first writing style, adds the characteristic of the intended reader and the first writing style to the positive example list and adds the characteristic of the intended reader and the second writing style to the negative example list; the first writing style characteristic is added as the preferred writing style characteristic; and the second writing style characteristic is added as the avoided writing style characteristic; and in the ninth step of the second phase, when there is a combination that does not create the fourth directive sentence, the first subcomponent proceeds to the second step of the second phase, and otherwise terminates the second phase.

11. An information processing method having a first phase and a second phase, wherein the first phase follows the second phase, and the second phase comprises steps 1 to 13, wherein in the first step of the second phase, the first component accepts an original document and sends it to the second component, and in the second step of the second phase, the second component accepts the original document and sends it to a fourth component, and in the third step of the second phase, the fourth component accepts the original document and converts it into an embedded representation using an embedding model, and in the fourth step of the second phase, the fourth component sends the embedded representation to the second component, and in the fifth step of the second phase, a second subcomponent selects two of n+1 typical documents, creates all unique combinations, and designates one of the combinations as a first document and the other as a second document, the second subcomponent is included in the second component, and n-1 of the n+1 representative documents are classified into different clusters and have embedded representations that are closest to the center of gravity of their respective clusters, and are all classified into a cluster different from the original; and two of the n+1 representative documents are classified into the same cluster as the original and have embedded representations selected in order of proximity to the center of gravity of the cluster; in the sixth step of the second phase, the first subcomponent selects one of the combinations that has not yet been selected, creates a fourth instruction statement, and sends it to the third component, and the fourth instruction statement includes a fourth instruction, the first document, and the second document; and the fourth instruction includes a procedure for extracting and listing first stylistic features from the first document, and a procedure for extracting and listing second stylistic features from the second document, In the seventh step of the second phase, the third component receives the fourth instruction sentence and generates the first stylistic features and the second stylistic features using the large-scale language model;In the eighth step of the second phase, the third component sends the first stylistic feature and the second stylistic feature to the second component; in the ninth step of the second phase, the third subcomponent creates a question and sends it to the first component; the third subcomponent is included in the second component; the question asks whether the first stylistic feature or the second stylistic feature is a style that is appropriate for the characteristics of the intended reader; in the tenth step of the second phase, the first component provides the question and waits for input of an answer; in the eleventh step of the second phase, the first component accepts the answer and sends it to the second component; 9. The information processing method of claim 8, wherein in the twelfth step of the second phase, the third subcomponent accepts the answer, and when the answer is characteristic of the first writing style, adds the characteristic of the intended reader and the first writing style to the positive example list and adds the characteristic of the intended reader and the second writing style to the negative example list; the first writing style characteristic is added as the preferred writing style characteristic; and the second writing style characteristic is added as the avoided writing style characteristic; and in the thirteenth step of the second phase, when there is a combination that does not create the fourth directive sentence, the first subcomponent proceeds to the sixth step of the second phase; otherwise, the second phase is terminated.

Citation Information

Patent Citations

  • Large model machine translation method based on RAG

    CN117993396A

  • LLM-based text translation method and apparatus, and program product

    CN118153587A