Information processing device, method, and program
The described technology addresses the limitations of existing software development methods by using a language model to convert natural language requirements into object-oriented software design, enhancing efficiency and quality.
Patent Information
- Application Number
- JP2024084520
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-24
- Publication Date
- 2025-12-05
AI Technical Summary
Existing technologies for object-oriented software development from natural language are restricted by pre-created basic ontologies, limiting the flexibility and efficiency of software development.
An information processing device and method that utilizes a language model to convert natural language requirement specifications into object-oriented software design, allowing for the generation of model data and source code through a series of instruction sentences and sentence patterns.
Facilitates efficient object-oriented software development by automating the analysis and design process, reducing developer workload and improving the quality and maintainability of software.
Smart Images

Figure 2025177566000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, a method, and a program. [Background technology]
[0002] Patent Document 1 discloses a technique for developing an object-oriented language program from natural language using an ontology engineering method. For example, the technique disclosed in Patent Document 1 uses a basic ontology that manages classes, methods, etc. provided by an object-oriented language by attaching vocabulary labels to them. Then, a class name is defined as a vocabulary whose nouns included in the input natural language match a vocabulary label included in the basic ontology. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-287695 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the technology disclosed in Patent Document 1 has no choice but to generate class names within the range of vocabulary labels defined in a pre-created basic ontology, which places restrictions on the natural language sentences that developers can write for software development, posing a challenge to development efficiency.
[0005] In view of the above-mentioned problems, the object of the present disclosure is to provide an information processing device, method, and program for supporting efficient object-oriented software development from natural language text written for software development. [Means for solving the problem]
[0006] The information processing device according to the present disclosure includes: a first input means for inputting a first instruction sentence for converting a requirement specification sentence, in which the software requirement specification is written in a natural language, into a set of sentences in accordance with a predetermined sentence pattern, and a first input text including the requirement specification sentence, into a predetermined language model; a first acquisition means for acquiring the set of sentences converted from the requirement specification sentence based on the first instruction sentence using the language model; a second input means for inputting a second input text including a second instruction sentence for converting the set of sentences into model data by an object-oriented software design method and the set of sentences to the language model; a second acquiring means for acquiring model data converted from the set of sentences based on the second instruction sentence using the language model; Equipped with.
[0007] The information processing method according to the present disclosure includes: The computer a first instruction sentence for converting a requirement specification sentence, in which the software requirement specification is written in a natural language, into a set of sentences in accordance with a predetermined sentence pattern, and a first input text including the requirement specification sentence are input to a predetermined language model; obtaining the set of sentences converted from the requirement specification sentence based on the first instruction sentence using the language model; inputting a second input text including a second instruction sentence for converting the set of sentences into model data using an object-oriented software design method and the set of sentences into the language model; The language model obtains model data converted from the set of sentences based on the second instruction sentence.
[0008] The information processing program according to the present disclosure includes: a first input process for inputting a first instruction sentence for converting a requirement specification sentence, in which the software requirement specification is written in a natural language, into a set of sentences in accordance with a predetermined sentence pattern, and a first input text including the requirement specification sentence, into a predetermined language model; a first acquisition process for acquiring the set of sentences converted from the requirement specification sentence based on the first instruction sentence using the language model; a second input process for inputting a second input text including a second instruction sentence for converting the set of sentences into model data using an object-oriented software design method and the set of sentences to the language model; a second acquisition process of acquiring model data converted from the set of sentences based on the second instruction sentence using the language model; to be executed by the computer. [Effects of the Invention]
[0009] The present disclosure makes it possible to support efficient object-oriented software development from natural language text written for software development. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram illustrating a configuration of an information processing device according to the present disclosure. [Figure 2] 1 is a flowchart illustrating a flow of an information processing method according to the present disclosure. [Figure 3] 1 is a block diagram showing the overall configuration of a software development support system including an information processing device according to the present disclosure. [Figure 4] FIG. 1 is a block diagram illustrating a configuration of an information processing device according to the present disclosure. [Figure 5] FIG. 1 is a diagram illustrating an example of a requirement specification statement according to the present disclosure. [Figure 6] 1 is a flowchart illustrating a flow of an object-oriented software development support method according to the present disclosure. [Figure 7] 10A-10C illustrate example prompts for analyzing sentence data according to the present disclosure. [Figure 8] FIG. 10 is a diagram illustrating an example of a set of sentences converted from a requirement specification sentence according to the present disclosure. [Figure 9]10A-10C illustrate example prompts for model transformation according to the present disclosure. [Figure 10] FIG. 10 is a diagram illustrating an example of object-oriented model data converted from a set of statements according to the present disclosure. [Figure 11] 10A-10C illustrate example prompts for code conversion according to the present disclosure. [Figure 12] FIG. 10 is a diagram illustrating an example of source code converted from model data into a specific programming language according to the present disclosure. [Figure 13] FIG. 1 is a diagram illustrating an example of a requirement specification statement according to the present disclosure. [Figure 14] 1 is a flowchart illustrating a flow of an object-oriented software development support method according to the present disclosure. [Figure 15] FIG. 10 is a diagram illustrating an example of a set of sentences converted from a requirement specification sentence according to the present disclosure. [Figure 16] FIG. 2 is a diagram illustrating an example of object-oriented first model data converted from a set of statements according to the present disclosure. [Figure 17] 10A-10C illustrate example layering prompts according to the present disclosure. [Figure 18] FIG. 2 is a diagram showing an example of second model data hierarchically organized by superordinately conceptualizing classes from the first model data according to the present disclosure. [Figure 19] FIG. 10 is a diagram illustrating an example of source code converted from second model data into a specific programming language according to the present disclosure. [Figure 20] FIG. 1 is a block diagram showing a hardware configuration of an information processing device according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each drawing, the same or corresponding elements are designated by the same reference numerals, and for clarity of explanation, duplicate explanations will be omitted as necessary.
[0012] (Embodiment 1) FIG. 1 is a block diagram showing the configuration of an information processing device 1. The information processing device 1 is a computer device that uses a predetermined language model to generate model data for object-oriented software design from a requirement specification document in which the software requirement specifications are written in natural language. The information processing device 1 may also be called a software development support device. The information processing device 1 includes a first input unit 11, a first acquisition unit 12, a second input unit 13, and a second acquisition unit 14. The first input unit 11, the first acquisition unit 12, the second input unit 13, and the second acquisition unit 14 may be used as a means for inputting and a means for acquiring information or data, respectively.
[0013] The first input unit 11 inputs a first input text including a first instruction sentence and a requirement specification sentence to a predetermined language model. Here, the "requirement specification sentence" is a sentence (text data) in which the software requirement specifications are written in a natural language. For example, the requirement specification sentence may be written by a software developer. Alternatively, the requirement specification sentence may be text data extracted and generated from a large amount of electronic data for software development. Furthermore, the "first instruction sentence" is a sentence for converting the requirement specification sentence into a set of sentences in accordance with a predetermined sentence pattern. Note that the "sentence" is text data including one or more sentences.
[0014] Furthermore, a "language model" is a computer program or information system that receives text data (input text) that expresses a question or instruction in natural language, and outputs text data that has been generated, converted, processed, summarized, or otherwise processed by a predetermined calculation on the input text. The language model corresponds to a natural language model in an AI (Artificial Intelligence) model. Note that the language model is executed inside the information processing device 1 or on an external server connected to the information processing device 1, and is capable of accepting input text. Furthermore, an "instruction sentence" is text data that instructs the language model on how to process the requirement specification sentence. Therefore, the language model converts the requirement specification sentence into a set of sentences according to a predetermined sentence pattern based on the first instruction sentence, and outputs the converted set of sentences as a processing result.
[0015] The first acquisition unit 12 acquires a set of sentences converted from the requirement specification sentence based on the first instruction sentence using the language model. That is, the first acquisition unit 12 acquires the set of sentences as output data from the language model. Here, the "set of sentences" is text data. The set of sentences is, for example, a set of simple sentences obtained by dividing and converting nouns, predicates, attributes, etc. included in the requirement specification sentence according to a predetermined sentence pattern.
[0016] The second input unit 13 inputs a second input text including a second instruction sentence and a set of sentences to the language model. Here, the "second instruction sentence" is a sentence for converting the set of sentences into model data using an object-oriented software design method. The "set of sentences" included in the second input text is text data acquired by the first acquisition unit 12 described above.
[0017] The second acquisition unit 14 acquires model data converted from the set of sentences based on the second instruction sentence using the language model. Here, "model data" refers to software design information generated by object-oriented analysis. The model data includes text data in a predetermined format that is the result of the object-oriented analysis. Furthermore, the model data may also include various graphic data that is the result of the object-oriented analysis.
[0018] FIG. 2 is a flowchart showing the flow of the information processing method. First, the first input unit 11 inputs a first input text including a first instruction sentence and a required specification sentence to a predetermined language model (S11). Here, the required specification sentence is a sentence in which the software required specifications are written in a natural language. The first instruction sentence is a sentence for converting the required specification sentence into a set of sentences according to a predetermined sentence pattern. Next, the first acquisition unit 12 acquires a set of sentences converted from the required specification sentence based on the first instruction sentence using the language model (S12).
[0019] Next, the second input unit 13 inputs a second input text including a second instruction sentence and a set of sentences to the language model (S13). Here, the set of sentences is acquired in step S12. The second instruction sentence is an instruction sentence for converting the set of sentences into model data using an object-oriented software design method. Then, the second acquisition unit 14 acquires model data converted from the set of sentences based on the second instruction sentence using the language model (S14).
[0020] In this way, the information processing device 1 uses a language model to convert requirement specification statements into a set of statements, and then converts the set of statements into model data that has undergone object-oriented analysis. Then, by inputting appropriate instruction statements into the language model for each conversion, the characteristics of the language model can be utilized to obtain model data, which is design information based on object-orientation. Therefore, software developers can efficiently develop desired software using the obtained model data. Therefore, the information processing device 1 according to the present disclosure can support efficient software development based on object-orientation from natural language statements written for software development.
[0021] The information processing device 1 includes a processor, a memory, and a storage device (not shown). The storage device stores a computer program that implements the process of the information processing method shown in FIG. 2, for example. The processor then loads the computer program from the storage device into the memory and executes the computer program. This allows the processor to implement the functions of a first input unit 11, a first acquisition unit 12, a second input unit 13, and a second acquisition unit 14.
[0022] Alternatively, each component of the information processing device 1 may be realized by dedicated hardware. Furthermore, some or all of the components of each device may be realized by general-purpose or dedicated circuits, processors, etc., or a combination of these. These may be configured by a single chip, or by multiple chips connected via a bus. Some or all of the components of each device may be realized by a combination of the above-mentioned circuits, etc., and programs. Furthermore, a CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), quantum processor (quantum computer control chip), etc., may be used as the processor.
[0023] Furthermore, when some or all of the components of the information processing device 1 are realized by multiple information processing devices, circuits, etc., the multiple information processing devices, circuits, etc. may be centrally or distributed. For example, the information processing devices, circuits, etc. may be realized as a client-server system, a cloud computing system, or the like, in a form in which each is connected via a communication network. Furthermore, the functions of the information processing device 1 may be provided in a SaaS (Software as a Service) format.
[0024] (Embodiment 2) 3 is a block diagram showing the overall configuration of a software development support system 1000 including an information processing device 100. The software development support system 1000 is an information system for supporting software development by automatically generating object-oriented software model data and source code from text written in a natural language using a language model. The software development support system 1000 includes the information processing device 100, an LLM (Large Language Model) server 200, and a developer terminal 300. The information processing device 100, the LLM server 200, and the developer terminal 300 are each connected to each other via a network N so as to be able to communicate with each other. Here, the network N is a network of wired and wireless communication lines.
[0025] Developer terminal 300 is an information processing device operated by a developer who develops software based on object-orientation. Developer terminal 300 may be a general-purpose PC (Personal Computer) or the like. Therefore, developer terminal 300 performs processing in response to keyboard and mouse operations by the developer, and communicates with information processing device 100 via network N as appropriate.
[0026] The LLM server 200 is a server computer on which a predetermined LLM runs. The LLM is an example of the language model in the first embodiment described above. The LLM is a trained model trained by repeatedly performing deep learning on a predetermined natural language model using a large amount of data set. The number of deep learning training sessions, the number of data sets used for training, and the number of trained parameters for the LLM are large compared to when AI models first began to be used. For this reason, the LLM is sometimes referred to as a large-scale language model. The LLM is a computer program that receives input text (prompt) written in a specific format, performs processing based on instruction sentences included in the prompt, and outputs the processing results. Here, the prompt includes text data to be processed and instruction sentences describing processing for the text data in a specific format.
[0027] For example, the first instruction sentence is an instruction to convert natural language text data into a set of sentences according to a predetermined sentence pattern, such as the five basic sentence patterns of English grammar. In this case, the LLM converts the natural language text data included in the input prompt into a set of sentences according to a predetermined sentence pattern, in accordance with the first instruction sentence included in the prompt. For example, if the text data is "The user logs in and reserves a ticket," the LLM extracts the subject (user), predicate (log in, reserve), object (ticket), etc. from the text data and converts it into two sets of sentences: "The user logs in" and "The user reserves a ticket."
[0028] Next, the second instruction sentence is assumed to be an instruction to convert the set of statements into model data using an object-oriented software design methodology. In this case, the LLM extracts object-oriented software components (classes, methods, members (attributes, properties), arguments, return values, etc.) from the set of statements converted above in accordance with the second instruction sentence included in the input prompt, and generates model data by organizing their relationships (inheritance, possession, correspondence, etc.). For example, in the case of the set of converted statements above, the LLM converts them into model data in which "User" and "Ticket" are defined as classes, the "User" class has methods for "Login" and "Make a Reservation," and the "Make a Reservation" method is defined with the "Ticket" class as an argument.
[0029] Then, the third instruction sentence is an instruction to convert the model data into source code based on the specifications of a specific programming language. In this case, the LLM converts the converted model data into source code based on the specifications of the specific programming language in accordance with the third instruction sentence included in the input prompt. This makes it possible to automatically generate object-oriented software source code from natural language sentences using the LLM.
[0030] Therefore, LLM can be said to function as a text data analyzer and a model generator. The text data analyzer analyzes natural language text data, generates multiple sentences based on multiple basic sentence patterns in the grammar of a specific natural language, and generates a set of sentences. The model generator converts the set of sentences into model data, which is information that defines the software components based on object-orientation and the relationships between each component.
[0031] When the LLM server 200 receives input text (prompt) from a requestor via the network N, it inputs the prompt to the LLM and returns output data, which is the processing result of the LLM, to the requestor via the network N. The requestor is, for example, the information processing device 100 or the developer terminal 300.
[0032] The information processing device 100 is an example of the information processing device 1 described above. The information processing device 100 is a computer device that generates object-oriented software model data and source code from a document written in a natural language using LLM. Specifically, the information processing device 100 receives a requirements specification document and an instruction for object-oriented software development support processing from a developer terminal 300 via a network N. The information processing device 100 then communicates with an LLM server 200 to generate object-oriented software model data and source code from the requirements specification document. Specifically, the information processing device 100 generates prompts for LLM according to each stage of object-oriented software development, transmits (inputs) the prompts to the LLM server 200, and receives (acquires) output data from the LLM server 200, repeating this process multiple times.
[0033] 4 is a block diagram showing the configuration of the information processing device 100. The information processing device 100 includes a reception unit 121, a generation unit 122, an input unit 123, an acquisition unit 124, and an output unit 125. The reception unit 121, the generation unit 122, the input unit 123, the acquisition unit 124, and the output unit 125 may be used as a means for receiving, a means for generating, a means for inputting, a means for acquiring, and a means for outputting information or data, respectively.
[0034] The receiving unit 121 receives a requirement specification document and instructions for object-oriented software development support processing from the developer terminal 300 via the network N. The receiving unit 121 may also receive various prompts, which will be described later, from the developer terminal 300. The receiving unit 121 may also receive from the developer terminal 300 a set of corrected statements, corrected model data, or a part of the set of statements or model data as processing targets.
[0035] The generation unit 122 generates a sentence analysis prompt including a first instruction sentence and a requirement specification sentence. The sentence analysis prompt is an example of the first input text described above. The generation unit 122 generates a model conversion prompt including a second instruction sentence and a set of sentences. The model conversion prompt is an example of the second input text described above. Here, the set of sentences included in the model conversion prompt is a set of sentences converted from the requirement specification sentence by the LLM based on the sentence analysis prompt described above. The generation unit 122 generates a code conversion prompt including a third instruction sentence and model data. The code conversion prompt is an example of a third input text. Here, the model data included in the code conversion prompt is data converted from the set of sentences by the LLM based on the code conversion prompt described above. Furthermore, the "third instruction sentence" is a sentence for converting the model data into source code based on the specifications of a specific programming language.
[0036] The input unit 123 is an example of the first input unit 11 and the second input unit 13 described above. The input unit 123 inputs the sentence analysis prompt generated by the generation unit 122 to the LLM. The input unit 123 inputs the model conversion prompt generated by the generation unit 122 to the LLM. The input unit 123 is also an example of a third input means. The input unit 123 inputs a third input text including a third instruction sentence and model data to the language model. That is, the input unit 123 inputs the code conversion prompt generated by the generation unit 122 to the LLM. Specifically, the input unit 123 inputs each prompt to the LLM by transmitting each prompt to the LLM server 200 via the network N.
[0037] The acquisition unit 124 is an example of the first acquisition unit 12 and the second acquisition unit 14 described above. The acquisition unit 124 acquires, as output data, a set of sentences converted by the LLM in response to a sentence analysis prompt. Here, the set of sentences is converted from the requirement specification sentence by the LLM in accordance with a plurality of basic sentence patterns in the grammar of a specific natural language. The acquisition unit 124 acquires, as output data, model data converted by the LLM in response to a model conversion prompt. The acquisition unit 124 is also an example of a third acquisition means. The acquisition unit 124 acquires source code converted from the model data by the LLM based on the code conversion prompt.
[0038] The output unit 125 outputs the output data acquired by the acquisition unit 124. For example, the output unit 125 may display the output data on a display device built into or connected to the information processing device 1. Specifically, the output unit 125 may transmit the output data to the developer terminal 300 via the network N, thereby causing the developer terminal 300 to display the output data.
[0039] (Example of a use case for object-oriented software development support processing) For example, suppose a developer is developing a web application for an online bookstore. The online bookstore will allow users to search for and browse books, add them to a cart, and order them. The developer of the web application will then input a document describing requirements for developing object-oriented software into the developer terminal 300. Figure 5 shows an example of a requirements document 40.
[0040] 6 is a flowchart showing the flow of the object-oriented software development support method. First, in response to a developer's operation, the developer terminal 300 transmits a request including the input requirement specification statement and instructions for the object-oriented software development support process to the information processing device 100 via the network N. In response, the information processing device 100 receives the request from the developer terminal 300 via the network N.
[0041] Specifically, the receiving unit 121 receives a requirement specification statement and an instruction for object-oriented software development support processing from the developer terminal 300 (S101). Then, the generating unit 122 generates a sentence analysis prompt using the received requirement specification statement in accordance with the received instruction (S102).
[0042] FIG. 7 is a diagram showing an example of a prompt 41 for analyzing text data. The prompt 41 for analyzing text data is text data including an instruction sentence 411 and a requirement specification sentence 412. The instruction sentence 411 is an example of the first instruction sentence described above, but is not limited to this. For example, the part of the instruction sentence 411 that reads "five basic sentence patterns in English grammar" may be replaced with another natural language or number of sentence patterns, as long as the sentence patterns are multiple basic sentence patterns in the grammar of a specific natural language. For example, analysis of sentence patterns in the grammar of a specific natural language may use techniques such as Japanese syntax analysis or morphological analysis (parts of speech and dependency relationships). The requirement specification sentence 412 is an example of text data that begins with the notation "#requirement specification sentence" and then includes the requirement specification sentence 40 described above from the next line onward.
[0043] Next, the input unit 123 transmits a text data analysis prompt 41 to the LLM server 200 (S103). In response, the LLM server 200 inputs the received text data analysis prompt 41 to the LLM. The LLM extracts nouns, predicates, attributes, etc. from the requirement specification text 412 based on the instruction sentence 411 included in the text data analysis prompt 41. At this time, the LLM may aggregate the extracted nouns, predicates, attributes, etc. by eliminating duplications, including spelling variations. The LLM then generates multiple simple sentences from the extracted or aggregated nouns, predicates, attributes, etc. For this purpose, the LLM converts the requirement specification text 412 into a set of sentences. The LLM server 200 transmits an output message including the set of sentences converted by the LLM to the information processing device 100 via the network N.
[0044] In response to this, the acquisition unit 124 acquires an output message including a set of sentences from the LLM server 200. Then, the output unit 125 displays the output message on the developer terminal 300 (S104). Specifically, the output unit 125 transmits the acquired output message to the developer terminal 300 via the network N. Then, the developer terminal 300 displays the received output message on the screen.
[0045] 8 is a diagram showing an example of a set of sentences 422 converted from the requirement specification document 40. The output data 42 is an example of an output message 421 and a set of sentences 422 displayed on the screen of the developer terminal 300. The output message 421 is text data added by the LLM when converting the requirement specification document 40 into a set of sentences. The set of sentences 422 is an example of text data showing a set of multiple simple sentences converted from the requirement specification document 40 by the LLM.
[0046] Next, the generation unit 122 generates a model conversion prompt using the acquired set of sentences 422 (S105). FIG. 9 is a diagram showing an example of a model conversion prompt 43. The model conversion prompt 43 is text data including an instruction sentence 431 and a set of sentences 432. The instruction sentence 431 is an example of the second instruction sentence described above, but is not limited to this. The set of sentences 432 is an example of text data that begins with the notation "#set of sentences" and then contains the set of sentences 422 described above from the next line onwards.
[0047] Next, the input unit 123 transmits a model conversion prompt 43 to the LLM server 200 (S106). In response, the LLM server 200 inputs the received model conversion prompt 43 into the LLM. Based on the instruction sentence 431 included in the model conversion prompt 43, the LLM organizes a set of sentences 432 by abstracting, hierarchically classifying, and aggregating similar nouns, predicates, attributes, etc. The LLM then associates the organized nouns, predicates, attributes, etc. with each other to generate a class structure. In other words, the LLM converts the set of sentences into model data, which is information defining the components of object-oriented software and the relationships between each component. Furthermore, the model data may include information defining the relationships between each component, in which components extracted from the set of sentences are classified as higher-level concepts by the LLM, which determines that the components have a high degree of similarity in logical structure, are hierarchically structured. Furthermore, the model data may include a class structure that associates each word included in the set of sentences with each other. Furthermore, the model data includes classes, methods, attributes, arguments, and return values as components. The model data may also include, as relationships between components, inheritance relationships between classes, ownership relationships between classes, correspondences between classes and methods, and correspondences between methods and arguments and return values. The LLM server 200 transmits an output message including the model data converted by the LLM to the information processing device 100 via the network N.
[0048] In response to this, the acquisition unit 124 acquires an output message including the model data from the LLM server 200. Then, the output unit 125 displays the output message on the developer terminal 300 (S107). Specifically, the output unit 125 transmits the acquired output message to the developer terminal 300 via the network N. Then, the developer terminal 300 displays the received output message on the screen.
[0049] FIG. 10 shows an example of object-oriented model data 442 converted from a set of statements. Output data 44 is an example of an output message 441 and model data 442 displayed on the screen of the developer terminal 300. The output message 441 is text data added by the LLM during conversion to model data. The model data 442 is text data converted from the set of statements 432 by the LLM, showing information defining the components of object-oriented software and the relationships between each component. In other words, the model data 442 shows an example of data converted from the requirements specification document 40 using an object-oriented software design method. The model data 442 may also include graphical information such as a class diagram. In this case, for example, the model conversion prompt may include a Mermaid syntax specification. Relationships r11 and r12 are examples of correspondences between classes and methods. Relationships r21 and r22 are examples of ownership relationships (has-a relationships) between classes.
[0050] Next, the generation unit 122 generates a code conversion prompt using the acquired model data 442 (S108). FIG. 11 is a diagram showing an example of a code conversion prompt 45. The code conversion prompt 45 is text data including an instruction sentence 451 and model data 452. The instruction sentence 451 is an example of the third instruction sentence described above, and is not limited to this. In particular, the instruction sentence 451 may describe a specific language name, such as "Java (registered trademark)," instead of describing "an object-oriented programming language." The same applies to the following explanation. The model data 452 is an example of text data that begins with the notation "#model data" and then describes the above-described model data 442 from the next line onwards.
[0051] Next, the input unit 123 transmits the code conversion prompt 45 to the LLM server 200 (S109). In response, the LLM server 200 inputs the received code conversion prompt 45 into the LLM. The LLM converts the model data 442 into source code based on the specifications of the specified programming language, based on the instruction sentence 451 included in the code conversion prompt 45. The LLM server 200 transmits an output message including the source code converted by the LLM to the information processing device 100 via the network N.
[0052] In response to this, the acquisition unit 124 acquires an output message including the source code from the LLM server 200. Then, the output unit 125 displays the output message on the developer terminal 300 (S110). Specifically, the output unit 125 transmits the acquired output message to the developer terminal 300 via the network N. Then, the developer terminal 300 displays the received output message on the screen.
[0053] FIG. 12 is a diagram showing an example of source code 452 converted from model data 452 into a specific programming language. Output data 46 is an example of an output message 461 and source code 462 displayed on the screen of the developer terminal 300. The output message 461 is text data added by the LLM during conversion into source code. The source code 462 is an example of Java (registered trademark), an object-oriented programming language, but the programming language to be converted is not limited to this. The same applies to the following explanation. The source code 462 shows an example in which members (attributes), methods, etc. are implemented for each of the classes User, Book, Cart, and Order. In other words, the source code 462 shows an example generated based on the requirement specification document 40.
[0054] 7, the set of statements 432 in FIG. 9, and the model data 452 in FIG. 11 described above are not limited to the above. For example, the beginnings of the requirement specification statement 412, the set of statements 432, and the model data 452 may be other notations conforming to the LLM prompt notation provided by the LLM server 200. Alternatively, the beginnings of the requirement specification statement 412, the set of statements 432, and the model data 452 may be omitted. In this case, the requirement specification statement 412 may be the same as the requirement specification statement 40. The set of statements 432 may be the same as the set of statements 422. The model data 452 may be the same as the model data 442.
[0055] As described above, the technology disclosed herein can significantly automate and streamline analysis and design work in the early stages of software development by gradually combining natural language processing using LLM with object-oriented analysis processing. This allows developers to obtain software designs with a certain degree of granularity simply by describing their software ideas (requirements specifications) in natural language. This significantly reduces the burden on software developers.
[0056] In other words, the technology disclosed herein can use a language model to automatically generate at least a portion of object-oriented software design and code from requirement specifications expressed in natural language. This allows developers to add detailed specifications and modify or extend the generated source code based on the generated model data and source code. This improves the efficiency of object-oriented software development.
[0057] In particular, by using the technology disclosed herein, it is possible to automatically infer what classes and methods are required. This allows for a significant reduction in the amount of work required in the early stages of design. This allows developers to focus on implementing more essential logic. Furthermore, it is possible to narrow the gap between the natural language requirements specification and the software implementation, which may improve the quality and maintainability of the software. In addition, this embodiment can achieve the same effects as the first embodiment described above.
[0058] (Embodiment 3) In the third embodiment, the first model data converted by the LLM is further superordinately conceptualized (converted) by the LLM to obtain second model data. That is, the input unit 123 inputs a fourth input text, including a fourth instruction sentence and the first model data, to a predetermined language model. Here, the "first model data" is the model data acquired by the acquisition unit 124. The "fourth instruction sentence" is a sentence for superordinately conceptualizing components that the language model determines to have a high degree of similarity in logical structure among the components included in the first model data, and reconverting the components to form a hierarchical structure. The acquisition unit 124 then acquires, as model data, second model data reconverted from the first model data based on the fourth instruction sentence using the language model to which the fourth input text is input. This improves the accuracy of object-oriented analysis of the model data. Note that other configurations of the third embodiment are the same as those of the second embodiment, and therefore redundant explanations and illustrations will be omitted as appropriate. The following description will focus on the differences from the second embodiment.
[0059] (Example of a use case for superordinate conceptualization) For example, suppose a developer is developing a web application for an online bookstore. The online bookstore allows staff to register books, and allows clients to search for and browse books, add them to carts, and order them. The web application developer then inputs a document describing requirements for developing object-oriented software into the developer terminal 300. Figure 13 shows an example of requirements document 50. The example of requirements document 50 indicates that the processes that staff and clients can perform overlap in part.
[0060] 14 is a flowchart showing the flow of the object-oriented software development support method. First, the developer terminal 300 transmits a request including the above-described requirement specification document 50 and instructions for the object-oriented software development support process to the information processing device 100 via the network N. In response to this, the information processing device 100 executes steps S101 to S106 of FIG. 6 (S201). At this time, in step S102, the generation unit 122 generates a sentence analysis prompt using the requirement specification document 50. Then, in step S104, the acquisition unit 124 acquires a set of sentences converted from the requirement specification document 50 by the LLM, and the output unit 125 displays an output message on the developer terminal 300.
[0061] FIG. 15 is a diagram showing an example of a set of sentences 522 converted from the requirement specification document 50. The output data 52 is an example of an output message 521 and a set of sentences 522 displayed on the screen of the developer terminal 300. The output message 521 is text data added by the LLM when converting it into a set of sentences. The set of sentences 522 is an example of text data showing a set of multiple simple sentences converted from the requirement specification document 50 by the LLM. The set of sentences 522 shows that the sentences "The client is..." and "The staff is..." are different from the set of sentences 422 in FIG. 8 described above.
[0062] Next, in step S105, the generation unit 122 generates a prompt for model conversion using the acquired set of sentences 522. Then, in step S106, the input unit 123 sends the prompt for model conversion to the LLM server 200. Then, the acquisition unit 124 acquires an output message including the first model data from the LLM server 200. Then, the output unit 125 displays the output message on the developer terminal 300 (S202).
[0063] FIG. 16 shows an example of object-oriented first model data 542 converted from the set of statements 522. The output data 54 is an example of an output message 541 and the first model data 542 displayed on the screen of the developer terminal 300. The output message 541 is text data added by the LLM during conversion to the first model data. The first model data 542 is text data converted from the set of statements 522 by the LLM, and indicates information defining the components of object-oriented software and the relationships between the components. In other words, the first model data 442 shows an example of conversion from the requirements specification document 50 using an object-oriented software design method. Compared to the model data 442 in FIG. 10 described above, the first model data 442 shows that the classes "Client" and "Staff" have been added instead of the class "User," and that their methods and relationships have been changed.
[0064] Next, the generation unit 122 generates a layering prompt using the acquired first model data 542 (S203). FIG. 17 is a diagram showing an example of a layering prompt 55. The layering prompt 55 is text data including an instruction sentence 551 and first model data 552. The instruction sentence 551 is an example of the fourth instruction sentence described above, but is not limited to this. The first model data 552 is an example of text data that begins with the notation "#model data" and then contains the above-mentioned first model data 542 from the next line onwards.
[0065] Next, the input unit 123 transmits the layering prompt 55 to the LLM server 200 (S204). In response, the LLM server 200 inputs the received layering prompt 55 to the LLM. Based on the instruction sentence 551 included in the layering prompt 55, the LLM generates a superclass that superconceptualizes multiple classes from the first model data 552, and converts each original class into a subclass by modifying methods, attributes, relationships, etc. to convert it into second model data. The LLM server 200 transmits an output message including the second model data converted by the LLM to the information processing device 100 via the network N.
[0066] In response to this, the acquisition unit 124 acquires an output message including the second model data from the LLM server 200. Then, the output unit 125 displays the output message on the developer terminal 300 (S205).
[0067] 18 is a diagram showing an example of second model data 562 hierarchically created by superimposing classes from the first model data 552. Output data 56 is an example of an output message 561 and model data 562 displayed on the screen of the developer terminal 300. The output message 561 is text data added by the LLM when converting to the second model data. The second model data 562 is text data indicating information hierarchically created by the LLM from the first model data 552 by superimposing classes.
[0068] Specifically, the second model data 562 indicates that a superclass "User" has been added, which is a higher-level conceptualization of the classes "Client" and "Staff." Accordingly, the second model data 562 indicates that the relationships r13, r31, and r32 have been changed from the first model data 552. Specifically, the relationship r13 indicates that the methods "Create an Account," "Log in," "Search for Books," and "View Detailed Information," which were common to the classes "Client" and "Staff," have been made methods of the superclass "User." Accordingly, the relationship r13 indicates that these methods have been deleted from the subclasses "Client" and "Staff." Furthermore, the relationships r31 and r32 are examples of inheritance relationships (is-a relationships) between classes. Specifically, the relationship r31 indicates that the subclass "Client" inherits from the superclass "User." Furthermore, the relationship r32 indicates that the subclass "Staff" inherits from the superclass "User."
[0069] Next, the information processing device 100 executes steps S108 to S110 in FIG. 6 described above (S206). In step S108, the generation unit 122 generates a code conversion prompt using the acquired second model data 562. In step S109, the input unit 123 transmits the code conversion prompt to the LLM server 200. In step S110, the acquisition unit 124 acquires an output message including the source code from the LLM server 200. Then, the output unit 125 displays the output message on the developer terminal 300.
[0070] FIG. 19 is a diagram showing an example of source code converted from the second model data 562 into a specific programming language. Output data 57 is an example of an output message 571 and source code 572 displayed on the screen of the developer terminal 300. The output message 571 is text data added by the LLM during conversion to source code. The source code 572 shows an example in which the superclass C1 "User" is inherited by the subclass C11 "Client" and the subclass C12 "Staff", and members (attributes), methods, etc. are implemented. In other words, the source code 572 shows an example in which a class is super-conceptualized and converted (from the second model data) from the first model data generated based on the requirements specification document 50.
[0071] In this way, the technology according to the present disclosure makes it possible to obtain second model data by further superordinate conceptualization (conversion) of the first model data converted by the LLM using the LLM. This improves the accuracy of object-oriented analysis of the model data. In addition, this embodiment can achieve various effects similar to those of the first and second embodiments described above.
[0072] After converting and displaying the first model data into source code, the information processing device 100 may generate a layering prompt using the first model data, input the layering prompt to the LLM, and acquire the second model data. For example, a developer may view source code displayed on the developer terminal 300 and attempt to superconceptualize a class by transmitting a superconceptualization request to the information processing device 100 via the developer terminal 300. In this case, the information processing device 100 may perform processes subsequent to the generation of the layering prompt in response to the superconceptualization request. In this way, the developer can view the model data and source code and perform reconversion, such as superconceptualization, for a desired version of software. The developer may also modify the model data via the developer terminal 300 and specify the modified model data to transmit a superconceptualization request. In this case, the information processing device 100 may perform processes subsequent to the generation of the layering prompt using the specified model data. As described above, the technology disclosed herein can support more efficient software development based on object-orientation.
[0073] (Other embodiments) The information processing device described above may have a built-in language model such as an LLM. Furthermore, the language models according to the present disclosure do not need to be identical at each stage. The technology according to the present disclosure may use, for example, a first language model that converts a requirement specification statement into a set of statements, a second language model that converts the set of statements into model data in which object-oriented software design has been performed, and a third language model that converts the model data into source code. In this case, the first language model may be a trained model that inputs a requirement specification statement and outputs a set of statements converted from the requirement specification statement. Furthermore, the second language model may be a trained model that inputs a set of statements and outputs model data in which object-oriented software design has been performed from the set of statements. Furthermore, the third language model may be a trained model that inputs model data and outputs source code. In these cases, instruction sentences in the input text to each language model can be omitted. Furthermore, a portion of the first to third language models may be common.
[0074] 20 is a block diagram showing the hardware configuration of the above-described information processing device 100, etc. The information processing device 100 includes a memory 101, a processor 102, and a network interface 103.
[0075] The memory 101 is configured by a combination of volatile memory and nonvolatile memory. The volatile memory is, for example, a volatile storage device such as RAM (Random Access Memory), and is a storage area for temporarily holding information when the processor 102 is operating. The nonvolatile memory is, for example, a nonvolatile storage device such as a hard disk or flash memory. The memory 101 stores at least a computer program that implements the processing of an information processing method (a software development support method based on object-orientation) in the information processing device 100 according to the present disclosure. Note that the memory 101 may include storage located away from the processor 102. In this case, the processor 102 may access the memory 101 via an I / O (Input / Output) interface (not shown).
[0076] The processor 102 is a control device that controls each component of the information processing device 100. The processor 102 reads and executes software (computer programs) from the memory 101. As a result, the processor 102 realizes the functions of the reception unit 121, the generation unit 122, the input unit 123, the acquisition unit 124, and the output unit 125. That is, the processor 102 performs processing of the information processing method in the information processing device 100 according to the present disclosure. The processor 102 may be, for example, a microprocessor, an MPU (Multi Processing Unit), or a CPU (Central Processing Unit). The processor 102 may also include multiple processors.
[0077] The network interface 103 may be used to communicate with a network node. The network interface 103 may include, for example, a network interface card (NIC) conforming to the IEEE 802.3 series. IEEE stands for Institute of Electrical and Electronics Engineers. The network interface 103 may also include a wireless local area network (LAN), a wired LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0078] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0079] Each drawing is merely an example for describing one or more embodiments. Each drawing may relate not only to one particular embodiment, but also to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessary to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.
[0080] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. (Appendix A1) a first input means for inputting a first instruction sentence for converting a requirement specification sentence, in which the software requirement specification is written in a natural language, into a set of sentences in accordance with a predetermined sentence pattern, and a first input text including the requirement specification sentence, into a predetermined language model; a first acquisition means for acquiring the set of sentences converted from the requirement specification sentence based on the first instruction sentence using the language model; a second input means for inputting a second input text including a second instruction sentence for converting the set of sentences into model data by an object-oriented software design method and the set of sentences to the language model; a second acquiring means for acquiring model data converted from the set of sentences based on the second instruction sentence using the language model; An information processing device comprising: (Appendix A2) The model data is data converted from the set of statements into information that defines the components of object-oriented software and the relationships between each component. 10. The information processing device according to claim 1, (Appendix A3) The model data includes information in which, among the components extracted from the set of sentences, those determined by a language model to have a high degree of similarity in logical structure are superordinated, and the components are hierarchically structured to define the relationships. 10. The information processing device according to claim 9, wherein the information processing device is a (Appendix A4) The model data includes a class structure that associates each word included in the set of sentences with each other. 10. The information processing device according to claim 9, wherein the information processing device is a device for processing information. (Appendix A5) The model data is The components include classes, methods, attributes, arguments, and return values; The relationships between components include inheritance relationships between classes, ownership relationships between classes, correspondence between classes and methods, and correspondence between methods and arguments / return values. 10. The information processing device according to claim 9, wherein the information processing device is a device for processing information. (Appendix A6) a third input means for inputting a third input text including a third instruction sentence for converting the model data into source code based on the specifications of a specific programming language and the model data to the language model; and a third acquisition means for acquiring the source code converted from the model data based on the third instruction sentence using the language model. An information processing device according to any one of appendices A1 to A5. (Appendix A7) the second input means inputs to the language model a fourth instruction sentence for reconverting the components so as to hierarchically structure those components that the language model determines to have a high degree of similarity in logical structure among the components included in the first model data acquired by the second acquisition means into a higher-level concept, and a fourth input text including the first model data; The second acquisition means acquires, as the model data, second model data reconverted from the first model data based on the fourth instruction sentence by the language model. An information processing device according to appendix A1 or A2. (Appendix A8) The set of sentences is converted from the requirement specification sentence by the language model according to a plurality of basic sentence patterns in the grammar of a specific natural language. An information processing device according to any one of appendices A1 to A7. (Appendix B1) The computer a first instruction sentence for converting a requirement specification sentence, in which the software requirement specification is written in a natural language, into a set of sentences in accordance with a predetermined sentence pattern, and a first input text including the requirement specification sentence are input to a predetermined language model; obtaining the set of sentences converted from the requirement specification sentence based on the first instruction sentence using the language model; inputting a second input text including a second instruction sentence for converting the set of sentences into model data using an object-oriented software design method and the set of sentences into the language model; obtaining model data converted from the set of sentences based on the second instruction sentence using the language model; Information processing methods. (Appendix C1) a first input process for inputting a first instruction sentence for converting a requirement specification sentence, in which the software requirement specification is written in a natural language, into a set of sentences in accordance with a predetermined sentence pattern, and a first input text including the requirement specification sentence, into a predetermined language model; a first acquisition process for acquiring the set of sentences converted from the requirement specification sentence based on the first instruction sentence using the language model; a second input process for inputting a second input text including a second instruction sentence for converting the set of sentences into model data using an object-oriented software design method and the set of sentences to the language model; a second acquisition process of acquiring model data converted from the set of sentences based on the second instruction sentence using the language model; An information processing program that causes a computer to execute the above.
[0081] Some or all of the elements (e.g., configurations and functions) described in Appendix A2 to Appendix A8 that are dependent on Appendix A1 {e.g., device} may also be dependent on Appendix B1 {e.g., method} and Appendix C1 {e.g., program} in the same dependency relationship as Appendix A2 to Appendix A8. Some or all of the elements described in any appendix may be applied to various hardware, software, recording means for recording software, systems, and methods. [Explanation of symbols]
[0082] 1. Information processing equipment 11 First input section 12 First Acquisition Section 13 Second input section 14 Second Acquisition Section 1000 Software Development Support System N Network 100 Information processing device 121 Reception 122 Generation part 123 Input section 124 Acquisition Department 125 Output section 200 LLM servers 300 Developer Terminals 40 Requirements specification document 41 Prompts for Analyzing Textual Data 411 Instruction text 412 Requirement specification document 42 Output Data 421 Output Message 422 sentence set 43 Model conversion prompt 431 Instruction text 432 sentence sets 44 Output Data 441 Output Message 442 model data 45 Code conversion prompt 451 Instruction text 452 model data 46 Output Data 461 Output Message 462 Source Code 50 Requirement specification documents 52 Output Data 521 Output Message 522 sentence set 54 Output Data 541 Output Message 542 model data 55 Hierarchical prompts 551 Instruction text 552 model data 56 Output Data 561 Output Message 562 model data 57 Output Data 571 Output Messages 572 Source Code r11 related r12 related r13 related r21 related r22 related r31 related r32 related C1 Superclass C11 subclass C12 subclass 101 Memory 102 processors 103 Network Interface
Claims
1. a first input means for inputting a first instruction sentence for converting a requirement specification sentence, in which the software requirement specification is written in a natural language, into a set of sentences in accordance with a predetermined sentence pattern, and a first input text including the requirement specification sentence, to a predetermined language model; a first acquisition means for acquiring the set of sentences converted from the requirement specification sentence based on the first instruction sentence using the language model; a second input means for inputting a second input text including a second instruction sentence for converting the set of sentences into model data by an object-oriented software design method and the set of sentences to the language model; a second acquiring means for acquiring model data converted from the set of sentences based on the second instruction sentence using the language model; An information processing device comprising:
2. The model data is data converted from the set of statements into information that defines the components of object-oriented software and the relationships between each component. The information processing device according to claim 1 .
3. The model data includes information in which, among the components extracted from the set of sentences, those determined by the language model to have a high degree of similarity in logical structure are superordinated, and the components are hierarchically structured to define the relationships. The information processing device according to claim 2 .
4. The model data includes a class structure that associates each word included in the set of sentences with each other.
4. The information processing device according to claim 2 or 3.
5. The model data is The components include classes, methods, attributes, arguments, and return values; The relationships between components include inheritance relationships between classes, ownership relationships between classes, correspondence between classes and methods, and correspondence between methods and arguments / return values.
4. The information processing device according to claim 2 or 3.
6. a third input means for inputting a third input text including a third instruction sentence for converting the model data into source code based on the specifications of a specific programming language and the model data to the language model; and a third acquisition means for acquiring the source code converted from the model data based on the third instruction sentence using the language model.
3. The information processing device according to claim 1 or 2.
7. the second input means inputs to the language model a fourth instruction sentence for reconverting components so as to hierarchically structure those components that the language model determines to have a high degree of similarity in logical structure among the components included in the first model data acquired by the second acquisition means, and a fourth input text including the first model data; The second acquisition means acquires, as the model data, second model data reconverted from the first model data by the language model based on the fourth instruction sentence.
3. The information processing device according to claim 1 or 2.
8. The set of sentences is converted from the requirement specification sentence by the language model according to a plurality of basic sentence patterns in the grammar of a specific natural language.
3. The information processing device according to claim 1 or 2.
9. The computer a first instruction sentence for converting a requirement specification sentence, in which the software requirement specification is written in a natural language, into a set of sentences in accordance with a predetermined sentence pattern, and a first input text including the requirement specification sentence are input to a predetermined language model; obtaining the set of sentences converted from the requirement specification sentence based on the first instruction sentence using the language model; inputting a second input text including a second instruction sentence for converting the set of sentences into model data using an object-oriented software design method and the set of sentences to the language model; obtaining model data converted from the set of sentences based on the second instruction sentence using the language model; Information processing methods.
10. a first input process for inputting a first instruction sentence for converting a requirement specification sentence, in which the software requirement specification is written in a natural language, into a set of sentences in accordance with a predetermined sentence pattern, and a first input text including the requirement specification sentence, into a predetermined language model; a first acquisition process for acquiring the set of sentences converted from the requirement specification sentence based on the first instruction sentence using the language model; a second input process for inputting a second input text including a second instruction sentence for converting the set of sentences into model data by an object-oriented software design method and the set of sentences to the language model; a second acquisition process for acquiring model data converted from the set of sentences based on the second instruction sentence using the language model; An information processing program that causes a computer to execute the above.
Citation Information
Patent Citations
Program development support device, program development support method, program, and computer-readable storage medium recording the program
JP2004287695A