AI Model-Based Test Case Generation Method, Device, System, and Medium
The AI-driven method automates the generation of testing cases by parsing textual and graphical data, addressing inefficiencies and reliability issues in existing AI-based methods, thereby enhancing test case generation efficiency and accuracy.
Patent Information
- Application Number
- CN202510535065.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In the prior art, testers need to manually input the flowchart into the OCR module to extract text, resulting in insufficient efficiency and reliability of AI model generation test cases, and errors are not easily discovered.
The OCR module automatically extracts the image content in the requirement document, distinguishes the document type based on title keywords and business keywords, automatically generates test cases, and records reasoning information to improve reliability.
Automatic test case generation is realized, the generation efficiency is improved, and the reliability of test cases is enhanced by recording inference information, making it easier for testers to make judgments.
Smart Images

Figure CN120066974B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a test case generation method, device, system, and medium based on an AI model. Background Art
[0002] Currently, property systems involve many business modules, such as financial modules and property owner management modules. When developing business modules, testers need to construct corresponding test cases to test each business function. Manually constructing test cases can no longer meet current requirements. Some AI models have emerged on the market that can automatically generate test cases according to the test requirements input by testers, effectively improving the generation efficiency of test cases.
[0003] However, the test requirements input into the AI model are usually in text form. The test documents constructed by testers not only include text-based requirement descriptions but also some flowcharts related to functional logic. Although relevant text can be extracted from the flowcharts using Optical Character Recognition (OCR), testers need to manually input the images into the OCR module for text extraction. Therefore, before inputting test requirements into the AI model, testers still need to manually perform a large number of information processing operations. Once improper information processing leads to incorrect test cases output by the AI model, it is not easy to detect, and the generation efficiency and reliability of test cases cannot be guaranteed. Summary of the Invention
[0004] The present invention aims to at least solve one of the technical problems existing in the prior art. For this purpose, the present invention provides a test case generation method, device, system, and medium based on an AI model, which can automatically generate test cases based on requirement documents and improve the generation efficiency and reliability of test cases.
[0005] In a first aspect, an embodiment of the present invention provides a test case generation method based on an AI model, which is applied to a use case generation tool. The use case generation tool includes an OCR module and an AI model for generating test cases. The method includes:
[0006] Obtain a requirement document, at least one title keyword, and business keyword input by a user. When a target title is traversed in the requirement document based on the title keyword, create a first document and a second document based on the target title, where the document types of the first document and the second document are different;
[0007] When a target image is traversed, extract first content from the target image based on the OCR module, and input the first content into the first document, where any one of the business keywords is recorded in the image path of the target image;
[0008] When a new target title is traversed based on the title keyword, the first document is input into the AI model to obtain a target test case and first inference information. The target test case is written into the second document, and the first inference information is written into the first document, where the first inference information is used to indicate the derivation process of the target test case;
[0009] Create a new first document and a new second document based on the new target title, and continue to traverse the requirement document.
[0010] According to some embodiments of the present invention, the use case generation tool is preset with an extraction identifier, and the initial value of the extraction identifier is empty. When a target title is traversed in the requirement document based on the title keyword, creating the first document and the second document based on the target title includes:
[0011] Start traversing after setting the extraction identifier to false, where when the extraction identifier is false, the traversal keyword is the title keyword;
[0012] When the target title is traversed based on the traversal keyword, set the extraction identifier to true;
[0013] In response to the extraction identifier changing from false to true, create the first document and the second document based on the target title traversed this time, and switch the traversal keyword to the service keyword;
[0014] When the extraction identifier is true, input the image path of the traversed target image into the OCR module, so that the OCR module obtains the target image based on the image path and extracts the first content.
[0015] According to some embodiments of the present invention, based on the first content extracted from the target image by the OCR module, inputting the first content into the first document includes:
[0016] Create a content queue based on the target title;
[0017] When the extraction identifier is true, input the image path of the traversed target image into the OCR module, so that the OCR module obtains the target image based on the image path and extracts the first content, and write the first content into the content queue;
[0018] When any target paragraph is traversed, and the target paragraph does not record the target title, and the extraction identifier is true, extract the body text of the target paragraph as the second content and write it into the content queue;
[0019] When a new target title is traversed based on the title keyword, change the extraction flag from true to false;
[0020] In response to the extraction flag changing from true to false, write the content queue to the first document. After completing the writing of the first document, change the extraction flag from false to true again, empty the content queue, and then continue to traverse the requirement document.
[0021] According to some embodiments of the present invention, extracting first content from the target image based on the OCR module includes:
[0022] Performing image recognition on the target image through the use case generation tool, and determining a plurality of recognition logic boxes from the target image, wherein each of the recognition logic boxes is connected to another recognition logic box by at least one directed line segment;
[0023] Extracting target strings corresponding to the respective recognition logic boxes from the target image based on the OCR module;
[0024] Constructing the first content based on a logical relationship between the plurality of target strings and the respective recognition logic boxes indicated by the directed line segments.
[0025] According to some embodiments of the present invention, extracting target strings corresponding to the respective recognition logic boxes from the target image based on the OCR module includes:
[0026] Performing a line-by-line traversal of the target image based on the OCR module to obtain a plurality of recognition text blocks, and extracting recognition strings corresponding to the respective recognition text blocks, wherein a distance between two adjacent recognition text blocks is greater than a preset threshold;
[0027] Determining a first text block and a second text block from the plurality of recognition text blocks, and determining a first serial number of each of the recognition text blocks, wherein the first text block has a character count greater than 1, and the second text block has a character count equal to 1;
[0028] Determining a second serial number of each of the recognition logic boxes and the target line segments line by line through the use case generation tool, wherein when a shape of the recognition logic box is a preset target shape, the directed line segments connected to the recognition logic box are determined as target line segments, and the recognition logic box corresponding to the target shape is used to indicate conditional judgment;
[0029] Based on the first serial number and the second serial number, determine the target string from the recognition string of the first text block corresponding to the recognition logic box, and determine the second text blocks associated with each of the target line segments.
[0030] According to some embodiments of the present invention, constructing the first content based on the logical relationships between the multiple target strings and the respective recognition logic boxes indicated by the directed line segments includes:
[0031] Based on any first logic box, determine the corresponding second logic box based on the corresponding directed line segment, where the first logic box is connected to the starting point of the directed line segment, and the second logic box is connected to the ending point of the directed line segment;
[0032] When the shape of the first logic box is the target shape, write the target strings of the respective second logic boxes into a second paragraph, where the second paragraph is a subordinate paragraph of the first paragraph, and the first paragraph records the target string corresponding to the first logic box;
[0033] Alternatively, when the shape of the first logic box is not the target shape, construct respective third paragraphs based on the target strings of the respective second logic boxes, where the third paragraphs are sibling paragraphs of the first paragraph.
[0034] According to some embodiments of the present invention, writing the first inference information into the first document includes:
[0035] Determine the respective second inference information corresponding to each of the target strings in the first inference information;
[0036] When the target string corresponds to any of the second paragraphs, and the second inference information does not indicate the target string of the associated first paragraph or the second text block, generate an inference exception information.
[0037] In a second aspect, an embodiment of the present invention provides a test case generation device based on an AI model, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the method for generating a test case based on an AI model as described in the first aspect above.
[0038] In a third aspect, an embodiment of the present invention provides a test case generation system based on an AI model, including the test case generation device based on an AI model as described in the second aspect above.
[0039] Fourthly, an embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions for executing the test case generation method based on an AI model as described in the first aspect above.
[0040] The test case generation method based on an AI model according to an embodiment of the present invention has at least the following beneficial effects: obtaining a requirement document, at least one title keyword, and a service keyword input by a user; when a target title is traversed in the requirement document based on the title keyword, creating a first document and a second document based on the target title, where the document types of the first document and the second document are different; when a target image is traversed, extracting first content from the target image based on the OCR module and inputting the first content into the first document, where any one of the service keywords is recorded in the image path of the target image; when a new target title is traversed based on the title keyword, inputting the first document into the AI model to obtain a target test case and first inference information, writing the target test case into the second document, and writing the first inference information into the first document, where the first inference information is used to indicate the derivation process of the target test case; creating a new first document and a new second document based on the new target title and continuing to traverse the requirement document. According to the technical solution of the embodiment of the present invention, it is possible to automatically traverse the requirement document, automatically distinguish different requirement modules by the target title and create corresponding first and second documents, record the first content identified by the OCR module from the flowchart in the first document, automatically generate test cases after inputting the first document into the AI model, improve the generation efficiency of test cases, and record the inference information of the test cases in the first document, providing a basis for the tester to judge the reliability of the test cases. Description of the Drawings
[0041] Figure 1 is a schematic diagram of the principle provided by an embodiment of the present invention;
[0042] Figure 2 is a flowchart of the test case generation method based on an AI model provided by another embodiment of the present invention;
[0043] Figure 3 is a diagram for prompting the extraction principle of the OCR module provided by another embodiment of the present invention;
[0044] Figure 4 is a structural diagram of the test case generation device based on an AI model provided by another embodiment of the present invention. Detailed Embodiments
[0045] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0046] In the description of the present invention, it should be understood that with respect to the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., is based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as limiting the present invention.
[0047] In the description of the present invention, the meaning of "a number of" is one or more, the meaning of "a plurality of" is two or more. Understandings such as "greater than", "less than", "exceeding", etc. do not include the recited number, and understandings such as "above", "below", "within", etc. include the recited number. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0048] In the description of the present invention, unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.
[0049] An embodiment of the present invention provides a method, apparatus, system, and medium for generating test cases based on an AI model. The method for generating test cases based on the AI model includes: obtaining a requirement document, at least one title keyword, and a business keyword input by a user. When a target title is traversed in the requirement document based on the title keyword, a first document and a second document are created based on the target title, where the document types of the first document and the second document are different; when a target image is traversed, first content is extracted from the target image based on the OCR module, and the first content is input into the first document, where any one of the business keywords is recorded in the image path of the target image; when a new target title is traversed based on the title keyword, the first document is input into the AI model to obtain a target test case and first inference information, the target test case is written into the second document, and the first inference information is written into the first document, where the first inference information is used to indicate the derivation process of the target test case; new first and second documents are created based on the new target title, and the requirement document is continuously traversed. According to the technical solution of the embodiment of the present invention, by automatically traversing the requirement document, different requirement modules can be automatically distinguished by the target title and corresponding first and second documents can be created. The first document records the first content recognized by the OCR module from the flowchart. After the first document is input into the AI model, test cases are automatically generated, improving the generation efficiency of the test cases. At the same time, the inference information of the test cases is recorded in the first document, providing a basis for the testers to judge the reliability of the test cases in terms of the inference process.
[0050] The following is a further elaboration of the technical solution of the embodiment of the present invention based on the principle schematic diagram shown in the appendix Figure 1 as shown.
[0051] Refer to Figure 2 , Figure 2 which is a flowchart of a method for generating test cases based on an AI model provided by an embodiment of the present invention, applied to a use case generation tool. The use case generation tool includes an OCR module and an AI model for generating test cases. The method for generating test cases based on the AI model includes but is not limited to the following steps:
[0052] S10. Obtain a requirement document, at least one title keyword, and a business keyword input by a user. When a target title is traversed in the requirement document based on the title keyword, create a first document and a second document based on the target title, where the document types of the first document and the second document are different;
[0053] S20. When a target image is traversed, extract first content from the target image based on the OCR module, and input the first content into the first document, where any one of the business keywords is recorded in the image path of the target image;
[0054] S30. When a new target title is traversed based on the title keywords, input the first document into the AI model to obtain the target test cases and the first inference information, write the target test cases into the second document, and write the first inference information into the first document, where the first inference information is used to indicate the derivation process of the target test cases.
[0055] S40. Create a new first document and a new second document based on the new target title, and continue to traverse the requirement document.
[0056] It should be noted that the OCR module in this embodiment is a functional module constructed based on conventional OCR technology. The technology of OCR for extracting text is well-known to those skilled in the art. The embodiments of the present invention do not involve the improvement of OCR technology. Instead, when text can be extracted, the process of further processing and sending the text is improved to obtain the OCR module in this embodiment. For the specific principle, refer to the subsequent description.
[0057] It should be noted that the AI model can be any model that can present the inference process and generate test cases. The principle of the AI model for generating test cases in this embodiment will not be elaborated here. For example, large models such as OpenAI GPT or DEEPSEEK can be used. The inference process of the AI model can be explicitly displayed in the form of text, so that the test case generation tool can obtain the relevant text as the first inference information. Of course, the inference process of the AI model can also be recorded in the relevant work logs. The test case generation model can access the work logs of the AI model to extract the inference process related to each first content as the first inference information.
[0058] It should be noted that the requirement document in this embodiment records multiple titles and flowcharts. As Figure 1 shown, the titles in the requirement document can include "Requirement Background", "Function Description", "Trigger", etc. There can be at least one image after each title, such as a flowchart, an example diagram, etc. To improve the flexibility of test cases, not all content in the requirement document needs to be used as test cases. There are also some content and images for explaining to testers. These information do not need to be input into the AI model. Usually, the content described in the requirement document is organized. The content following a title is usually related to this title. Therefore, testers can know which content under which titles in the requirement document is useful for generating test cases. Thus, as Figure 1 shown, after importing the requirement document into the test case generation tool in this embodiment, input at least one title keyword and at least one business keyword in the text title option. As Figure 1The title keywords in it include "function description" and "trigger", which makes the explanatory text under the above "requirement background" title not be written into the first document and thus not input into the AI model. Similarly, the naming method of each image in the requirement document usually conforms to the specification and does not use randomly generated characters by the computer. Therefore, in this embodiment, the required images can be described by business keywords. For example Figure 1 The input business keyword shown is "business process". File names are usually referenced in the image address of the flowchart. When the image naming conforms to the rules, the corresponding flowchart can be matched based on "business process", while some sample diagrams are not extracted because they do not have the above business keywords, avoiding the OCR module from extracting useless strings.
[0059] It should be noted that the document types of the first document and the second document in this embodiment are different. As Figure 1 shown, the first document is used to record strings and inference information, so the first document uses a Word document. The second document is used to record test cases, and test cases are usually in JSON format, so the second document uses an Excel document.
[0060] It should be noted that according to the description of the above embodiment, the requirement document is usually a document with a standardized format. As Figure 1 shown, in the requirement document, the content of the current part is first described by the title, and then the relevant images and text descriptions are recorded under the title. When constructing test cases, each title corresponds to a test object. For example Figure 1 shown, the content corresponding to the "function description" title is the specific function, and tests need to be executed according to the content of the function description. The content corresponding to the "trigger" title is the specific trigger condition, etc., and tests need to be executed according to the content of the trigger description. Based on this, this embodiment uses the title as the basis for dividing test case objects. First, the requirement document is traversed line by line based on the title keywords. When any target title is detected, the corresponding first document and second document are created according to the target title, and the content corresponding to each title is the object for one operation in this embodiment.
[0061] Exemplarily, as Figure 1 shown, the titles seen are "function description" and "trigger". When traversing the requirement document, when the target title "function description" is traversed, the first document and the second document are created with the name "function description". When step S40 is executed, since the test cases corresponding to "function description" have been generated, when the new target title "trigger" is traversed, the first document and the second document are created with the name "trigger". Each target title corresponds to a group of independent first documents and second documents, and the first documents corresponding to different target titles are not shared with each other, nor are the second documents.
[0062] It should be noted that after creating the first document and the second document, the images obtained by continuing to traverse the requirements document are the content under the target title. For example Figure 1 as shown in Figure 1 , after traversing to "Function Description" and before traversing to the new target title "Trigger", all the images traversed can be determined to belong to "Function Description". According to the description of the above embodiments, not all images are flowcharts. Therefore, in this embodiment, the images traversed are screened based on business keywords. When the image path records a business keyword, it is determined as the target image, and after extracting the first content through the OCR module, it is input into the first document.
[0063] Exemplarily, as Figure 1 shown in Figure 1 , taking "Business Process" as an example of the business keyword, when the corresponding image path 1 records "Business Process" after traversing to the image, it is determined as the target image and input into the OCR module for content extraction. If the corresponding image path does not record "Business Process" but records keywords such as "Business Description" or "Principle Description", since they do not belong to the business keywords input by the user, they are not determined as the target images. It should be noted that the testers are professional personnel with development capabilities, so they can fill in accurate business keywords according to actual needs. The target images in this embodiment are obtained through precise matching and do not perform related operations such as fuzzy matching.
[0064] It should be noted that after inputting the target image into the OCR module, relevant text can be extracted from the flowchart and output in the form of a string. Therefore, the first content in this embodiment is the process steps recorded in the target image. The number of target images can be arbitrary, and step S20 can be executed each time a target image is traversed.
[0065] It should be noted that after extracting the first content of the target image in the above manner, this embodiment uses traversing to a new target title as the trigger condition for ending a content collection. This embodiment generates a test case based on each target title. Therefore, the first documents for different target titles are independent of each other. After matching to a new target title, the current first document is input into the AI model to obtain the test case, and new first and second documents are created based on the new target title, so that the first content obtained by continuing to traverse is written into the new first document.
[0066] It should be noted that since the content recorded in the first document is the business logic indicated by each flowchart, test cases can be generated for the relevant business logic, such as Figure 1As shown, in this embodiment, the generated target test cases in JSON format are written into a second document in Excel format for subsequent viewing and management. This embodiment only completes the generation of the target test cases and does not involve subsequent processes such as the running of the test cases, so no more elaboration will be made here.
[0067] It should be noted that the AI model can also output the reasoning process based on the first document. Since the first document has been used to generate the target test model at this time, adding text content to the first document will not affect the target test cases. For example, Figure 1 As shown, in this embodiment, the first reasoning information is input into the first document that is no longer input into the AI model, enabling testers to view the first reasoning information in the first document to determine whether there are logical errors in the construction process of the target test cases and improving the reliability of the test cases.
[0068] It should be noted that after creating the new first document and second document, the steps based on the above embodiment can be repeatedly executed until the requirement document is traversed, and no more repetition will be made here.
[0069] In addition, in one embodiment, the use case generation tool is preset with an extraction identifier, and the initial value of the extraction identifier is empty. Step S10 specifically includes but is not limited to the following steps:
[0070] S11, start traversing after setting the extraction identifier to false. Among them, when the extraction identifier is false, the traversal keyword is the title keyword;
[0071] S12, when the target title is traversed based on the traversal keyword, set the extraction identifier to true;
[0072] S13, in response to the extraction identifier changing from false to true, create the first document and the second document based on the target title traversed this time, and switch the traversal keyword to the business keyword;
[0073] S14, when the extraction identifier is true, input the image path of the traversed target image into the OCR module so that the OCR module can obtain the target image based on the image path and extract the first content.
[0074] It should be noted that in this embodiment, an extraction identifier is preset in the use case generation tool, and its initial value is empty (NULL). After starting the traversal, the extraction identifier is set to false. In this embodiment, the extraction identifier is used as the basis for selecting the traversal keyword. When the extraction identifier is false, the traversal keyword uses the title keyword, which is used to identify the target title during the traversal. When the extraction identifier is true, the traversal keyword uses the business keyword, which is used to identify the target image during the traversal and serves as a judgment condition for content extraction, instructing the use case generation tool to extract the target image for OCR recognition during the traversal.
[0075] It should be noted that the value of the extraction identifier is also used to trigger the generation of the first document and the second document. When the extraction identifier is false, the target title can be matched based on the traversal keyword. At this time, the use case generation tool does not need to extract the target title, but sets the extraction identifier to true. When the extraction identifier changes from false to true, it triggers the use case generation tool to generate the first document and the second document, and names the documents using the matched target title.
[0076] It should be noted that when the extraction identifier is true, the traversal keyword is the business keyword. After traversing to the target image, it is further combined with the value of the extraction identifier to determine whether the image is extracted. If the extraction identifier is true and the target image is traversed, and both conditions are met, the corresponding image path is input to the OCR module for content extraction. When the extraction identifier is false, regardless of whether the traversal keyword is the business keyword, content extraction will not be performed due to the value of the extraction identifier, effectively ensuring that the content of the target image is extracted at the correct time and position.
[0077] In addition, in one embodiment, step S20 specifically includes but is not limited to the following steps:
[0078] S21, creating a content queue based on the target title;
[0079] S22, when the extraction identifier is true, input the image path of the traversed target image to the OCR module, so that the OCR module obtains the target image based on the image path and extracts the first content, and writes the first content into the content queue;
[0080] S23, when traversing to any target paragraph, and the target paragraph does not record the target title, and the extraction identifier is true, extract the text of the target paragraph as the second content and write it into the content queue;
[0081] S24, when traversing to a new target title based on the title keyword, change the extraction identifier from true to false;
[0082] S25. In response to the extraction flag changing from true to false, write the content queue to the first document. After completing the writing of the first document, change the extraction flag from false to true again. After emptying the content queue, continue traversing the requirement document.
[0083] It should be noted that the use case generation tool in this embodiment includes an OCR module and an AI model, and creates a first document and a second document internally. The input of the OCR module and the AI model is data transmission, so the consumed read and write resources are less. For the first document and the second document, read and write threads need to be allocated. During traversal, according to the description of the above embodiment, each target title may include multiple target images and text contents. If the first document is written each time the first content is extracted, it will increase the number of file accesses, resulting in frequent allocation of read and write threads. The time consumed by allocating threads and accessing documents is much greater than the time for writing information. Therefore, in this embodiment, a content queue is created, and all the traversed first contents and second contents are written into the content queue. Trigger a write to the first document when a new target title is traversed, which can effectively save system resources and improve the efficiency of data processing.
[0084] It should be noted that when the extraction flag is true, the first content extracted by the OCR module is written into the content queue. By using the first-in-first-out characteristic of the content queue, it is ensured that when the content queue is written to the first document, the multiple first contents written can reflect the traversal order.
[0085] It should be noted that this embodiment can not only extract the first content in the flowchart. When traversing a target paragraph, if the target paragraph does not record the target title, it can be determined that the target paragraph is not a title but a text description related to the content indicated by the target title. Combined with the extraction flag being true, the Python paragraph.text.strip() function can be used to extract the text body as the second content, so that the text records in the requirement document can also be incorporated into the reasoning process of the AI model. The content of the target paragraph is usually a relevant supplementary description of the flowchart, so it can improve the reasoning accuracy of the AI model and the reliability of the target test case.
[0086] It is worth noting that when traversing to a new target title, this embodiment changes the extraction flag from true to false, triggers the writing of the content queue to the first document, and changes the extraction flag from false to true again after the writing is completed. At this time, it can further trigger the creation of a new first document and a new second document based on the new target title. Therefore, this embodiment uses the change of the parameter value of the extraction flag as the trigger condition to realize the automatic cycle of document creation, content extraction and use case generation, and can trigger the generation of corresponding target test cases based on each target title during the traversal process.
[0087] In addition, in one embodiment, step S20 specifically includes but is not limited to the following steps:
[0088] S25, perform image recognition on the target image through a use case generation tool, and determine multiple recognition logic boxes from the target image, where each recognition logic box is connected to another recognition logic box by at least one directed line segment;
[0089] S26, extract the target strings corresponding to each recognition logic box from the target image based on the OCR module;
[0090] S27, construct the first content based on the logical relationship between the multiple target strings and the various recognition logic boxes indicated by the directed line segments.
[0091] It should be noted that the use case generation tool in this embodiment can also pre-deploy common image recognition functions. The image recognition function in this embodiment only needs to have simple graphic recognition. Flowcharts usually use standard graphics to represent the logic of each logic box. For example, common rectangular boxes are used to represent regular steps, and diamond boxes are used to represent judgment steps, etc. In this embodiment, image recognition is performed based on the target image, multiple recognition logic boxes are determined from the target image, and at the same time, the directed line segments connecting each recognition logic box are recognized, so as to determine whether there is a sequential order between two adjacent recognition logic boxes, providing a basis for the string arrangement of the first content.
[0092] Exemplarily, as Figure 3 shown, through simple image recognition of the flowchart, 6 rectangular recognition logic boxes and 1 diamond recognition logic box can be determined therefrom. The direction of the directed line segment can be determined according to the arrow direction, and no more details will be elaborated here.
[0093] It should be noted that for a flowchart, the content recorded in each recognition logic box is a process step. After the target image is input into the OCR module, the OCR module can extract the target string corresponding to each recognition logic box. For example, the target string extracted from the first recognition logic box as Figure 3 shown is "AAAAAA", and so on.
[0094] It should be noted that after obtaining the target strings of each recognition logic box, the directed line segments of the flowchart can indicate the logical relationship between each recognition logic box. Based on the logical relationship, the target strings can be logically arranged so that the first content not only records all the image content but also can be arranged according to the logical relationship indicated in the flowchart, that is, the first content in this embodiment can represent the process logic of the target image.
[0095] Exemplarily, as Figure 3As shown, taking the extracted target strings "AAAAAA", "BBB", "CCC", "DDDDD", and "FFF" as examples, according to the identified directed line segments, the steps can be determined to be "AAAAAA", "BBB", and "CCC" in sequence, while "DDDDD" and "FFF" are two branches after the execution judgment of step "CCC", and the constructed first content can be referred to Figure 3 as shown.
[0096] In addition, in an embodiment, step S26 specifically includes but is not limited to the following steps:
[0097] S261, traversing the target image line by line based on the OCR module to obtain multiple recognized text blocks, where the distance between two adjacent recognized text blocks is greater than a preset threshold, and extracting the recognized strings corresponding to each recognized text block;
[0098] S262, determining a first text block and a second text block from the multiple recognized text blocks, and determining the first serial number of each recognized text block, where the number of characters in the first text block is greater than 1, and the number of characters in the second text block is equal to 1;
[0099] S263, determining the second serial number of each recognized logic box and the target line segment line by line through a use case generation tool. When the shape of the recognized logic box is a preset target shape, the directed line segments connected to the recognized logic box are determined as target line segments, and the recognized logic box corresponding to the target shape is used to indicate conditional judgment;
[0100] S264, based on the first serial number and the second serial number, determining the recognized string of the first text block corresponding to the recognized logic box as the target string, and determining the second text block associated with each target line segment.
[0101] It should be noted that when the OCR module performs text extraction, it usually outputs strings in units of lines, while the text of the flowchart is recorded in each recognized logic box. The text within the same recognized logic box is relatively close, and the text between different recognized logic boxes is relatively far due to the existence of directed line segments. Based on this, in this embodiment, a preset threshold is configured in the OCR module. When the distance between strings is less than the preset threshold, it can be determined that the strings belong to the same recognized text block, and the characters extracted from the same recognized text block can be arranged into a recognized string.
[0102] Exemplarily, such as Figure 3As shown, the OCR module detected "AAA" when recognizing the first line and also detected "AAA" when recognizing the second line. The distance between the two lines is less than the preset threshold and they are grouped into the same recognized text block. When recognizing the third line, "BBB" was detected. Since the distance between "BBB" and the "AAA" in the first line is greater than the preset threshold, "BBB" is grouped into the second recognized text block, and so on.
[0103] It should be noted that in this embodiment, the first text block and the second text block are distinguished by the number of characters. The first text block is used to represent the steps in the flowchart, and the second text block is used to represent the basis for judgment. For example Figure 3 As shown, the second text block includes "Yes" in the third line and "No" in the fourth line, which is used for the judgment of the content "CCC" in the first text block of the diamond box in the third line.
[0104] It should be noted that in this embodiment, each recognized text block is numbered in the traversal order. As Figure 3 shown, traversing from top to bottom, the serial numbers of each recognized text block are 1 to 9 in sequence. Similarly, the test case generation tool can identify each recognition logic box, and also numbers the recognition logic box and the target line segment according to the traversal order from top to bottom and from left to right. The target line segment in this embodiment is the directed line segment connected to the recognition logic box for conditional judgment. For example Figure 3 the two directed line segments connected after the diamond box in, because the target line segment usually records the basis for conditional judgment. For example Figure 3 the "Yes" or "No" in, or it can also be the value of a specific parameter.
[0105] It should be noted that the OCR module can only extract text. In order to endow the target string extracted by the OCR module with the logic represented by the flowchart, in this embodiment, according to the principle of pairing the first serial number and the second serial number, the recognition logic box corresponding to each recognized string is determined. For example Figure 3 in, the serial number 1 is used to correspond to the first recognition logic box and the recognized string "AAAAAA", and the serial number 4 is used to correspond to the first target line segment and the corresponding recognized string "Yes", thereby establishing the relationship between the logic of the flowchart and the recognized string, and ensuring that the first content can carry the process logic of the target image.
[0106] In addition, in one embodiment, step S27 specifically includes but is not limited to the following steps:
[0107] S271, based on any first logic box, determine the corresponding second logic box based on the corresponding directed line segment, where the first logic box is connected to the starting point of the directed line segment, and the second logic box is connected to the ending point of the directed line segment;
[0108] S272. When the shape of the first logical box is the target shape, write the target strings of the respective second logical boxes into the second paragraph, where the second paragraph is a subordinate paragraph of the first paragraph, and the first paragraph records the target string corresponding to the first logical box.
[0109] S273. When the shape of the first logical box is not the target shape, construct respective third paragraphs based on the target strings of the respective second logical boxes, where the third paragraphs are sibling paragraphs of the first paragraph.
[0110] It should be noted that according to the description of the above embodiments, the recognition logical boxes determined by the application example generation tool and the recognition text blocks extracted by the OCR module can be matched using serial numbers. In this embodiment, this correspondence is used to sort the respective recognition text blocks. Since the flowchart is in order, in this embodiment, the recognition logical boxes are determined as the first logical box one by one, and those connected by directed line segments are determined as the second logical box. For example Figure 3 the serial number of the second logical box corresponding to the first logical box with serial number 1 in [example] is 2, and the serial numbers of the second logical boxes corresponding to the first logical box with serial number 3 are 4 and 7.
[0111] It should be noted that when the first logical box is the first recognition logical box, directly write the target string of the first logical box into the first paragraph. When the first logical box is not the first recognition logical box, relevant target strings have been written when the first logical box is determined as the second logical box. Therefore, the first paragraph can be directly determined without repeated writing.
[0112] It should be noted that when the first logical box is the target shape, such as the rhombus in the above example, write the target strings of the second logical boxes into the second paragraph, and determine the second paragraph as a subordinate paragraph of the first paragraph. When the first logical box is not the target shape, write the target strings of the second logical boxes into the third paragraph and configure it as a sibling paragraph of the first paragraph, so as to determine the logical relationship between the recognition logical boxes using graphic recognition and adjust the paragraph levels between different target strings based on this logical relationship, enabling the AI model to determine the attribution relationship between different paragraph strings according to the paragraph levels.
[0113] Exemplarily, as Figure 3 shown, the serial number of the first logical box is 3, the first paragraph records "CCC", the second logical boxes include the recognition logical boxes corresponding to serial numbers 4 and 7. Therefore, two second paragraphs are constructed under the first paragraph to record "DDDDD" and "FFF" respectively; continue traversing such that the recognition logical box corresponding to serial number 4 is the first logical box, and the serial number of the second logical box is 6. Since it is not a rhombus, a sibling third paragraph is constructed and records "EEEE". The first content obtained based on the above logic can be referred to Figure 3 as shown, and will not be repeated here.
[0114] In addition, in one embodiment, step S30 specifically includes but is not limited to the following steps:
[0115] S31. Determine the corresponding second inference information for each target string in the first inference information;
[0116] S32. When the target string corresponds to any second paragraph and the second inference information does not indicate the target string of the associated first paragraph or second text block, generate inference exception information.
[0117] It should be noted that in this embodiment, the target test cases are obtained through content inference by the AI model. During the inference process, each target string must be analyzed. In this embodiment, the second inference information is queried based on the target string in the first inference information, and exception recognition is performed for each second inference information to realize the self-check of the AI model inference process.
[0118] It should be noted that if the target string corresponds to the first paragraph, its inference process does not depend on the judgment of the previous target string. Therefore, as long as the second inference information is recorded, it can be determined that the inference is normal. The correctness and error of the specific inference need to be judged by the tester based on the first inference information. This embodiment only judges whether the context connection is complete. When the target string corresponds to the second paragraph, its inference process necessarily introduces the target string corresponding to the associated first paragraph. For example Figure 3 As shown, in the second inference information corresponding to the target string "DDDDD", it must be recorded that "when CCC is judged to be yes". If the relevant information is not recorded in the second inference information, it can be determined that the precondition is not introduced in the inference process, and an error may occur in the inference process, and inference exception information is generated to prompt the tester for manual judgment.
[0119] As Figure 4 shown Figure 4 is the structural diagram of a test case generation device based on an AI model provided by an embodiment of the present invention. The present invention also provides a test case generation device based on an AI model, including:
[0120] A processor 401, which can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0121] The memory 402 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 402 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 402 and are called by the processor 401 to execute the method for generating test cases based on the AI model in the embodiments of this application;
[0122] The input / output interface 403 is used to implement information input and output;
[0123] The communication interface 404 is used to implement communication and interaction between this device and other devices. It can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);
[0124] The bus 405 transmits information between various components of the device (such as the processor 401, the memory 402, the input / output interface 403, and the communication interface 404);
[0125] Among them, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404 achieve communication connections with each other inside the device through the bus 405.
[0126] The embodiments of this application also provide a system for generating test cases based on an AI model, including the device for generating test cases based on an AI model as described above.
[0127] The embodiments of this application also provide a storage medium. The storage medium is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned method for generating test cases based on an AI model.
[0128] A memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. The device embodiments described above are merely illustrative, where the units described as separate components may or may not be physically separated, and may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0129] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0130] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. A test case generation method based on an AI model, characterized in that, Applied to a use case generation tool, the use case generation tool includes an OCR module and an AI model for generating test cases, and the method includes: Obtain a requirement document, at least one title keyword, and a business keyword input by a user. When a target title is traversed in the requirement document based on the title keyword, create a first document and a second document based on the target title, where the document types of the first document and the second document are different; When a target image is traversed, extract first content from the target image based on the OCR module, and input the first content into the first document, where any of the business keywords is recorded in the image path of the target image; When a new target title is traversed based on the title keyword, input the first document into the AI model to obtain a target test case and first inference information, write the target test case into the second document, and write the first inference information into the first document, where the first inference information is used to indicate the derivation process of the target test case; Create a new first document and a new second document based on the new target title, and continue to traverse the requirement document; The use case generation tool is preset with an extraction flag, and the initial value of the extraction flag is empty. When a target title is traversed in the requirement document based on the title keyword, creating a first document and a second document based on the target title includes: Set the extraction flag to false and then start traversing. When the extraction flag is false, the traversal keyword is the title keyword; When the target title is traversed based on the traversal keyword, set the extraction flag to true; In response to the extraction flag changing from false to true, create the first document and the second document based on the target title traversed this time, and switch the traversal keyword to the business keyword; When the extraction flag is true, input the image path of the traversed target image into the OCR module, so that the OCR module obtains the target image based on the image path and extracts the first content.
2. The test case generation method based on the AI model according to claim 1, wherein Extracting first content from the target image based on the OCR module and inputting the first content into the first document includes: Create a content queue based on the target title; When the extraction flag is true, input the image path of the traversed target image into the OCR module, so that the OCR module obtains the target image based on the image path and extracts the first content, and write the first content into the content queue; When any target paragraph is traversed, and the target paragraph does not record the target title, and the extraction flag is true, extract the text of the target paragraph as second content and write it into the content queue; When a new target title is traversed based on the title keyword, change the extraction flag from true to false; In response to the extraction identifier changing from true to false, write the content queue to the first document. After completing the writing of the first document, change the extraction identifier from false to true again. After emptying the content queue, continue to traverse the requirements document.
3. The test case generation method based on the AI model according to claim 1, wherein Based on the OCR module, extracting first content from the target image includes: Performing image recognition on the target image through the use case generation tool, and determining a plurality of recognition logic frames from the target image, wherein each of the recognition logic frames is connected to another recognition logic frame by at least one directed line segment; Based on the OCR module, extracting the target string corresponding to each of the recognition logic frames from the target image; Constructing the first content based on the logical relationship between the plurality of target strings and each of the recognition logic frames indicated by the directed line segments.
4. The test case generation method based on the AI model according to claim 3, wherein Based on the OCR module, extracting the target string corresponding to each of the recognition logic frames from the target image includes: Performing line-by-line traversal on the target image based on the OCR module to obtain a plurality of recognition text blocks, and extracting the recognition string corresponding to each of the recognition text blocks, wherein the distance between two adjacent recognition text blocks is greater than a preset threshold; Determining a first text block and a second text block from the plurality of recognition text blocks, and determining the first serial number of each of the recognition text blocks, wherein the number of characters in the first text block is greater than 1, and the number of characters in the second text block is equal to 1; Determining the second serial number of each of the recognition logic frames and the target line segments line by line through the use case generation tool, wherein when the shape of the recognition logic frame is a preset target shape, each of the directed line segments connected to the recognition logic frame is determined as a target line segment, and the recognition logic frame corresponding to the target shape is used to indicate conditional judgment; Based on the first serial number and the second serial number, determining the recognition string of the first text block corresponding to the recognition logic frame as the target string, and determining the second text block associated with each of the target line segments.
5. The test case generation method based on the AI model according to claim 4, wherein Constructing the first content based on the logical relationship between the plurality of target strings and each of the recognition logic frames indicated by the directed line segments includes: Based on any first logic frame, determining a corresponding second logic frame based on the corresponding directed line segment, wherein the first logic frame is connected to the starting point of the directed line segment, and the second logic frame is connected to the ending point of the directed line segment; When the shape of the first logic frame is the target shape, writing the target string of each of the second logic frames to a second paragraph, wherein the second paragraph is a subordinate paragraph of the first paragraph, and the first paragraph records the target string corresponding to the first logic frame; Or, when the shape of the first logic frame is not the target shape, constructing a respective third paragraph based on the target string of each of the second logic frames, wherein the third paragraph is a sibling paragraph of the first paragraph.
6. The test case generation method based on the AI model according to claim 5, wherein, Writing the first inference information to the first document includes: Determine the respective second inference information corresponding to each of the target strings in the first inference information; When the target string corresponds to any one of the second paragraphs and the second inference information does not indicate the target string of the associated first paragraph or the second text block, generate inference exception information.
7. A test case generation device based on an AI model, characterized in that, Comprising at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the AI model-based test case generation method according to any one of claims 1 to 6.
8. A test case generation system based on an AI model, characterized in that, Comprising the AI model-based test case generation device according to claim 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the AI model-based test case generation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Aerial document analysis and test case generation system driven by large model agent
CN117909243A
Apparatus for document structure information extraction and document merging using artificial intelligence
KR102538108B1
Cited By
AI-based test demand analysis and use case generation method and system
CN122045069A
AI-based test requirement analysis and use case generation method and system
CN122045069B