Financial Requirement Item Generation Method and Device
Through word segmentation and clustering methods based on financial vocabulary collection and natural language processing technology, financial demand items are automatically generated, which solves the problem of low efficiency in generating financial demand items, realizes personalized and efficient demand item generation, and improves the optimization efficiency and reliability of financial software products.
Patent Information
- Application Number
- CN202110891478.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-04
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-08-04
AI Technical Summary
The generation efficiency of financial demand items in the existing technology is low, and it cannot adapt to the rapid optimization and launch of financial software products on the basis of ensuring personalization. The labor cost is high and the degree of automation is insufficient.
Based on preset financial vocabulary collection and natural language processing technology, financial demand items are automatically extracted and generated through word segmentation, clustering and template generation methods, including the use of conditional random field CRF algorithm, FudanNLP toolkit and LP clustering algorithm, reducing the data volume and improving the accuracy and automation of clustering processing.
It improves the personalization, accuracy and reliability of financial demand item generation, enhances the degree of automation, improves the online efficiency and reliability of financial software products, and improves the user experience of developers.
Smart Images

Figure CN113535125B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, particularly to the financial technology field, and specifically to a method and device for generating financial requirement items. Background Art
[0002] In the Internet era, in order to quickly adapt to business requirements, financial institutions such as banks have gradually transformed the project development process from the traditional method to agile development to meet the market requirements for rapid product launch. Each project consists of multiple requirement items, and each requirement item is jointly developed and implemented by one or more product applications. During the agile iteration process of the project, a large number of usability problems that do not affect the basic functions of the product usually occur in the testing process, and these problems are not required to be forced to be optimized within an extremely short iteration time. After the product is promoted, in order to facilitate project management and close the loop of the original requirements in the development link, it is usually artificially written into the requirement items of the next round of iteration projects again to further optimize the product and continuously update it. This method weakens the value of the testing problems found in the early stage and also seriously consumes the labor cost of the requirement items output in the iteration project.
[0003] In terms of the output of financial requirement items, the current requirements can be roughly divided into business transformation requirements and technical transformation requirements. The current common method is as follows: for financial business transformation requirements, the requester collects the problems feedback by customers during the use of the product through online and offline forms, and summarizes them in combination with the previous testing problems, and manually writes the requirement items. For technical transformation requirements, the requester solicits the opinions of developers and testers, and manually writes the technical transformation requirements in combination with the actual situation of financial production. However, no matter which type it is, due to the extremely high requirements for labor cost in each step, there are problems such as too low efficiency in generating financial requirement items, so it is impossible to fully adapt to the agile process of rapid product optimization and launch while ensuring the personalization of financial requirement items. Summary of the Invention
[0004] Aiming at the problems in the prior art, this application provides a method and device for generating financial requirement items, which can effectively improve the personalization, accuracy and reliability of generating financial requirement items, and can be specifically applicable to the financial industry, and can effectively improve the automation degree and efficiency of the financial requirement item generation process, thereby improving the efficiency and reliability of launching and optimizing financial software products according to financial requirement items.
[0005] To solve the above technical problems, this application provides the following technical solutions:
[0006] In a first aspect, this application provides a method for generating financial requirement items, including:
[0007] Extract keywords for generating financial requirement items from the tokenized problem dataset based on a preset financial vocabulary set, so as to obtain a target problem dataset composed of each of the keywords, wherein the problem dataset includes test questions in multiple iterative processes corresponding to each development project;
[0008] Cluster each test question in the target problem dataset to obtain a corresponding set of key problem words;
[0009] Input the set of key problem words into a preset text template for financial requirement items to generate corresponding financial requirement items.
[0010] Further, before extracting keywords for generating financial requirement items from the tokenized problem dataset based on a preset financial vocabulary set, it further includes:
[0011] Obtain test questions in multiple iterative processes corresponding to each development project and generate a corresponding problem dataset;
[0012] Preprocess the problem dataset;
[0013] Perform word segmentation on the preprocessed problem dataset based on a preset conditional random field (CRF) algorithm to obtain a tokenized problem dataset.
[0014] Further, the preprocessing of the problem dataset includes:
[0015] Clean the data of the problem dataset;
[0016] Format the problem dataset after data cleaning so that the problem dataset contains the correspondence between the unique identifier of each test question and the test question content, wherein the test question content is divided by each attribute, and each attribute includes: project name, business type, business scenario, problem description, and involved application.
[0017] Further, the extraction of keywords for generating financial requirement items from the tokenized problem dataset based on a preset financial vocabulary set includes:
[0018] Retrieve a preset financial vocabulary set;
[0019] Extract financial vocabulary in the financial vocabulary set for generating financial requirement items from the tokenized problem dataset;
[0020] Use a preset FudanNLP toolkit to label the part-of-speech of each of the financial vocabulary and extract the financial vocabulary with the part-of-speech of noun and verb as keywords to form a target problem dataset composed of each of the keywords.
[0021] Further, before clustering each test question in the target question dataset to obtain a corresponding set of key question words, it further includes:
[0022] Based on a preset professional vocabulary division rule, the keywords corresponding to each test question in the target question dataset are divided into professional nouns, non-professional nouns, professional verbs, and non-professional verbs;
[0023] In the order of decreasing weight values, weight assignments are respectively performed on the keywords belonging to professional nouns, non-professional nouns, professional verbs, and non-professional verbs.
[0024] Further, the clustering of each test question in the target question dataset to obtain a corresponding set of key question words further includes:
[0025] Based on a preset LP clustering algorithm, each test question in the target question dataset after keyword weight assignment is clustered to obtain a corresponding set of key question words.
[0026] Further, it further includes:
[0027] Output the financial requirement item.
[0028] In a second aspect, the present application provides a financial requirement item generation device, including:
[0029] A data extraction module, configured to extract keywords for generating financial requirement items from a tokenized question dataset based on a preset financial vocabulary set, so as to obtain a target question dataset composed of each of the keywords, where the question dataset includes test questions in multiple iteration processes corresponding to each development project; a data clustering module, configured to cluster each test question in the target question dataset to obtain a corresponding set of key question words; a template generation module, configured to input the set of key question words into a preset financial requirement item text template to generate a corresponding financial requirement item.
[0030] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the financial requirement item generation method described above.
[0031] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the financial requirement item generation method described above.
[0032] As can be seen from the above technical solutions, a method and device for generating financial requirement items provided by this application, the method includes: extracting keywords for generating financial requirement items from a tokenized problem dataset based on a preset financial vocabulary set to obtain a target problem dataset composed of each of the keywords, where the problem dataset includes test problems in multiple iterative processes corresponding to each development project; clustering each test problem in the target problem dataset to obtain a corresponding set of key problem words; inputting the set of key problem words into a preset financial requirement item text template to generate corresponding financial requirement items. By extracting keywords for generating financial requirement items from a tokenized problem dataset based on a preset financial vocabulary set, on the basis of effectively reducing the data volume of the tokenized problem dataset, the target problem dataset can be specifically applicable to the financial industry, and thus the reliability, accuracy, and applicability of subsequent clustering processing on the target problem dataset can be effectively improved; by clustering each test problem in the target problem dataset to obtain a corresponding set of key problem words, the accuracy, automation degree, and intelligence degree of finding similar test problems in the target problem dataset can be effectively improved; by inputting the set of key problem words into a preset financial requirement item text template, the personalization, accuracy, and reliability of generating financial requirement items can be effectively improved, and the automation degree and efficiency of the financial requirement item generation process can be effectively improved, thereby improving the efficiency and reliability of launching and optimizing financial software products according to financial requirement items, and effectively improving the user experience of financial software product developers. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0034] Figure 1 It is a schematic diagram of the relationship between the financial requirement item generation device and the client device in the embodiments of this application.
[0035] Figure 2 It is the first flowchart of the financial requirement item generation method in the embodiments of this application.
[0036] Figure 3 It is the second flowchart of the financial requirement item generation method in the embodiments of this application.
[0037] Figure 4 It is the third flowchart of the financial requirement item generation method in the embodiments of this application.
[0038] Figure 5 It is the fourth process schematic diagram of the financial demand item generation method in the embodiments of the present application.
[0039] Figure 6 It is the fifth process schematic diagram of the financial demand item generation method in the embodiments of the present application.
[0040] Figure 7 It is the sixth process schematic diagram of the financial demand item generation method in the embodiments of the present application.
[0041] Figure 8 It is the seventh process schematic diagram of the financial demand item generation method in the embodiments of the present application.
[0042] Figure 9 It is the structural schematic diagram of the financial demand item generation device in the embodiments of the present application.
[0043] Figure 10 It is the structural schematic diagram of the financial demand item generation system provided by the application example of the present application.
[0044] Figure 11 It is the structural schematic diagram of the data preprocessing device provided by the application example of the present application.
[0045] Figure 12 It is the structural schematic diagram of the clustering analysis device provided by the application example of the present application.
[0046] Figure 13 It is the structural schematic diagram of the demand item generation device provided by the application example of the present application.
[0047] Figure 14 It is the structural schematic diagram of the electronic device in the embodiments of the present application. Detailed implementation manners
[0048] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0049] It should be noted that the financial demand item generation methods and devices disclosed in the present application can be used in the financial technical field and can also be used in any field other than the financial technical field. The application fields of the financial demand item generation methods and devices disclosed in the present application are not limited.
[0050] The existing financial demand item generation method has the problem that it is impossible to meet the financial demand item generation efficiency requirements on the basis of ensuring the personalization of financial demand items, that is, in the process of demand item output, there is no method that can automatically generate demand items without manpower. How to find a solution that can automatically summarize and generate demand items from known problems is a technical problem that needs to be solved urgently in this field. The embodiment of the present application provides a financial demand item generation method, which extracts keywords for generating financial demand items from a segmented problem data set based on a preset financial vocabulary set to obtain a target problem data set composed of each of the keywords, wherein the problem data set contains test questions in multiple iterations corresponding to each development project; clusters each test question in the target problem data set to obtain a corresponding key question word set; inputs the key question word set into a preset financial demand item text template to generate corresponding financial demand items, and The financial vocabulary set extracts keywords for generating financial demand items from the segmented question data set, which can effectively reduce the data volume of the segmented question data set, so that the target question data set can be specifically applicable to the financial industry, thereby effectively improving the reliability, accuracy and applicability of subsequent clustering processing of the target question data set; by clustering each test question in the target question data set to obtain the corresponding key question word set, the accuracy, automation and intelligence of finding similar test questions in the target question data set can be effectively improved; by inputting the key question word set into a preset financial demand item text template, the personalization, accuracy and reliability of generating financial demand items can be effectively improved, and the automation and efficiency of the financial demand item generation process can be effectively improved, thereby improving the efficiency and reliability of launching and optimizing financial software products according to financial demand items, and effectively improving the user experience of financial software product developers.
[0051] In one or more embodiments of the present application, the Conditional Random Fields (CRF) algorithm is a conditional probability distribution model of another set of output sequences under a given set of input sequences, and is widely used in natural language processing.
[0052] In one or more embodiments of the present application, the FudanNLP toolkit is a toolkit developed for Chinese natural language processing, and also includes machine learning algorithms and data sets for implementing these tasks, including Chinese word segmentation, part-of-speech tagging, named entity recognition, dependency syntax analysis, keyword extraction, time phrase recognition, text classification, news clustering, hierarchical classification, online learning and other functions.
[0053] In one or more embodiments of the present application, the LP (Layer-Partition) clustering algorithm is based on the idea of partitioning and hierarchical clustering. Each time it calculates the cluster distance, it depends on the previous calculation result to find the current optimal solution, avoiding comparing the similarities between all clusters and improving the overall clustering speed.
[0054] Based on the above content, the present application also provides a financial demand item generation device for implementing the financial demand item generation method provided in one or more embodiments of the present application. The financial demand item generation device can be a server. Refer to Figure 1 , the financial demand item generation device can communicate with each client device in sequence by itself or through a third-party server, etc. The financial demand item generation device can receive a financial demand item generation request sent by the client device, extract keywords for generating financial demand items from the segmented problem data set based on a preset financial vocabulary set, so as to obtain a target problem data set composed of each of the keywords, where the problem data set includes test problems in multiple iterative processes corresponding to each development project; cluster each test problem in the target problem data set to obtain a corresponding set of key problem words; input the set of key problem words into a preset financial demand item text template to generate corresponding financial demand items. By extracting keywords for generating financial demand items from the segmented problem data set based on a preset financial vocabulary set, the financial demand item generation device can also send each financial demand item to a preset display device for display, or send a notification message including the specific content of each financial demand item to the client device of developers, etc.
[0055] In another practical application scenario, the part of the foregoing financial demand item generation device for generating financial demand items can be executed in a server as described above, or all operations can be completed in the user terminal device. Specifically, it can be selected according to the processing capacity of the user terminal device and the limitations of the user usage scenario, etc. The present application does not make any limitations in this regard. If all operations are completed in the user terminal device, the user terminal device may further include a processor for specific processing of financial demand item generation.
[0056] It can be understood that the mobile terminal may include any mobile device capable of loading applications, such as a smart phone, a tablet electronic device, a network set-top box, a portable computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, smart watches, smart bracelets, etc.
[0057] The above-mentioned mobile terminal may have a communication module (i.e., a communication unit), which can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side. In other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.
[0058] Any suitable network protocol can be used for communication between the above-mentioned server and the mobile terminal, including network protocols that have not been developed as of the filing date of this application. The network protocol may, for example, include TCP / IP protocol, UDP / IP protocol, HTTP protocol, HTTPS protocol, etc. Of course, the network protocol may also, for example, include RPC protocol (Remote Procedure Call Protocol) and REST protocol (Representational State Transfer) used on top of the above-mentioned protocols.
[0059] Specifically, detailed descriptions will be given respectively through the following various embodiments and application examples.
[0060] To solve the problems that the existing financial requirement item generation methods cannot meet the requirements of financial requirement item generation efficiency on the basis of ensuring the personalization of financial requirement items, etc., an embodiment of a financial requirement item generation method is provided in this application. See Figure 2 , the financial requirement item generation method executed by the financial requirement item generation device specifically includes the following contents:
[0061] Step 100: Extract keywords for generating financial requirement items from the segmented problem data set based on a preset financial vocabulary set to obtain a target problem data set composed of each of the keywords, where the problem data set includes test problems in multiple iteration processes corresponding to each development project.
[0062] In step 100, the financial vocabulary set can be pre-stored locally in the financial requirement item generation device or in a database accessible to the financial requirement item generation device. The financial vocabulary set is used to store various financial vocabularies such as pre-set banking professional vocabularies. Specifically, it can be pre-input by the user into the local financial requirement item generation device or a database accessible to the financial requirement item generation device for storage. In one or more embodiments of the present application, the problem data set containing test problems in multiple iteration processes corresponding to each development project means that the problem data set is used to store the correspondence between the unique identifiers of test problems in multiple iteration processes corresponding to each development project and the test problem contents.
[0063] Before the problem data set is subjected to word segmentation processing, the test problem content corresponding to the unique identifier of each test problem stored in the problem data set is the complete test problem content; after the problem data set is subjected to word segmentation processing, the test problem content corresponding to the unique identifier of each test problem stored in the problem data set is the content composed of each divided (for example, separated by punctuation between each vocabulary) vocabulary.
[0064] After extracting keywords from the problem data set, the test problem content corresponding to the unique identifier of each test problem stored in the problem data set is the content composed of each financial vocabulary (i.e., keyword). And at this time, the problem data set is confirmed as the target problem data set that needs to be subjected to the clustering process in step 200. That is to say, this target problem data set is still used to store the correspondence between the unique identifiers of test problems in multiple iteration processes corresponding to each development project and the test problem contents. However, at this time, the test problem content is no longer the complete test problem content or each vocabulary after word segmentation, but each financial vocabulary involved in the original complete test problem content.
[0065] Step 200: Cluster each test problem in the target problem data set to obtain the corresponding set of key problem words.
[0066] In step 200, clustering each test problem in the target problem data set means clustering each test problem composed of each financial vocabulary in the target problem data set to improve the accuracy of finding similar test problems in the target problem data set.
[0067] In one or more embodiments of the present application, the set of key problem words is also used to store the correspondence between the unique identifiers of test problems in multiple iteration processes corresponding to each development project and the content of the test problems. However, the content of the test problems is composed of various financial terms (i.e., key words), and the number of unique identifiers of the test problems is significantly less than the number of unique identifiers of the test problems in the target problem dataset because, in step 200, similar test problems are clustered, resulting in a corresponding reduction in the total number of unique identifiers of the corresponding test problems.
[0068] Step 300: Input the set of key problem words into a preset financial requirement item text template to generate corresponding financial requirement items.
[0069] In step 300, a financial requirement item text template can be set, and it supports manual modification by the requester. By assembling the set of key problem words, the corresponding requirement item description is output to complete the output of the complete requirement item.
[0070] As can be seen from the above description, the financial requirement item generation method provided by the embodiments of the present application extracts the key words for generating financial requirement items from the tokenized problem dataset based on a preset set of financial terms. On the basis of effectively reducing the data volume of the tokenized problem dataset, the target problem dataset can be specifically applicable to the financial industry, thereby effectively improving the reliability, accuracy, and applicability of subsequent clustering processing of the target problem dataset; by clustering each test problem in the target problem dataset to obtain the corresponding set of key problem words, the accuracy, automation degree, and intelligence level of finding similar test problems in the target problem dataset can be effectively improved; by inputting the set of key problem words into a preset financial requirement item text template, the personalization, accuracy, and reliability of generating financial requirement items can be effectively improved, and the automation degree and efficiency of the financial requirement item generation process can be effectively improved, thereby improving the efficiency and reliability of launching and optimizing financial software products according to financial requirement items, and effectively improving the user experience of financial software product developers.
[0071] In order to make the target problem dataset specifically applicable to the financial industry, in an embodiment of the financial requirement item generation method provided by the present application, refer to Figure 3 Before step 100 of the financial requirement item generation method, the following specific content is further included:
[0072] Step 010: Obtain the test problems in multiple iteration processes corresponding to each development project and generate a corresponding problem dataset.
[0073] In step 010, test problems in multiple iterations of multiple projects are imported, and then usability problems with the status of uncompleted modification in the iteration process are screened out to form a cleaned data source to be processed, that is, an initial problem data set.
[0074] Step 020: Preprocess the problem data set.
[0075] Step 030: Perform word segmentation on the preprocessed problem data set based on the preset conditional random field (CRF) algorithm to obtain a word-segmented problem data set.
[0076] Specifically, use the CRF algorithm for basic Chinese word segmentation, set a value set {B, E, M, S} to calculate the annotation probability between words, set the vector F(x, y) and the weight vector w, the observation sequence x, and set the recurrence function Calculate and output the y set as the optimal path output, where Identify professional words in combination with banking vocabulary, and screen out keywords that can be used to generate banking requirement items.
[0077] As can be seen from the above description, the financial requirement item generation method provided by the embodiment of the present application can provide an accurate and effective data basis for extracting keywords for generating financial requirement items in the word-segmented problem data set by performing word segmentation on the preprocessed problem data set based on the preset conditional random field (CRF) algorithm. Furthermore, on the basis of effectively reducing the data volume of the word-segmented problem data set, the target problem data set can be specifically applicable to the financial industry, thereby effectively improving the reliability, accuracy, and applicability of subsequent clustering processing of the target problem data set.
[0078] In order to improve the effectiveness and efficiency of data preprocessing, in an embodiment of the financial requirement item generation method provided by the present application, see Figure 4 , the specific content of step 020 of the financial requirement item generation method includes the following:
[0079] Step 021: Clean the problem data set.
[0080] Step 022: Format the problem data set after data cleaning so that the problem data set contains the corresponding relationship between the unique identifier of each test problem and the test problem content. Among them, the test problem content is divided by each attribute, and each attribute includes: project name, business type, business scenario, problem description, and involved application.
[0081] Specifically, the cleaned data is constructed according to attributes such as project name, business type, business scenario, problem description, and involved applications, and then the data after attribute construction is output in a regularized manner and transformed into a modeling data source that can be recognized by modeling.
[0082] As can be seen from the above description, the financial demand item generation method provided by the embodiments of the present application formats the problem data set after data cleaning, so that the corresponding relationship between the unique identifier of each test problem and the test problem content is included in the problem data set, which can effectively improve the effectiveness and efficiency of data preprocessing, and further can effectively improve the efficiency and reliability of tokenizing the preprocessed problem data set.
[0083] In order to improve the effectiveness and application reliability of keywords in the target problem data set, in an embodiment of the financial demand item generation method provided by the present application, see Figure 5 , step 100 of the financial demand item generation method specifically includes the following contents:
[0084] Step 110: Retrieve a preset financial vocabulary set.
[0085] Step 120: Extract the financial vocabulary in the financial vocabulary set for generating financial demand items from the tokenized problem data set.
[0086] Step 130: Use a preset FudanNLP toolkit to label the part-of-speech of each financial vocabulary, and extract the financial vocabulary with the part-of-speech of noun and verb as keywords to form a target problem data set composed of each keyword.
[0087] Specifically, first use the CRF algorithm for basic Chinese word segmentation. This algorithm is based on a sequence labeling model and can well handle problems such as Chinese ambiguity generated when describing test problems, combine with banking vocabulary to identify professional vocabulary, screen out keywords that can be used to generate banking demand items, and use the FudanNLP toolkit to label the part-of-speech.
[0088] As can be seen from the above description, the financial demand item generation method provided by the embodiments of the present application can effectively reduce the data volume of the target problem data set by using a preset FudanNLP toolkit to label the part-of-speech of each financial vocabulary and extracting the financial vocabulary with the part-of-speech of noun and verb as keywords, improve the effectiveness and application reliability of keywords in the target problem data set, and further can improve the personalization, accuracy and reliability of generating financial demand items, and can be specifically applicable to the financial industry, and can effectively improve the automation degree and efficiency of the financial demand item generation process.
[0089] To improve the reliability and accuracy of clustering each test question in the target problem dataset, in an embodiment of the financial demand item generation method provided in this application, refer to Figure 6 Between step 100 and step 200 of the financial demand item generation method, the following specific content is further included:
[0090] Step 140: Based on a preset professional vocabulary division rule, divide the keywords corresponding to each test question in the target problem dataset into professional nouns, non-professional nouns, professional verbs, and non-professional verbs.
[0091] Step 150: Assign weight values to the keywords belonging to professional nouns, non-professional nouns, professional verbs, and non-professional verbs in descending order of weight value.
[0092] Specifically, retain the professional nouns, non-professional nouns, professional verbs, and non-professional verbs in each problem description to construct a keyword set; assume the number of terms in the keyword set is n, and a certain keyword is t i , then each test question consists of multiple keywords. A question S can be represented as S(t 1, t 2,..., t n ) using the VSM method to construct a vector space model. If t i appears, it is recorded as 1; if it does not appear, it is recorded as 0, so that each question forms an n-dimensional space vector. However, if only the calculation method of 0 and 1 is used for vector calculation, it is difficult to distinguish the similarity in detail. Therefore, weight values w1 to w4 are set in descending order for professional nouns, non-professional nouns, professional verbs, and non-professional verbs, with default values of 100, 50, 10, and 5 respectively, to construct the text vector S corresponding to the test question set S c = S(w 11 , w 21 , w 32 , w 43 ,...... w nn ).
[0093] From the above description, it can be seen that in the financial demand item generation method provided in the embodiment of this application, by dividing the keywords corresponding to each test question in the target problem dataset into professional nouns, non-professional nouns, professional verbs, and non-professional verbs, and assigning weight values to the keywords belonging to professional nouns, non-professional nouns, professional verbs, and non-professional verbs respectively, it can provide an effective and reliable data basis for clustering each test question in the target problem dataset, and thus can effectively improve the reliability and accuracy of clustering each test question in the target problem dataset.
[0094] In order to improve the effectiveness of the critical problem word set, in an embodiment of the financial demand item generation method provided in this application, refer to Figure 7 , step 200 of the financial demand item generation method specifically includes the following content:
[0095] Step 210: Based on a preset LP clustering algorithm, cluster each test problem in the target problem dataset after keyword weight assignment to obtain a corresponding critical problem word set.
[0096] Specifically, use the Layer-Partition (abbreviated as LP) clustering algorithm to construct an algorithm model. This clustering algorithm inherits the idea of partition-based and hierarchical clustering. Each time the cluster distance is calculated, it depends on the previous calculation result to find the current optimal solution, avoiding comparing the similarities between all clusters, and improving the overall clustering speed. Set the test problem set S = {S1, S2,..., S m}, calculate the similarity Sim(S i , S j ). The similarity calculation uses the cosine theorem. Set the distance threshold as α. Initially, each S i is used as a single cluster T i . Arbitrarily select a T i , and calculate the distance between each S j and T i in turn. If the distance is less than α, then classify S j into T i until all the remaining S j are greater than α. Then, select the T j that is least similar to the clustering starting point as the starting point, and repeat the distance calculation step until all clusters have participated in the clustering. Set each professional term as the maximum keyword and output the clustering result around the maximum keyword.
[0097] As can be seen from the above description, in the financial demand item generation method provided in the embodiment of this application, by clustering each test problem in the target problem dataset after keyword weight assignment based on a preset LP clustering algorithm, the effectiveness of the critical problem word set can be effectively improved, and the accuracy, automation degree, and intelligence degree of finding similar test problems in the target problem dataset can be effectively improved.
[0098] In order to improve the convenience and efficiency for users such as R & D personnel to obtain the automatically generated financial demand items, in an embodiment of the financial demand item generation method provided in this application, refer to Figure 8 , after step 300 of the financial demand item generation method, it specifically further includes the following content:
[0099] Step 400: Output the financial demand item.
[0100] Specifically, each financial requirement item generated in step 300 can be sent to a preset display device for display, or a notification message containing the specific content of each financial requirement item can be sent to a client device such as a developer, etc.
[0101] As can be seen from the above description, the financial requirement item generation method provided by the embodiments of the present application can effectively improve the convenience and efficiency for users such as R & D personnel to obtain the automatically generated financial requirement items by outputting the financial requirement items, can further improve the user experience of users such as R & D personnel, and further improve the efficiency and reliability of launching and optimizing financial software products according to financial requirement items, effectively improving the user experience of developers of financial software products.
[0102] At the software level, in order to solve the problems that the existing financial requirement item generation methods cannot meet the requirements of financial requirement item generation efficiency on the basis of ensuring the personalization of financial requirement items, etc., the embodiments of the present application provide an embodiment of a financial requirement item generation device for executing all or part of the content in the financial requirement item generation method, see Figure 9 , the financial requirement item generation device specifically includes the following:
[0103] A data extraction module 10, configured to extract keywords for generating financial requirement items from a tokenized problem data set based on a preset financial vocabulary set, so as to obtain a target problem data set composed of each of the keywords, where the problem data set includes test problems in multiple iteration processes corresponding to each development project;
[0104] A data clustering module 20, configured to cluster each test problem in the target problem data set to obtain a corresponding set of key problem words;
[0105] A template generation module 30, configured to input the set of key problem words into a preset financial requirement item text template to generate corresponding financial requirement items.
[0106] The embodiments of the financial requirement item generation device provided by the present application can specifically be used to execute the processing flow of the embodiments of the financial requirement item generation method in the above embodiments, and its functions will not be elaborated here, and reference can be made to the detailed description of the above method embodiments.
[0107] As can be seen from the above description, the financial requirement item generation device provided by the embodiments of the present application can extract keywords for generating financial requirement items from the tokenized question dataset based on a preset financial vocabulary set. On the basis of effectively reducing the data volume of the tokenized question dataset, the target question dataset can be specifically applicable to the financial industry, thereby effectively improving the reliability, accuracy, and applicability of subsequent clustering processing of the target question dataset. By clustering each test question in the target question dataset to obtain the corresponding key question word set, the accuracy, automation level, and intelligence level of finding similar test questions in the target question dataset can be effectively improved. By inputting the key question word set into a preset financial requirement item text template, the personalization, accuracy, and reliability of generating financial requirement items can be effectively improved, and the automation level and efficiency of the financial requirement item generation process can be effectively improved. Furthermore, the efficiency and reliability of launching and optimizing financial software products according to financial requirement items can be improved, and the user experience of financial software product developers can be effectively improved.
[0108] To further illustrate the present solution, the application example of the present application provides a method for generating financial requirement items based on a test question set and the LP clustering algorithm, which relates to the field of product requirements. To solve the problem that there is no method for automatically generating requirement items without human effort in the process of outputting requirement items in the field of product requirements, the application example of the present application provides a method for constructing an LP clustering analysis model based on natural language processing based on a test question dataset, summarizing and discovering requirement keywords in the same business scenario, constructing a personalized requirement item generation model, and finally generating requirement items by designing a template for keyword input, aiming to liberate the human cost invested in outputting requirement items during the project iteration process.
[0109] A method for automatically generating requirement items based on a test question set provided by the application example of the present application mainly includes the following steps:
[0110] Step 1): A data preprocessing device preprocesses the test question set, parses it into a standardized question format including project name, business type, business scenario, question description, and involved application, and provides the new dataset to the clustering analysis device.
[0111] Step 2): Clustering analysis device. The device uses the word segmentation algorithm of the conditional random field (CRF) algorithm as the basic word segmentation algorithm, combines it with banking professional vocabulary to form a special test question set word segmentation tool, and uses the FudanNLP toolkit to annotate the part of speech. Compared with the traditional word segmentation method, this method can identify banking professional nouns and provide part-of-speech annotation functions. Through part-of-speech annotation, it is convenient to organize the word order of the keywords finally generated by the requirement items. The LP clustering algorithm is used to construct a clustering analysis model to find the key question word sets of similar test question scenarios. This clustering algorithm inherits the idea of partition-based and hierarchical clustering, avoids comparing the similarities between all clusters, can improve the overall clustering speed, and is convenient for providing keywords to the requirement item generation device.
[0112] Step 3): Requirement item generation device. Set a special template for requirement items, support manual modification by the requirer, output the corresponding requirement item description by assembling the key question word sets, and complete the output of the complete requirement items.
[0113] See Figure 10 , the financial requirement item generation system provided by the application example of this application for implementing the financial requirement item generation method specifically includes the following:
[0114] Data preprocessing device, clustering analysis device and requirement item generation device.
[0115] Among them, the data preprocessing device is connected to the clustering analysis device; the clustering analysis device is connected to the requirement item generation device.
[0116] (1) Data preprocessing device: Used to clean and process the original test question data set, including test questions in multiple iterative processes of multiple projects, screen out usability problems, and clean and construct attributes for the data. The attributes include project name, business type, business scenario, problem description, and involved applications, etc. After reconstruction, it is provided to the clustering analysis device.
[0117] (2) Clustering analysis device: Used to input the preprocessed data source, construct a model using a clustering algorithm, and output the key question word sets corresponding to various business scenarios. First, use the CRF algorithm for basic Chinese word segmentation. This algorithm is based on a sequence labeling model and can handle problems such as Chinese ambiguity generated when describing test questions well. Combine it with banking vocabulary to identify professional vocabulary, screen out keywords that can be used to generate banking requirement items, and use the FudanNLP toolkit to annotate the part of speech. On the basis of performing Chinese word segmentation and annotating the part of speech of nouns, verbs, etc., it is also necessary to increase the word segmentation weights of banking professional vocabulary and nouns, use the LP clustering algorithm to construct a clustering analysis model to find the key question word sets of similar test question scenarios, and provide them to the requirement item generation device.
[0118] (3) Requirement Item Generation Device: It is used to take the keyword set obtained by clustering analysis as the input data source, set the requirement item text template to automatically fill in the keyword set to form a complete requirement item, support manual modification and saving, and finally output the complete requirement item generated based on the test question set.
[0119] See Figure 11 , the data preprocessing device 1 specifically includes the following:
[0120] Data acquisition unit 11, data cleaning unit 12, attribute construction unit 13, and data transformation unit 14.
[0121] (1) Data acquisition unit 11: It is used to import test questions in multiple iterative processes of multiple projects.
[0122] (2) Data cleaning unit 12: It screens out the usability problems with the status of uncompleted modification in the iterative process and forms the data source to be processed after cleaning.
[0123] (3) Attribute construction unit 13: It is used to construct the cleaned data according to attributes such as project name, business type, business scenario, problem description, and involved application.
[0124] (4) Data transformation unit 14: It is used to perform regular output on the data after attribute construction and transform it into a modeling data source that can be recognized by modeling.
[0125] Among them, the examples of the test questions are shown in Table 1:
[0126] Table 1 Example Table of Test Questions
[0127]
[0128] See Figure 12 , the clustering analysis device 2 specifically includes the following:
[0129] Word segmentation unit 21, data screening unit 22, weight calculation unit 23, clustering calculation unit 24, and result output unit 25.
[0130] (1) Word segmentation unit 21: It uses the CRF algorithm for basic Chinese word segmentation, sets the value set {B, E, M, S} for calculating the annotation probability between words, sets the vector F(y, x) and the weight vector w, the observation sequence x, and sets the recurrence function Calculate and output the y set as the optimal path output, where Identify professional terms by combining banking vocabulary, screen out keywords that can be used to generate banking requirements, and use the FudanNLP toolkit to label parts of speech. On the basis of Chinese word segmentation and labeling parts of speech such as nouns and verbs, it is also necessary to filter out useless words such as special symbols, adjectives, and auxiliary words to facilitate dimensionality reduction of the word vector model, and the order of the keywords finally generated by the requirements items can be conveniently organized through part-of-speech labeling.
[0131] (2) Data screening unit 22: Retain professional nouns, non-professional nouns, professional verbs, and non-professional verbs in each problem description to construct a keyword set.
[0132] (3) Weight calculation unit 23: Assume that the number of terms in the keyword set is n, and a certain keyword is t i , then each test question is composed of multiple keywords. A question S can be represented as S(t 1, t 2,..., t n ) using the VSM method to construct a vector space model. If t i appears, it is recorded as 1, and if it does not appear, it is recorded as 0, so that each question forms an n-dimensional space vector. However, if only 0 and 1 are used in vector calculation, it is difficult to distinguish the similarity in detail. Therefore, weight values w1 to w4 are set in descending order according to professional nouns, non-professional nouns, professional verbs, and non-professional verbs, with default values of 100, 50, 10, and 5 respectively, to construct the text vector S corresponding to the test question set S c = S(w 11 , w 21 , w 32 , w 43 ,...... w nn ).
[0133] (4) Clustering calculation unit 24: Use the Layer-Partition (abbreviated as LP) clustering algorithm to construct an algorithm model. This clustering algorithm inherits the idea of partition-based and hierarchical clustering. Each time the cluster distance is calculated, it depends on the previous calculation result to find the current optimal solution, avoiding comparing the similarity between all clusters and improving the overall clustering speed. Set the test question set S = {S1, S2,..., S m} and calculate the similarity Sim(S i , S j ) between each question. The similarity calculation uses the cosine theorem, and the distance threshold is set to α. Initially, each S i is used as a single cluster T i . Arbitrarily select a T i , and calculate the distance between each S j and T i in turn. If the distance is less than α, then S jClassified as T i until all the remaining S j is greater than α, and then select the previous T that is least similar to the clustering starting point j as the starting point, repeat the distance calculation step until all clusters have participated in the clustering. Set each professional term as the maximum keyword and output the clustering results around the maximum keyword.
[0134] (5) Result output unit 25: Used to send the clustering results to the text template.
[0135] See Figure 13 , the requirement item generation device 3 specifically includes the following:
[0136] Text template setting unit 31, text layout unit 32, personalized setting unit 33, and requirement item publishing unit 34.
[0137] (1) Text template setting unit 31: Preset multiple text templates for generating requirement items, and can be filled and output according to the reserved templates when the requirement items are published;
[0138] (2) Text layout unit 32: Typeset the generated requirement items, indent two spaces at the beginning, and add a line break character at the end to perform a simple beautification function;
[0139] (3) Personalized setting unit 33: Supports functions such as manually entering new templates, modifying existing text templates, and deleting templates, and supports modification and saving after the system automatically generates requirement items.
[0140] (4) Requirement item publishing unit 34: Supports generating a complete set of requirement items according to keywords. Tables 1 and 2 show the comparison results of the original problem and the finally generated requirement items. After clustering, the requirement items are automatically filled and generated using the template "'Channel class', 'Noun', initiate 'Scenario class' 'Verb', 'Noun' needs to support 'Verb'".
[0141] Among them, the example of the output requirement items can be seen in Table 2:
[0142] Table 2 Example Table of Output Requirement Items
[0143]
[0144] The financial requirement item generation method provided by the application example of this application, based on the test problem set generated during the project iteration process, constructs a clustering model based on natural language processing technology, discovers the keywords in the business scenario and problem description, and automatically summarizes and generates the product requirements for the next iteration project. Its advantages are as follows:
[0145] Based on the original data of the test question set, it is possible to save manpower from collecting the discovered product defects again, enhance the utilization value of usability issues in test questions, and play a greater role throughout the project lifecycle.
[0146] Through the clustering algorithm, it is possible to self-learn and discover the keywords in the questions, cluster and mine the scenarios involved in the keywords, greatly saving the time cost of collecting and summarizing similar questions manually.
[0147] The automatic generation of requirement items relying on text templates can save the time for the requester to write requirement items, and personalized tools can be used to generate templates to maintain the diversity of different requirement generations.
[0148] From the hardware level, in order to solve the problems that the existing financial requirement item generation methods cannot meet the requirements of financial requirement item generation efficiency on the basis of ensuring the personalization of financial requirement items, etc., this application provides an embodiment of an electronic device for implementing all or part of the content in the financial requirement item generation method. The electronic device specifically includes the following content:
[0149] Figure 14 It is a schematic block diagram of the system composition of the electronic device 9600 according to an embodiment of this application. As Figure 14 shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It should be noted that this Figure 14 is exemplary; other types of structures can also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0150] In one embodiment, the financial requirement item generation function can be integrated into the central processing unit. Among them, the central processing unit can be configured to perform the following controls:
[0151] Step 100: Extract keywords for generating financial requirement items from the segmented question data set based on a preset financial vocabulary set to obtain a target question data set composed of each of the keywords, where the question data set includes test questions in multiple iteration processes corresponding to each development project.
[0152] Step 200: Cluster each test question in the target question data set to obtain a corresponding set of key question words.
[0153] Step 300: Input the set of key question words into a preset financial requirement item text template to generate corresponding financial requirement items.
[0154] As can be seen from the above description, the electronic device provided by the embodiments of the present application extracts keywords for generating financial requirement items from the tokenized question dataset based on a preset financial vocabulary set, which can effectively reduce the data volume of the tokenized question dataset and make the target question dataset specifically applicable to the financial industry. Furthermore, it can effectively improve the reliability, accuracy, and applicability of subsequent clustering processing of the target question dataset; by clustering each test question in the target question dataset to obtain a corresponding set of key question words, the accuracy, automation level, and intelligence level of finding similar test questions in the target question dataset can be effectively improved; by inputting the set of key question words into a preset financial requirement item text template, the personalization, accuracy, and reliability of generating financial requirement items can be effectively improved, and the automation level and efficiency of the financial requirement item generation process can be effectively improved. Furthermore, the efficiency and reliability of launching and optimizing financial software products according to financial requirement items can be improved, and the user experience of financial software product developers can be effectively improved.
[0155] In another embodiment, the financial requirement item generation device can be separately configured from the central processing unit 9100. For example, the financial requirement item generation device can be configured as a chip connected to the central processing unit 9100 to implement the financial requirement item generation function through the control of the central processing unit.
[0156] As Figure 14 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include Figure 14 all the components shown in Figure 14 ; in addition, the electronic device 9600 may further include
[0157] components not shown in Figure 14 ; reference can be made to the prior art.
[0158] Among them, the memory 9140 can be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. It can store the above information related to failures, and can also store programs for executing relevant information. And the central processing unit 9100 can execute the program stored in the memory 9140 to implement information storage or processing, etc.
[0159] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.
[0160] The memory 9140 may be a solid-state memory, for example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be a memory that stores information even when power is off, can be selectively erased and has more data. An example of this memory is sometimes referred to as an EPROM, etc. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 that is used to store application programs and function programs or the processes for operating the electronic device 9600 by the central processing unit 9100.
[0161] The memory 9140 may also include a data storage unit 9143 that is used to store data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers of the electronic device for communication functions and / or for performing other functions of the electronic device (such as a messaging application, an address book application, etc.).
[0162] The communication module 9110 is a transmitter / receiver 9110 that transmits and receives signals via the antenna 9111. The communication module (transmitter / receiver) 9110 is coupled to the central processing unit 9100 to provide input signals and receive output signals, which may be the same as in the case of a conventional mobile communication terminal.
[0163] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module (transmitter / receiver) 9110 is also coupled to the speaker 9131 and the microphone 9132 via the audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby implementing normal telecommunication functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 9130 is also coupled to the central processing unit 9100, so that it is possible to record on the local machine through the microphone 9132 and play the sound stored on the local machine through the speaker 9131.
[0164] Embodiments of the present application further provide a computer-readable storage medium capable of implementing all steps in the financial requirement item generation method in the above embodiments. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, all steps of the financial requirement item generation method with the execution entity being a server or a client in the above embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented:
[0165] Step 100: Extract keywords for generating financial requirement items from the tokenized problem data set based on a preset financial vocabulary set to obtain a target problem data set composed of each of the keywords, where the problem data set includes test problems in multiple iterative processes corresponding to each development project.
[0166] Step 200: Cluster each test problem in the target problem data set to obtain a corresponding set of key problem words.
[0167] Step 300: Input the set of key problem words into a preset financial requirement item text template to generate corresponding financial requirement items.
[0168] As can be seen from the above description, the computer-readable storage medium provided by the embodiments of the present application can extract keywords for generating financial requirement items from the tokenized problem data set based on a preset financial vocabulary set, which can effectively reduce the data volume of the tokenized problem data set and make the target problem data set specifically applicable to the financial industry. Furthermore, it can effectively improve the reliability, accuracy, and applicability of subsequent clustering processing of the target problem data set; by clustering each test problem in the target problem data set to obtain a corresponding set of key problem words, it can effectively improve the accuracy, automation degree, and intelligence level of finding similar test problems in the target problem data set; by inputting the set of key problem words into a preset financial requirement item text template, it can effectively improve the personalization, accuracy, and reliability of generating financial requirement items, and can effectively improve the automation degree and efficiency of the financial requirement item generation process, thereby improving the efficiency and reliability of launching and optimizing financial software products according to financial requirement items, and effectively improving the user experience of financial software product developers.
[0169] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0170] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (devices), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0171] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0173] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for generating financial demand items, characterized in that, Including: Extracting keywords for generating financial requirement items from the segmented problem dataset based on a preset financial vocabulary set to obtain a target problem dataset composed of each of the keywords, where the problem dataset contains test problems in multiple iterative processes corresponding to each development project; the problem dataset is used to store the correspondence between the unique identifier of the test problems in multiple iterative processes corresponding to each development project and the content of the test problems. Clustering each of the test problems in the target problem dataset to obtain a corresponding set of key problem words. Inputting the set of key problem words into a preset text template for financial requirement items to generate corresponding financial requirement items. Wherein, before extracting keywords for generating financial requirement items from the segmented problem dataset based on a preset financial vocabulary set, it further includes: Obtaining test problems in multiple iterative processes corresponding to each development project and generating a corresponding problem dataset. Preprocessing the problem dataset. Performing word segmentation on the preprocessed problem dataset based on a preset conditional random field (CRF) algorithm to obtain a segmented problem dataset. Wherein, the preprocessing of the problem dataset includes: Performing data cleaning on the problem dataset. Formatting the problem dataset after data cleaning so that the problem dataset contains the correspondence between the unique identifier of each test problem and the content of the test problem, where the content of the test problem is divided by each attribute, and each attribute includes: project name, business type, business scenario, problem description, and involved application.
2. The method for generating financial demand items according to claim 1, wherein The extracting keywords for generating financial requirement items from the segmented problem dataset based on a preset financial vocabulary set includes: Retrieving a preset financial vocabulary set. Extracting financial vocabulary in the financial vocabulary set for generating financial requirement items from the segmented problem dataset. Using a preset FudanNLP toolkit to annotate the part-of-speech of each of the financial vocabulary and extracting the financial vocabulary with the part-of-speech of noun and verb as keywords to form a target problem dataset composed of each of the keywords.
3. The method for generating financial demand items according to claim 2, wherein Before clustering each of the test problems in the target problem dataset to obtain a corresponding set of key problem words, it further includes: Dividing the keywords corresponding to each of the test problems in the target problem dataset into professional nouns, non-professional nouns, professional verbs, and non-professional verbs based on a preset professional vocabulary division rule. Assigning weight values to the keywords belonging to professional nouns, non-professional nouns, professional verbs, and non-professional verbs in descending order of weight values.
4. The method for generating financial demand items according to claim 3, wherein, The clustering each of the test problems in the target problem dataset to obtain a corresponding set of key problem words further includes: Clustering each of the test problems in the target problem dataset after keyword weight assignment based on a preset LP clustering algorithm to obtain a corresponding set of key problem words.
5. The method for generating financial demand items according to any one of claims 1 to 4, characterized in that It further includes: Outputting the financial requirement items.
6. A financial demand item generation device, characterized in that Including: A data extraction module, which is used to extract keywords for generating financial requirement items from the tokenized problem dataset based on a preset financial vocabulary set, so as to obtain a target problem dataset composed of each of the keywords, wherein the problem dataset contains test questions in multiple iteration processes corresponding to each development project; the problem dataset is used to store the corresponding relationship between the unique identifier and the test question content of the test questions in multiple iteration processes corresponding to each development project; a data clustering module, which is used to cluster each test question in the target problem dataset to obtain a corresponding set of key problem words; a template generation module, which is used to input the set of key problem words into a preset text template of financial requirement items to generate corresponding financial requirement items. Wherein, before extracting keywords for generating financial requirement items from the tokenized problem dataset based on a preset financial vocabulary set, it further includes: Obtaining test questions in multiple iteration processes corresponding to each development project and generating a corresponding problem dataset; Preprocessing the problem dataset; Performing a tokenization process on the preprocessed problem dataset based on a preset conditional random field (CRF) algorithm to obtain a tokenized problem dataset; Wherein, the preprocessing of the problem dataset includes: Performing data cleaning on the problem dataset; Formatting the problem dataset after data cleaning so that the problem dataset contains the corresponding relationship between the unique identifier and the test question content of each test question, wherein the test question content is divided by each attribute, and each attribute includes: project name, business type, business scenario, problem description, and involved application.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the financial requirement item generation method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the financial requirement item generation method according to any one of claims 1 to 5.
9. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the financial requirement item generation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and terminal for automatically generating ideas based on knowledge network
CN106940726A
OCR-technology-based text automatic generation method and device, equipment and medium
CN111782772A