Software development demand analysis system fused with natural language processing
Through a software development requirement analysis system that integrates natural language processing and knowledge graph technology, the existing system is solved by the problem of difficulty in analyzing complex user feedback and insufficient real-time collaboration capabilities, efficient demand analysis and team collaboration are achieved, and project adaptability and competitiveness are improved.
Patent Information
- Application Number
- CN202411875211.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-09
AI Technical Summary
Existing software development requirements analysis systems are difficult to effectively analyze complex user feedback, especially natural language comments on social media, and have limited capabilities in real-time collaboration and multi-source data integration, resulting in low efficiency in team collaboration and decision support.
The software development requirements analysis system that integrates natural language processing and knowledge graph technology includes real-time collaboration module, voice to text module, predictive demand analysis module, demand generation and automation design module, emotional intelligence and user feedback analysis module, and knowledge graph and intelligent decision support module. Through the integration of these modules, automatic extraction and analysis of complex user feedback is realized, and real-time collaboration and multi-source data integration is supported.
It significantly improves the automation and accuracy of software development requirements analysis, enhances team communication and decision-making efficiency, and improves project adaptability and market competitiveness.
Smart Images

Figure CN119960730A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software development requirement analysis, and in particular to a software development requirement analysis system integrating natural language processing. Background Art
[0002] Software development requirements analysis is the process of determining what the software must do and how to implement it. It critically affects the software development cycle, cost, and final quality. As software projects grow in size and complexity, manual requirements analysis becomes increasingly time-consuming and error-prone. The integration of natural language processing can automatically process large amounts of text data, such as requirements documents, user feedback, and forum discussions, helping to quickly and accurately extract and analyze requirements-related information, thereby improving the efficiency and accuracy of requirements analysis.
[0003] Existing software development requirements analysis systems mainly rely on template or rule-driven methods to identify and analyze requirements. These systems usually include functions such as requirements collection, requirements document generation, and requirements verification. They standardize the structure of requirements documents by using predefined templates to make requirements clearer and more consistent. In addition, some systems identify important information in documents through keyword matching and simple text analysis techniques. These methods perform well when processing structured and semi-structured data, and can improve the automation level and accuracy of requirements analysis to a certain extent.
[0004] Although existing software development requirements analysis systems perform well in processing structured and semi-structured data and can improve the automation level and accuracy of requirements analysis to a certain extent, they still have some shortcomings. Existing software development requirements analysis systems often cannot effectively parse complex user feedback, such as natural language comments on social media, and it is difficult to automatically extract and utilize sentiment and semantic information from them. In addition, existing software development requirements analysis systems have limited capabilities in real-time collaboration and multi-source data integration, resulting in low efficiency in team collaboration and decision support. These shortcomings limit the flexibility and efficiency of software development projects in responding to changes in market and user needs. Summary of the invention
[0005] In view of the shortcomings of the existing technology, the present invention provides a software development requirements analysis system that integrates natural language processing. The present invention integrates natural language processing and knowledge graph technology to significantly improve the automation and accuracy of software development requirements analysis. The real-time collaboration module enhances team communication, and the intelligent reasoning and emotional intelligence modules optimize decision-making and quickly respond to user feedback, thereby improving the adaptability and market competitiveness of the project as a whole.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a software development requirements analysis system integrating natural language processing, comprising:
[0007] Real-time collaboration module, suitable for capturing and processing real-time audio and video data;
[0008] The speech-to-text module converts speech data into text using a bidirectional long short-term memory network;
[0009] Predictive demand analysis module, which uses the extreme gradient boosting machine model to predict the impact of demand changes;
[0010] Demand generation and automated design module, which uses conditional generative adversarial networks to generate demand text;
[0011] Emotional intelligence and user feedback analysis module, which uses long short-term memory networks to analyze the emotional tendencies of user feedback;
[0012] The knowledge graph and intelligent decision support module analyzes the relationship between entities based on the Neo4j graph database and page ranking algorithm.
[0013] Preferably, the speech-to-text module uses a bidirectional long short-term memory network, and its specific calculation formula includes:
[0014] Forward long short-term memory network calculation formula:
[0015]
[0016] Backward long short-term memory network calculation formula:
[0017]
[0018] Forward and backward hidden state merging formula:
[0019]
[0020] Among them, x t is the input feature vector at time step t, h t is the output feature vector at time step t.
[0021] Preferably, the predictive demand analysis module adopts an extreme gradient boosting machine model, whose objective function and parameters include:
[0022] Objective function:
[0023]
[0024] in, is the square error loss function, γ and λ are regularization parameters, T k is the number of trees in the model, w k is the leaf weight of the tree.
[0025] Preferably, the demand generation and automated design module uses a conditional generative adversarial network, and the objective functions of its generator and discriminator are:
[0026]
[0027] Among them, D(x|y) represents the probability that the discriminator D judges that the data x under given condition y is real data, and G(z|y) represents the ability of the generator G to generate data from noise z under condition y.
[0028] Preferably, the emotional intelligence and user feedback analysis module uses a long short-term memory network, and its key calculation formula includes:
[0029] The calculation formula of the forget gate is:
[0030] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0031] Among them, f t is the activation value of the forget gate, W f is the weight matrix of the forget gate, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b f is the bias vector of the forget gate, σ is the S-type activation function;
[0032] The calculation formula of the input gate is:
[0033] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0034] Among them, i t is the activation value of the input gate, W i is the weight matrix of the input gate, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b i is the bias vector of the input gate, σ is the S-type activation function;
[0035] The calculation formula of the output gate is:
[0036] o t =σ(W o ·[h t-1 ,x t ]+b o )
[0037] Among them, t is the activation value of the output gate, W o is the weight matrix of the output gate, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b o is the bias vector of the output gate, σ is the S-type activation function;
[0038] The unit state update formula is:
[0039]
[0040] in, is the candidate cell state at the current moment, W C is the weight matrix of the unit state, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b C is the bias vector of the unit state, and tanh is the hyperbolic tangent activation function;
[0041] The update formula of the unit state is:
[0042]
[0043] Among them, C t is the unit state at the current moment, f t is the activation value of the forget gate, C t-1 is the unit state at the previous moment, i t is the activation value of the input gate, is the candidate cell state at the current moment, * indicates the element-level multiplication operation;
[0044] The output update formula is:
[0045] h t =o t *tanh(C t )
[0046] Among them, h t is the hidden state vector at the current moment, o t The activation value of the output gate, C t is the cell state at the current moment, tanh is the hyperbolic tangent activation function, and * represents the element-level multiplication operation.
[0047] Preferably, the knowledge graph and intelligent decision support module use a page ranking algorithm to analyze the relationship between entities, and the calculation formula is:
[0048]
[0049] Among them, PR(u) is the page ranking algorithm value of node u, B u is the set of nodes pointing to node u, L(v) is the number of links pointed by node v, and d is the damping factor, which is set to 0.85.
[0050] Preferably, the real-time collaboration module further comprises:
[0051] A real-time video acquisition submodule, used to capture video streams through the user device camera;
[0052] A real-time audio acquisition submodule, used to capture audio streams through the user device microphone;
[0053] The data transmission submodule transmits audio and video data through the network protocol.
[0054] Preferably, the predictive demand analysis module further comprises:
[0055] Data preprocessing submodule, used to clean and convert historical demand change data;
[0056] A feature extraction submodule is used to extract feature vectors from preprocessed data;
[0057] The model training submodule is used to train the extreme gradient boosting machine model based on the extracted feature vectors.
[0058] Preferably, the emotional intelligence and user feedback analysis module further includes:
[0059] The data collection submodule is used to collect user feedback data from social media, customer service systems and other channels;
[0060] The text processing submodule is used to perform preprocessing on the user feedback text, such as word segmentation and stop word removal;
[0061] The sentiment classification submodule is used to perform sentiment classification based on the preprocessed text data.
[0062] Preferably, the knowledge graph and intelligent decision support module further includes:
[0063] The knowledge extraction submodule is used to automatically extract entities and their relationships from the requirements document;
[0064] Graph database management submodule, used to store and manage the extracted knowledge graph data;
[0065] Intelligent reasoning submodule, used for complex query and reasoning based on knowledge graph.
[0066] The present invention provides a software development requirements analysis system integrating natural language processing. It has the following beneficial effects:
[0067] 1. The present invention integrates natural language processing and knowledge graph technology to automatically extract and analyze key information in software development requirement documents and construct a dynamic view of entities and their relationships. This not only improves the level of automation in information processing and reduces the need for manual intervention, but also provides data-driven decision support through the intelligent reasoning submodule, significantly improving the speed and accuracy of decision-making.
[0068] 2. The present invention introduces a real-time collaboration module, including real-time capture and transmission of video and audio, so that distributed team members can communicate effectively and provide instant feedback. This not only enables smooth communication during remote work, but also improves the dynamic response capability of project development and the collaborative efficiency among team members through instant information exchange and problem solving.
[0069] 3. The present invention uses emotional intelligence and user feedback analysis modules to capture and analyze user emotions and opinions from various channels in real time, ensuring that user voices are heard and understood. It not only overcomes the lag and one-sidedness of traditional methods in processing user feedback, but also can predict changes in user needs, provide a scientific basis for product iteration and service improvement, and enhance customer satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a system structure diagram of the present invention. DETAILED DESCRIPTION
[0071] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0072] Please refer to the attached Figure 1 The embodiment of the present invention provides a software development requirements analysis system integrating natural language processing, including:
[0073] The real-time collaboration module is suitable for capturing and processing real-time audio and video data. This module realizes the capture and processing of audio and video data by integrating real-time communication technology. The real-time collaboration module enables remote teams to conduct seamless audio and video communication, greatly improving communication efficiency and collaboration effects. This is especially important for distributed software development teams, which can discuss and solve development problems in real time and reduce the risk of project delays;
[0074] The speech-to-text module uses a bidirectional long short-term memory network to convert speech data into text. This module uses a bidirectional long short-term memory network to convert captured speech data into text. By adding a reverse transfer path to the traditional long short-term memory network, the bidirectional long short-term memory network can simultaneously consider the preceding and following information of the speech data, thereby more accurately capturing the contextual characteristics of the language. Through more accurate speech-to-text processing, the meeting content and oral needs can be automatically recorded, reducing the workload and error rate of manual transcription;
[0075] The predictive demand analysis module uses the extreme gradient boosting model to predict the impact of demand changes. It improves the accuracy and stability of predictions by building multiple decision trees and combining their prediction results. The algorithm also includes regularization terms to avoid overfitting and optimize the generalization ability of the model. By accurately predicting the impact of demand changes on the project, the project team can better manage risks and make adjustment plans in advance. This not only improves the flexibility and adaptability of project management, but also effectively controls project costs and schedules.
[0076] The requirements generation and automated design module uses conditional generative adversarial networks to generate requirements texts. Conditional generative adversarial networks are used to automatically generate requirements texts that meet specific conditions. The generator tries to create texts that are sufficient to deceive the discriminator, while the discriminator tries to distinguish the generated texts from the real texts. In this way, the generator continuously learns and improves to generate texts that better meet actual requirements, which improves the efficiency of preparing requirements documents, especially in development environments where requirements change frequently or require rapid iterations.
[0077] The emotional intelligence and user feedback analysis module uses the long short-term memory network to analyze the emotional tendency of user feedback. The long short-term memory network can effectively process and memorize long-term dependent information through its unique gating mechanism, and is suitable for processing sequence data in emotional analysis. This enables the system to accurately extract emotional attitudes from user feedback and assist in decision-making. Through emotional analysis, the system can timely capture user satisfaction and demand changes, and provide data support for product iteration and service improvement. This helps to improve user satisfaction;
[0078] The knowledge graph and intelligent decision support module analyzes the relationship between entities based on the Neo4j graph database and page ranking algorithm. By constructing a knowledge graph, it represents and analyzes information such as demand, users, and market trends in the form of a graph, and uses the page ranking algorithm to evaluate the influence and importance of each entity. The knowledge graph and intelligent decision support module provides the project team with an intuitive and powerful tool for mining potential connections and insights between data. This in-depth analysis helps the team better understand the complex factors behind the demand and optimize the decision-making process, thereby improving the project success rate and the ability to meet user needs.
[0079] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the speech-to-text module uses a bidirectional long short-term memory network, and its specific calculation formula includes:
[0080] Forward long short-term memory network calculation formula:
[0081]
[0082] Backward long short-term memory network calculation formula:
[0083]
[0084] Forward and backward hidden state merging formula:
[0085]
[0086] Among them, x t is the input feature vector at time step t, h t is the output feature vector at time step t;
[0087] The traditional LSTM network is extended by using a bidirectional LSTM network, which can capture context information more comprehensively by processing the forward and reverse sequences of data simultaneously. The input feature vector x at each time step is t It is processed by both forward and backward LSTM networks, which integrates the preceding and succeeding information at each point in the time series. By capturing contextual information, the bidirectional LSTM network shows higher recognition accuracy when processing speech data with complex background noise and different accents. Automatic transcription of meetings and discussions reduces the need to manually record requirements, improves the efficiency and completeness of requirements capture, and enables development teams to respond to and implement these requirements more quickly.
[0088] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the predictive demand analysis module adopts an extreme gradient boosting machine model, whose objective function and parameters include:
[0089] Objective function:
[0090]
[0091] in, is the square error loss function, γ and λ are regularization parameters, T k is the number of trees in the model, w kThe leaf weights of the tree. The Extreme Gradient Boosting Machine is an optimized distributed gradient boosting library that can improve execution speed and efficiency through engineering optimization while retaining the advantages of the gradient boosting machine algorithm. By using the Extreme Gradient Boosting Machine algorithm, the system can accurately predict the potential impact of demand changes on the project. Because the Extreme Gradient Boosting Machine model uses the learning results of multiple trees in combination, it usually has better prediction performance and accuracy than a single decision tree or other basic models. By accurately predicting the impact of demand changes, project managers can allocate resources more reasonably, avoid resource waste, and avoid or prepare for potential risks in advance, thereby optimizing the execution efficiency and success rate of the entire project. Predictive demand analysis provides decision makers with data-supported insights to help them make more informed decisions based on quantitative prediction results.
[0092] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the demand generation and automated design module uses a conditional generative adversarial network, and the objective functions of its generator and discriminator are:
[0093]
[0094] Among them, D(x|y) represents the probability that the discriminator D judges that the data x under the given condition y is the real data, and G(z|y) represents the ability of the generator G to generate data from the noise z under the condition y;
[0095] Model building steps:
[0096] Data preparation: First, collect and prepare training data, which includes requirement text and its related conditional labels, such as requirement type, priority, etc.
[0097] Generator construction: The generator G accepts random noise z and condition y, and generates data to try to imitate the real demand text. The goal of the generator is to generate data that is sufficient to deceive the discriminator.
[0098] Discriminator construction: The discriminator D receives real data or generated data and judges the authenticity of the data given the condition y. The goal of the discriminator is to correctly distinguish between real data and generated data.
[0099] Training process:
[0100] Alternately train the generator and discriminator, the generator tries to minimize log(1-D(G(z|y))), while the discriminator tries to maximize logD(x|y);
[0101] Through this adversarial training, the generator continuously learns how to improve the data it generates to better fool the discriminator, while the discriminator learns how to recognize the generated data more effectively.
[0102] By taking conditions into account during the generation process, cGANs can generate highly customized requirement texts that directly correspond to specific project requirements or user preferences, improving the relevance and practicality of the requirements. Automated requirement generation reduces the time and labor cost of manually writing requirement documents, which is particularly useful in development environments where requirements change frequently. It speeds up the requirements collation and approval process and improves the responsiveness and flexibility of the development team.
[0103] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the emotional intelligence and user feedback analysis module uses a long short-term memory network, and its key calculation formula includes:
[0104] The calculation formula of the forget gate is:
[0105] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0106] Among them, f t is the activation value of the forget gate, W f is the weight matrix of the forget gate, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b f is the bias vector of the forget gate, σ is the S-type activation function, and the forget gate is used to decide which information is discarded or retained;
[0107] The calculation formula of the input gate is:
[0108] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0109] Among them, i t is the activation value of the input gate, W i is the weight matrix of the input gate, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b i is the bias vector of the input gate, σ is the S-type activation function, and the input gate determines which part of the new information is worth updating to the unit state;
[0110] The calculation formula of the output gate is:
[0111] o t =σ(W o ·[h t-1,x t ]+b o )
[0112] Among them, t is the activation value of the output gate, W o is the weight matrix of the output gate, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b o is the bias vector of the output gate, σ is the S-type activation function, and the output gate determines which part of the unit state will be output to the outside;
[0113] The unit state update formula is:
[0114]
[0115] in, is the candidate cell state at the current moment, W C is the weight matrix of the unit state, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b C is the bias vector of the unit state, and tanh is the hyperbolic tangent activation function;
[0116] The update formula of the unit state is:
[0117]
[0118] Among them, C t is the unit state at the current moment, f t is the activation value of the forget gate, C t-1 is the unit state at the previous moment, i t is the activation value of the input gate, is the candidate cell state at the current moment, * indicates the element-level multiplication operation;
[0119] The output update formula is:
[0120] h t =o t *tanh(C t )
[0121] Among them, h t is the hidden state vector at the current moment, o t The activation value of the output gate, C tis the cell state at the current moment, tanh is the hyperbolic tangent activation function, and * represents the element-level multiplication operation. The long short-term memory network is a special type of recurrent neural network, which is particularly suitable for processing and predicting important events with long intervals and delays in sequence data. Since the long short-term memory network can effectively process long sequence data and remember long-term dependent information, it can more accurately identify and classify user emotions and feelings when analyzing complex user feedback or text data. Accurate sentiment analysis enables the development team to gain insight into changes in user emotions and adjust products or services in time to better meet user needs.
[0122] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the knowledge graph and intelligent decision support module use a page ranking algorithm to analyze the relationship between entities, and the calculation formula is:
[0123]
[0124] Among them, PR(u) is the page ranking algorithm value of node u, B u is the set of nodes pointing to node u, L(v) is the number of links pointed to by node v, d is the damping factor, which is set to 0.85. The page ranking algorithm can quantify each entity, helping project managers and team members understand the priority and criticality of each requirement or task. This method provides data-driven support for resource allocation and task scheduling, and enhances the accuracy and effectiveness of decision-making. By identifying key requirements and central tasks, the team can prioritize the most critical parts for project success, thereby promoting project progress more efficiently, reducing unnecessary work and iterations, and not only improving the team's collaborative efficiency, but also promoting knowledge sharing and reuse, which is especially important in large or long-term projects.
[0125] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the real-time collaboration module further comprises:
[0126] The real-time video acquisition submodule is used to capture the video stream through the camera of the user device. The real-time video acquisition submodule uses the camera on the computer or smart phone to capture the video stream. It involves accessing the media hardware interface of the device, such as using the user media interface obtained by the web real-time communication protocol, which allows web applications to directly access the camera and microphone of the device. By capturing the real-time video stream, this submodule can provide a face-to-face communication experience for remote team meetings or virtual collaboration, thereby enhancing the effectiveness of communication and reducing misunderstandings and information loss, while improving the dynamics and participation of collaboration;
[0127] The real-time audio acquisition submodule is used to capture the audio stream through the microphone of the user device. The real-time audio acquisition submodule captures the audio stream through the microphone of the user device. Similar to video acquisition, audio acquisition also uses the web real-time communication protocol to obtain the user media interface to ensure the real-time capture and processing of audio signals. Real-time audio acquisition makes remote communication smoother and more natural, especially when dealing with multi-party calls and large virtual meetings. High-quality audio capture reduces communication barriers and improves the efficiency of meetings;
[0128] The data transmission submodule transmits audio and video data through the network protocol. The data transmission submodule is responsible for transmitting the captured audio and video data to other users or servers through the network protocol. The effective data transmission mechanism ensures the real-time and high-quality audio and video data, no matter how geographically dispersed the users are. In addition, efficient data transmission also reduces network delays and freezes, providing a smoother and more stable communication experience;
[0129] Through the collaborative work of these three sub-modules, the real-time collaboration module not only improves the efficiency and interactivity of remote team collaboration, but also improves the quality of communication between team members through high-quality audio and video transmission, thereby helping to enhance the collaboration effect and productivity of the entire project team.
[0130] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the predictive demand analysis module further comprises:
[0131] The data preprocessing submodule is used to clean and convert historical demand change data. The data preprocessing submodule is responsible for converting the original historical demand change data into a format suitable for analysis and modeling;
[0132] The data preprocessing submodule includes the following steps:
[0133] Data cleaning: Identify and address outliers, duplicate records, or incomplete entries in the data.
[0134] Missing value handling: Apply different strategies to fill or ignore missing values, such as using the mean, median, or inferring missing values from related data.
[0135] Data standardization: Normalize the data to ensure that data features in different ranges have a balanced impact on the model.
[0136] Through thorough data preprocessing, the interference of abnormal data on the model is reduced, and the performance of the model can be optimized so that it can be better generalized to new data, significantly improving the data quality of model training, thereby increasing the accuracy and reliability of the model.
[0137] The feature extraction submodule is used to extract feature vectors from preprocessed data, identify which data attributes have the greatest impact on predicting demand changes, and build new features that can enhance the model's predictive capabilities, such as by combining existing data fields or applying mathematical transformations. Through effective feature extraction, the predictive model can more accurately capture the key factors that affect demand changes, thereby providing more accurate prediction results;
[0138] A model training submodule for training an extreme gradient boosting machine model based on the extracted feature vectors;
[0139] The following steps are involved:
[0140] Model training: Use techniques such as cross-validation to evaluate the performance of the model on different subsets of data to avoid overfitting;
[0141] Parameter tuning: Adjust the learning rate of the extreme gradient boosting machine, the depth of the tree, the number of trees, etc. to find the best model configuration.
[0142] Through precise tuning and training, the model training submodule is able to produce a highly optimized prediction model that can accurately predict the potential impact of requirement changes on software development projects.
[0143] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the emotional intelligence and user feedback analysis module further includes:
[0144] The data collection submodule is used to collect user feedback data from social media, customer service systems and other channels. The data collection submodule is responsible for collecting user feedback data from various channels, including social media platforms, customer service systems, forums and comment areas. This module uses API calls, web crawlers or integrated data streaming services to automatically collect text data. In addition, it also involves cooperation with third-party data providers to obtain more comprehensive user feedback data. By systematically collecting and aggregating user feedback from multiple channels, the data collection submodule provides a rich data resource. The ability to access a large amount of real-time data enables the system to respond quickly to market changes and provide support for data-based decision-making;
[0145] The text processing submodule is used to perform preprocessing on user feedback text, such as word segmentation and stop word removal. The purpose of preprocessing is to reduce data dimension and noise, and improve the efficiency and accuracy of subsequent analysis. The text processing submodule lays the foundation for efficient and accurate sentiment analysis by standardizing and simplifying input data. Processed text is easier to analyze and interpret, which helps improve the accuracy of sentiment classification, making sentiment analysis results more reliable and useful;
[0146] The sentiment classification submodule is used to perform sentiment classification based on preprocessed text data. The sentiment classification submodule uses natural language processing technology and machine learning algorithms to analyze the preprocessed text and determine its sentiment tendency. Accurate sentiment analysis results can help companies better understand user needs, optimize products and services, and thus enhance user satisfaction.
[0147] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the knowledge graph and intelligent decision support module further includes:
[0148] The knowledge extraction submodule is used to automatically extract entities and their relationships from the requirements document;
[0149] The knowledge extraction submodule is responsible for automatically extracting key entities and their relationships from software development requirement documents. This is achieved through natural language processing technology, including text analysis, entity recognition, relationship extraction and other steps;
[0150] Text analysis: Perform grammatical and semantic analysis on documents to identify vocabulary and sentence structure.
[0151] Entity recognition: Use named entity recognition technology to identify key entities in the document, such as functions, users, tasks, etc.
[0152] Relation extraction: Analyze the semantic relationships between entities and build associations between these entities.
[0153] By automatically extracting key information from requirement documents and building entity relationships, the knowledge extraction submodule not only saves a lot of time in manual document processing and analysis, but also improves information utilization and accuracy.
[0154] The graph database management submodule is used to store and manage the extracted knowledge graph data. By using Neo4j, the ability to store a large number of nodes and edges and efficiently query these data is optimized. Entities and relationships can be stored as nodes and edges in the graph. Each node and edge can contain multiple attributes. The graph database management submodule makes the data structure of the entire system more intuitive and easy to operate by efficiently storing and processing complex entity relationship networks, thereby improving the speed and flexibility of data query, thereby accelerating the process of decision support and data analysis;
[0155] Intelligent reasoning submodule, used for complex query and reasoning based on knowledge graph.
[0156] The intelligent reasoning submodule includes the following steps:
[0157] Logical reasoning: Using graph query and reasoning algorithms, such as shortest path, graph traversal, and pattern matching, to discover non-obvious connections between data;
[0158] Decision support: Provides decision support based on real-time data analysis, such as impact analysis, demand fulfillment strategies, and resource optimization recommendations.
[0159] In-depth analysis and understanding of large data sets can help project teams optimize resource allocation and develop more effective project strategies. In addition, the in-depth insights provided by intelligent reasoning can also help improve product quality and customer satisfaction.
[0160] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A software development requirements analysis system integrating natural language processing, characterized by: include: Real-time collaboration module, suitable for capturing and processing real-time audio and video data; The speech-to-text module converts speech data into text using a bidirectional long short-term memory network; Predictive demand analysis module, which uses the extreme gradient boosting machine model to predict the impact of demand changes; Demand generation and automated design module, which uses conditional generative adversarial networks to generate demand text; Emotional intelligence and user feedback analysis module, which uses long short-term memory networks to analyze the emotional tendencies of user feedback; The knowledge graph and intelligent decision support module analyzes the relationship between entities based on the Neo4j graph database and page ranking algorithm.
2. The software development requirements analysis system integrating natural language processing according to claim 1 is characterized in that: The speech-to-text module uses a bidirectional long short-term memory network, and its specific calculation formula includes: Forward long short-term memory network calculation formula: Backward long short-term memory network calculation formula: Forward and backward hidden state merging formula: Among them, x t is the input feature vector at time step t, h t is the output feature vector at time step t.
3. The software development requirements analysis system integrating natural language processing according to claim 1 is characterized in that: The predictive demand analysis module adopts an extreme gradient boosting machine model, whose objective function and parameters include: Objective function: in, is the square error loss function, γ and λ are regularization parameters, T k is the number of trees in the model, w k is the leaf weight of the tree.
4. The software development requirements analysis system integrating natural language processing according to claim 1 is characterized in that: The demand generation and automated design module uses a conditional generative adversarial network, and the objective functions of its generator and discriminator are: Among them, D(x|y) represents the probability that the discriminator D judges that the data x under the given condition y is the real data, and G(z|y) represents the ability of the generator G to generate data from the noise z under the condition y.
5. The software development requirements analysis system integrating natural language processing according to claim 1 is characterized in that: The emotional intelligence and user feedback analysis module uses a long short-term memory network, and its key calculation formula includes: The calculation formula of the forget gate is: f t =σ(W f ·[h t-1 ,x t ]+b f ) Among them, f t is the activation value of the forget gate, W f is the weight matrix of the forget gate, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b f is the bias vector of the forget gate, σ is the S-type activation function; The calculation formula of the input gate is: i t =σ(W i ·[h t-1 ,x t ]+b i ) Among them, i t is the activation value of the input gate, W i is the weight matrix of the input gate, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b i is the bias vector of the input gate, σ is the S-type activation function; The calculation formula of the output gate is: the t =σ(W o ·[h t-1 ,x t ]+b o ) Among them, t is the activation value of the output gate, W o is the weight matrix of the output gate, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b o is the bias vector of the output gate, σ is the S-type activation function; The unit status update formula is: in, is the candidate cell state at the current moment, W C is the weight matrix of the unit state, h t-1 is the hidden state vector at the previous moment, x t is the input vector at the current moment, b C is the bias vector of the unit state, and tanh is the hyperbolic tangent activation function; The update formula of the unit state is: Among them, C t is the unit state at the current moment, f t is the activation value of the forget gate, C t-1 is the unit state at the previous moment, i t is the activation value of the input gate, is the candidate cell state at the current moment, * indicates the element-level multiplication operation; The output update formula is: h t =o t *tanh(C t ) Among them, h t is the hidden state vector at the current moment, o t The activation value of the output gate, C t is the cell state at the current moment, tanh is the hyperbolic tangent activation function, and * represents the element-level multiplication operation.
6. The software development requirements analysis system integrating natural language processing according to claim 1 is characterized in that: The knowledge graph and intelligent decision support module uses a page ranking algorithm to analyze the relationship between entities, and its calculation formula is: Among them, PR(u) is the page ranking algorithm value of node u, B u is the set of nodes pointing to node u, L(v) is the number of links pointed by node v, and d is the damping factor, which is set to 0.
85.
7. The software development requirements analysis system integrating natural language processing according to claim 1 is characterized in that: The real-time collaboration module further comprises: A real-time video acquisition submodule, used to capture video streams through the user device camera; A real-time audio acquisition submodule, used to capture audio streams through the user device microphone; The data transmission submodule transmits audio and video data through network protocols.
8. The software development requirements analysis system integrating natural language processing according to claim 1 is characterized in that: The predictive demand analysis module further includes: Data preprocessing submodule, used to clean and convert historical demand change data; A feature extraction submodule is used to extract feature vectors from preprocessed data; The model training submodule is used to train the extreme gradient boosting machine model based on the extracted feature vectors.
9. The software development requirements analysis system integrating natural language processing according to claim 1 is characterized in that: The emotional intelligence and user feedback analysis module further includes: The data collection submodule is used to collect user feedback data from social media, customer service systems and other channels; The text processing submodule is used to perform preprocessing on the user feedback text, such as word segmentation and stop word removal; The sentiment classification submodule is used to perform sentiment classification based on the preprocessed text data.
10. The software development requirements analysis system integrating natural language processing according to claim 1, characterized in that: The knowledge graph and intelligent decision support module further includes: The knowledge extraction submodule is used to automatically extract entities and their relationships from the requirements document; The graph database management submodule is used to store and manage the extracted knowledge graph data; the intelligent reasoning submodule is used to perform complex queries and reasoning based on the knowledge graph.
Citation Information
Cited By
Demand processing method and system based on natural language processing and data mining
CN120743229A